
Visual Intelligence Hackathon (Real Time Video Anomaly Detection)
Description
A car stopped in a parking lot? Normal.The same car stopped on a highway? Now you probably want to know.That tiny difference is exactly where most computer vision systems start struggling.On 5th September, we’re bringing together AI/ML engineers, researchers, students, and multimodal AI builders for a full day to work on one problem:Can we teach a small vision language model to spot what actually matters in live drone video, in real time?And no, this isn’t a “prompt something cool and demo it at 6 PM” kind of hackathon.You’ll be working with real drone footage, limited GPU capability, and the kind of messy visual context that makes this problem genuinely hard.The ChallengeA drone flying over a city might see thousands of perfectly ordinary moments.And then there’s the one that matters.A vehicle has broken down on a highway. Or traffic is slowly building up. Or smoke appears where there shouldn’t be any or anything unusual.The difficult part is that an object itself usually isn’t the anomaly. The context is.Traditional object detectors can tell you what they see. The challenge is building something that can understand whether what it sees is unusual enough to require attention.Vision-language models can reason about that context, but large models are too slow and expensive to continuously run across live video feeds.So we’re making the problem harder:Make it work with a small model. Make it work in real time. And make it economical enough that it could eventually run across many drone feeds at once.What You’ll Spend the Day DoingWe’re not dropping a problem statement at 9 AM and leaving you alone with Stack Overflow.The morning starts with a state-of-the-art session on video anomaly detection, related work across robotics and computer vision, and demos from the FlytBase team.Then we build.You might fine-tune a small vision-language model, distill a larger model, build a lightweight detection + verification pipeline, implement recent research, or try something none of us thought of.The approach is open.The constraint is what makes it interesting.What You’ll GetThe complete problem statement, evaluation criteria, and submission format will be revealed on the day.Who Should Join?This is for you if you’re working with, learning, or seriously curious about:Computer Vision · Multimodal AI · Vision-Language Models · Video Understanding · Model Fine-Tuning · Distillation · Efficient Inference · Anomaly Detection · Open-Set RecognitionAI/ML engineers, researchers, students, and builders are all welcome.You don’t need to walk in knowing the answer.You should walk in ready to code, experiment, train, fail a few times, and keep going.One important thing: arrive with your coding setup and model access already tested. If you plan to fine-tune, have your training environment ready too.We’d rather spend the day fighting the actual problem than fighting CUDA.The Day9:00 AM – 9:30 AM Breakfast + meet the people you’ll be building alongside9:30 AM – 11:00 AM State of the Art: Video Anomaly Detection + related robotics/CV work + FlytBase demos11:00 AM – 6:00 PM Build, train, test, break things, try again6:00 PM – 7:00 PM Selected demos + resultsDate: 5th September 2026 Time: 9:00 AM – 7:00 PMIf you spend your weekends reading papers, fine-tuning models, experimenting with vision systems, or wondering what happens when multimodal AI leaves the benchmark and meets messy real-world video...you’ll probably want to be in this room.Register and come build with the AI Hackers Collective. Join the WhatsApp group for all the updates: https://chat.whatsapp.com/EYb3jCdtcS31Ul1jWe0wWF?s=cl&p=i&mlu=4
Event location
Let your network know you`re going
Share this event to start conversations, invite colleagues, and connect before it begins.
