The Working Notes: EP5 - Can We Trust How AI Understands Video?
Can an AI model actually reason about video, or is it just describing what's in frame? In this episode of The Working Notes, Research Scientist Thong and our Forward Deploy Engineer Zhen Wei walk through four kinds of reasoning our models perform on real video: where something is, what happened over time, who it is, and whether it matters. We put it through a worst-case identity test: twenty competitors, near-identical vehicles, high speed, the model has to figure out who's who from pixels alone. The result: 90.9% top-1 accuracy, up from a 39.6% baseline. The gain didn't come from a bigger model, it came from fixing how the evaluation itself was run. Then we go further: can a model tell the difference between a door opening at 9am and the same door opening at 3am with someone concealing their face? That's judgment, not recognition, and it's the hardest reasoning type of all. š Read the full research: https://reka.ai/labs/research/beyond-recognition-how-our-models-reason-about-video š¬ Join the conversation on Discord: https://link.reka.ai/discord š More from Reka Labs: https://reka.ai/labs #PhysicalAI #WorldModels #AI #RekaLabs