Ambient @ EgoLongQA 2026: Distilling Long-Video perception into a Sub-2B Model
The system is a single 2B vision-language model that answers multiple-choice questions about ten-minute egocentric videos in one greedy forward pass, and reaches 89% of the accuracy of the large agentic pipeline using 1.1% of its parameters.