This work introduces MissionBench, a benchmark for mission-level evaluation of MLLMs in aerial 3D environments, and shows that mission-level competence requires coordinating multiple capabilities beyond spatial perception, including multi-step planning and adaptive reasoning.
Suman Navaratnarajah, Taehyoung Kim, Jona Ruthardt et al.· arXiv.org· 0 citations
VideoRouter (VR) is proposed that rethinks long-video understanding as coordinating complementary evidence views rather than selecting a single subset of frames, and introduces a verification-guided router to determine which view is better supported by the selected evidence and select the final answer.
Ziling Huang, Yuki M. Asano, Shin'ichi Satoh· 0 citations
This paper proposes TWIST, a twin-expert stepwise tuning module that modifies the decoder of the language model using one frozen module pre-trained on image understanding tasks and another learnable one for visual grounding tasks, which allows the MLLM to retain previously learned knowledge and skills, while acquiring...
AritraBhowmik, MohammadMahdiDerakhshani, Dennis C. Koelma et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.