This paper studies \textbf{thinking--answer consistency} in vision-language models. We focus on Visual Intention Grounding, where a model infers a target object based on a human intention query and predicts a bounding box. We reveal that previous IoU-based reinforcement learning (RL) frameworks suffer from ``thinking d...
Peng-Zhan Sun, Shiu-hong Kao, Shi-Jie Li et al.· 1 citation
AI companions are envisioned as always-on assistants that support users in daily life. With this regard, we introduce BuddyVQA, a benchmark for companion-style question answering (QA) on egocentric streaming video. BuddyVQA contains 21.6K questions linked to 6K highlight moments across 1,012 long, egocentric videos. It...
Hangyu Qin, Jun-Bin Xiao, Sheng Zhang et al.· 0 citations
This work proposes SmartRes, a framework that performs efficiency optimization in the pixel space via dynamic resolution routing and introduces a margin-regularized routing objective that increases foreground-background logit separation and improves foreground recall.
H. Sun, Wang-Bo Zhao, Fanyue Wei et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.