Preprint
Jul 2026
See2Think: Do Multimodal Models Really Use Intermediate Visual States?
Evaluating representative proprietary and open-source multimodal models, it is found that visual reasoning is strongly model- and environment-dependent, with no single setting consistently dominating across tasks.
Siyu Yan, Zhuoran Yan, Haiying Xu et al.
· 0 citations