Preprint
Jul 2026
Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models
For the coarse attributes the authors study, MLLMs encode the visual evidence but cannot reliably control their reliance on it, indicating that for the coarse attributes they study, MLLMs cannot reliably control their reliance on it.
Jiaang Li, Chengzu Li, Zhaochong An et al.
· 0 citations