Reading Right, Answering Wrong: How Visual Configuration Changes Affect Evidence Use in VLMs
Findings show that configuration changes can affect how models use information they can still read, and attention interventions in LLaVA-NeXT suggest that configuration changes can weaken the use of readable information during answering.
Dingyang Lin, Yingfeng Luo, Cheng-Long Wang et al.
· 0 citations