Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence
ReaLVR is proposed, which brings visual-evidence supervision to the model's own free-running latent trajectories, and is the first to scale visual reasoning in latent space, showing that the framework continues to deliver robust improvements at frontier model scales up to 235B.
Xi Xiao, Tian-Chen Zhao, Youngeun Kim et al.
· 0 citations