Conversational MLLM agents are increasingly expected to assist in professional workflows, from AI research and engineering design to product management and business operations. Yet this capability remains underexplored: existing benchmarks largely focus on informal, everyday interactions and personal-life scenarios fea...
Ji-Ke Zhong, Ritwick Chaudhry, Xuan-Bai Chen et al.· 0 citations
Diffusion language models (DLMs) enable fast generation by predicting multiple tokens in parallel, but their practical adoption remains limited by a persistent quality gap relative to comparably sized autoregressive (AR) models. We attribute this gap to a computation-difficulty mismatch: within a partially observed seq...
Lian-Cheng Fang, Zhuo-Wei Li, Youngeun Kim et al.· 0 citations
ReaLVR is proposed, which brings visual-evidence supervision to the model's own free-running latent trajectories, and is the first to scale visual reasoning in latent space, showing that the framework continues to deliver robust improvements at frontier model scales up to 235B.
Xi Xiao, Tian-Chen Zhao, Youngeun Kim et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.