Conversational MLLM agents are increasingly expected to assist in professional workflows, from AI research and engineering design to product management and business operations. Yet this capability remains underexplored: existing benchmarks largely focus on informal, everyday interactions and personal-life scenarios fea...
Ji-Ke Zhong, Ritwick Chaudhry, Xuan-Bai Chen et al.· 0 citations
A benchmark with tab-and-blank interlocking pieces where geometric constraints provide strong local compatibility requirements that, combined with visual content, yield unambiguous ground truth is introduced, establishing scalable geometric reasoning as an open challenge for vision-language models.
This work proposes to model objects as a stronger semantic unit for visual prediction, encouraging the encoder to learn the global context and semantics among visual elements, and shows that an object-centric objective reduces pixel-averaging shortcuts and yields more globally coherent and context-consistent representa...
The results suggest context learning hinges on not only content acquisition but also specification acquisition, and designs a deliberately simple intervention PSCI (private specification-contract induction) which extracts local specifications and enforces them through adversarial checking and repair.
Jike Zhong, Ming Li, Yuxiang Lai et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.