Skip to content

Author

Mohit Bansal

We have 2 of 36 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

V-Co: A Closer Look at Visual Representation Alignment via Co-Denoising

V-Co, a systematic study of visual co-denoising in a unified JiT-based framework, outperforms the underlying pixel-space diffusion baseline and strong prior pixel-diffusion methods while using fewer training epochs, offering practical guidance for future representation-aligned generative models.

Han Lin, Xichen Pan, Zun Wang et al. · 0 citations
Jul 2026

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA

VideoTreeSearch (VTS) is proposed, a framework that casts grounded LVQA as iterative self-correcting search over an adaptive temporal tree, and trains an agent to navigate the tree through four discrete operations: zoom_in, zoom_out, shift, and answer.

Ce Zhang, Ziyang Wang, Yu-Lu Pan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.