Skip to content

Author

Yu-Bo Zhu

We have 6 of 17 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Video generation begins in text space by authoring a cinematic screenplay, then materializes into pixels. As contemporary video generators scale to 30 seconds and faithfully follow complex conditions, the textual prompt largely directs the production, planning how actions, camera trajectories, lighting, and sound unfol...

Yu-Bo Zhu, Ya-Wen Shao, Zi-Yun Dai et al. · 0 citations
Preprint Aug 2026

VIVAS: Vitalizing Visual Perception in VLM Pre-training via Vision-language Unified Autoregressive Supervision

VIVAS is proposed, a framework built upon the unified token space paradigm, which introduces a dense-structural-semantic vision tokenizer, which expands the textual vocabulary into a unified vision-language vocabulary by incorporating a visual vocabulary.

Zhe-Han Kan, Yu-Bo Zhu, Xing-Hua Jiang et al. · 0 citations
Preprint Aug 2026

UVU: Improving Multimodal Understanding via Vision-Language Unified Autoregressive Paradigm

Despite remarkable advancements in multimodal large language models (MLLMs), their fine-grained visual understanding is constrained by a primary reliance on sparse textual supervision. Existing efforts to introduce visual supervision typically do so during post-training, when visual representations have already been la...

Zhe-Han Kan, Xing-Hua Jiang, Yu-Bo Zhu et al. · 0 citations
Preprint Aug 2026

When Does Visual Generation Help Visual Understanding in Unified Multimodal Models?

VGAU-Diag is introduced, a fine-grained evaluation framework for vision generation-assisted understanding that stratifies samples by difficulty, enables unified evaluation of multiple reasoning paradigms, and uses Oracle-Ass Reference Protocols.

Yu-Bo Zhu, Zhe-Han Kan, Jing-Yi Yang et al. · 1 citation
Preprint Aug 2026

PURPOSE: Poisoning Conflict Resolution in RAG via Proxy-Fact-Grounded Updates

PURPOSE is proposed, a strict black-box poisoning attack that reframes the injection as an update that minimizes conflict, rather than as a counter-claim, and identifies non-contradicting injection as a practical mode to enhance poisoning attack.

Zijian Wang, Yubo Zhu, M. Dong et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.