Unified multimodal models (UMMs) can perform both understanding and generation, raising a central question: can visual generation improve understanding? Existing evaluations provide mixed evidence, but confound task difficulty, reasoning paradigms, and the closed-loop interaction between generation and understanding. We introduce VGAU-Diag, a fine-grained evaluation framework for vision generation-assisted understanding. It stratifies samples by difficulty, enables unified evaluation of multiple reasoning paradigms, and uses Oracle-Assisted Reference Protocols. Our analysis shows that generated visual aids help on easier instances but become unreliable as reasoning complexity increases. Oracle-assisted diagnosis further reveals that the main bottleneck often lies on the visual-understanding side rather than the visual-generation side, as current UMMs struggle to leverage even faithful visual aids. We also show that effective visual generation should target visual-understanding bottlenecks rather than add more reasoning steps, and identify a three-stage transition from task-irrelevant noise, to misleading plausible guidance, and finally to useful assistance. These findings would be useful to guide the development of better UMMs.
Yubo Zhu, Zhehan Kan, Jing-Yi Yang et al.· 0 citations
KFS-RAG is proposed, a defense that mitigates information leakage by reformulating the retrieved context by identifying a small set of influential keywords from the retrieved context via an attention rollout plus a causal perturbation mechanism.
Ziliang Zhang, Yubo Zhu, Wei Tong et al.· 0 citations
PURPOSE is proposed, a strict black-box poisoning attack that reframes the injection as an update that minimizes conflict, rather than as a counter-claim, and identifies non-contradicting injection as a practical mode to enhance poisoning attack.
Zijian Wang, Yubo Zhu, M. Dong et al.· 0 citations