Diffusion multimodal large language models (dMLLMs) frequently produce long-form outputs marred by semantic drift and repetition, with quality generally degrading as output length increases. We identify two structural deficiencies in existing decoding methods as primary drivers of these failures: confidence-based scori...
Yikai Zhao, Qiyan Zhao, Jia-Quan Zhang et al.· 0 citations
IGFD is proposed, a training-free decoding strategy that ranks candidates using token confidence, neighborhood uncertainty, and structural commitment risk and consistently outperforms existing decoding strategies across the majority of benchmarks and diffusion MLLM backbones under identical decoding budgets.
Xing-You Fang, Jin Zhong, Xiaosong Yuan et al.· 0 citations
Systematic skill-bank audits reveal two recurring reliability failures: coalition pollution, where bank-level gains conceal negative coalition-level skill contributions, and cross-domain utility reversal, where source-beneficial skills reverse their effects after transfer.
Qiyan Zhao, Xiaofeng Zhang, Bo Liu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.