Before the Token Commits: Trajectory-Level Benchmarking of Visual Hallucinations in Diffusion VLMs
Multimodal diffusion language models generate responses by iteratively unmasking tokens, making each answer the endpoint of a multi-step trajectory rather than an immediate commitment. Hallucination benchmarks built for autoregressive models evaluate only the final output, and therefore cannot determine whether an unsu...