Multimodal diffusion language models generate responses by iteratively unmasking tokens, making each answer the endpoint of a multi-step trajectory rather than an immediate commitment. Hallucination benchmarks built for autoregressive models evaluate only the final output, and therefore cannot determine whether an unsu...
Ya-Dong Wang, Si-Ping Yue, Yu Tian et al.· 0 citations
LSTR (Latent Sparse Transcoder Reasoning), a framework that turns sparse transcoders from post-hoc diagnostic tools into in-loop, intervenable transition components for latent reasoning, and suggests that sparse latent transitions can preserve the compression benefits of latent reasoning while making the resulting traj...
Yadong Wang, Hao-Dong Chen, Yu Tian et al.· 0 citations
Recall-Anchored Distillation (RAD), a base-anchored self-distillation objective that preserves out-of-distribution generation behavior by aligning the adapted model with the original base model's soft continuation distribution on unlabeled OOD text, is introduced.
Hao-Dong Chen, Yadong Wang, Shengtao Wen et al.· 0 citations
Pre-norm is the standard normalization placement in modern Transformers because it facilitates joint optimization of full-depth models. We ask whether this preference persists when depth is introduced through a curriculum. In curriculum depth growth, each appended block receives the boundary representation produced by...
Sheng Ren, Yadong Wang, Naiqiang Tan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.