We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch with a frozen vision-language understanding module built on the LLaDA2.0-Mini diffusion language model backbone. Instead of relying heavily on paired image-text data from the beginning, we first build a strong...
Chu-Yan Chen, Hao-Xin Chen, Kun Chen et al.· 1 citation
This work proposes Hyper-Spherical Quantization (HSQ), which decouples semantic content from feature magnitude via angular routing, preventing code assignment from being dominated by scale rather than meaning.
Tianren Ma, Lin Long, Chu-Yan Chen et al.· arXiv.org· 0 citations
Chunked Muon (CMuon) is introduced, a simple yet highly effective strategy that partitions these matrices into independent sub-components prior to orthogonalization, effectively overcoming the late-stage convergence plateaus of vanilla Muon.
Chu-Yan Chen, Peng Sun, Kun Yuan· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.