Autoregressive models via next token prediction have emerged as a promising way towards unified models across multiple modalities for general artificial intelligence. Nevertheless, exploiting this idea to achieve single image to 3D generation remains a challenging problem. In this paper, we introduce a new Autoregressi...
Haibo Yang, Yang Chen, Yingwei Pan et al.· IEEE Transactions on Pattern...· 0 citations
This work constructs TextMuSS-10M, a large-scale synthetic scene text dataset spanning 10 scripts and 229 languages and proposes ScriptMoE, a script-aware Mixture-of-Experts (MoE) architecture that achieves the highest accuracy and is simpler than per-language experts, lighter than VLMs, and more accurate than both.
Xing-Song Ye, Yong-Kun Du, Jia-Xin Zhang et al.· 1 citation
Perceptual Flow Matching supervises flow matching in a perceptual feature space using pretrained perceptual models, which substantially improves the few-step generation capability of flow-matching models, reducing the number of sampling steps from 35-50 to 4-8 while preserving generation quality.
Chuyang Zhao, Yifei Song, Hongfa Wang et al.· arXiv.org· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.