Preprint
Jul 2026
Scaling Native Multimodal Pre-Training From Scratch
This empirical research establishes the essential groundwork for predictably scaling multimodal foundation models by modeling the influence of data composition on compute laws and allocation exponents and derive an efficiency frontier specifying precise configurations of model size, token count, and data mixture.
Haoyuan Wu, Aoqi Wu, Hai Wang et al.
· 1 citation