Reinforcement learning with verifiable rewards (RLVR) has become a standard recipe for post-training vision-language models (VLMs), but it typically assumes a static training environment. As the actor improves, fixed tasks drift out of its learning frontier: many become trivial, others remain unsolvable; and the learni...
Meng Lu, Li-Geng Zhu, Olivia C. Xiao et al.· 0 citations
This work takes an RSI-inspired approach at the harness layer, scaling auto-research loops across increasingly numerous and diverse environments for harness rollouts, and yields reusable improvements that transfer beyond their development setting.
Hao-Zhe Liu, Tian Ye, Sen-Sen Gao et al.· 4 citations
DC-Gen, a general framework that accelerates text-to-image diffusion models by leveraging a deeply compressed latent space, uses an efficient post-training pipeline to preserve the quality of the base model to reduce the latency of 4K image generation.
Parason is introduced, which reveals and learns both forms of parallelism in LLM reasoning, and identifies Trial Parallelism as the majority of parallelizable reasoning computation, and it becomes increasingly dominant on hard problems.
Zhengyang Zhang, Zijian Zhang, Jiaxuan Gao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.