Across seven benchmarks and four base VLMs, VLM-in-Sandbox achieves the highest sample-weighted average accuracy among Vanilla VLM, Append-only Sandbox, and the proposed method, and is identified as a central abstraction for sandboxed VLM agents.
He-Xiong Yang, Mingrui Chen, Jie Cao et al.· 0 citations
Educational foundation models must solve problems, understand curriculum structure, diagnose learner difficulties, and provide appropriate instructional support. Existing educational language models often focus on either problem solving or tutoring, with training mixtures organized by source or task rather than capabil...
Hao Liang, Qi-Han Lin, Mei-Yi Qiang et al.· 0 citations
DataFlex-RL, an evaluation platform for comparing choices under a common GRPO recipe, is introduced, finding that changing the data policy measurably changes the training process but does not produce a reproducible improvement over uniform training.
Hao Liang, Ming-Rui Chen, Hengyi Feng et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.