Multi-agent systems (MAS) split a task across specialized roles and are promising on complex tasks, yet a prevailing approach relies on inference-time orchestration alone. General-purpose APIs are costly and hard to customize, while small models with role prompts rarely develop stable role competence or reliable collab...
Qi-Yong Zhong, Mao Zheng, Ming-Yang Song et al.· 0 citations
Experiments on reasoning, code-generation, scientific-knowledge, scientific-knowledge, and tool-use benchmarks show that these implementations can be executed through the same verl-based backend while retaining their method-specific objectives and task-dependent performance profiles.
Jie Sun, Mao Zheng, Mingyang Song et al.· arXiv.org· 0 citations
Self-Evaluation Elicitation (SEE) is introduced, a method that surfaces a latent ability to predict how a judge will score its own output through a short cycle comprising a calibration-coupled reinforcement learning phase that improves the answer and predicts the judge, followed by a masked distillation phase that shar...