Reinforcement learning from verifiable rewards (RLVR) usually optimizes answer correctness, yet useful language-model behavior also requires high-quality reasoning and concise responses. Existing multi-reward post-training methods typically scalarize rewards or combine specialists without explicitly protecting a reward...
Doseok Jang, Jon Ander Campos, You-Ran Qi· 0 citations
Building2Building (B2B), a large-scale suite of realistic HVAC control environments built on EnergyPlus, a state-of-the-art building simulator, is introduced, defining benchmark tasks targeting key open challenges in RL, including goal adaptation, dynamics adaptation, action-space shifts, and cross-domain transfer.
Vincent Taboga, Justine Veilleux, Doseok Jang et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.