Alignment-Guided Flow Transformer (AGFT) is presented, a novel framework that explicitly enforces tri-modal alignment through a dedicated alignment loss, bridging the representational gap across modalities and enhancing task adaptation.
Sheng-Chao Hu, Peng Wang, Qi-Yang Zhou et al.· 0 citations
This paper addresses the problem of safe offline reinforcement learning, which involves training a policy to satisfy safety constraints using an offline dataset. This problem is inherently challenging as it requires balancing three highly interconnected and competing objectives: satisfying safety constraints, maximizin...
Sheng-Chao Hu, Peng Wang, Ji-Feng Hu et al.· 0 citations
FailForge is proposed, an agentic framework that converts failed rollouts into training signal, and recovers over 26% of previously failed instances at marginal additional cost, and training Qwen3.5-4B on the augmented corpus improves the SWE-bench Verified resolve rate by 6.6 points over a strong RFT baseline.
Dongyi Lv, E. Fushun, Aichen Cai et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.