Time-series foundation models offer a unified approach to forecasting across heterogeneous domains. Textual context and auxiliary observations provide complementary information about temporal dynamics, yet reusable multimodal predictive representations remain underexplored. We introduce Pythia, a foundation world model...
Xilin Dai, Hong-Zhou Chen, Yi-Fan Hu et al.· 0 citations
Agentic time series forecasting concerns systems whose underlying mechanisms evolve, making the relative effectiveness of numerical models, reasoning strategies, and intervention rules inherently time-varying. Consequently, a time series agent must adapt the forecasts it produces and the orchestration policy that deter...
Yi-Fan Hu, Xilin Dai, Zhi-Yuan Qu et al.· 1 citation
Overall, Code2Skill transforms procedural knowledge embedded in repositories into grounded, verifiable, and transferable agent skills, providing initial evidence that the pipeline can expand with the growing volume of AI-generated software.
Yong-Qi Tong, Pan Wang, Hang Wang et al.· 1 citation
The proposed ARC (Advantage Regularization via Conditioning), a training recipe that restores fairer relative comparison through strategy-conditioned rollout grouping, together with hybrid rewards and entropy regularization, is proposed.
Yong-Qi Tong, Tan Li Hui Faith, C. Marcus et al.· 0 citations
ACA-RL supports a new mission for NLP evaluation: measuring whether models can recognize when a task is underdetermined and handle uncertainty, not only whether they can answer fully specified questions.
Yong-Qi Tong, Zhenyu Zhang, Zimou Liu et al.· 1 citation
This work proposes \methodname, a stability-guided active-set controller for controlled objective admission, a stability-guided active-set controller for controlled objective admission in reward-vector RLHF, which positions objective-entry timing as a concrete control variable in reward-vector RLHF.
Yong-Qi Tong, Z. Zhang, Ruirui Wang et al.· 1 citation
DARC is proposed, a diagnosis-guided recovery harness that profiles task-family failure modes, prunes mismatched interventions from a shared recovery library, and freezes a verifier-selected success-cost policy for deployment, providing a practical route toward more reliable agents in domains where compiler-like feedba...
Pan Wang, Yihao Hu, Hang Wang et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.