Policies with similar mean returns can differ sharply in rare failures, yet estimating lower-tail conditional value-at-risk (CVaR) accurately can require many costly rollouts. When different conditional components of a stochastic workflow can be queried separately, we ask how to allocate a fixed evaluation budget to es...
P. Bourigault, Xiao-Tong Ji, Matthieu Zimmer et al.· 0 citations
Self-Distillation Fine-Tuning (SDFT) enables a language model to act as its own teacher: by conditioning on a demonstration, the model produces an implicit reward via pointwise mutual information, which guides on-policy learning without external supervision. However, SDFT operates at training time: it requires gradient...
Su-Ee Tan, Xiao-Tong Ji, Rasul Tutunov et al.· 0 citations
On-policy self-distillation fine-tuning (SDFT) learns new skills from demonstrations while reducing forgetting, but it always distils toward the full demonstration-conditioned teacher. This fixes teacher influence at the full-teacher endpoint, providing no control over how much demonstration information should be trans...
Ahmed Khaled Khamis, Xiao-Tong Ji, Hassan Jaber et al.· 0 citations
A counterexample-guided reinforcement learning method that navigates safe exploration in autonomous systems without prior knowledge, even when safety and optimality conflict, and a novel belief-based regularization method to address the distributional shift between online and offline learning and to balance optimizatio...
Xiao-Tong Ji, Antonio Filieri· ACM Transactions on Autonomo...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.