TACIT-Switch: Cost-Aware Model Escalation for LLM Agents from Censored Supervision
This method learns permanent handoff policies from accumulated trajectory evidence and Teacher-Annotated Censored Intervention Times (TACIT) and represents each annotation as an interval-censored observation on a cumulative-risk scale and achieves the highest held-out success among learned policies on both ALFWorld and DABench.
Ji'an Lei, Jian-Hao Huang
· 0 citations