The increasing deployment of AI agents in long-horizon tasks yields massive execution logs. Diagnosing failures within these records is crucial for reliability, as it transforms outcome-level signals into actionable interventions. The sheer scale of the data renders human review impractical, driving the need for automa...
Harsh Raj, David Lee, Anas Mahmoud et al.· 0 citations
Reliable Enterprise Agent Deployment (READY), a framework for qualifying AI agents for deployment on enterprise workflows, and provides a basis for comparing agent systems, setting oversight requirements, and making evidence-based deployment decisions.
Veronica Chatrath, Bryan Zhu, Jingxuan Fan et al.· 0 citations
CRAFT is introduced, a method that converts any rubric based evaluation dataset into a model specific diagnosis of weak capabilities, and yields both a sharper picture of what a model cannot do and measurably better models after finetuning on that diagnosis.
Vipul Gupta, Zihao Wang, Razvan-Gabriel Dumitru et al.· 0 citations
This work proposes Rubric Dropout, a one-line fix borrowed from neuron dropout that randomly drops a subset of the rubric's criteria before computing the reward, so the policy never optimizes the same rubric twice.
Minglai Yang, Xinyu Guo, Utkarsh Tyagi et al.· 0 citations
This work introduces an interaction-centric taxonomy that localizes failures to the interactions in which they originate and identifies the responsible component, and organizes 41 failure modes by assigning each to an edge between two components and a fault side indicating where the repair belongs.
Harsh Raj, Vipul Gupta, Anas Mahmoud et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.