Outcome Monitors, which detect violations of outcome contracts mined from task-disjoint traces or derived from public schemas, are introduced, which detect violations of outcome contracts mined from task-disjoint traces or derived from public schemas.
Abstract
When a tool call times out, the agent sees the failure and can route around it. A cached error page or negative price can instead arrive in the expected format and be consumed as fact. We introduce Outcome Monitors, which detect violations of outcome contracts mined from task-disjoint traces or derived from public schemas. On a violation, the monitor preserves the result and issues a nonbinding receipt naming the violated property and public recovery tools. In frozen, prespecified evaluations with injected failures, Outcome Monitors raise ToolMaze completion from 10.9% to 28.1% across four models in two provider families and replicate in a third. In tau-bench retail, completion improves by 14.0 and 12.0 points on two tiers. In separate ToolMaze controls, removing the recovery-tool list eliminates the measured gain and restoring it recovers the effect; diagnostic detail and timing produce no detectable differences. Gains concentrate where the fault blocks completion. On a suite transcribed from a published incident taxonomy, detection outside the mined vocabulary falls to 46%, though delivery continues and completion is unchanged. Recovery tools are the active receipt content in these controls; extending detection beyond the contract vocabulary remains open.
A client that receives isError:true knows that something went wrong. It may still have no machine-readable basis for deciding whether to fix an argument, authenticate, wait, choose another tool, or stop. This paper studies what deterministic software can learn from a completed MCP failure result alone; request argument...
Metis, a multi-provider runtime that converts provider streams into typed events before admitted calls reach external effects is presented, a multi-provider runtime that converts provider streams into typed events before admitted calls reach external effects.
Overall, TraceGate shows that rethinking debugging through controlled observability, rather than relying solely on stronger models or larger prompts, can make LLM-assisted repair more effective, efficient and controllable.
Nicolas Schuler, MateVincenzoScotti, RaffaelaMirandola· 0 citations
An executed exploit is reported showing a specification-level authorization defect that produced no high or medium impact finding, its consequence for repair metrics, where it biases both transition counts upward and can confound comparison between methods producing differently sized .patches.
Staley Ian· International Journal of Inn...· 0 citations
This work measures a correct-invocation rate that separates the two, under both a clean teacher-forced context and the model's own free-running context, on five open-weight models over contamination-free multi-step tasks.
Afiya Noorain, Subhranshu Mohanty, Amritesh Banerjee et al.· 0 citations
A wrong answer does not show whether the model lacked the needed information or held it and failed to use it. On an entity-obligation binding task, a language model can emit an incorrect prompt-supplied binding while a linear probe can recover the correct one from its frozen hidden state. We measure how often this occu...
Manas Venkata Sai Ravulapalli, S. Chadha, Abhinav M. Hari· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.