Skip to content
Preprint

Outcome Monitors: Recovery Affordances for Silent Tool Failures

Aug 2026 · 0 citations · 34 references
Computer Science

TL;DR

Outcome Monitors, which detect violations of outcome contracts mined from task-disjoint traces or derived from public schemas, are introduced, which detect violations of outcome contracts mined from task-disjoint traces or derived from public schemas.

Abstract

When a tool call times out, the agent sees the failure and can route around it. A cached error page or negative price can instead arrive in the expected format and be consumed as fact. We introduce Outcome Monitors, which detect violations of outcome contracts mined from task-disjoint traces or derived from public schemas. On a violation, the monitor preserves the result and issues a nonbinding receipt naming the violated property and public recovery tools. In frozen, prespecified evaluations with injected failures, Outcome Monitors raise ToolMaze completion from 10.9% to 28.1% across four models in two provider families and replicate in a third. In tau-bench retail, completion improves by 14.0 and 12.0 points on two tiers. In separate ToolMaze controls, removing the recovery-tool list eliminates the measured gain and restoring it recovers the effect; diagnostic detail and timing produce no detectable differences. Gains concentrate where the fault blocks completion. On a suite transcribed from a published incident taxonomy, detection outside the mined vocabulary falls to 46%, though delivery continues and completion is unchanged. Recovery tools are the active receipt content in these controls; extending detection beyond the contract vocabulary remains open.

View source

Similar papers

Review Aug 2026

Can MCP Clients Decide What to Do After Failure? A Result-Only Actionability Audit

A client that receives isError:true knows that something went wrong. It may still have no machine-readable basis for deciding whether to fix an argument, authenticate, wait, choose another tool, or stop. This paper studies what deterministic software can learn from a completed MCP failure result alone; request argument...

Rishabh Mehan · 1 citation
Preprint Aug 2026

Metis: Typed Runtime Mediation for Tool-Using Software Agents

Metis, a multi-provider runtime that converts provider streams into typed events before admitted calls reach external effects is presented, a multi-provider runtime that converts provider streams into typed events before admitted calls reach external effects.

Jun'an Yu · 0 citations
Open access Aug 2026

Three Measurement Hazards in Analyzer-in-the-Loop Repair of Smart Contracts

An executed exploit is reported showing a specification-level authorization defect that produced no high or medium impact finding, its consequence for repair metrics, where it biases both transition counts upward and can confound comparison between methods producing differently sized .patches.

Staley Ian · 0 citations
Preprint Aug 2026

Invocation-Level Reliability of Tool-Using Agents

This work measures a correct-invocation rate that separates the two, under both a clean teacher-forced context and the model's own free-running context, on five open-weight models over contamination-free multi-step tasks.

Afiya Noorain, Subhranshu Mohanty, Amritesh Banerjee et al. · 0 citations
#machine learning Preprint Sep 2026

Legible Failures: Detecting and Repairing In-Context Binding Errors

A wrong answer does not show whether the model lacked the needed information or held it and failed to use it. On an entity-obligation binding task, a language model can emit an incorrect prompt-supplied binding while a linear probe can recover the correct one from its frozen hidden state. We measure how often this occu...

Manas Venkata Sai Ravulapalli, S. Chadha, Abhinav M. Hari · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.