Skip to content

Policy Loopholes in Agent Evaluation: When Policy Ambiguity Masquerades as Agent Error

Sep 2026 · 0 citations · 32 references
Computer Science

TL;DR

Auditing two $\tau^2$-bench domains, a taxonomy of policy loopholes is developed and it is shown that affected tasks produce unreliable scores: they lower scores across different models in different ways and make every model less consistent across repeated trials.

Abstract

Agent benchmarks evaluate policy compliance but assume each policy determines a unique correct action. Natural-language policies can violate this assumption through silence, ambiguity, or contradiction, admitting multiple defensible readings that a single gold trajectory cannot capture. Auditing two $\tau^2$-bench domains, we develop a taxonomy of such policy loopholes and show that affected tasks produce unreliable scores: they lower scores across different models in different ways and make every model less consistent across repeated trials. A cross-domain comparison reveals that exploitability requires both policy ambiguity and tool permissiveness: when policy complexity exceeds what tools can enforce, agents resolve gaps inconsistently and scores become unreliable. Policy specification quality sets the ceiling on evaluation quality. Benchmark developers should audit policies before collecting gold annotations.

View source

Similar papers

Preprint Aug 2026

ReguSim: Evaluating LLM Agent Rule Grounding in Financial Compliance

ReguSim, a controlled financial-compliance environment, and ReguBench, a target-marked monitoring benchmark, are introduced to separate four artifacts: stated reasoning, attempted action, execution enforcement, and monitor evidence to frame financial compliance evaluation as an audit of rule-grounded actions and eviden...

Yiyan Luo, Yihang Jiang, Qijun Xie et al. · 1 citation
#artificial intelligence Preprint Sep 2026

AgentBoundary: Counterfactual Evaluation of Safety in Tool-Using LLM Agents

Safety alignment for large language models (LLMs) in conversational settings is largely framed around whether to answer or refuse a request. In agentic settings, however, the same models must decide whether to act as permission-critical evidence emerges during execution. This creates a distinct challenge: apparent risk...

Tian-Zhuo Yang, Zi-Rui Mi, Yan-Tao Huang et al. · 0 citations
Preprint Aug 2026

RePolicy: Reinforcement Learning for Safety-Policy Invocation in Agent Safeguards

Safeguarding language model agents requires assessing complete execution trajectories under context-dependent safety policies. Existing policy-aware safeguards mainly rely on prompting or supervised fine-tuning, limiting their ability to adapt to unseen trajectories and changing policy contexts. We propose RePolicy, an...

Hou-Cheng Jiang, Bo-Xuan Zhang, Qi-Yong Zhong et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Sapien: A Stateful Policy Engine for Autonomous AI Agents

Contextual security defenses prevent AI agents from taking rogue actions by synthesizing a task-specific policy and enforcing it on the agent's tool calls. In multi-step tasks, however, which actions are valid often depends on what the agent has already done and learned. We present Sapien, a policy engine for enforcing...

CO Tiffany, Wen Zhang, E. Bagdasarian et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Auditing Action Settlement in LLM Agent Environments: Order, Progress, and Replay

Concurrent actions in large language model (LLM) agent environments require arbitration even when each proposal is individually valid. We implement a typed snapshot-settlement contract and audit three distinct properties: order sensitivity, useful progress, and replay consistency. Five settlement policies are tested in...

Hao-Tian Chen, Bo-Wen Ye, Yu-Ning Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Auditability Is Not One Property: Rule Overlap, Behavioural Agreement, and Composition in Reinforcement Learning

This work contributes an evidence-bounded audit and composition protocol, not a claim of universal interpretability or autonomous skill generation, and defines auditability as six separately testable predicates: trace integrity, lossless coding, rule coverage, behavioral agreement, composition quality, and value-model...

Hung-Ming Liu · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.