Skip to content

Despite Instructions: Frontier Agents Improvise Covert Channels at Test Time

Sep 2026 · 0 citations · 42 references
Computer Science

TL;DR

It is found that model pairs can learn to communicate the secret using only one bit of feedback indicating whether the receiver inferred it correctly, during inference with fixed parameters and no supplied codebook or encoding examples.

Abstract

In security-sensitive applications, language-model agents are often required to coordinate without disclosing confidential information. Yet repeated interactions may also let ordinary messages acquire shared private meaning. We study a repeated game with pairs of models in which the sender model observes one of four secret states and selects one of four summaries of the same public report, while the receiver model tries to infer the secret state. We find that model pairs can learn to communicate the secret using only one bit of feedback indicating whether the receiver inferred it correctly. This learning occurs during inference with fixed parameters and no supplied codebook or encoding examples. The effect also persists when agents generate their own free-form updates in a simulated incident-response task. Across ten independent games, pairs of GPT-5.6 Sol agents reach 98.8% final accuracy, compared with 25% chance, despite explicit instructions prohibiting disclosure and a monitor that screens each message without access to the agents'interaction histories. The same interactions that help agents cooperate can therefore allow confidential information to pass through messages intended for legitimate coordination.

View source

Similar papers

Preprint Aug 2026

Confident at the moment of action: belief miscalibration in LLM play under hidden information

This work tests a hidden-information chess variant where royal status can be secretly, repeatedly relocated between pieces, and where an agent's stated probability distribution over the opponent's hidden royal piece -- elicited every turn, separately from the move it chooses -- is scored against ground truth recoverabl...

Bhushan Kashinath Joshi · 1 citation
#machine learning Preprint Sep 2026

Be Careful Who You Trust: Coordination Dynamics under Corrupted Communication in LLM Multi-Agent Games

Large language models are increasingly used as interacting agents, but it remains unclear how robust their coordination is when public communication is unreliable. We study this question in iterated $N$-player Stag Hunt games played by homogeneous LLM groups under controlled programmatic action inversion, which changes...

Xuan-Yi Liu, Niall J. Dalton, Hairil Amin et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Communication Shapes Collective Inference in Self-Adapting LLM Societies: Evidence from Mafia

When does communication help a group identify hidden adversaries, and how does its value change as the group adapts? In Mafia, an informed minority hides inside an uninformed majority whose only evidence is open play. The zero-information game, where each day's vote eliminates a random player, is exactly solved and sco...

Hao-Nan Huang, Joey Xiao · 0 citations
#artificial intelligence Preprint Sep 2026

Disclosure-Gated User Simulation for Companion-Agent Evaluation

A disclosure gate conditioning information release on the companion agent's behaviour is answered, with a disclosure gate conditioning information release on the companion agent's behaviour: its state is a ladder of five ordered gates, merged onto three observable depth layers.

Yaozhong Liu, Yu He · 0 citations
Conference Open access Sep 2026

Disclosure Failure in Compromised AI Coding Agents

Preliminary evidence is found that a single sentence added to the developer's own prompt substantially improves disclosure, and a benchmark scored on attack success alone cannot rank agents on the risk a developer actually carries.

Aaron Su, Rong-Xing Lu · 0 citations
Preprint Aug 2026

Knowledge-Verified Emergent Deception in LLM Agents Under Conflicting Incentives

A knowledge-verified benchmark that first confirms through a neutral probe that an agent knows a user's entitlement, and then evaluates whether it makes false claims once an incentive to deny that entitlement is introduced, which reduces the confound between lying and not knowing and enables more rigorous auditing and...

Zhe-Yuan Liu, Wei-Liang Zhao, Xiangchi Yuan et al. · 1 citation

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.