Skip to content

Author

Vasilis Syrgkanis

We have 3 of 179 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection

Chain-of-thought (CoT) monitoring is a safety strategy where the reasoning of a large language model"actor"is inspected by a"monitor"(often another language model) for signs of unsafe planning, deception, or misalignment. We find that planting harmful but benign-sounding reasoning in the actor's context can steer it to perform adversarial actions while evading monitors, an attack we term"plan injection". We initially discover this attack in the multiple-choice question-answering monitorability setting proposed by Lanham et al. (2023), using the investigator-agent elicitation framework of Li et al. (2025). We generalize the attack and show that the discovered behavior scales to harder tasks (achieving 25-33% monitor evasion rates across different monitorability benchmarks) and larger models such as DeepSeek-R1. Across the settings we study, actor models not only follow injected plans but also paraphrase them as their own reasoning, without explicit attribution to the injections. Finally, we find cases where extra monitor resources cause harm - giving the monitor access to the injected plan drops detection by as much as 50% in the Bio-Math task and in a case study on monitor reasoning budget, we find transcripts where additional thinking tokens are spent rationalizing the injected plan rather than flagging it.

Keertana Chidambaram, Andrew Ilyas, Vasilis Syrgkanis · 0 citations
#machine learning Review Jul 2026

CausalSmith: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference

A framework for automated theoretical research in causal inference built on the Lean proof assistant, where a proof is checked by a program rather than read by a referee is presented, where a proof is checked by a program rather than read by a referee.

Jiyuan Tan, Vasilis Syrgkanis · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.