Skip to content
Preprint

Capability Is Not Propensity: Measuring Pressure-Robust Cooperative Behavior in Civic LLM Agents

Aug 2026 · 0 citations · 39 references
Computer Science

TL;DR

This work argues that Cooperative AI evaluations should separate what models can do under benign instructions from what they tend to do under realistic civic pressure, and introduces DiffCoop-Civic, a 10-scenario pilot evaluation suite spanning preference understanding, evidence and persuasion, commitment design, asymmetric information, and dissent preservation.

Abstract

Cooperative capabilities in language models are dual-use. The same social reasoning that supports civic deliberation can also enable strategic omission, false consensus, and manipulative framing. We argue that Cooperative AI evaluations should separate what models can do under benign instructions from what they tend to do under realistic civic pressure. We introduce DiffCoop-Civic, a 10-scenario pilot evaluation suite spanning preference understanding, evidence and persuasion, commitment design, asymmetric information, and dissent preservation. Across seven models from four model families, subtle omission pressure produces a near-uniform shift: manipulative enablement rises by 1.17 points and dissent preservation falls by 1.67 points on a 5-point scale. Overt false-consensus pressure behaves differently: it triggers refusal or redirection in some aligned API models, but direct compliance in several open-weight models. A lightweight Pareto-Trace prompting intervention improves pressure robustness without simply relying on hard refusal. An anonymous reproducibility package is available at https://anonymous.4open.science/r/diffcoop-civil-771C.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Persuaded, Not Informed: Incentive-Misaligned Witnesses Defeat In-Context Grounding

This work identifies a failure mode not addressed by a stronger model: when the context contains an assertion by a party with an incentive toward optimism - here the sales representative, a witness recorded in the CRM - the model treats the assertion as evidence and clears deals the company's own records deem unaccepta...

Rahul Balakavi · 0 citations
Preprint Aug 2026

Group Perspective Matters: Regulating Debate Relationships Can Mitigate Blind Conformity in Multi-Agent Debate

This paper proposes a novel framework for Dynamicallyynamically regulatingdebate Relationships (DEAR) from the group perspective, and formulate the execution of the two RL-Agents as a sequential decision-making process, jointly optimizing via multi-agent reinforcement learning.

Hao Wu, Shoucheng Song, Chang Yao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Before the Script, Set the Stage: How Worldview Simulation Amplifies Psychologically Grounded Persuasion in Multi-Turn Jailbreaking

BluePRINT is introduced, a safety-evaluation framework separating a factorized social-influence strategy space from WORLDVIEWSIM, a cross-turn situational context module, and Monte Carlo Tree Search optimizes turn-level combinations of 18 theory-grounded influence factors across a four-turn trajectory.

Si-Yu Chen, Hao-Ran Wang, Xiaojian Li et al. · 0 citations
Conference Open access 2026

Bias and Fairness in LLM-Based Recruitment: A Systematic Review

A PRISMA 2020-guided systematic literature review draws on 82 studies selected from 493 records retrieved from Scopus and Web of Science and reveals a structural disconnect in the fairness-in-NLP and HCAI governance literature.

Asmae El Moutafail, Khalid Belkhoutout · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.