Skip to content
Preprint

Personalized Privacy Control in LLMs via Attention Head Intervention

Aug 2026 · 0 citations · 18 references
Computer Science

TL;DR

The introduction of personalized privacy, which incorporates user-specific disclosure preferences into privacy control, and a robust inference-time attention head intervention method that adjusts disclosure behavior toward policy-consistent responses are proposed.

Abstract

The rise of agentic AI enables LLMs to access diverse user data, raising critical privacy concerns. Prior work on contextual privacy studies whether LLMs regulate information disclosure according to context-dependent norms. However, acceptable disclosure boundaries may vary across users even within the same context. To address this limitation, we introduce \textit{personalized privacy}, which incorporates user-specific disclosure preferences into privacy control. We further present P3Bench~(\textbf{P}ersonalized \textbf{P}rivacy \textbf{P}reservation \textbf{Bench}mark), a novel benchmark extending contextual privacy policies with personalized disclosure policies. Experiments show that prompt-based policies fail to reliably enforce personalized privacy policies, with Qwen2.5-7B and Gemma3-4B showing average policy ignorance ratios of 51.25\% and 74.28\%, respectively. Finally, to address this problem, we propose \textsc{Repair}, a robust inference-time attention head intervention method that adjusts disclosure behavior toward policy-consistent responses. Our method significantly improves adherence to user-specific privacy preferences by reducing cases where the model fails to follow the given policy.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

EP-Mem: Elastic Privacy Memory for Social Relationship-Aware LLM Agents

Large language model (LLM) agents face critical privacy risks when acting as delegates in human-agent-human communication. To prevent such breaches, agents must understand users'social relationships and adhere to context-dependent social information disclosure boundaries. Current studies on agent memory privacy focus o...

Feng-Zhou Sun, Yuan Zhang, Xin-Tong Yu et al. · 0 citations
Preprint Aug 2026

Beyond Direct Identifiers: Probabilistic Privacy Risk Estimation for Privacy-Conscious LLM Query Delegation

This work investigates a probabilistic variant of PCD, where an LLM-driven probabilistic estimation of k-anonymity is augmented with an LLM-driven probabilistic estimation of k-anonymity, and proposes k-anonymity as a useful auxiliary metric for tackling PCD.

Si-Yan Li, Yu Zhou, Julia Hirschberg · 2 citations
Preprint Aug 2026

What to Remember, What to Reveal: Privacy-Aware Memory for Conversational Agents

Long-term memory enables personalized conversational agents to retain user information across sessions. However, existing memory architectures primarily optimize for utility while neglecting the risks of unnecessarily storing and reusing private attributes such as personally identifiable information (PII). Addressing p...

Wen-Jie Wang, Wen-He Si, Xinyue Xu et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Confuse the Model, Control the Flow: Understanding and Mitigating Privacy Leakage from LLM Agents with Information Flow Control

Personal AI agents built on large language models (LLMs) are increasingly given access to a user's private data and communications in order to provide personalized assistance. This access creates a persistent privacy risk: the agent must decide whether a given sensitive information should be disclosed to a particular p...

Minsun Shim, Ramisha Raida Karim, Ruthwik Jakkula et al. · 1 citation
Preprint Aug 2026

CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment

This work introduces CIDER, a dataset of 14,850 human annotations from 169 users, forming 1,650 contextual disclosure boundary sets across 60 interpersonal communication scenarios involving information sharing that violates privacy norms, and reveals both the promise and limitations of inference-time personalization fo...

Bingcan Guo, Er-Yue Xu, Jijie Zhou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.