Skip to content
Preprint

CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment

Aug 2026 · 0 citations · 67 references
Computer Science

TL;DR

This work introduces CIDER, a dataset of 14,850 human annotations from 169 users, forming 1,650 contextual disclosure boundary sets across 60 interpersonal communication scenarios involving information sharing that violates privacy norms, and reveals both the promise and limitations of inference-time personalization for privacy preference modeling.

Abstract

Aligning large language models (LLMs) with human privacy preferences requires capturing individuals'disclosure boundaries beyond general privacy norms. However, a gap remains in eliciting such nuanced preferences to evaluate alignment in realistic settings. We introduce CIDER, a dataset of 14,850 human annotations from 169 users, forming 1,650 contextual disclosure boundary sets across 60 interpersonal communication scenarios involving information sharing that violates privacy norms. Each boundary represents a real user's disclosure decisions over 9 sharing variants in a scenario, for a given communication role and AI-mediated condition. We formulate a task in which models predict a user's disclosure decision from historical boundaries, with varying levels of contextual information. Across 12 open and proprietary models, in-context personalization improves prediction accuracy by up to 11.41 percentage points using only 6 historical examples. Larger models such as GPT-5.4 (with medium reasoning effort) and Claude Sonnet 4.6 are better at leveraging semantic context to understand user-specific, context-dependent disclosure preferences for more accurate predictions, while smaller models tend to rely on structured heuristics based on disclosure granularity and identifiability. Personalization generally improves prediction accuracy, but the improvement is often accompanied by imbalanced shifts in false-positive and false-negative rates across models, with only Claude Sonnet 4.6 achieving balanced improvements in both. Our findings reveal both the promise and limitations of inference-time personalization for privacy preference modeling and position CIDER as a resource for advancing personalized privacy alignment.

View source

Similar papers

Preprint Aug 2026

Beyond Direct Identifiers: Probabilistic Privacy Risk Estimation for Privacy-Conscious LLM Query Delegation

This work investigates a probabilistic variant of PCD, where an LLM-driven probabilistic estimation of k-anonymity is augmented with an LLM-driven probabilistic estimation of k-anonymity, and proposes k-anonymity as a useful auxiliary metric for tackling PCD.

Si-Yan Li, Yu Zhou, Julia Hirschberg · 2 citations
#human-computer interacti... Preprint Sep 2026

Alignment and Divergence between Humans and AI in Interpersonal Privacy Decisions

AI assistants increasingly mediate interpersonal communication on behalf of their primary user, but they risk violating the privacy expectations of third-party information owners. Resolving these tensions requires understanding how humans anticipate interpersonal privacy boundaries. Therefore, we conducted a dyadic stu...

Han-Xiang Zeng, Shu-Ning Zhang, Xin-Yuan Zhou et al. · 1 citation
#artificial intelligence Preprint Sep 2026

EP-Mem: Elastic Privacy Memory for Social Relationship-Aware LLM Agents

Large language model (LLM) agents face critical privacy risks when acting as delegates in human-agent-human communication. To prevent such breaches, agents must understand users'social relationships and adhere to context-dependent social information disclosure boundaries. Current studies on agent memory privacy focus o...

Feng-Zhou Sun, Yuan Zhang, Xin-Tong Yu et al. · 0 citations
Preprint Aug 2026

Personalized Privacy Control in LLMs via Attention Head Intervention

The introduction of personalized privacy, which incorporates user-specific disclosure preferences into privacy control, and a robust inference-time attention head intervention method that adjusts disclosure behavior toward policy-consistent responses are proposed.

Junseok Kim, Nakyeong Yang, Kyomin Jung · 0 citations
Preprint Sep 2026

Can Prompt Anonymity Protect Your Identity From LLM Providers?

User conversations with large language models (LLMs) often contain highly sensitive personal information that can be exploited by LLM providers to create detailed user dossiers, enable targeted advertising, and train more powerful models. To protect user privacy, anonymizing LLM proxies have emerged as a practical solu...

Dzung Pham, Dillon Sheils, Naina Singh et al. · 0 citations
Book Open access Aug 2026

Asking for Privacy: Contrasting Consumer Questions with Questions in Privacy Policies

Privacy policies are text documents intended to inform consumers about the data practices of apps and websites, but they can be challenging to read. Some organizations try to address the obstacles by organizing their privacy policies as question-answer pairs. We use a combination of automated and manual analysis method...

Tian-Yang Zhao, Younes Karimi, Thomas B. Norton et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.