Skip to content

Detecting Pretraining Data in Large Language Models from a Free-Energy Perspective

Sep 2026 · 0 citations · 54 references
Computer Science

TL;DR

This work introduces an inclined boundary that evaluates prediction loss relative to predictive entropy, and shows that entropy correction can preserve the expected membership signal while reducing its variance, thereby improving standardized member--non-member separation.

Abstract

Detecting pretraining data in large language models is challenging because high likelihood can reflect either training exposure or strong generalization. In the joint space of prediction loss and predictive entropy, a likelihood-only detector uses a horizontal boundary and can mistake predictable non-members for members. Motivated by this, we introduce an inclined boundary that evaluates prediction loss relative to predictive entropy. Our analysis shows that entropy correction can preserve the expected membership signal while reducing its variance, thereby improving standardized member--non-member separation. We further extend the mean--variance analysis to the more general setting with a nonzero mean entropy gap. Interestingly, this entropy-adjusted score admits a Helmholtz free-energy interpretation, leading to Energy Transfer Detection (ETD), which views pretraining data detection from a macroscopic residual free-energy transfer perspective. Extensive experiments show that ETD achieves the best average detection performance, improving average AUROC by up to 3.5\% and TPR@5\%FPR by up to 5.1\%, while remaining robust across diverse settings.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Relative Generalization Invariance of LLM Pretraining

Large Language Model (LLM) pretraining performance is jointly shaped by three components of the training triplet: the optimizer, model architecture, and training data stream. However, how these components influence performance in distinct ways remains unclear. We take a first step toward isolating their effects by stud...

Feng-Zhuo Zhang, Shu-Che Wang, Sheng-Gui Li et al. · 0 citations
#machine learning Preprint Sep 2026

Entropy Regularization: A Free Correction to Cross-Entropy for Verified Demonstrations

Large language models are often post-trained on expert demonstrations using cross-entropy (CE), even when the downstream objective is not to imitate the demonstrated solution but to produce any output accepted by a verifier. This mismatch is seen in verifiable domains with multiple correct solutions, such as mathematic...

Mihir Dhanakshirur, Adam Ousherovitch, A. Tewari · 0 citations
#machine learning Preprint Sep 2026

Does a Shared Temperature Imply a Shared Angular Scale in Probabilistic Contrastive Learning?

In probabilistic contrastive learning, a shared temperature is commonly interpreted as a shared similarity scale, but this interpretation does not hold for high-dimensional distributional class representations. We study the exact von Mises-Fisher (vMF) probabilistic score used by ProCo when representation dimension and...

Ning-Kang Peng, Qian-Feng Yu, Jing Mao et al. · 0 citations
Preprint Aug 2026

Cross-Domain Generalization in Machine Unlearning via Label-Conditioned Energy Magnitude Regularization

This paper studies what happens to the rest of the model when a class is forgotten, using a label-conditioned energy-based model (EBM) that assigns per-class energies, making the effect directly observable.

Syed Ali Ahmed, Syed Bilal Ahsan, Muhammad Zaigham Zaheer National University of Computer et al. · 0 citations
Preprint Aug 2026

Respect Your Zero-Shot Uncertainty: Conservative Calibration for Test-Time-Adapted Vision-Language Models

It is shown that TTA can increase confidence and reduce entropy even when the top-1 prediction and its correctness remain unchanged, a failure mode the authors term prediction-preserving sharpening, and proposed Zero-Shot-Anchored Entropy Calibration (ZAEC), a label-free post-hoc method that uses zero-shot entropy as a...

Jing-Yan Jiang, Yaru Sun, Xiao Chen et al. · 0 citations
#machine learning Preprint Sep 2026

Understanding and Exploiting Anisotropy in Post-Training

LLM post-training combines supervised fine-tuning (SFT), a mode-covering forward-KL objective, with reinforcement learning (RL), a mode-seeking reverse-KL objective. Frequency-weighted likelihood training leaves a well-known signature: \emph{anisotropy}, in which a few residual channels carry disproportionately large a...

Samyak Jha, Harshvardhan Saini, Yi-Zhen Liao et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.