Skip to content
Preprint

Prediction of Prediction (PoP): Inter-Layer Activation Fusion for Single-Pass Hallucination Detection in Large Language Models

Aug 2026 · 0 citations · 23 references
Computer Science

Abstract

Autoregressive large language models (LLMs) routinely generate factually incorrect outputs with high decoding confidence, limiting their deployment in high-stakes workflows. Existing output-stage uncertainty metrics can fail when models are overconfident on false assertions, while multi-sample verification pipelines introduce substantial memory and latency overhead. This work evaluates whether internal hidden-state transition dynamics during generation can signal factual errors without auxiliary decoding calls. We introduce Prediction of Prediction (PoP), a mechanism that captures layer-transition uncertainty by fusing intermediate hidden representations across depth during a single forward pass. Evaluated on the TruthfulQA benchmark using autoregressive transformer backbones, PoP achieves an area under the receiver operating characteristic curve (AUROC) of 75.5% for factual-correctness classification. The mechanism operates within the base forward pass, adding less than 1.2% runtime latency and requiring zero additional generation passes. The numerical results are reported from the author-verified experimental implementation and are bounded by the evaluation scope described below.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Domain-Specific Hallucination Detection in Large Language Models

Large language models generate fluent text that can contain unfaithful claims -- a phenomenon known as hallucination. We present a multi-signal detection pipeline combining fine-tuned DeBERTa-v3 classification, Monte Carlo (MC) Dropout uncertainty quantification, and temperature-scaled calibration for response-level ha...

Varun Teja Chundru, Debasmita Biswas · 0 citations
#artificial intelligence Preprint Sep 2026

Look Before You Leap: Factual Decoding with Internal Attribution Signals

This work identifies a factual-salient layer span within LLMs whose derived signal is selectively elevated for factual tokens and exhibits anomalous spikes at hallucination-prone steps, and proposes DescaPE, a decoding framework that leverages internal model signals to suppress hallucination-prone trajectories at infer...

Hayeong Ryu, Jungmin Yun, B. Lim et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ActMap: Single-Pass Uncertainty Quantification from Generation-Time Activation Maps

ActMap is introduced, a white-box representation that compresses the generation-time hidden-state trajectory into a fixed tensor of temporal-statistic channels that preserves structure across transformer depth and pooled hidden coordinates, and supports abstention, routing, and selective verification from a single gene...

Jacopo Dardini, Roberta Calegari · 0 citations
#artificial intelligence Preprint Sep 2026

The Alignment Paradox: How Post-Training Amplifies Confident Hallucinations in Language Models

Bounded margins mitigate confident hallucinations during post-training, implemented through an entropy-dependent margin bound in direct preference optimization (DPO) and shown to mitigate confident hallucinations during post-training.

Qing-Jia Huang, Ya-Kai Li, Jian-Guo Wu et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Dynamic Alignment Compensation for Hallucination Mitigation in Large Vision-Language Models

Experiments show that DAC consistently reduces hallucinations while maintaining strong overall performance, and combines Layer-wise Semantic Compensation to mitigate inter-layer degradation with Sequential Semantic Correction to constrain temporal drift.

Kai-Rong Yu, Zi-Xin Zhu, Le Yu et al. · 0 citations
Preprint Aug 2026

Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic Critique

The Latent Critic is introduced, a lightweight low-rank adapter that operates concurrently with a frozen base LLM's generation to actively restructure the transformer's residual stream---amplifying latent grounding signals and translating them into localized, natural language feedback within a single sequence.

S. Vijayvargiya, R. Lokesh · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.