Skip to content

NeuroActiSep: Detecting Factual Hallucinations from Feed-Forward Neurons in a Single Pass

Sep 2026 · 0 citations · 23 references
Computer Science

TL;DR

This work proposes a method to rank feed-forward neurons at the final prompt token using a custom neuron selection dataset, and transfers the selected neuron identities to train hallucination classifiers on other factual question answering datasets.

Abstract

Hallucination in large language models reduces their reliability and slows adoption. Various white-box studies have used internal representations to detect patterns of truthfulness and factuality. A less-studied approach is to identify feed-forward neurons correlated with hallucination. We propose a method to rank feed-forward neurons at the final prompt token using a custom neuron selection dataset. We transfer the selected neuron identities to train hallucination classifiers on other factual question answering datasets. Our work provides empirical evidence that probes trained using the features from the selected neurons perform on par with probes trained on internal states. We also analyze the distribution of selected neurons and the effect of layer depth on detection performance.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Hallucination Neurons and Where to Find Them: An Investigation into the existence of Hallucination Neurons

Interpretable machine learning for Large Language Models (LLMs) increasingly relies on sparse probing methods that identify small sets of neurons claimed to detect and causally influence behaviors such as factuality recall, safety alignment, and hallucination. These claims have important implications for model auditing...

Huseyin Cavus, Sebin Sabu, J. Spear et al. · 0 citations
#machine learning Preprint Sep 2026

Attention Dispersion as a Diagnostic Signal for Hallucination in Large Language Models

Large Language Models (LLMs) frequently exhibit hallucinations, presenting a major barrier to reliability in complex reasoning tasks. While traditional detection methods rely on output-based confidence metrics, these logits are often miscalibrated by modern alignment techniques. In this paper, we investigate the tempor...

S. More, Tanuja S. Pawar · 0 citations
#artificial intelligence Preprint Sep 2026

Domain-Specific Hallucination Detection in Large Language Models

Large language models generate fluent text that can contain unfaithful claims -- a phenomenon known as hallucination. We present a multi-signal detection pipeline combining fine-tuned DeBERTa-v3 classification, Monte Carlo (MC) Dropout uncertainty quantification, and temperature-scaled calibration for response-level ha...

Varun Teja Chundru, Debasmita Biswas · 0 citations
Open access Sep 2026

Mitigating Hallucinations in Finance-Based Multi-Agent Large Language Model Systems

Large language models are increasingly used for financial question answering, while they are prone to generating hallucinated content. In this research, we propose a multi-signal framework for hallucination detection and mitigation. Our framework combines six signals (entailment, semantic similarity, claim verification...

Rashmi Nagpal, Unyimeabasi Usua, Kailey Simons et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.