Skip to content

What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness

Jul 2026 · arXiv.org · Vol abs/2607.08046 · 2 citations · ⚡ 1 influential · 23 references
Computer Science

TL;DR

Puzzling internal representations as a practical tool for calibrating, auditing, and triaging language model forecasters and reasoning models more broadly is established.

Abstract

Large language models fine-tuned for forecasting can be accurate yet poorly calibrated, and their chain-of-thought (CoT) reasoning may not faithfully reflect the evidence behind a forecast. We ask whether internal representations offer a more direct window into both. Working with Eternis-Forecaster 8B on OpenForesight, we train representation-pooling probes on intermediate activations and find they achieve substantially better calibration; a result that also holds for GLM-4.7-Flash and GLM-4.5-Air. We then assess CoT faithfulness through evidence ablation and diversionary injection: removing an influential source in the prompt often changes the model's forecast while leaving the reasoning trace untouched. The same probes function as lie detectors: their activations track behavioral shifts far better than the reasoning trace does, and they also predict the direction of change in 84% of cases, including when the CoT conceals the perturbation's influence. Finally, forced answering reveals that forecasts are largely fixed before reasoning begins: a single pre-reasoning pass recovers the committed answer and confidence, and routing questions by the spread of this pre-set answer distribution saves 30-47% of generated tokens, with no loss of accuracy. Together, these results establish probing internal representations as a practical tool for calibrating, auditing, and triaging language model forecasters and reasoning models more broadly.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Counterfactual Tests for Measuring Chain-of-Thought Faithfulness in Visual Language Models

The analysis shows that CoTs do not reliably track visual evidence that influences model predictions, and it is found that Predict-then-Explain explanations align more strongly with perturbation-induced probability shifts than pre-answer CoTs, while binary vCT scores are often nearly saturated.

Bayar Menzat, Max Süss, Rui-Zhi Wang et al. · 0 citations
#machine learning Preprint Sep 2026

Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning

This work operationalizes the importance of a reasoning step as its advantage: the change in expected reward, e.g., producing the correct final answer, from including that step, estimated via Monte Carlo rollouts, and evaluates whether LLM judges can identify high-advantage steps and finds that sufficiently capable LLM...

Kevin Du, A. Hoyle, L. Ruis et al. · 0 citations
#artificial intelligence Preprint Sep 2026

The Corroboration Illusion: When More News Makes LLM Forecasts Less True

Large language models (LLMs) are increasingly used to forecast real-world events by retrieving and reasoning over news. We show that this dependence on an open, crawlable news corpus creates a new attack surface: an adversary who can merely publish articles--without access to the retriever, the model, or the user's que...

Yuan-Lai Lu, Yu-Kuan Zhang · 0 citations
#natural language process... Preprint Sep 2026

An Empirical Study of Counterfactual Self-Explanations in LLMs

This work evaluates ten instruction-tuned models from the LLaMA-3 and Qwen-2.5 families, measuring faithfulness, minimality, and alignment with human-annotated rationales and shows that model scale is the strongest determinant of explanation quality.

Giannis Kalyvas, Giorgos Filandrianos, Orfeas Menis-Mastromichalakis et al. · 0 citations
#artificial intelligence Preprint Sep 2026

The Missing"I Don't Know": Why Three Reasoning-Reliability Findings Converge on Calibrated Abstention

It is argued these findings converge on a single intervention: calibrated abstention is what each independently identifies as the missing capability, even though the unavailability they document, a capability gap, a policy gap, and a recursion-theoretic gap, has a different source in each case.

Srijith Ravikumar · 0 citations
#artificial intelligence Preprint Sep 2026

From Concept Alignment to Causal Grounding: An Intervention Test of Chain-of-Thought Faithfulness

Chain-of-thought (CoT) can sound plausible yet be unfaithful to the model's underlying reasoning. Most prior work probes CoT faithfulness through input--output behavior or input attributions, leaving internal computation largely underexplored. We instead cast faithfulness as internal concept grounding: Does a large lan...

Qianli Wang, Yilong Wang, Dennis Wei et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.