Skip to content

Calibrating Semantic Uncertainty from Observable Language-Model Probabilities

Jul 2026 · arXiv.org · Vol abs/2607.17447 · 3 citations · 37 references
Computer Science Mathematics

TL;DR

The proposed map turns semantic uncertainty in generative systems into an identifiable and testable statistical measurement problem and, when its acceptance conditions hold, yields an auditable posterior estimate.

Abstract

As generative artificial intelligence enters scientific and professional work, its uncertainty must be defined on the states that matter for inference and decision-making. Language models assign probabilities to words, whereas applications require uncertainty over meaningful states such as diagnoses, hypotheses or operational conditions. We introduce a \emph{semantic map}: a prespecified, testable bridge from probabilities over verbal responses to a posterior over declared finite states. The language distribution remains unrestricted; held-out calibration connects it to a reference posterior. We derive posterior-error bounds and conditions for existence, conditional uniqueness, presentation stability and stable inverse recovery. This distinction matters because language probabilities depend on prompt wording, while the target posterior should not change under information-equivalent rewording. Experiments use professional market text compiled from Federal Reserve economic and financial series, together with controlled simulations having exact posteriors. Across two fitted language models, language-derived probabilities outperform printed numerical confidence, recover held-out posteriors with valid uncertainty coverage, remain largely stable under paraphrase and respond appropriately to altered evidence. \textbf{Prompt engineering optimises a wording-dependent response; robust scientific use requires validated stability of application-relevant meaning.} The proposed map turns semantic uncertainty in generative systems into an identifiable and testable statistical measurement problem and, when its acceptance conditions hold, yields an auditable posterior estimate.

View source

Similar papers

Preprint Aug 2026

DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning

DirEAG is proposed, a Dirichlet Evidence Aggregation method that converts each elicited answer-confidence observation into calibrated soft evidence over generated candidate answers and an additional null state, allowing the model to represent cases where none of the candidates is correct.

Haorui Xu, Yu-Zhou Zhu, Li-Yuan Gao · 0 citations
#natural language process... Preprint Sep 2026

From Token Probabilities to Semantic Constraints: Towards Declarative Probabilistic Evaluation of Language Models

A declarative probabilistic framework for pre-training evaluation that makes the semantic structure of model behavior explicit and provides new formal tools for relating evaluation to learning, and links evaluation and learning through a shared semantics.

Kyle Richardson, C. Anderson, Pranav Balakrishnan et al. · 1 citation
#machine learning Review Sep 2026

Complex Problem Solving in Large Language Models: A Statistical Control Survey and Diagnostic Framework

This framework organizes existing methods around five components: explicit state representation, transition structuring, validation and constraint enforcement, search and rollback, and uncertainty management, and yields a diagnostic hypothesis: interventions should be most effective when they target the error or uncert...

Jia-Zhang Cai, Tao Wang, Rui-Dong Zhang et al. · 0 citations
Preprint Aug 2026

Attention-Path Fragility as an Uncertainty Signal in Large Language Models

It is proposed that a model's uncertainty about a token is reflected not only in the breadth of its output distribution but also in whether a confident prediction is \emph{fragile} under perturbation of its attention pathways, a training-free estimator that masks attention heads and measures the BALD mutual information...

Minsoo Kim, Sungyoung Ji, Kisung Moon et al. · 0 citations
#machine learning Preprint Sep 2026

ABSOL: Aggregated Bayesian Subsampling Orchestrated with LLMs

A hybrid LLM-guided Bayesian network structure-learning framework that uses LLMs as bounded semantic guides, and shows that language-derived semantic knowledge can substantially improve scalable probabilistic structure learning when used as bounded guidance within a statistically grounded reasoning pipeline.

Jackson Hassell, Chen Shen, Estevam Hruschka · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.