Skip to content
Review Open access

The Physical Fidelity Gap as an Evidence-Traceability Problem in AI Uncertainty Quantification: A Structured Review

Aug 2026 · Italian National Conference on Sensors · Vol 26, pp. 5447 · 0 citations · 51 references
Medicine

TL;DR

This study conducted a structured 15-dimensional coding review of studies and summarizes the above evidence discontinuity as the physical fidelity gap (PFG), which refers to incomplete or unverifiable evidential links between AI uncertainty claims publicly reported in the literature and the relevant physical, measurement, or operational conditions.

Abstract

Highlights What are the main findings? A total of 566 studies were included, of which 556 formed stable dominant AI uncertainty claim units and constituted the common analytic set for 15-dimensional coding. By evidence-traceability tier, 347 studies (62.4%) were classified as Weak, 133 (23.9%) as Medium, and 76 (13.7%) as Strong. What are the implications of the main findings? Specific risk claims, Multiple uncertainty entry points, Multiple physical information types, and Sampling/ensemble approximation more often co-occurred with higher evidence traceability. PFG provides a scope-bounded reference for locating breakpoints in the claim–evidence chain and indicating directions for evidence strengthening; its interpretation is limited to the traceability of evidence publicly reported in the literature. Abstract Artificial intelligence (AI) increasingly produces uncertainty outputs for sensing and measurement tasks, but the evidence supporting these outputs may not maintain a traceable correspondence with the relevant real-world conditions. This study conducted a structured 15-dimensional coding review of 566 studies, of which 556 formed stable dominant AI uncertainty claim units and entered the common analytic set. Each final code was linked to row-level audit evidence. The unidimensional distributions first showed breakpoint states with nonzero frequencies at the relevant nodes of the claim–condition–test–uncertainty-response chain, thereby confirming observable evidence discontinuities in the current corpus. The studies were then stratified using three sequential, non-compensatory evidence questions. Among the 556 studies, 347 (62.4%) were classified as Weak, 133 (23.9%) as Medium, and 76 (13.7%) as Strong. Descriptive cross-dimensional comparisons showed that specific OOD/drift risk or Error/quality estimation claims, Multiple uncertainty entry points, Multiple physical information types, and Sampling/ensemble approximation more often co-occurred with higher evidence traceability; Prediction reliability/confidence claims, a standalone Uncertainty proxy/score, and Latency/real-time inference constraints more often co-occurred with lower evidence traceability. This study summarizes the above evidence discontinuity as the physical fidelity gap (PFG), which refers to incomplete or unverifiable evidential links between AI uncertainty claims publicly reported in the literature and the relevant physical, measurement, or operational conditions. PFG provides a scope-bounded reference for locating links in the evidence chain that may need strengthening. The breakpoints observable in the current corpus indicate that the public evidence still has room for improvement in forming continuous, verifiable claim–condition–test–response correspondences; the tiered comparison indicates that subsequent work can strengthen the evidence chain by clarifying claim conditions, incorporating the relevant conditions into empirical testing, and reporting identifiable uncertainty responses and extended corroboration.

Read PDF

Similar papers

#artificial intelligence Review Sep 2026

What Can Component-Replacement Evidence Establish? A Critical Scoping Review of Local Decisions in LLM Agents

Background. A component replacement in a language-model agent changes an execution trajectory, potentially altering later observations, resource use, and recovery opportunities. Different evidence is needed to assess its task-level benefit and the contribution of local decision quality. Methods. This critical scoping r...

Shu-Yang Zhang, Jian-Shuo Chang · 0 citations
#artificial intelligence Preprint Sep 2026

When Agent Metrics Measure Different Things: An Evidence-Grounded Audit of the Praxa AI Pipeline

Agent evaluations can be numerically correct while measuring a different construct from the one implied by their labels. We present a retrospective measurement audit of selected Praxa AI implementation files, historical evaluation artifacts, and operational records. A 139-case offline routing report contains 112 passes...

Stefan Creadore, Peyton Woakz · 0 citations
Preprint Aug 2026

TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs

A probe corpus of 42 retracted, fraudulent, and pseudoscientific papers is paired with a methodology for eliciting and scoring single-shot model engagement with each paper's framing, indicating an urgent need for guardrail infrastructure for scientific deployment of language models.

V. Rodionov, Shamil Assylbekov · 0 citations
#artificial intelligence Preprint Sep 2026

Do Frontier Models Seek Safety Evidence Before Acting?

SAFE, a controlled benchmark in which models make deployment decisions with optional evidence that varies in retrieval cost, probability, severity, and presentation, is introduced and suggests that deployment-time safety depends not only on how models respond to known risks, but also on whether they acquire the evidenc...

Omer Tafveez · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.