Skip to content
Review Open access

Why Large Language Models Cannot Be Certified for Safety-Critical Systems

Aug 2026 · American Impact Review · Vol 1, pp. e2026066 · 0 citations · 26 references

TL;DR

This review synthesizes three literatures that rarely meet: statistical learning theory on hallucination, empirical measurements of LLM error rates in high-stakes domains, and the standards and regulatory documents that define certification.

Abstract

Large language models (LLMs) are moving rapidly into domains governed by functional-safety certification: aviation, road vehicles, medicine, and critical infrastructure. Certification regimes such as IEC 61508, DO-178C, and ISO 26262 combine system-level risk targets with process-, traceability-, configuration-, and evidence-based obligations, in mixes that differ by regime but share a demand for verifiable, bounded behavior. This review synthesizes three literatures that rarely meet: statistical learning theory on hallucination, empirical measurements of LLM error rates in high-stakes domains, and the standards and regulatory documents that define certification. The convergent finding is that the mismatch between the two worlds is structural rather than incidental. Impossibility results establish a nonzero error floor for calibrated probabilistic generators under stated conditions; reported benchmark metrics in legal, medical, and agentic tasks are not directly convertible to certification targets, and no published deployment supplies the system-level hazard and exposure model that a compliance demonstration would require; and every major mitigation family (retrieval augmentation, guardrails, formal verification, uncertainty quantification, and neurosymbolic hybrids) is documented to narrow, but not close, the gap. We formalize error compounding over execution horizons, tabulate the documented limits of each mitigation, and examine why plausibility cannot substitute for assurance. Two coherent exits emerge: statistical acceptance criteria in standards for bounded tasks, and architectures that confine the LLM to an untrusted-proposer role inside a deterministic, independently verifiable execution envelope. Certifying the envelope rather than the model is, on the current evidence, the most defensible path consistent with both the mathematics and the standards.

Read PDF

Similar papers

#natural language process... Preprint Aug 2026

Beyond Semantic Accuracy: Consequence-Aware Evaluation for Safety-Critical Language Understanding

A semantic-safety gap is uncovered in air traffic control (ATC), and conventional scores give substantially higher performance estimates than consequence-aware evaluation, even for models that appear reliable under standard met- rics.

Yujing Chang, Thinh Pham, Van-Phat Thai et al. · 0 citations
#software testing Preprint Sep 2026

A Large-Scale Empirical Study of Quality Assurance Practices and Gaps in AI Agents

A large-scale empirical study of quality assurance (QA) practices in 157 open-source LLM-based agent projects with at least 100 GitHub stars highlights the need to move beyond feature-level testing toward systematic end-to-end validation that ensures agent workflows remain within intended boundaries when interacting wi...

Wu-Yang Dai, Moses Openja, Jiho Shin et al. · 0 citations
Open access Sep 2026

Development of a Framework for Evaluating Large Language Model Safety and Reliability: a Proof-of-Concept Evaluation

Large language models (LLMs) are entering clinical decision support faster than methodology can characterise their safety. Aggregate accuracy treats all errors as interchangeable and cannot support safe deployment under Software as a Medical Device (SaMD) and EU AI Act frameworks. To develop and demonstrate a framework...

Fang-Yan Liu, Zhi Liu, Xiaolu Fei et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Beyond the Answer Key: Robustness Evaluation of Large Language Models for Step-Level Mathematical Verification

Evaluating state-of-the-art open LLMs reveals a significant robustness gap, and shows that reliable process-level verification remains challenging, and evaluator robustness should be measured separately from solver accuracy, even in a simple algebraic domain with exact ground truth.

Fatemeh Mazdarani, Carlos Toxtli · 2 citations

Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models

A large-scale assessment of the effectiveness and robustness of these automated pipelines is conducted by evaluating five widely used benchmark suites across 26 open-source SLMs under a unified judging rubric, which reveals a capability-safety confound that mixes model capability with apparent safety.

Nyamtulla Shaik, Feng-Jun Li, Bo Luo · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.