Skip to content

Beyond Semantic Accuracy: Consequence-Aware Evaluation for Safety-Critical Language Understanding

Aug 2026 · 0 citations · 64 references
Computer Science

TL;DR

A semantic-safety gap is uncovered in air traffic control (ATC), and conventional scores give substantially higher performance estimates than consequence-aware evaluation, even for models that appear reliable under standard met- rics.

Abstract

Can language models be trusted in safety- critical operations? In such settings, strong per- formance on semantic metrics does not guaran- tee operational reliability: a misread altitude, a dropped execution condition, or a confused call- sign may score well under standard F1 yet carry sharply asymmetric operational consequences. We study this problem in air traffic control (ATC), where controller-pilot communication demands near-zero error tolerance, and use consequence-aware evaluation to test whether semantic scores misstate operational reliabil- ity. The framework is instantiated in a con- trolled diagnostic ATC benchmark grounded in aviation standards and feedback from 40 air traffic controllers across three countries. Evaluating 8 models, we uncover a system- atic semantic-safety gap: conventional scores give substantially higher performance estimates than consequence-aware evaluation, even for models that appear reliable under standard met- rics. Risk-aware fine-tuning narrows but does not close this gap, showing that consequence- aware evaluation is a necessary complement to standard NLP metrics before any real safety- critical deployment claim

View source

Similar papers

#artificial intelligence Preprint Aug 2026

Validity-Aware Jailbreak Evaluation for Large Language Models

This work proposes Sequential Epistemic and Action-Level Validation (SEAV), a verification-centric jailbreak evaluation framework that decomposes responses into ordered steps and evaluates both validity and correctness, and shows that enforcing correctness substantially reshapes measured robustness.

Qilong Wu, Sahil Wadhwa, Pranab Mohanty et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Beyond the Answer Key: Robustness Evaluation of Large Language Models for Step-Level Mathematical Verification

Evaluating state-of-the-art open LLMs reveals a significant robustness gap, and shows that reliable process-level verification remains challenging, and evaluator robustness should be measured separately from solver accuracy, even in a simple algebraic domain with exact ground truth.

Fatemeh Mazdarani, Carlos Toxtli · 2 citations

Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models

A large-scale assessment of the effectiveness and robustness of these automated pipelines is conducted by evaluating five widely used benchmark suites across 26 open-source SLMs under a unified judging rubric, which reveals a capability-safety confound that mixes model capability with apparent safety.

Nyamtulla Shaik, Feng-Jun Li, Bo Luo · 1 citation
Review Open access Aug 2026

Why Large Language Models Cannot Be Certified for Safety-Critical Systems

This review synthesizes three literatures that rarely meet: statistical learning theory on hallucination, empirical measurements of LLM error rates in high-stakes domains, and the standards and regulatory documents that define certification.

Akbar Sayakov · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond Routine Compliance: Cunning Data Cultivates Safety Vigilance in Large Language Models

Experiments show that Cunning training improves robustness to out-of-distribution jailbreak attacks and strengthens subsequent safety fine-tuning, and suggest that cunning data can strengthen model vigilance and complement conventional safety alignment.

You-Jia Wang, Lin Xu, Yang Sun et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.