A semantic-safety gap is uncovered in air traffic control (ATC), and conventional scores give substantially higher performance estimates than consequence-aware evaluation, even for models that appear reliable under standard met- rics.
Abstract
Can language models be trusted in safety- critical operations? In such settings, strong per- formance on semantic metrics does not guaran- tee operational reliability: a misread altitude, a dropped execution condition, or a confused call- sign may score well under standard F1 yet carry sharply asymmetric operational consequences. We study this problem in air traffic control (ATC), where controller-pilot communication demands near-zero error tolerance, and use consequence-aware evaluation to test whether semantic scores misstate operational reliabil- ity. The framework is instantiated in a con- trolled diagnostic ATC benchmark grounded in aviation standards and feedback from 40 air traffic controllers across three countries. Evaluating 8 models, we uncover a system- atic semantic-safety gap: conventional scores give substantially higher performance estimates than consequence-aware evaluation, even for models that appear reliable under standard met- rics. Risk-aware fine-tuning narrows but does not close this gap, showing that consequence- aware evaluation is a necessary complement to standard NLP metrics before any real safety- critical deployment claim
This work proposes Sequential Epistemic and Action-Level Validation (SEAV), a verification-centric jailbreak evaluation framework that decomposes responses into ordered steps and evaluates both validity and correctness, and shows that enforcing correctness substantially reshapes measured robustness.
Qilong Wu, Sahil Wadhwa, Pranab Mohanty et al.· 0 citations
Evaluating state-of-the-art open LLMs reveals a significant robustness gap, and shows that reliable process-level verification remains challenging, and evaluator robustness should be measured separately from solver accuracy, even in a simple algebraic domain with exact ground truth.
A large-scale assessment of the effectiveness and robustness of these automated pipelines is conducted by evaluating five widely used benchmark suites across 26 open-source SLMs under a unified judging rubric, which reveals a capability-safety confound that mixes model capability with apparent safety.
Nyamtulla Shaik, Feng-Jun Li, Bo Luo· Lecture notes in computer sc...· 1 citation
This review synthesizes three literatures that rarely meet: statistical learning theory on hallucination, empirical measurements of LLM error rates in high-stakes domains, and the standards and regulatory documents that define certification.
Akbar Sayakov· American Impact Review· 0 citations
Experiments show that Cunning training improves robustness to out-of-distribution jailbreak attacks and strengthens subsequent safety fine-tuning, and suggest that cunning data can strengthen model vigilance and complement conventional safety alignment.
You-Jia Wang, Lin Xu, Yang Sun et al.· 0 citations
AtmosCoder-Bench is introduced, an execution-grounded benchmark that makes the calculation process visible, and finds that multiple-choice formats inflate measured accuracy by at least 12 percentage points.
Mao-Hao Ran, Chendong Ma, Yanting Zhang et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 24, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.