Skip to content

When 99.98% is too Good to be True: Preventing Rule-Induced Overfitting in Embodied Clinical AI for Surgical Readmission Prediction

Jul 2026 · 2026 ITU Kaleidoscope - AI and Frontier Technologies for Good (ITU K) · pp. 1-8 · 0 citations · 28 references

Abstract

Embodied Artificial Intelligence (AI) systems are increasingly used to support clinical decision-making, telemedicine follow-up, and resource allocation, particularly in remote and resource-constrained healthcare settings. In these deployments, predictive models are embedded within clinical workflows and operate under human oversight, making safety, transparency, and reliability essential for regulatory-compliant use. A key but underexplored factor affecting trustworthiness is how supervision signals are derived from unstructured clinical text. This paper analyses the impact of text-derived label construction on embodied clinical AI for postoperative risk monitoring. Using surgical readmission prediction from free-text clinical notes, we compare two weak supervision strategies: a myopic keyword-based heuristic and a context-aware labeling framework that accounts for negation, temporal scope, and clinical severity. Although the naive approach achieves near-perfect accuracy (up to 99.98%), we show that this performance is driven by rule-induced target leakage, where models learn documentation artifacts rather than clinically meaningful deterioration signals. We propose an auditable, context-aware labeling protocol aligned with the requirements of deployable clinical decision support systems. While trading inflated accuracy for more realistic performance, the proposed approach improves discrimination and patient-level calibration-properties essential for safe human-in-the-loop operation in telemedicine and remote care. These findings highlight that trustworthy embodied AI depends not only on model sophistication, but also on clinically grounded and transparent supervision mechanisms.

View source

Similar papers

Open access 2026

Evaluating Uncertainty Quantification in Clinical Machine Learning: Calibration, Robustness, and Decision Utility under Distribution Shift

A rigorous empirical framework is presented for comparing three uncertainty quantification approaches on two clinical prediction tasks, in-hospital mortality and 30-day readmission, using 74,829 ICU admissions from the MIMIC-IV database to support a more demanding evaluation standard for UQ in clinical machine learning...

Isaac Tosin Adisa, Francis Mawutor Amuyao, Ezekiel Olaoluwa Joaquim · 0 citations
Open access Aug 2026

An Explainable AI-Driven Framework for Integrating Diverse Health Data to Enhance Predictive Accuracy and Clinical Interpretability

A framework through XAI to incorporate various health data sources such as electronic health records, medical imaging, laboratory reports, and wearable sensor information, which can be integrated in the context of achieving higher predictive performance in disease prediction and treatment stratification, as well as dec...

M. Aparna, S. Lahane, Dr. Bharti A. Dixit · 0 citations
Sep 2026

[Accuracy alone is not enough: cognitive safety of artificial intelligence in medicine.]

Cognitive safety is proposed here as a longitudinal property of the clinician-AI-organization sociotechnical system: its capacity to support or improve clinical performance without eroding independent hypothesis generation, uncertainty calibration, reasoned dissent, metacognitive control, and resilient performance when...

S. Corrao · 0 citations
Review Open access Jul 2026

From Algorithm to Bedside: A Clinician's Framework for AI in Practice

This article proposes seven questions that clinicians can run through to evaluate any clinical AI tool in the time it takes to read an abstract, alongside a traffic-light schema for matching oversight to risk and a short list of demands clinicians should make of vendors and institutions.

Alaa Abdelqader, M. Alkhateeb, Abdullah Al-Marrawi et al. · 0 citations
Review Open access Aug 2026

AI safety evaluation in an underrepresented population: real-world performance of clinical decision support and frontier language models on Medicaid patient messaging triage

No evaluated tool or combination was sufficiently accurate to enable physician-unassisted triage in this setting of patient-initiated text messages in a multi-state Medicaid population.

S. Basu, Sadiq Y. Patel, Parth Sheth et al. · 0 citations
#artificial intelligence Review Aug 2026

AI Morbidity and Mortality: A Framework for Clinical AI Failure Review

AI Morbidity and Mortality (AI M&M) is intended to complement, rather than replace, model monitoring, patient safety reporting, and regulatory oversight by converting individual AI-in-workflow failures into actionable institutional learning.

Paulius Mui, Dean F. Sittig, Steven Labkoff et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.