In this simulation, AI-generated notes scored higher on documentation quality, varied less, and carried fewer clinically significant errors than notes written on the same consultations by junior-to-middle-grade clinicians.
Abstract
Background Ambient AI documentation tools, known as scribes, are entering routine clinical practice at scale, but the evidence comparing the notes they produce against clinician-written notes is dominated by single-site, single-language studies that rely on human review to find errors, a method known to miss most documentation errors. Methods We conducted a paired simulation across five countries and languages (Cambridge/English, Barcelona/Spanish, Milan/Italian, Paris/French, Cologne/German; 385 paired consultations, 770 notes). From each actor-performed consultation, an AI scribe (Heidi) and a junior-to-middle-grade clinician independently produced a note. Notes were scored on the PDQI-9 by evaluators blinded to authorship. Documentation errors were identified by two methods of deliberately different sensitivity - clinician adjudication, and a calibrated automated reviewer externally validated against a blinded ten-clinician panel - then graded for clinical risk by a three-model panel. The co-primary outcomes were PDQI-9 total and Critical+High error burden, the latter reported under both detection arms. The analysis plan was registered before any pooling across sites. Results AI notes scored higher than clinician notes on the PDQI-9 (40.6 vs 35.6; difference +5.08, 95% CI 4.6-5.6; Cohen dz=0.55), consistently across all five sites (dz 0.41-0.75), and were less dispersed (5.7% of AI vs 27.8% of clinician notes fell below the study pre-specified low-score threshold (<32)). On the principal safety outcome - the paired probability that a note carried [≥]Critical+High error - clinician notes were affected more often under both detection arms: 61.0% versus 24.4% by the calibrated reviewer (relative risk 2.50, 95% CI 2.09-3.00) and 21.8% versus 6.2% by clinician adjudication (relative risk 3.50, 95% CI 2.32-5.27). The difference was largest for omissions. Unaided clinician review identified roughly 12% of the errors the calibrated reviewer retained, and a smaller fraction in AI notes than in clinician notes. Conclusions In this simulation, AI-generated notes scored higher on documentation quality, varied less, and carried fewer clinically significant errors than notes written on the same consultations by junior-to-middle-grade clinicians. The magnitude of the safety difference depends on the sensitivity of error detection, so we report both detection regimes and bound rather than point-estimate the absolute error rate. Extension to live practice, consultant-authored documentation, and notes as filed after clinician editing remains to be established.
Objectives Safety claims for ambient artificial intelligence (AI) scribes rest on automated judges that detect documentation errors and grade clinical risk. Expert reviewers are under-sensitive and disagree with one another, so no gold standard exists and validation cannot mean accuracy. We tested whether such judges are a defensible instrument: reproducible, within the envelope of expert disagreement, and non-differential across arms. Methods Pre-registered, blinded validation study nested in a multi-country simulation of ambient AI documentation (English setting), reported per GRRAS. Ten external clinicians independently adjudicated a stratified sample of 434 pipeline flags, retained and screen-discarded, blinded to note authorship, identification source, the pipeline's verdict and severity tier. Agreement used Gwet's AC1; proportions carry Wilson intervals. Three propositions were pre-specified: envelope parity, non-differential behaviour across arms, and concordance on consensus cases. Results All ten reviewers completed: 565 adjudications across 434 items, 131 of them double-rated. Inter-clinician agreement on genuineness was fair (raw 59%, 95% CI 50 to 67; AC1 0.24), leaving no human consensus to serve as truth. Judge-clinician agreement was 64% (95% CI 60 to 68), overlapping that interval. Behaviour was near-symmetric on contrast-critical metrics: kept-precision 74% for AI against 81% for clinician notes, and severity signed gap +0.06 against -0.09 tiers. One sub-metric was asymmetric: removed-confirmed 56% against 42%, so the screen over-removes more on clinician notes, a direction conservative to the parent contrast. On 77 consensus items the pipeline concurred on 70% (95% CI 59 to 79). Latent-class triangulation placed the genuine-error rate among flagged candidates at 68% (94% credible interval 48 to 83). Conclusions The judges behave as a consistent, near-non-differential, clinician-equivalent instrument. This licenses a directional AI-versus-clinician contrast under a non-differential misclassification argument, subject to its conditions. It is not a claim of accuracy, which moderate consensus concordance and fair reliability preclude, and the genuine-error rate is best reported as an interval.
H. Bergman, V. Liu, B. Austin et al.· medRxiv· 0 citations
Background Prediction models are central to advancing precision oncology, yet many fail to translate into clinical practice due to methodological flaws and inadequate validation. This review provides a practical, clinician-oriented guide to the statistical principles and advanced methods for developing, validating, and interpreting robust prediction models. Methods This narrative review used a targeted literature search of PubMed, Embase, and Web of Science to identify methodological papers, reporting guidelines, and representative oncology prediction model studies, with a focus on literature published between January 1, 2005, and February 28, 2025. Landmark methodological papers published before 2005 were also included when directly relevant. Rather than performing a systematic review or meta-analysis, we synthesized key statistical principles and illustrative examples to guide clinicians and researchers through model development, validation, interpretation, and clinical translation. Findings A multifaceted evaluation encompassing discrimination, calibration, clinical utility, and external validation is essential for prediction models. Over-reliance on discrimination metrics such as the area under the receiver operating characteristic curve (AUC), while neglecting calibration and clinical utility, can lead to misleading conclusions about a model’s value. Rigorous external validation in geographically or temporally distinct cohorts is the most direct test of generalizability, and performance degradation should be interpreted through root-cause analysis rather than treated simply as model failure. Key challenges include managing overfitting, selecting appropriate modeling and validation strategies for different oncology scenarios, addressing special settings such as rare tumors and real-world data, and improving the interpretability of complex “black-box” models. Conclusion Building a trustworthy prediction model requires a combination of advanced computational methods and rigorous statistical principles. To bridge the gap from model development to clinical impact, researchers must prioritize comprehensive validation, transparent reporting, scenario-appropriate modeling decisions, and critical assessment of a model’s real-world utility.
Xuexing Wang, Youxian Dou, Yufeng Wang et al.· Frontiers in Oncology· 1 citation
Aim/Background: Diagnostic uncertainty persists as a major driver of preventable patient harm in clinical practice. This feasibility study examined whether a structured data-submission protocol — the Rapid Dx Analyzer and Clinical Decision Tool (R-DA) — could improve the reliability of a commercial large language model (LLM) in generating differential diagnoses for complex clinical presentations. Methods: A single-user, retrospective case series was conducted. Twenty non-consecutive clinical cases were selected from an institutional database using pre-defined complexity criteria (involvement of two or
more organ systems, three or more differential diagnoses, ambiguous or conflicting data, or time-todiagnosis exceeding 48 hours). For each case, a structured prompt was submitted to Gemini 3.0 Pro via its web interface. The primary outcome was diagnostic concordance — defined as the model\'s top-ranked output matching the confirmed clinical diagnosis (established by biopsy, surgical findings, or definitive clinical course). The clinician providing input was not blinded to the final diagnosis.
Results: Concordance between the R-DA-generated output and the confirmed diagnosis was observed in all 20 cases (100%; 95% CI [Clopper-Pearson]: 83.2%–100%). This result should be interpreted with caution given the small sample size and the lack of independent adjudication.
Conclusion: These preliminary findings suggest that structured prompt engineering may meaningfully improve LLM-assisted diagnostic reasoning. The R-DA protocol warrants further investigation through prospective, multi-center trials with blinded adjudication. This study does not support claims of specialistlevel performance, but provides a hypothesis-generating foundation for future validation work
Aya Kawssan, A. Bazzal, Aktham Abdelhadi et al.· Journal of Medical Research...· 0 citations
Background/Objectives: The hospital discharge report is a critical document for care continuity that generates a substantial administrative burden for clinicians. Generative artificial intelligence (AI) offers the potential to reduce this burden while improving documentary quality. This study aims to compare, under real-world conditions with a GDPR-oriented architecture based on prior local anonymisation, the quality of AI-assisted discharge reports (IAIA) against those drafted by the responsible physician (INF). Methods: A retrospective, paired, expert-evaluation study was conducted at a Spanish university hospital. One hundred and twenty consecutive clinical cases from nine departments were included (240 reports total). Each case was independently evaluated by one of ten primary care physicians using a structured rubric covering 13 clinical dimensions (ordinal scale 1–3) and a global rating scale (1–10). The Wilcoxon signed-rank test was applied to all paired comparisons; effect size was estimated using the paired rank-biserial correlation (r). Results: IAIA achieved a significantly higher overall mean rating than INF (8.14 vs. 7.30 out of 10; p < 0.0001; r ≈ 0.76, large effect). IAIA was nominally superior in 9 of 13 clinical dimensions; after Bonferroni correction for the 13 per-dimension comparisons, six of these differences remained statistically significant, with the largest gains in family history, principal diagnosis hierarchy, and structured listing of secondary diagnoses. INF retained an advantage only in allergies and intolerances (2.69 vs. 2.46; p = 0.002), where IAIA tended to use generic formulas. Three dimensions showed no significant difference (prior treatment, physical examination, procedures). Conclusions: AI-assisted discharge reports received higher expert-rated documentary quality scores in a non-blinded paired evaluation across most evaluated dimensions. The physician-written report retained an advantage only in the safety-critical allergy domain, where allergy information must not be inferred by the model but sourced from verified structured fields or explicitly flagged as pending physician validation, supporting the need for a supervised hybrid model in which AI generates the initial draft while the clinician mandatorily validates sensitive content. Prior local anonymisation constitutes a GDPR-oriented approach to generative AI deployment in European hospital settings, substantially reducing the risk of disclosure of identifiable clinical information.
Daniela Velásquez-Villegas, Toni Alonso Solís, Alex Trejo-Omeñaca et al.· Healthcare· 0 citations
BACKGROUND
The creation of diagnostic criteria for a disease is always challenging in medicine. Consensus meetings and expert's opinions are sometimes responsible for establishing useful and precise diagnostic criteria but, unfortunately, not all the consensus meetings are based on evidence-based conclusions.
METHODS
A perspective analysis of the historical evolution of diagnostic criteria for Vogt-Koyanagi-Harada.
RESULTS
We examine the historical evolution of diagnostic criteria for Vogt-Koyanagi-Harada disease, an autoimmune condition targeting melanocyte-containing tissues, primarily affecting the choroid and potentially involving the skin, ears, and meninges. Early diagnostic criteria proposed in the late twentieth century were considered inadequate, leading to an international consensus conference in 1999 and the publication of revised diagnostic criteria in 2001. However, these criteria contained major conceptual flaws, notably the combination of acute and chronic clinical features that rarely occur simultaneously. In addition, important diagnostic tools such as indocyanine green angiography were largely excluded due to prevailing local practices and skepticism regarding their use. These limitations hindered accurate diagnosis and delayed recognition of VKH as a disease with distinct acute-onset and chronic forms requiring different diagnostic frameworks and management strategies. Subsequent studies corrected some deficiencies but often generated complex criteria difficult to apply in routine practice.
CONCLUSION
We would like to emphasize the risks of poorly structured consensus processes and propose safeguards to ensure scientifically rigorous and clinically useful recommendations.
C. Herbort, Ioannis Papasavvas, A. A. El-Asrar et al.· Journal of Ophthalmic Inflam...· 0 citations
Bottom-up incremental scoring showed the closest alignment with human assessment in clinical AI evaluation, underscoring the need for standardised prompt architectures in clinical AI evaluation.
What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.