Skip to content
Review Open access

From Output Errors to Workflow Harm: A Practitioner-Audit Method for LLM-Mediated Research

Aug 2026 · medRxiv · 0 citations
Medicine

TL;DR

TRACE (Tracking Reliability of AI-generated Conversational Evidence), a practitioner-audit framework for evaluating the downstream workflow reliability of conversational AI, is presented, suggesting taxonomy legibility under standardized conditions even where human judgment diverged.

Abstract

Objective. Formal large language model (LLM) evaluations score isolated prompts, but clinicians and health-informatics researchers meet model failures inside multi-step workflows where erroneous output can alter procedures or contaminate documents. We present TRACE (Tracking Reliability of AI-generated Conversational Evidence), a practitioner-audit framework for evaluating the downstream workflow reliability of conversational AI. Materials and Methods. A method paper with an empirical demonstration: 45 documentation-positive incidents recorded by one clinician-informatician across scholarly, clinical informatics, and clinical-adjacent workflows over seven weeks, coded with a consequence-based severity rubric, an error definition, a taxonomy crosswalk, and a Response-Audit Scorecard. Three reviewer-authors independently coded a 16-incident subsample; three vendor-blinded AI comparators applied the taxonomy to all 45 incidents. Results. Four categories tied as most frequent: verification failure, factual numerical error, tool-behavior misunderstanding, and citation or reference formatting (n=7 each). Four workflow-harm patterns recurred: procedural propagation, documentary contamination, trust-calibration disruption, and user-borne corrective burden, and one incident carried an estimated $2500 impact. Category agreement across three human reviewer-authors was low (Fleiss {kappa}=0.155), whereas three AI comparators agreed substantially (Fleiss {kappa}=0.632), suggesting taxonomy legibility under standardized conditions even where human judgment diverged. Discussion. Category assignment is comparatively legible, whereas severity and claimed-verification remain judgment-dependent. The claimed-verification gap is a measurable failure mode distinct from hallucination, sycophancy, and over-refusal. Conclusion. Practitioner audits with structured response scoring complement benchmarks by documenting workflow harm as an applied evaluation unit for clinical informatics and public-health work; this is a pilot that motivates, not estimates, error rates or cross-model comparisons.

Read PDF

Similar papers

Review Open access Sep 2026

Cross-System Legibility of a Practitioner-Derived Workflow-Error Taxonomy for Conversational AI: A Three-Comparator Agreement Study

Practitioner-derived taxonomies of conversational artificial intelligence (AI) workflow errors show low inter-rater agreement among human coders, leaving open whether the instrument is ill-specified or the judgments are inherently difficult. We delivered a locked eight-category workflow-error taxonomy verbatim, under s...

D. Austria, B. McCollister, J. Lindsey et al. · 0 citations
Open access Sep 2026

A framework for developing, validating, and utilizing automated measures of order errors using the retract-and-reorder methodology

The methodology described can be applied to develop automated measures to detect a range of order error types, examine the epidemiology, and investigate the root causes of order errors in near-real-time, as well as rigorously evaluate the impact of preventive interventions.

P. Spector, A. Grauer, Jerard Z. Kneifati-Hayek et al. · 0 citations
#generative ai Review Sep 2026

Vibe coding and the illusion of competence: why laboratory medicine specialists shouldn’t trust their AI-generated code for clinical deployment

An effective approach is that laboratory specialists use generative AI to prototype clinical concepts in a sandbox environment, followed by collaboration with qualified software developers and regulatory experts who translate them into production-grade systems.

S. De Bruyne, Stef Rommes, T. Fiers et al. · 0 citations

PRISMA-LLM: An Empirical Reporting Framework for AI-Assisted Systematic Reviews

Large language models (LLMs) and AI-enabled software increasingly participate in systematic-review decisions, yet the information needed to audit these workflows is reported inconsistently. We analyze SciLitBench, a corpus of 888 review-automation papers with 14,726 annotations, to characterize changes in methods, revi...

Miguel Zabaleta, Bai-Han Lin · 0 citations
Review Open access Aug 2026

The CODEX action incubator: a consensus-driven approach to identify and implement diagnostic excellence measures in the context of artificial intelligence

Abstract Diagnostic errors are a substantial source of patient harm. As artificial intelligence (AI) integrates into clinical workflows, opportunities are emerging to assess their impacts on diagnostic excellence (DxEx). The Coordinating Center for Diagnostic Excellence (CODEX) at the University of California San Franc...

B. Rosner, Molly Hammer, Aaron Tabacco et al. · 0 citations
Review Open access Aug 2026

Are automated documentation-error judges fit to measure ambient AI scribes? A pre-registered, blinded human-validation study

The judges behave as a consistent, near-non-differential, clinician-equivalent instrument, which licenses a directional AI-versus-clinician contrast under a non-differential misclassification argument, subject to its conditions.

H. Bergman, V. Liu, B. Austin et al. · 1 citation · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.