TRACE (Tracking Reliability of AI-generated Conversational Evidence), a practitioner-audit framework for evaluating the downstream workflow reliability of conversational AI, is presented, suggesting taxonomy legibility under standardized conditions even where human judgment diverged.
Abstract
Objective. Formal large language model (LLM) evaluations score isolated prompts, but clinicians and health-informatics researchers meet model failures inside multi-step workflows where erroneous output can alter procedures or contaminate documents. We present TRACE (Tracking Reliability of AI-generated Conversational Evidence), a practitioner-audit framework for evaluating the downstream workflow reliability of conversational AI. Materials and Methods. A method paper with an empirical demonstration: 45 documentation-positive incidents recorded by one clinician-informatician across scholarly, clinical informatics, and clinical-adjacent workflows over seven weeks, coded with a consequence-based severity rubric, an error definition, a taxonomy crosswalk, and a Response-Audit Scorecard. Three reviewer-authors independently coded a 16-incident subsample; three vendor-blinded AI comparators applied the taxonomy to all 45 incidents. Results. Four categories tied as most frequent: verification failure, factual numerical error, tool-behavior misunderstanding, and citation or reference formatting (n=7 each). Four workflow-harm patterns recurred: procedural propagation, documentary contamination, trust-calibration disruption, and user-borne corrective burden, and one incident carried an estimated $2500 impact. Category agreement across three human reviewer-authors was low (Fleiss {kappa}=0.155), whereas three AI comparators agreed substantially (Fleiss {kappa}=0.632), suggesting taxonomy legibility under standardized conditions even where human judgment diverged. Discussion. Category assignment is comparatively legible, whereas severity and claimed-verification remain judgment-dependent. The claimed-verification gap is a measurable failure mode distinct from hallucination, sycophancy, and over-refusal. Conclusion. Practitioner audits with structured response scoring complement benchmarks by documenting workflow harm as an applied evaluation unit for clinical informatics and public-health work; this is a pilot that motivates, not estimates, error rates or cross-model comparisons.
Practitioner-derived taxonomies of conversational artificial intelligence (AI) workflow errors show low inter-rater agreement among human coders, leaving open whether the instrument is ill-specified or the judgments are inherently difficult. We delivered a locked eight-category workflow-error taxonomy verbatim, under s...
D. Austria, B. McCollister, J. Lindsey et al.· medRxiv· 0 citations
The methodology described can be applied to develop automated measures to detect a range of order error types, examine the epidemiology, and investigate the root causes of order errors in near-real-time, as well as rigorously evaluate the impact of preventive interventions.
P. Spector, A. Grauer, Jerard Z. Kneifati-Hayek et al.· JAMIA Open· 0 citations
An effective approach is that laboratory specialists use generative AI to prototype clinical concepts in a sandbox environment, followed by collaboration with qualified software developers and regulatory experts who translate them into production-grade systems.
S. De Bruyne, Stef Rommes, T. Fiers et al.· Clinical Chemistry and Labor...· 0 citations
Large language models (LLMs) and AI-enabled software increasingly participate in systematic-review decisions, yet the information needed to audit these workflows is reported inconsistently. We analyze SciLitBench, a corpus of 888 review-automation papers with 14,726 annotations, to characterize changes in methods, revi...
Abstract Diagnostic errors are a substantial source of patient harm. As artificial intelligence (AI) integrates into clinical workflows, opportunities are emerging to assess their impacts on diagnostic excellence (DxEx). The Coordinating Center for Diagnostic Excellence (CODEX) at the University of California San Franc...
B. Rosner, Molly Hammer, Aaron Tabacco et al.· Diagnosis· 0 citations
The judges behave as a consistent, near-non-differential, clinician-equivalent instrument, which licenses a directional AI-versus-clinician contrast under a non-differential misclassification argument, subject to its conditions.
H. Bergman, V. Liu, B. Austin et al.· medRxiv· 1 citation· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.