Skip to content
Review Open access

Cross-System Legibility of a Practitioner-Derived Workflow-Error Taxonomy for Conversational AI: A Three-Comparator Agreement Study

Sep 2026 · medRxiv · 0 citations
Medicine

Abstract

Practitioner-derived taxonomies of conversational artificial intelligence (AI) workflow errors show low inter-rater agreement among human coders, leaving open whether the instrument is ill-specified or the judgments are inherently difficult. We delivered a locked eight-category workflow-error taxonomy verbatim, under standardized conditions, to three frontier large language model comparators from distinct developer lineages, which coded a documented 45-incident error corpus. On the same 16 incidents coded by three human reviewer-authors, comparator category agreement was substantial (Fleiss {kappa}=0.625) against slight human agreement ({kappa}=0.155); across the full corpus it was stable ({kappa}=0.632), and a 10-category refinement did not reduce it ({kappa}=0.690). Severity and a claimed-verification flag remained only fair in both arms and on both samples. Substantial cross-system consistency provides a legibility signal consistent with recoverable category distinctions, but cannot separate instrument clarity from shared model priors, and is not a validation of any coding.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.