Statistical Methods for Assessing Diagnostic Agreement
Abstract
With the rise of artificial intelligence (AI), an increasing number of AI-based diagnostic tools are being developed. Before clinical implementation, these tools must be validated against existing gold standards. This requires trials that quantify the agreement between AI predictions and reference measurements. However, designing such agreement studies poses methodological challenges that differ substantially from classical superiority trials. This paper aims to provide statistical methods for assessing diagnostic agreement. Methods were categorized according to the measurement scale of the data (nominal, ordinal, continuous)—with a separate group for methods that apply across several scales—and according to the number of raters or measurements involved. A decision tree is provided as a simplified educational framework for method selection rather than as a general method-selection algorithm: design features such as repeated measurements, clustering, spectrum effects, dependence between raters, and an imperfect reference method are not encoded in it and are discussed separately, together with the circularity and confounding issues specific to the validation of AI-based tools. For each method, we summarized assumptions, appropriate use cases, interpretation of results, and available open-source software for sample size calculation and analysis, and we illustrate the sample size calculations in three fully worked examples covering binary, ordinal, and continuous outcomes. We further distinguish conditional inference about one fixed, frozen model version from the broader generalization to a class of algorithms or to future model versions, which require additional sources of algorithmic and dataset variability to be represented in the design and analysis. We conclude by outlining open methodological questions—including Bayesian approaches to agreement estimation, methods for complex AI outputs, agreement models for clustered and repeated-measures designs, and the limited software support for Gwet’s AC1/AC2 sample size planning—that warrant further work as diagnostic technologies and statistical methodology continue to evolve.