Reference Standards in Artificial Intelligence Detection of Apical Periodontitis: A Narrative Review of What Reported Accuracy Measures
Abstract
The accuracy reported for detecting apical periodontitis using artificial intelligence (AI) is generally interpreted as accuracy against the disease; however, this is usually not the case. Compared with histopathology, single-view periapical radiography detects 16 to 27% of confirmed lesions and 38% with parallax, with a pooled digital sensitivity of 61.0% and a negative predictive value of 41.6%. Two blinded specialists reading 1717 radiographs disagreed on 22% of teeth. Detection depended on lesion size and site: simulated cancellous-bone lesions were invisible below 1 mm in incisors and 3 mm in molars. Of the 44 primary studies identified, 34 (77%) used a label that was a reading of the index image or was unascertainable; eight (18%) used an independent reference—six (14%) suited to contemporaneous detection and two to later treatment outcomes. No cone-beam computed tomography (CBCT) input study had a reference independent of its volume. One commercial platform reported 92.3% sensitivity against same-image consensus and 47.9% against blinded CBCT. The reported accuracy is thus in agreement with a noisy rater; recognising this changes what meta-analyses pool. A model that reproduces human labels would have a lower sensitivity against disease than reported; for real-world models, the direction depends on error dependence that the study design cannot identify.