Diagnostic Accuracy of Cybersight AI for Glaucoma Detection: Sensitivity, Specificity, and ROC Analysis
Abstract
Glaucoma is a chronic, progressive optic neuropathy and a leading cause of irreversible blindness worldwide. Early detection is crucial but is often limited by nonspecific clinical signs and the need for specialist expertise. Artificial intelligence (AI), particularly deep learning, has shown promise for glaucoma detection from ophthalmic images. This study independently validated an existing third-party AI tool, Cybersight AI, powered by the Visulytix Pegasus engine, for detecting primary glaucoma. No new algorithm was developed or proposed. A cross-sectional diagnostic accuracy study was conducted on 229 eyes from 121 patients, including 133 eyes with primary glaucoma and 96 controls, at Hue Central Hospital between December 2024 and September 2025. Fundus photographs were analyzed using the Cybersight AI platform, which estimated the vertical cup-to-disc ratio and generated a binary glaucoma referral recommendation. AI outputs were compared with diagnoses established by ophthalmologists using an expert multimodal clinical reference standard. Diagnostic performance was assessed using sensitivity, specificity, positive and negative predictive values, and the area under the receiver operating characteristic curve (AUC). At its default operating point, Cybersight AI achieved 51.9% sensitivity (95% CI 43.5–60.2) and 90.6% specificity (95% CI 83.1–95.0), with an AUC of 0.744 (95% CI 0.682–0.807). Positive and negative predictive values were 88.5% and 57.6%, respectively. The Youden-optimal cut-off was 10%, yielding 60.9% sensitivity and 85.4% specificity. Cybersight AI demonstrated moderate discrimination, high specificity, and modest sensitivity at its default binary setting in this Vietnamese tertiary-care cohort. It may serve as a screening or triage adjunct, although context-specific thresholding and further external validation are required before broad clinical deployment. Limitations include the single-center convenience sample, within-patient clustering, and comparison of a fundus-only index test against a multimodal reference standard. Future studies should prioritize multicenter validation in Vietnamese settings and comparison with expert fundus assessment alone.