Evaluating artificial intelligence in the diagnosis of hip fractures: an analysis of sensitivity, specificity, positive and negative predictive values, and accuracy.
The Milvue SmartUrgences AI software may provide useful decision support during acute radiographic assessment, but AI output should complement rather than replace clinical and radiological assessment.
Abstract
Objectives
To evaluate the diagnostic performance of Milvue SmartUrgences AI software for detecting hip fractures on radiographs obtained in the emergency department (ED).
Methods
This retrospective diagnostic accuracy study included consecutive patients aged ≥ 60 years undergoing hip and pelvis radiography at two EDs between June 2023 and September 2024. The index test was Milvue SmartUrgences, which classified radiographs as "YES," "NO," or "DOUBT." The reference standard was established by a consensus panel of three musculoskeletal radiologists and two senior orthopaedic trauma surgeons, blinded to AI output. Sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), accuracy, and F1-score were calculated. The primary analysis included patients with definitive "YES" or "NO" classifications; "DOUBT" classifications were evaluated separately.
Results
Among 539 included patients, 487 received definitive AI classifications and were included in the primary analysis. The mean age was 83.4 years (SD 8.4), and 66% were female. The reference standard identified 191 hip fractures. Sensitivity was 97.4% (95% CI 94.0-98.9), and specificity was 86.5% (95% CI 82.1-89.9). PPV was 82.3% (95% CI 76.8-86.7), NPV 98.1% (95% CI 95.6-99.2), and accuracy 90.8% (95% CI 87.9-93.0). The F1-score was 89.2%.
Conclusions
Milvue SmartUrgences demonstrated high sensitivity and negative predictive value for hip fracture detection in a consecutive emergency department cohort. The software may provide useful decision support during acute radiographic assessment, but AI output should complement rather than replace clinical and radiological assessment.
PURPOSE
To estimate diagnostic performance of a deep learning algorithm for extremity fracture detection on radiographs in patients aged ≥ 2 years using a refined reference standard. Secondary, to compare positive predictive value (PPV) by the algorithm's built-in confidence level (high vs. low) and diagnostic performa...
Sabine Morris Delhez, Thomas Breiner Enøe, O. Gerke et al.· Emergency Radiology· 0 citations
BackgroundSeveral commercial artificial intelligence (Al) algorithms are available for detecting intracranial hemorrhage (ICH), but independent clinical validation remains limited.PurposeTo compare three commercially available Al algorithms for ICH detection on non- contrast head computed tomography (NCHCT).Material an...
Kristoffer Järlevi, Daniel Lindbom, Andreas Lindholm et al.· Acta Radiologica· 0 citations
AI models demonstrate high diagnostic accuracy in imaging-based ONFH diagnosis, however, the current evidence is constrained by the limited number of included studies, predominantly retrospective designs, and a lack of adequate external validation, and should therefore be interpreted with caution.
FeiLong Lu, Li-Rong Wang, Wen-Bin Zhang et al.· Journal of Medical Internet...· 0 citations
Background/Aim: AI in pediatric dentistry is used for the purpose of making an accurate diagnosis and assisting clinicians, dentists, and pediatric dentists in clinical decision making, developing preventive strategies, and is also a time-saving procedure. The current state-of-the-art artificial intelligence-based natu...
Mds Chinta Brahmaja, Mds Challagulla Anusha, Mds A. Ratnaditya et al.· Genetics and Molecular Resea...· 0 citations
Multimodal large language models can interpret medical images, but their performance for pediatric elbow radiographs remains uncertain. We evaluated the diagnostic performance of GPT-5.2 Instant as accessed through the ChatGPT web interface during the defined study period.
In this prospective, single-cente...
O. Taş, Mehmet Yorgun, R. Aktaş et al.· BMC Medical Imaging· 0 citations
Background Radiographs sometimes do not depict femoral-neck fractures, particularly radiograph-negative or indeterminate femoral-neck fractures, leading to delayed treatment and complications. Purpose To develop and externally evaluate a deep learning model, OccuNet, for detecting femoral-neck fractures on pelvic or hi...
Writing as a participant and researcher, PhD student JS Tan SM ’22 has co-authored a new book about the rise of tech worker protests and the employer backlash that followed.
Requirements in large systems rarely exist in isolation. Their meaning depends on the wider project context - other requirements, policies, decisions, tests, and implementation details. That becomes especially important when AI is used for review, because spotting a possible conflict or gap is only the beginning. ReqSpace explores how AI, visualisation, and connected project context can help reviewers understand those findings, trace the relationships behind them, and focus on the questions that…
AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.