Skip to content
#software testing Open access

Evaluating artificial intelligence in the diagnosis of hip fractures: an analysis of sensitivity, specificity, positive and negative predictive values, and accuracy.

Sep 2026 · Skeletal Radiology · 0 citations · 18 references
Medicine

TL;DR

The Milvue SmartUrgences AI software may provide useful decision support during acute radiographic assessment, but AI output should complement rather than replace clinical and radiological assessment.

Abstract

Objectives

To evaluate the diagnostic performance of Milvue SmartUrgences AI software for detecting hip fractures on radiographs obtained in the emergency department (ED).

Methods

This retrospective diagnostic accuracy study included consecutive patients aged ≥ 60 years undergoing hip and pelvis radiography at two EDs between June 2023 and September 2024. The index test was Milvue SmartUrgences, which classified radiographs as "YES," "NO," or "DOUBT." The reference standard was established by a consensus panel of three musculoskeletal radiologists and two senior orthopaedic trauma surgeons, blinded to AI output. Sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), accuracy, and F1-score were calculated. The primary analysis included patients with definitive "YES" or "NO" classifications; "DOUBT" classifications were evaluated separately.

Results

Among 539 included patients, 487 received definitive AI classifications and were included in the primary analysis. The mean age was 83.4 years (SD 8.4), and 66% were female. The reference standard identified 191 hip fractures. Sensitivity was 97.4% (95% CI 94.0-98.9), and specificity was 86.5% (95% CI 82.1-89.9). PPV was 82.3% (95% CI 76.8-86.7), NPV 98.1% (95% CI 95.6-99.2), and accuracy 90.8% (95% CI 87.9-93.0). The F1-score was 89.2%.

Conclusions

Milvue SmartUrgences demonstrated high sensitivity and negative predictive value for hip fracture detection in a consecutive emergency department cohort. The software may provide useful decision support during acute radiographic assessment, but AI output should complement rather than replace clinical and radiological assessment.

Read PDF

Similar papers

Review Open access Sep 2026

Diagnostic accuracy & clinical importance of AI confidence for extremity fracture detection: 2,508-patient retrospective cohort.

PURPOSE To estimate diagnostic performance of a deep learning algorithm for extremity fracture detection on radiographs in patients aged ≥ 2 years using a refined reference standard. Secondary, to compare positive predictive value (PPV) by the algorithm's built-in confidence level (high vs. low) and diagnostic performa...

Sabine Morris Delhez, Thomas Breiner Enøe, O. Gerke et al. · 0 citations
Review Open access Sep 2026

Retrospective comparison of three commercial artificial intelligence algorithms for detection of intracranial hemorrhage (ICH) in the emergency radiology department.

BackgroundSeveral commercial artificial intelligence (Al) algorithms are available for detecting intracranial hemorrhage (ICH), but independent clinical validation remains limited.PurposeTo compare three commercially available Al algorithms for ICH detection on non- contrast head computed tomography (NCHCT).Material an...

Kristoffer Järlevi, Daniel Lindbom, Andreas Lindholm et al. · 0 citations
Review Open access Aug 2026

Diagnostic Accuracy of Medical Imaging–Based Artificial Intelligence for Osteonecrosis of the Femoral Head: Systematic Review and Meta-Analysis

AI models demonstrate high diagnostic accuracy in imaging-based ONFH diagnosis, however, the current evidence is constrained by the limited number of included studies, predominantly retrospective designs, and a lack of adequate external validation, and should therefore be interpreted with caution.

FeiLong Lu, Li-Rong Wang, Wen-Bin Zhang et al. · 0 citations
Open access Sep 2026

“COMPARISON OF DIAGNOSTIC AND TREATMENT RELIABILITY OF ARTIFICIAL INTELLIGENCE (AI) SOFTWARE WITH HUMAN EVALUATION USING PANORAMIC RADIOGRAPHS”

Background/Aim: AI in pediatric dentistry is used for the purpose of making an accurate diagnosis and assisting clinicians, dentists, and pediatric dentists in clinical decision making, developing preventive strategies, and is also a time-saving procedure. The current state-of-the-art artificial intelligence-based natu...

Mds Chinta Brahmaja, Mds Challagulla Anusha, Mds A. Ratnaditya et al. · 0 citations
Review Open access Sep 2026

Diagnostic performance of GPT-5.2 ınstant for pediatric elbow fracture detection on radiographs: a prospective single-center diagnostic accuracy study

Multimodal large language models can interpret medical images, but their performance for pediatric elbow radiographs remains uncertain. We evaluated the diagnostic performance of GPT-5.2 Instant as accessed through the ChatGPT web interface during the defined study period. In this prospective, single-cente...

O. Taş, Mehmet Yorgun, R. Aktaş et al. · 0 citations
Sep 2026

Improving Recognition Performance and Reader Efficiency for Femoral-Neck Fracture Detection on Pelvic or Hip Radiographs Using Artificial Intelligence.

Background Radiographs sometimes do not depict femoral-neck fractures, particularly radiograph-negative or indeterminate femoral-neck fractures, leading to delayed treatment and complications. Purpose To develop and externally evaluate a deep learning model, OccuNet, for detecting femoral-neck fractures on pelvic or hi...

Xin-Xiao Lin, Yu-Rui Qian, Shi-Hao Zhou et al. · 1 citation

Related blog posts

MIT News · Artificial Intelligence Oct 2, 2026

Documenting the tech worker movement

Writing as a participant and researcher, PhD student JS Tan SM ’22 has co-authored a new book about the rise of tech worker protests and the employer backlash that followed.

GPT-Lab Sep 23, 2026

Requirements Don’t Live in Isolation: What We’re Exploring with Req-Space

Requirements in large systems rarely exist in isolation. Their meaning depends on the wider project context - other requirements, policies, decisions, tests, and implementation details. That becomes especially important when AI is used for review, because spotting a possible conflict or gap is only the beginning. ReqSpace explores how AI, visualisation, and connected project context can help reviewers understand those findings, trace the relationships behind them, and focus on the questions that…

GPT-Lab Sep 17, 2026

Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering

AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.