YOLO11-based deep learning system for automated tubal patency classification in hysterosalpingography: a comparative study for clinical decision support.
While the results are promising for a novel application domain, the model's failure on clinically critical minority classes (Bilateral Blockage, Bilateral Patency) means it is not yet suitable for unsupervised clinical use.
Abstract
In the United States, around 500,000 hysterosalpingography (HSG) procedures are performed annually. One fluoroscopic procedure that is frequently used to evaluate tubal patency in infertile women is hysterosalpingography. Clinical decision-making depends on the fast and accurate classification of tubal patency results, but this procedure still depends on radiologist skill, which varies greatly across clinical situations. In order to automatically classify tubal patency categories in HSG pictures, this study suggests using YOLO11, a cutting-edge deep learning architecture. Three baseline models YOLOv8n, ResNet50, and EfficientNetB0 were used to train and evaluate YOLO11s using a publicly accessible clinical dataset of 892 real HSG images annotated by three board-certified radiologists and arranged into four pathological categories: bilateral patency, bilateral blockage, bilateral partial patency, and unilateral patency. Standardized training techniques were used to conduct experiments on GPU-accelerated infrastructure. With an inference speed of 19.98 ms per image and an overall test accuracy of 76%, YOLO11s demonstrated clinically relevant performance for the Unilateral Patency class (F1-score = 0.83, recall = 0.96). YOLO11s demonstrated competitive accuracy with much fewer parameters than ResNet50 (5.4 M vs. 25.6 M), outperforming ResNet50 (75.0%) and YOLOv8n (67.12%), matching EfficientNetB0 (76.19%) within 0.2%. The main factor restricting performance on minority classes was found to be class imbalance, with Unilateral Patency accounting for 66% of training images. While the results are promising for a novel application domain, the model's failure on clinically critical minority classes (Bilateral Blockage, Bilateral Patency) means it is not yet suitable for unsupervised clinical use. The proposed system should be considered as an exploratory research baseline requiring further development, class-imbalance mitigation, and prospective clinical validation before any clinical decision support application.
An advanced system capable of automatically detecting complicated appendicitis from ultrasound images was developed and was explained with gradient-weighted class activation mapping (Grad-CAM), which creates a heatmap of the regions responsible for the model's prediction of the infected areas.
Fahad Ahammed, Omar Faruq Shikdar, Navid Zaman et al.· 0 citations
Foreign body aspiration (FBA) of non-high-density objects (NHDFBs) in children is a critical pediatric emergency, posing risks of airway obstruction and requiring prompt diagnosis, which currently relies on clinician experience with low-dose computed tomography (LDCT). This study aimed to develop and validate a deep learning (DL) model for the automated detection of tracheobronchial NHDFBs in pediatric LDCT scans. A retrospective, multicenter cohort of 600 children with suspected FBA was utilized, with bronchoscopic confirmation as the gold standard. A ResUnet-based model was trained and evaluated on internal and external validation sets, with its performance systematically compared against junior and senior radiologists. The DL model achieved foreign body detection rates comparable to senior radiologists across training, internal test and external validation cohorts without significant intergroup differences, whereas both outperformed junior radiologists significantly (all
p
< 0.05). The DL model exhibited drastically shorter reading time (17.5 ± 2.4 s) than senior radiologists (80.2 ± 13.1 s) and junior radiologists (109.3 ± 16.8 s, all
p
< 0.05). The model maintained stable high diagnostic performance in all cohorts. Its sensitivity was significantly higher than junior radiologists in both validation sets (all
p
< 0.05), while sensitivity, specificity, PPV and NPV showed no statistical disparities between the DL model and senior radiologists. The DL model yielded AUC values of 0.93, 0.89 and 0.85 in the three cohorts, which were statistically equivalent to those of senior radiologists (0.94, 0.95, 0.90), and substantially superior to junior radiologists (0.82, 0.83, 0.79). The developed DL model achieves expert-level accuracy with superior efficiency for detecting pediatric NHDFBs on LDCT, demonstrating strong potential as a rapid, objective decision-support tool to enhance diagnostic workflows, particularly in settings with limited specialist availability.
Junzhong Liu, Qi Wang, Haogang Li et al.· Scientific Reports· 0 citations
Timely identification of children with ileocolic intussusception likely to fail air-enema reduction is critical to avoid delays and bowel perforation. However, even expert sonographers show inter-observer variability. We developed and prospectively validated a Vision Transformer (ViT) deep learning system to predict reduction failure from static B-mode ultrasound images. This multicenter bidirectional cohort study included 5602 children (4-60 months) who underwent air-enema reduction at 14 Chinese tertiary hospitals (retrospective cohort: 2019-2024). After data augmentation, 10,151 images (8122 training, 2029 validation) were used to train a ViT model for binary classification ("success" vs. "failure"). External validation was performed on a prospective cohort of 190 patients (March-June 2025), with three junior and three senior sonographers independently predicting outcomes. The study was approved by the Ethics Committee of Yijishan Hospital of Wannan Medical University (approval No. 2025-04) and registered with ChiCTR2500098673. The model achieved high internal performance (failure: accuracy 0.880, precision 0.969; success: accuracy 0.970, precision 0.898). In the prospective cohort, the ViT model achieved 93.7% overall accuracy, significantly higher than senior (74.7%) and junior (60.7%) sonographers (p < 0.05). This study innovatively applies ViT to assess pediatric ileocolic intussusception severity, providing an objective, accurate tool to support clinical decision-making and reduce treatment risks.
Jie Liu, Yue Wang, Danping Zeng et al.· npj Digital Medicine· 0 citations
The feasibility of developing a deep learning model for the detection of intussusception using a smaller dataset of POCUS images is demonstrated and fine-tuning models were best adapted to the screening nature of POCUS images.
A. Thyagachandran, Brian Lefchak, H. Murthy et al.· Frontiers in Radiology· 0 citations
Objective To develop and validate an interpretable multimodal machine learning model that integrates transvaginal ultrasound (TVUS) deep learning features with clinical risk factors for differentiating benign from malignant endometrial thickening in postmenopausal women, with the aim of reducing unnecessary invasive procedures. Methods In this retrospective single-centre diagnostic study, 601 postmenopausal women with histopathologically confirmed endometrial thickening (509 benign, 92 malignant) were identified consecutively and allocated to a training set (n = 420) and a stratified held-out test set (n = 181) at a 7:3 ratio. Deep learning features were extracted from manually segmented mid-sagittal TVUS images via four pretrained ResNet architectures and compressed into a single imaging score (DL_score) through sequential mRMR filtering and LASSO regression. Independent clinical predictors were identified by multivariate logistic regression. Ten fusion classifiers were trained on the combined feature set and evaluated by AUC, calibration, decision curve analysis, and SHAP-based interpretability. Results In the test set, all image-based models outperformed the clinical-only model (AUC, 0.797; 95% CI, 0.707–0.888), with ResNet152 achieving the highest single-modality AUC (0.850; 95% CI, 0.782–0.917). Among the combined models, logistic regression (LR) performed best, with an AUC of 0.906 (95% CI, 0.842–0.970), an accuracy of 84.6%, a sensitivity of 75.9%, and a specificity of 92.1%. The LR model was well calibrated (Brier score, 0.08; Hosmer–Lemeshow P = 0.60) and, on DCA, provided the greatest net benefit across the clinically relevant range of threshold probabilities. SHAP analysis identified the DL_score, postmenopausal bleeding, and BMI as the three most influential predictors; Grad-CAM activation maps indicated that the DL_score captured spatially localised information related to endometrial bulk and junction integrity. Conclusion The combined image–clinical logistic regression model demonstrated promising discriminatory and calibration performance for risk stratification of postmenopausal endometrial thickening, achieving improved specificity over conventional thickness-based criteria. This interpretable multimodal approach may serve as a useful adjunct to support clinical decision-making in identifying postmenopausal women at lower risk of malignancy, potentially reducing the burden of unnecessary invasive procedures.
Ling Li, Dan Yang, Yan Zhai et al.· Frontiers in Oncology· 0 citations
Benign Disease: New Technologies
Hiatal hernia diagnosis using barium swallow studies often requires multiple image acquisitions to visualize the esophagogastric junction adequately. Repeated acquisitions may increase radiation exposure and barium ingestion. Interpretation is observer-dependent. The aim of this study was to train an artificial intelligence model for detecting and classifying hiatal hernias.
A limited dataset of 70 anonymized barium swallow images centered on the esophagogastric junction was retrospectively analyzed and classified as hiatal hernia (n=45) or normal (n=25). Images were standardized using region-of-interest cropping, autocontrast adjustment, resizing to 512×512 pixels, and data augmentation to mitigate small-sample limitations. The dataset was divided into training (70%), validation (15%), and testing (15%) subsets.
Three pretrained convolutional neural networks (ResNet18, DenseNet121, EfficientNetB0) were fine-tuned using transfer learning. A structured grid search (72 trials) optimized learning rate, dropout, weight decay, and training epochs under identical cross-validation conditions. The primary evaluation metric was mean area under the curve (AUC). Secondary metrics included F1-score, accuracy, sensitivity, and specificity. Grad-CAM visualization was applied to assess anatomical regions influencing predictions.
Despite the limited dataset and use of focused images, all architectures demonstrated strong internal performance. DenseNet121 achieved mean AUC and F1 values of 1.00 in cross-validation, while EfficientNetB0 showed the strongest out-of-fold performance (AUC 0.9733; F1 0.9556). ResNet18 achieved mean cross-validation AUC 0.9956 and F1 0.9882.
Under a unified grid-search comparison framework, ResNet18 demonstrated the most consistent global performance (mean AUC 0.7860; F1 0.7662; accuracy 0.7571; sensitivity 0.6215; specificity 1.0000), supporting its selection as the final model. Grad-CAM confirmed consistent attention to the esophagogastric junction region. The prototype web application generated automated binary classification with confidence scoring.
Deep learning models can support detection and classification of hiatal hernias using focused AP barium swallow images, even when trained on a limited dataset. This training approach may enable automated identification and classification while potentially reducing radiation exposure and barium ingestion during diagnostic studies.
B. Borráez-Segura, Ricardo Arango-Slingsby, Lucas Dueñas-Ramirez et al.· Diseases of the esophagus· 0 citations