The feasibility of developing a deep learning model for the detection of intussusception using a smaller dataset of POCUS images is demonstrated and fine-tuning models were best adapted to the screening nature of POCUS images.
Abstract
Objective Intussusception is a pediatric emergency with delays in diagnosis. We sought to develop a novel deep learning model for the detection of target sign on point-of-care ultrasound (POCUS) images. Materials and methods POCUS image/video clip media files obtained during emergency department (ED) visits were included for study. Images underwent preprocessing to enhance and resize the region of interest (ROI), ImageNet high dimensional feature extraction and further training using various machine learning models, including classical machine learning, ensemble learning, transfer learning, and fine-tuning. Outputs were analyzed on a patient case, individual image frame and relative threshold bases. Results A POCUS image database of originally 785 media files from 49 patients, 8 of whom were positive for intussusception, was converted to 1,582 intussusception and 1,965 normal images for training. The output results show that the fine-tuning models performed better than the classical machine, ensemble and transfer learning models across three different analyses, and that the threshold-based approach to intussusception cases resulted in the greatest predictive performance. Discussion Intussusception is a potential candidate for development of deep learning tools due to limited capacity for pediatric focused imaging and accurate diagnosis. Prior studies are limited and have used either large private datasets or formal radiology studies. Modeling involved converting dynamic video files into several static images. Fine-tuning models were best adapted to the screening nature of POCUS images. Conclusion Our work demonstrated the feasibility of developing a deep learning model for the detection of intussusception using a smaller dataset of POCUS images.
Timely identification of children with ileocolic intussusception likely to fail air-enema reduction is critical to avoid delays and bowel perforation. However, even expert sonographers show inter-observer variability. We developed and prospectively validated a Vision Transformer (ViT) deep learning system to predict reduction failure from static B-mode ultrasound images. This multicenter bidirectional cohort study included 5602 children (4-60 months) who underwent air-enema reduction at 14 Chinese tertiary hospitals (retrospective cohort: 2019-2024). After data augmentation, 10,151 images (8122 training, 2029 validation) were used to train a ViT model for binary classification ("success" vs. "failure"). External validation was performed on a prospective cohort of 190 patients (March-June 2025), with three junior and three senior sonographers independently predicting outcomes. The study was approved by the Ethics Committee of Yijishan Hospital of Wannan Medical University (approval No. 2025-04) and registered with ChiCTR2500098673. The model achieved high internal performance (failure: accuracy 0.880, precision 0.969; success: accuracy 0.970, precision 0.898). In the prospective cohort, the ViT model achieved 93.7% overall accuracy, significantly higher than senior (74.7%) and junior (60.7%) sonographers (p < 0.05). This study innovatively applies ViT to assess pediatric ileocolic intussusception severity, providing an objective, accurate tool to support clinical decision-making and reduce treatment risks.
Jie Liu, Yue Wang, Danping Zeng et al.· npj Digital Medicine· 0 citations
An advanced system capable of automatically detecting complicated appendicitis from ultrasound images was developed and was explained with gradient-weighted class activation mapping (Grad-CAM), which creates a heatmap of the regions responsible for the model's prediction of the infected areas.
Fahad Ahammed, Omar Faruq Shikdar, Navid Zaman et al.· 0 citations
Objective Deep learning (DL), a core branch of artificial intelligence (AI), has revolutionized medical image analysis. Driven by advances in computational power and access to large-scale datasets, DL excels at extracting hierarchical features from complex, unstructured data. Accurate interpretation of imaging features is essential for diagnosing and treating middle and inner ear diseases. This narrative review aims to summarize current literature on the application of DL in otological imaging. Methods A narrative literature search was conducted using the PubMed and MEDLINE databases. Keywords related to AI or DL in otology were used to identify relevant articles published up to the time of writing. Results DL models have achieved favorable results in classifying common middle ear diseases (e.g., otitis media) using otoscopic images and in segmenting major inner ear structures (e.g., cochlea, ossicular chain) in 3D volumetric data. While 2D CNNs are mature for otoscopic classification, 3D U-Net and UNETR architectures dominate CT and MRI analysis. Models also show value in low-dose CT reconstruction and multimodal diagnosis. Conclusion DL has demonstrated strong potential to improve clinical decision-making and healthcare efficiency in otology. However, the field faces challenges related to data scarcity for rare diseases, poor segmentation performance for tiny structures (e.g., stapes), and a lack of integration with clinical workflows. Future efforts should focus on standardizing data and optimizing network structures for specific modalities.
Sining Wu, Yating Liao, Zhenhua Li· Frontiers in Neurology· 0 citations
We developed and evaluated a deep learning (DL) model for image-based detection of acute pancreatitis (AP) on abdominal contrast-enhanced CT (CECT). A total of 552 patients from two university centers (January 2010–January 2026) were included. The internal dataset comprised 207 patients with clinically and radiologically confirmed AP (499 scans) and 250 control patients with suspected AP (368 scans). An independent external validation cohort included 95 patients. Convolutional neural network–based models were trained using monophasic and biphasic CECT data. The final model was evaluated on a 20% patient-level hold-out test set from the internal cohort and on the external cohort. Performance was assessed using the F1 score and area under the receiver operating characteristic curve (AUROC). A single-input multiphase model incorporating arterial and portal venous phase scans achieved the best performance, with ensembling applied for biphasic studies. On the internal hold-out test set (n = 116), the model achieved an F1 score of 0.83 (95% confidence interval 0.75–0.89) and an AUROC of 0.89 (0.82–0.95). Performance remained robust on external validation (n = 95), with an AUROC of 0.99 (0.96–1.00) and an F1 score of 0.92 (0.86–0.97). DL enabled accurate CECT-based identification of AP in this retrospective multicenter cohort, with performance maintained in an independent external dataset. Prospective validation using broader and independently adjudicated clinical populations remains necessary. The model showed promising performance for CECT-based acute pancreatitis detection but was not designed or tested as a triage system. Diagnostic uncertainty in acute pancreatitis often arises from nonspecific abdominal symptoms and inter-reader variability in CECT interpretation. The DL model achieved high internal accuracy (AUROC 0.89) and maintained robust performance in an independent external validation cohort (AUROC 0.99). Diagnostic uncertainty in acute pancreatitis often arises from nonspecific abdominal symptoms and inter-reader variability in CECT interpretation. The DL model achieved high internal accuracy (AUROC 0.89) and maintained robust performance in an independent external validation cohort (AUROC 0.99).
Oleksandra Seidel, M. Theis, Sebastian Nowak et al.· European Radiology Experimen...· 0 citations
Objective To investigate the performance and clinical application potential of an improved YOLOv5-based deep learning model for automated detection and classification of pulmonary adenocarcinoma in situ (AIS) and minimally invasive adenocarcinoma (MIA), with particular emphasis on attention mechanisms and multiscale feature fusion for small pulmonary nodule detection. Methods CT images of 225 pathologically confirmed AIS or MIA lesions treated at the Affiliated People's Hospital of Fujian University of Traditional Chinese Medicine between December 2019 and June 2025 were retrospectively collected. In the deep learning workflow, the 225 lesions were first divided at the lesion level: 27 lesions (11 MIA and 16 AIS) were set aside, and the 218 original PNG images derived from these lesions were reserved as an independent test set that was excluded from all data augmentation and model training. The remaining lesions were converted into PNG format and further augmented using horizontal flipping, yielding a final dataset of 3,000 images for training and validation. Subsequently, the dataset was partitioned into training and validation sets at an 8:2 ratio. During this process, the origins of the augmented lesion images were manually verified to ensure that an original image and its corresponding augmented variants were not assigned to different subsets simultaneously. YOLOv5s was used as the baseline model. Local contrast-spatial attention modules (LCSA and LCSAv2), BiFPN, and ASFF were incorporated for attention enhancement and multiscale feature fusion. Five-fold cross-validation was performed. Mean average precision (mAP), precision, and recall were evaluated across IoU thresholds from 0.6 to 0.9, and ablation experiments were conducted to quantify the contribution of each module. Results In the baseline model selection stage, comparison under identical conditions showed that YOLOv5s achieved detection performance comparable to YOLOv11s while reducing GFLOPs by approximately 25%; therefore, YOLOv5s was selected as the improvement baseline. In five-fold cross-validation at IoU = 0.6, baseline YOLOv5s achieved an mAP@50 (AP at the matching IoU of 0.6) of 0.947, precision of 0.849, and recall of 0.932. After LCSA x ASFF was introduced, recall increased to 0.941 and mAP@50 was maintained at 0.946. Under the stringent localization criterion of IoU = 0.9, LCSA × ASFF achieved an mAP@50 of 0.878, a 3.8-percentage-point higher value than the baseline (0.846); precision (0.879) and recall (0.773) were also the highest among all models, with the smallest performance attenuation in the high-IoU range. Ablation experiments showed that the combined use of LCSA and ASFF produced stronger synergistic effects in the high-IoU range than either module alone, whereas LCSAv2 combined with ASFF did not show the expected synergistic gain. On the independent test set, baseline YOLOv5s achieved precision, recall, and mAP@50 values of 0.747, 0.715, and 0.753, respectively; the improved LCSA × ASFF model achieved corresponding values of 0.693, 0.753, and 0.788. Recall and mAP@50 increased simultaneously, indicating better overall detection performance than the baseline. The improved YOLOv5s-based deep learning model performed better in automated detection, and the LCSA × ASFF combination improved localization accuracy while maintaining recall through the synergy between local contrast enhancement and adaptive scale fusion. Conclusion The improved YOLOv5s model achieved automated detection and classification of small pulmonary nodules through end-to-end training, showing clear advantages in clinical workflow automation without manual intervention. The synergistic effect of local contrast enhancement and adaptive scale fusion effectively improved AIS/MIA localization precision and classification accuracy, supporting its potential as an assistive tool for early lung adenocarcinoma screening, pending external multi-centre validation.
Zhipeng Sun, Jinghui Chen, Lianxin Xie et al.· Frontiers in Medicine· 0 citations
Heart disease is significant health burden and a major cause of death worldwide, hence the need to develop automated diagnostic systems with correctness and computational efficiency in order to rescue lives through early clinical intervention. This paper introduces a fine-grained deep transfer learning architecture to classify cardiac disease images into multiple classes using two large convolutional backbones ResNet50 and DenseNet121, which are trained systematically with AdamW, stochastic gradient descent (SGD), and Lion under a single experimental environment. The suggested pipeline combines standardized image preprocessing, stratified division of data, transfer learning hierarchical feature extraction based on the feature, hyper parameter convergence analysis, and 50-epoch supervised fine-tuning, and then thorough performance and efficiency assessment. The experimental findings indicate the stable optimization dynamic in all six backbone optimizer setups, with the most consistent validation convergence in ResNet50 with Lion optimizer. The most successful configuration obtained nearly 93.9% accuracy of validation, precision, recall, and F1-score and a robust ROC-AUC of 0.995, which verified a high inter-class separability and stable threshold-independent result. The loss and validation accuracy curves also show that the convergence is quick and the ability to generalize is high. In terms of deployment, DenseNet121 was found to have much lower architectural complexity (8.0M parameters) and reduced inference latency (8.4 ms/image) than ResNet50 (25.6M parameters, 12.8 ms/image), and yet have a competitive level of classification. The graph driven analysis also shows that optimizer selection mostly influences the smoothness of convergence, predictive calibration, but backbone architecture influences a tradeoff between representational richness and computational efficiency. On the whole, the suggested framework is a clinically applicable and deployment-focused solution to smart computer aided diagnosis of cardiac diseases.