Jul 2026· Journal of the American Society of Echocardiography· 0 citations
Medicine
TL;DR
MitralVision reliably distinguishes clinically significant MR using single-view B-mode echocardiography without Doppler input for model inference and may support more standardized MR screening.
Abstract
INTRODUCTION
Mitral regurgitation (MR) is one of the most prevalent valvular heart diseases, and its diagnosis traditionally relies on Doppler echocardiography, which is subject to significant variability and technical challenges. We developed and externally validated MitralVision, a deep learning model for automated classification of clinically significant MR using single-view, B-mode echocardiographic loops.
Methods
MitralVision, a deep neural network, was trained on 28,487 apical four-chamber (A4C) B-mode echocardiographic cine loops from 11,244 studies across 20 U.S. states. The model was designed to differentiate clinically significant (moderate/severe) from non-significant (none/trace/mild) MR using grayscale cine loops without Doppler input. External validation was performed on 629 studies from 26 independent clinical sites using the original clinical interpretation as the reference standard. A separate board-certified Level III echocardiographer independently re-graded all external validation studies to assess interobserver variability.
Results
On external validation, MitralVision achieved an AUROC of 0.91, sensitivity 82.1%, specificity 84.3%, negative predictive value 91.6%, and positive predictive value of 69.3%, compared with the original clinical read. Interobserver agreement between the original clinical read and the additional expert reader was 75.2% for clinically significant MR, with exact agreement across five MR grades of 35.5%. When benchmarking the additional expert reader's interpretation, model AUROC was 0.89. The model demonstrated excellent calibration (Brier score 0.12; expected calibration error 0.03).
Conclusions
MitralVision reliably distinguishes clinically significant MR using single-view B-mode echocardiography without Doppler input for model inference and may support more standardized MR screening. This streamlined AI-based approach offers reproducible MR assessment and may be compatible with future workflow implementation in high-throughput or resource-limited settings.
BACKGROUND
Accurate assessment of left ventricular outflow tract (LVOT) gradients is critical for hypertrophic cardiomyopathy management, yet Doppler-based measurements are technically demanding and require expertise. The objective of this work was to develop a multi-view deep learning model capable of classifying LVOT obstruction (>20 mm Hg) using routine 2-dimensional echocardiographic windows without reliance on Doppler imaging.
METHODS
We trained and externally validated a cross-attention-based video-to-video fusion framework that integrated EchoPrime-derived video representations from 3 standard transthoracic echocardiographic views to classify LVOT gradients.
RESULTS
Training was performed on a derivation cohort (N=1833) from a tertiary care system in the United States, with model performance evaluated on an internally held-out test set (N=275) and a Korean external validation cohort (N=46). Single-view baselines showed limited discrimination (external area under the receiver operating curves, 0.47-0.70). Conversely, the domain-specific foundational model (EchoPrime) achieved superior single-view performance (area under the receiver operating curves, 0.75-0.80 internal; 0.79-0.83 external), highlighting the importance of echo-specific pretraining and temporal modeling. The proposed multi-view fusion further enhanced predictive performance, with the late fusion model reaching an area under the receiver operating curve of 0.84 on the external cohort with significant population-shift.
CONCLUSIONS
These results suggest LVOT physiology is encoded in routine 2-dimensional imaging and can be leveraged for clinically relevant gradient classification without Doppler input. The proposed artificial intelligence-guided strategy demonstrates substantial cost savings compared with the screen-all approach. By integrating complementary spatial-temporal information across multiple views, our approach generalizes robustly across populations and may enable real-time decision support, extend LVOT assessment to portable or resource-limited settings, and complement Doppler-based evaluation for longitudinal hypertrophic cardiomyopathy management.
O. Crystal, J. Farina, I. Scalia et al.· Circulation Cardiovascular I...· 0 citations
BACKGROUND
Echocardiography is essential for assessing cardiac structure and function, yet accurate interpretation requires specialized expertise, creating challenges in emergency care and in regions with limited access to experienced echocardiographers. Although artificial intelligence-based interpretation is increasingly studied, most existing models rely on still images or single views and do not reflect the video-based, multiview integration used in clinical practice.
OBJECTIVES
The authors aim to develop and evaluate a multiview video-language model that integrates cardiac motion and multiple standard echocardiographic views.
METHODS
We trained a vision-language model on 577,061 transthoracic echocardiography videos paired with Japanese clinical reports from 46,852 examinations acquired at a single tertiary center (2015-2023). For each examination, features from 5 standard views (parasternal long-axis and short-axis; apical 2-, 3-, and 4-chamber) were aggregated. Performance was assessed in a retrieval task selecting the correct report from 16,436 candidate reports in the test cohort based on video input; the primary metric was correct match retrieval probability within the top 10 candidates (ie, R@10).
RESULTS
The image-based baseline achieved an R@10 of 1.3%. Using video input increased R@10 to 4.2%. Integrating 5 views further improved R@10 to 6.0% (95% CI: 5.5%-6.4%). Gains were greatest for findings that depended on temporal dynamics and multiview assessment, including left ventricular wall motion abnormality and left ventricular dilation.
CONCLUSIONS
A multiview video-language framework improved report retrieval compared with conventional image-based approaches. These findings support the utility of video-based, multiview representation learning for echocardiographic report retrieval.
R. Takizawa, Chiemi Yamazaki, S. Kodera et al.· JACC: Asia· 0 citations
Aortic stenosis (AS) is a prevalent and progressive valvular heart disease requiring accurate severity assessment for optimal clinical decision-making. Transthoracic echocardiography (TTE) is the standard diagnostic modality; however, its interpretation remains operator-dependent and subject to inter-observer variability. In this study, we propose an anatomically guided BackMix-enhanced semi-supervised learning framework for automated echocardiographic view classification and AS severity assessment. The approach leverages Gradient-weighted Class Activation Mapping (Grad-CAM) to preserve diagnostically relevant anatomical regions during augmentation while modifying background areas to mitigate shortcut learning. A semi-supervised self-training strategy combined with an ensemble classification framework was used to exploit both labeled and unlabeled data. The framework was evaluated on the TMED2 dataset of echocardiography, comprising 24,964 TTE images. To assess generalizability, external validation was performed on an independent dataset of 300 cardiologist-labeled echocardiographic images, including apical four-chamber (A4C), parasternal long-axis (PLAX), and parasternal short-axis (PSAX) views. Experimental results demonstrated that the proposed method outperformed the baseline semi-supervised model, improving view classification accuracy by approximately 3% and AS severity classification accuracy by 4–6%, with the largest gain observed in moderate AS. Performance remained consistent on the external validation dataset, supporting the robustness of the proposed approach. Statistical analysis confirmed the significance of these improvements (p < 0.01). Grad-CAM evaluation further demonstrated improved localization of clinically relevant regions. These findings suggest that anatomically guided BackMix augmentation combined with semi-supervised ensemble learning can improve classification accuracy, robustness, and interpretability in echocardiographic analysis under limited annotation conditions, offering a promising approach for automated AS assessment across independent clinical datasets
Fatima Ezzahra Elkouahy, Badreddine Labakoum, H. Ouahid et al.· Journal of Electronics Elect...· 0 citations
Echocardiography is the primary imaging modality for congenital heart disease (CHD) assessment. However, the real-world clinical application of artificial intelligence in this domain is often hindered by hidden data leakage from homologous frames and the high diagnostic variance of single-frame static inference. To bridge the gap between idealized model evaluation and clinical reality, this study proposes a standardized, region-of-interest (ROI)-guided deep learning workflow for differentiating atrial septal defect (ASD), ventricular septal defect (VSD), and normal cases. Specifically, utilizing a retrospective cohort of 1,987 patients (11,638 images) from Tengzhou Central People’s Hospital, we implemented strict subject-disjoint partitioning to eliminate data leakage, and simultaneously introduced cross-frame case aggregation to emulate the multi-frame visual synthesis process of expert echocardiographers. Among the evaluated model configurations, fully fine-tuned EfficientNet-B0 achieved the highest internal held-out test performance, with an accuracy of 0.9950 (95% CI, 0.9875–1.0000), a macro-F1 of 0.9941, and a macro-AUROC of 0.9997; 2 of 401 test patients were misclassified. Five-fold patient-level cross-validation yielded an accuracy of $0.9945 \pm 0.0040$ , further supporting the stability of the model across patient partitions. These findings suggest that the proposed workflow has the potential to serve as an adjunctive tool for septal defect screening.
Tao Zhang, Peipei Zhang, Qing-Yuan Zhang et al.· IEEE Access· 0 citations
Background: Regurgitant valvular heart disease (rVHD) is a major cause of cardiovascular morbidity. Echocardiography is the diagnostic standard but is resource-intensive for large-scale screening. Electrocardiography (ECG) has shown promise for predicting incident rVHD, yet performance varies across phenotypes, particularly for aortic regurgitation (AR). Chest radiography (CXR) provides complementary structural and hemodynamic information. We hypothesized that a multimodal model integrating ECG and CXR would improve prediction of incident moderate-to-severe rVHD. Methods: In this retrospective multicenter study, we identified 212,888 paired ECG-CXR examinations from 116,380 patients across two Chinese centers. Baseline ECG and CXR were obtained within 60 days of echocardiography. Outcome was progression to moderate-to-severe AR, mitral regurgitation (MR), or tricuspid regurgitation (TR). We developed a multimodal neural network with pretrained unimodal encoders, token-level cross-modal fusion, and a class-specific gating mechanism that adaptively weighted ECG-only, CXR-only, and fused predictions. Performance was assessed using C-index, AUROC, AUPRC, decision curve analysis, net reclassification improvement (NRI), and Kaplan-Meier stratification. Results: Multimodal fusion consistently outperformed unimodal models across all phenotypes. For AR, C-index improved from 0.616 (ECG-only) to 0.713 (multimodal; AUROC 0.729, AUPRC 0.972). For MR, multimodal C-index was 0.801 (AUROC 0.814, AUPRC 0.972), versus 0.782 for ECG and 0.775 for CXR alone. For TR, multimodal and CXR-only models showed similar discrimination (C-index 0.802), but multimodal fusion yielded greater net benefit on decision curve analysis. NRI was positive across all time horizons (1-5 years) for all valve types. Grad-CAM interpretability analyses revealed that ECG attention localized to leads II, V-V (AR), leads I, II, aVF, V-V (MR), and inferior/right precordial leads (TR); CXR attention highlighted chamber-specific enlargement and pulmonary congestion patterns consistent with pathophysiology. Conclusion: A multimodal deep learning model integrating ECG and CXR significantly improved prediction of incident rVHD compared with ECG alone, with the greatest benefit observed for AR. The model leveraged complementary electrical and structural information, demonstrated biological plausibility through interpretability analyses, and provided consistent clinical utility. Given the widespread availability and low cost of both modalities, this approach offers a scalable tool for risk stratification in routine care. Prospective studies are warranted to validate clinical implementation.
S. Li, B. Zhang, L. Pan et al.· medRxiv· 0 citations
Left ventricular ejection fraction (LVEF) and global longitudinal strain (GLS) are essential for the diagnosis, clinical decision-making, and prognosis of cardiovascular disease. However, accurate assessments of LVEF and GLS by echocardiography are hampered by inter-observer variability, time-consuming, and labor-intensive. This study aimed to develop an automated method to accurately and rapidly assess LVEF and GLS. Based on the datasets of 500 patients (1,500 videos) from the internal center and 363 patients (1,089 videos) from four external centers, we successfully developed a dual-flow convolutional neural network called Echo-DFCNN, which allowed for synchronous acquisition of LVEF and GLS. We evaluated the performance of the Echo-DFCNN in a cardiac magnetic resonance (CMR) validation dataset composed of 67 patients. On the internal test dataset, the AI and manual measurements of LVEF demonstrated a median absolute error of 3.02% and a mean absolute error of 3.94%. AI-predicted LVEF showed good agreement with manually measured LVEF, with an ICC of 0.927, a bias of 0.89%, and a LOA of -10.91 to 12.69. For GLS, the median absolute error and mean absolute error between AI and manual measurements were 1.43% and 1.83%. AI-predicted GLS exhibited high agreement with manually measured GLS (ICC = 0.913; bias = -1.22%, LOA = -5.12 to 2.68). In addition, Echo-DFCNN maintained good performance when applied to external validation datasets. In the CMR validation dataset, the AI model showed good agreement with CMR measurements for both LVEF and GLS. Echo-DFCNN achieves simultaneous and precise assessment of LVEF and GLS in the study cohorts, demonstrating its potential for robust performance across a wide range of cardiac functions, different image qualities, and machine types.
Zisang Zhang, Ye Zhu, Shujun Chen et al.· PLOS Digital Health· 0 citations