Background Congenital heart disease (CHD) is a clinically important fetal anomaly. Early-pregnancy fetal cardiac microflow imaging (FCMI) can show low-velocity intracardiac flow, but brief shunt-related, regurgitant, and outflow-tract disturbances remain difficult to recognize when image quality, fetal position, and gestational age vary across examinations. Objective To evaluate uFlowAM for fetus-level detection and visualization of abnormal intracardiac microflow patterns on early-pregnancy FCMI. Methods This multicenter diagnostic accuracy study analyzed 650 early-pregnancy FCMI examinations from fetuses referred for suspected CHD or CHD risk assessment, including 500 examinations in the internal cohort and 150 in the external cohort. Standard four-chamber, left ventricular outflow tract (LVOT), and right ventricular outflow tract (RVOT) clips were processed by microflow extraction, cardiac-cycle alignment, signal normalization, and 16-frame windowing. uFlowAM used self-supervised training to learn a 256-dimensional representation of control fetal microflow from temporal-order discrimination and masked-frame reconstruction. Model training used no pixel-level or lesion-level labels. Control embeddings were grouped by view and cardiac phase to build a normal microflow template library, and an abnormality index (AbI) was calculated from latent-space Mahalanobis distances. The operating threshold was calibrated in internal validation and then kept fixed for external testing. Clinical utility was assessed in a 9-reader, 240-case multi-reader multi-case (MRMC) study. Results Using the fixed operating threshold (τ* = 2.15), uFlowAM achieved an area under the receiver operating characteristic curve (AUC) of 0.94 (95% CI, 0.92–0.96), sensitivity of 0.92, and specificity of 0.88 in the internal cohort. In the external cohort, AUC was 0.92 (95% CI, 0.88–0.95), with sensitivity of 0.90 and specificity of 0.86. Median reader-level AUC increased from 0.85 to 0.92 with uFlowAM assistance, weighted kappa for subtype agreement increased from 0.62 to 0.78, visibility scores increased from 2.8 ± 0.6 to 4.3 ± 0.5, and median reading time decreased from 78 s to 59 s. Mean inference time was 6.8 ± 1.3 s per case. Conclusions In this CHD-enriched referral/risk-assessment cohort, uFlowAM detected abnormal early-pregnancy fetal cardiac microflow patterns and improved reader consistency and efficiency on selected fetal cardiac microflow views. The framework should be considered an assistive second-reader tool for early fetal CHD assessment. It should not be used as a substitute for a complete fetal echocardiographic examination.
Yan Xia, Yarui Wei, Zhanru Lan et al.· Frontiers in Pediatrics· 0 citations
Specialized thoracic-surgery questions require the integration of multi-factor clinical relationships within text, yet general-purpose large language models (LLMs) may underperform on such exam-style benchmarks.
We constructed a DK-LLM agent by embedding curated medical textbook knowledge into a LangChain-based framework to support domain-specific reasoning. The model was tested on a 56-item thoracic-surgery examination question set in a restricted text-only setting without internet browsing or external tools and was compared with generic LLM baselines, three thoracic surgeons, and three non-expert engineers. Examination score and error patterns were assessed.
The knowledge-augmented DK-LLM configuration showed an 11.8-point examination-score advantage over the version without the local knowledge base. Commercial LLM-based agents outperformed the open-source baselines and non-expert participants on this question set, whereas experienced thoracic surgeons achieved the highest scores overall; in the ablation analysis, removing the local knowledge base reduced the examination score by 11.8 percentage points.
Embedding domain-specific knowledge into LLMs may improve performance on specialized exam-style thoracic-surgery questions on this text-only benchmark. However, the present 56-item evaluation does not establish clinical equivalence, diagnostic accuracy in practice, multimodal competence, or readiness for real-world clinical decision support.
Qian Li, Yongxin Li, Chao Ye et al.· Frontiers in Artificial Inte...· 0 citations