Skip to content

Author

M. Hirokawa

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

Comprehensive Echocardiography Interpretation Using Video and Multiview Vision-Language AI.

BACKGROUND Echocardiography is essential for assessing cardiac structure and function, yet accurate interpretation requires specialized expertise, creating challenges in emergency care and in regions with limited access to experienced echocardiographers. Although artificial intelligence-based interpretation is increasingly studied, most existing models rely on still images or single views and do not reflect the video-based, multiview integration used in clinical practice. OBJECTIVES The authors aim to develop and evaluate a multiview video-language model that integrates cardiac motion and multiple standard echocardiographic views. METHODS We trained a vision-language model on 577,061 transthoracic echocardiography videos paired with Japanese clinical reports from 46,852 examinations acquired at a single tertiary center (2015-2023). For each examination, features from 5 standard views (parasternal long-axis and short-axis; apical 2-, 3-, and 4-chamber) were aggregated. Performance was assessed in a retrieval task selecting the correct report from 16,436 candidate reports in the test cohort based on video input; the primary metric was correct match retrieval probability within the top 10 candidates (ie, R@10). RESULTS The image-based baseline achieved an R@10 of 1.3%. Using video input increased R@10 to 4.2%. Integrating 5 views further improved R@10 to 6.0% (95% CI: 5.5%-6.4%). Gains were greatest for findings that depended on temporal dynamics and multiview assessment, including left ventricular wall motion abnormality and left ventricular dilation. CONCLUSIONS A multiview video-language framework improved report retrieval compared with conventional image-based approaches. These findings support the utility of video-based, multiview representation learning for echocardiographic report retrieval.

R. Takizawa, Chiemi Yamazaki, S. Kodera et al. · 0 citations