Skip to content
Open access

Patient-specific multimodal learning with multi-view contrastive alignment for chest X-ray report generation

Nov 2024 · Bioinform. · Vol 42 · 6 citations · ⚡ 3 influential · 76 references
Medicine Computer Science

TL;DR

The proposed EVOKE surpasses recent state-of-the-art methods across multiple datasets, and introduces a multi-view contrastive learning method that captures semantic correspondences both among multi-view radiographs within a study and between these radiographs and their associated report, thereby improving visual representation learning.

Abstract

Abstract Motivation Radiology reports play a pivotal role in guiding treatment planning and enabling effective doctor-patient communication. However, their manual composition imposes a substantial workload on radiologists. Although automatic radiology report generation has emerged as a promising alternative, existing approaches predominantly rely on single-view chest X-rays and fail to adequately leverage patient-specific context, thereby limiting diagnostic accuracy. Results To address this challenge, we propose EVOKE, a novel chest X-ray report generation framework that incorporates multi-view contrastive learning and patient-specific knowledge. Specifically, we introduce a multi-view contrastive learning method that captures semantic correspondences both among multi-view radiographs within a study and between these radiographs and their associated report, thereby improving visual representation learning. We further present a knowledge-guided report generation module that integrates available patient-specific knowledge (i.e. indication, which includes symptom descriptions) to facilitate the generation of accurate and coherent radiology reports. To support research in multi-view report generation, we construct Multi-view CXR and Two-view CXR datasets using publicly available sources. Our proposed EVOKE surpasses recent state-of-the-art methods across multiple datasets, achieving a 2.9% F1 RadGraph improvement on MIMIC-CXR, a 5.0% BLEU-1 improvement on MIMIC-ABN, a 1.5% BLEU-4 improvement on Multi-view CXR, and an 8.2% F1,mic-14 CheXbert improvement on Two-view CXR. Availability Code is publicly available at https://github.com/mk-runner/EVOKE, with an archived release available on Zenodo (doi:10.5281/zenodo.21000219).

Read PDF

Similar papers

Sep 2026

Comorbidity-Aware Radiology Report Generation.

Chest X-ray report generation systems are valuable for assisting disease diagnosis and improving healthcare efficiency. However, existing methods still face two key challenges. First, multiple diseases often co-occur, leading to a combinatorial explosion of label combinations and sparse supervision for learning a gener...

Hong-Ze Zhu, Hong Liu, Ya-Wen Huang et al. · 0 citations
Open access Sep 2026

Dual-branch cross-modal architecture with global-to-local feedback for chest X-ray retrieval

Radiology reports are vital for accurate diagnosis and treatment planning, yet their manual generation is time-consuming and dependent on radiologist expertise, leading to delays and inconsistent clinical decisions. Medical image–text retrieval offers a scalable solution by enabling the retrieval of relevant prior...

Rezaul Abedin, Sofiane Laridi, Kam-Ming Mark Tam · 0 citations
Conference Aug 2026

Multimodal Radiology Assistant with Graph- Enhanced Reasoning and Uncertainty-Guided Report Generation

Automatic radiology report generation has become an active research area due to its potential to reduce radiologist workload and standardize reporting quality. However, state-of- the-art systems still suffer from hallucinated findings, limited clinical reasoning, and a lack of calibrated uncertainty estimates, all of w...

Lokesh P, Kamaleshwaran K, Naveenraj M et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MAC-RRG: Iterative Multi-Agent Collaboration for X-ray Radiology Report Generation

Despite the remarkable progress of LLM-based and knowledge graph-augmented Radiology Report Generation (RRG) methods, existing techniques still suffer from inherent defects. Conventional LLM-only models lack structured medical prior knowledge, resulting in frequent medical hallucinations and low diagnostic interpretabi...

Fu-Tian Wang, Yu-Han Qiao, Xiao Wang et al. · 0 citations
Preprint Aug 2026

PerFact: Perception-Derived Fact Prompting for 3D Brain MRI Report Generation

In a controlled study that fixes the backbone, data split, target reports, and adaptation while varying only the injected grounding, perception-derived facts outperform retrieved prior reports, retrieval becomes redundant once facts are present, and end-to-end predicted facts remain effective without any ground-truth a...

Jian-Yu Sun, Zhen-Xuan Zhang, Guang Yang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.