Sep 2026· Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence· 0 citations· 47 references
TL;DR
ClinAlign—a memory-based retrieval framework aligned with clinical workflow, drawing inspiration from clinical diagnostic workflows is proposed, which constructs a disease-aware visual memory bank and introduces Classification-Guided Prompt Augmentation (CGPA), where disease state predictions are converted into structured diagnostic prompts to provide explicit semantic guidance for textual memory retrieval.
Abstract
Automated radiology report generation aims to create clear and clinically correct diagnostic reports from medical images. Existing retrieval enhancement methods primarily focus on reusing textual knowledge, neglecting the crucial role of local visual pattern memory in clinical diagnosis. Furthermore, cross-modal retrieval lacking explicit clinical semantic constraints can easily introduce irrelevant pathological information, thereby reducing the clinical effectiveness of the generated reports. To address these challenges, we propose ClinAlign—a memory-based retrieval framework aligned with clinical workflow, drawing inspiration from clinical diagnostic workflows. Visually, we construct a disease-aware visual memory bank and enhance local patch representations through proposed Memory‑based Patch Pattern Augmentation (MPPA), thereby improving the perception and discrimination of pathological regions. On the textual side, we construct a disease-aware textual memory bank and introduce Classification-Guided Prompt Augmentation (CGPA), where disease state predictions are converted into structured diagnostic prompts to provide explicit semantic guidance for textual memory retrieval.Extensive experiments on two medical report generation benchmarks, MIMIC-CXR and IU X-Ray, demonstrate the effectiveness and practical value of our proposed method.
Automated radiology report generation (ARRG) has emerged as a promising application of artificial intelligence for reducing radiologists’ documentation workload and improving the consistency of clinical reporting. However, conventional image-to-text models often struggle to capture subtle abnormalities, establish meani...
P. Dayaker, M. Vignesh, I. Z. et al.· International journal of com...· 0 citations
Interventional radiology (IR) requires joint reasoning over procedural images and domain-specific clinical knowledge. Existing medical retrieval-augmented generation (RAG) methods are mainly text-oriented or designed for general medical vision-language tasks, and therefore remain limited in retrieving fine-grained visu...
Jing-Xiong Li, Cheng-Lu Zhu, Yuxuan Sun et al.· IEEE Transactions on Medical...· 0 citations
Generating clinically accurate radiology reports from chest X-rays demands both precise pathology recognition and coherent medical language. However, fine-tuning large vision-language models can be computationally challenging in deployment settings. We present a lightweight framework that improves report generation fro...
Taishi Nishizawa, Ayesha Issah, Ana Fuertes-Brito et al.· IEEE/ACM International Confe...· 0 citations
Experiments on BUS-CoT and IU X-ray datasets demonstrate consistent improvements in diagnostic accuracy, concept consistency, and report quality over strong general-purpose and medical MLLMs, indicating that concept-grounded reasoning better aligns generation with clinical decision processes.
Xin-Yue Xu, Hong-Bin Lin, Juan-Gui Xu et al.· 0 citations
Radiologists typically adopt a coarse-to-fine, dynamic focusing cognitive strategy when interpreting medical images, focusing on potential abnormal regions while integrating semantic context to compose reports. However, most existing medical report generation methods rely on fixed-resolution image encoding, which strug...