Dual-branch cross-modal architecture with global-to-local feedback for chest X-ray retrieval
Abstract
Radiology reports are vital for accurate diagnosis and treatment planning, yet their manual generation is time-consuming and dependent on radiologist expertise, leading to delays and inconsistent clinical decisions. Medical image–text retrieval offers a scalable solution by enabling the retrieval of relevant prior cases and reports. However, existing models often rely on static embeddings, which limits their ability to capture fine-grained, clinically meaningful cross-modal relationships. This limitation is particularly critical in chest X-ray interpretation, where nuanced textual descriptions must align with subtle visual cues. We introduce a novel dual-branch retrieval framework that distinguishes between shared semantics and complementary features through a Synergy branch and a Difference branch , respectively. These branches are stabilized through orthogonal regularization, ensuring minimal redundancy while keeping complementary diagnostic cues. A Global-to-Local Feedback mechanism guides fine-grained local attention using global context, enhancing interpretability and clinical relevance. Evaluated on two benchmark datasets, our model achieves state-of-the-art retrieval performance across both image $$\rightarrow$$ text and text $$\rightarrow$$ image tasks. Through hierarchical gated co-attention, our approach dynamically aligns image and text representations, addressing the limitations of static fusion and providing a foundation for interpretable, real-time retrieval systems that can accelerate and improve clinical decision support.