From Labels to Language: Zero-Shot Radiology Report Generation with Hybrid Retrieval Augmentation
Abstract
Generating clinically accurate radiology reports from chest X-rays demands both precise pathology recognition and coherent medical language. However, fine-tuning large vision-language models can be computationally challenging in deployment settings. We present a lightweight framework that improves report generation from a frozen MedGemma-4B model purely through hybrid retrieval augmentation. The only learned component is a lightweight MLP classifier on frozen BiomedCLIP-derived embeddings, leaving both foundation models unmodified. Our approach combines BiomedCLIP image embeddings with two complementary retrieval augmentation generation (RAG) mechanisms: FAISS-based nearest-neighbor search over training reports for visual similarity, and a label co-occurrence graph using normalized pointwise mutual information (NPMI) for concept-driven retrieval. At inference, predicted pathology labels from a lightweight MLP classification head serve as graph entry points, expanding via NPMI-weighted traversal to retrieve clinically related cases that FAISS alone can miss, particularly for rare findings where embedding-based recall collapses. On MIMIC-CXR-JPG with the CheXpert 14-label schema, our full pipeline improves RadGraph-F1 from 0.167 (zero-shot MedGemma) to 0.207 (dual-track RAG), all without LLM fine-tuning. Our results demonstrate that hybrid text and graph retrieval augmentation can be an effective, lightweight strategy for improving radiology report generation at deployment.