Skip to content

Learning Anatomy-Grounded CT Vision-Language Representations with Organ-Hierarchical Report Knowledge

Jul 2026 · arXiv.org · Vol abs/2607.10953 · 1 citation · 27 references
Computer Science

TL;DR

OKA-CT is proposed, an organ-hierarchical knowledge-augmented framework for CT-report VLP that achieves zero-shot abnormality diagnosis AUROCs on CT-RATE and RAD-ChestCT datasets and shows improved report-image alignment and stronger sensitivity to disease-associated anatomical regions.

Abstract

Medical vision-language pretraining (VLP) from paired CT images and radiology reports enables scalable representation learning, but most existing methods align either whole scans with entire reports or local image regions with text fragments. These formulations underuse a key property of radiology reports: findings are organized around anatomical structures, with abnormalities described by organs, disease concepts, locations, and severity-related attributes. We propose OKA-CT, an organ-hierarchical knowledge-augmented framework for CT-report VLP. OKA-CT first converts free-text reports into organ-conditioned knowledge using radiology report parsing and LLM-assisted semantic structuring. The extracted hierarchy is used across two learning stages. Stage~1 injects anatomy-grounded evidence into the CT visual representation through fine-grained organ-conditioned supervision, while Stage~2 uses organ-specific report evidence to guide structured report-CT contrastive learning, where hierarchy-derived semantic soft targets treat non-paired cases with shared organ-level findings as weak semantic positives rather than uniform negatives. A lightweight query-based global branch further aggregates disease-relevant volumetric evidence for whole-scan representation. On CT-RATE and RAD-ChestCT datasets, OKA-CT achieves zero-shot abnormality diagnosis AUROCs of 84.9 and 72.2, outperforming prior CT VLP baselines. Retrieval and patch-occlusion analyses further show improved report-image alignment and stronger sensitivity to disease-associated anatomical regions.

View source

Similar papers

#computer vision Review Sep 2026

NV-Reason-CT: 3D Visual Language Model for CT Analysis

The NV-Reason-CT model, a generative vision--language model for chest and abdominal CT combining native 3D visual encoding with radiologist-guided reasoning, and the model and training code are released to support reproducible research on explainable AI for volumetric medical imaging.

Andriy Myronenko, Dong Yang, Yu-Cheng Tang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

ARC-CT: Anatomy-Routed Contrastive Vision-Language Learning for 3D Chest CT

An Anatomy-Routed Contrastive Learning for 3D Chest CT (ARC-CT), a region-aware framework that addresses limitations using only reports extracted from reports by an LLM, with no manual annotations or bounding boxes, outperforms both comparable efficient baselines and larger transformer models.

H. Isik, Mehmet Alp Ozaydin, S. Kurugol et al. · 1 citation
Sep 2026

Concept-Enhanced Multi-Scale Cross-Modal Alignment for Medical Visual Representation Learning.

A Concept Clause Decomposition method is designed to extract semantically complete descriptions of pathological findings or radiology manifestations from medical reports as medical concept clauses, which are then utilized within a multi-granularity cross-modal alignment framework to enhance medical concept perception a...

Xiang-Min Kong, Xi-Bin Jia, Da-Wei Yang et al. · 0 citations
Preprint Aug 2026

MedPlex: Deep Vision-Language Co-Adaptation for Clinically Grounded Medical Segmentation

This work introduces MedPlex (Medical Plexus of Vision and Language), an end-to-end VLM framework that makes text guidance a continuous, clinically grounded component of segmentation learning, and introduces Bi-Fusion (Bidirectional Fusion), visual and textual representations evolve jointly across the encoding hierarch...

Rafi Ibn Sultan, Hui Zhu, Chengyin Li et al. · 0 citations
Sep 2026

Deep learning-based diagnostic report generation for low-resolution functional medical images via cross-modal visual and textual alignment

A unified framework for automatic report generation from SPECT bone scintigrams that integrates domain-adaptive representation learning, fine-grained image–text alignment, and anatomy-guided supervision is proposed, offering a valuable pathway to achieving trustworthy and intelligent diagnostic support within nuclear m...

Tao Song, Qiang Lin, Tong-Tong Li et al. · 0 citations
Preprint Sep 2026

MetaStructAtlas: A Grounded 3D Vision-Language Dataset and Benchmark for Functional and Structural Reasoning in Whole-Body PET/CT

The joint interpretation of metabolic function and anatomical structure is essential for clinical diagnosis in whole-body PET/CT. Although recent advances in 3D medical vision-language models have demonstrated remarkable progress, current efforts are limited to regional CT imaging, leaving a critical void in comprehens...

Chen-Guang Zheng, Le Xue, Yi-Chi Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.