Skip to content
Book Open access

Salient-Q: A Saliency-Guided Vision-Language Framework for Medical Image De-Identification

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · 0 citations · 26 references

TL;DR

Salient-Q, a saliency-guided vision-language framework that addresses challenges of PHI identification and localization through three architectural innovations, outperforms other baseline models in terms of PHI identification and localization.

Abstract

Medical image de-identification is critical for artificial intelligence research in the healthcare domain. It requires the detection and localization of Protected Health Information (PHI) burned into medical images. While vision-language models have shown strong visual understanding capabilities, recent studies reveal they suffer from attention dispersion, which correctly localizes but fails to perceive small visual details. Moreover, naive feature selection approaches that discard spatial context destroy positional information essential for accurate bounding box prediction. We propose Salient-Q, a saliency-guided vision-language framework that addresses these challenges through three architectural innovations: (1) a Saliency Module that learns per-token PHI probability and amplifies relevant features through adaptive soft gating; (2) a position-aware scout detector module with explicit 2D positional embeddings that preserves spatial relationships, and (3) a saliency-weighted cross-attention with coordinate supervision that aligns attention centers with ground-truth bounding box centers. We also developed a digital-decayed PHI synthesis pipeline and constructed the Decayed-PHI-50K dataset that was used to fine-tune and evaluate the developed models. Salient-Q outperforms other baseline models in terms of PHI identification and localization. Our code is available at: https://github.com/zjsuper/phi\_deidentification\_vlm.

Read PDF

Similar papers

Preprint Sep 2026

Beyond Classification: Structured Supervision Aligns Visual Evidence with Medical Semantics

This work evaluates three independent forms of structured supervision: topological priors via graph self-supervision, dense pixel-level constraints via segmentation, and cross-modal semantic grounding via image-text pairs, and finds that image-text alignment achieves the most superior performance.

He-Xiang Bai, Han-Yang Xu, Xiao-Xue Li et al. · 0 citations
Sep 2026

GAA-DETR: Query-level gaze alignment for end-to-end lesion detection.

Lesion detection is a fundamental task in medical image analysis. However, detectors trained only with bounding box supervision often develop boundary-biased attention and may overlook diagnostically relevant lesion content, which is closely related to false-positive predictions. To take a step toward addressing this l...

Yan Kong, Sheng Wang, Yuan Yin et al. · 0 citations
Preprint Aug 2026

DistMedVL: Distributional Vision-Language Alignment for Uncertainty-Aware Medical Image Segmentation

DistMedVL is proposed, a probabilistic vision-language framework that introduces a lightweight Probabilistic Cross-Modal Adapter (PCM-Adapter) upon frozen encoders to explicitly model representational uncertainty, exhibiting superior data efficiency, perturbation robustness and cross-dataset generalization.

Jiaxuan Li, Qing Xu, Xiang-Jian He et al. · 0 citations
Preprint Aug 2026

MedPixel: A Unified Pixel-Language Model for Medical Reasoning and Segmentation

This work presents MedPixel, a unified medical pixel-language model built around a shared language--mask interface that achieves strong performance in both pixel-level prediction and response generation, together with effective zero-shot transfer to external grounding benchmarks and robustness to imperfect spatial prom...

Haoyu Yang, Meixing Shi, Zeng-Jie Chen et al. · 0 citations
Preprint Aug 2026

LocAnyMed: Vision-Language Grounding for Multimodal Medical Images

LocAnyMed-CoT-20K is derived, a rationale-augmented subset that connects anatomical context, visual observations, and spatial conclusions through structured reasoning and further improves cross-source generalization through fine-tuning.

Zi-Han Wang, Tong Liu, Zhi-Wei Wang et al. · 0 citations
Open access Aug 2026

Medrecord-CLIP: enhancing fundus disease diagnosis via EHR-guided vision-language pre-training

This work proposes MedRecord-CLIP, a knowledge-enhanced foundation model featuring a diagnosis-guided cross-attention mechanism to adaptively extract and fuse salient patient history with diagnostic representations that highlights the critical value of integrating personalized clinical context to enhance the generaliza...

Lei Shi, Wenbin Zhai, Lei Yu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.