Skip to content
Open access

Semantic de-identification of burned-in PHI in DICOM medical images: a deep learning–NLP pipeline validated on clinical and phantom TMM datasets

Jul 2026 · Frontiers of Computer Science · Vol 8 · 0 citations · 33 references

TL;DR

A semantic de-identification pipeline integrating YOLOv11n-based text detection, domain-optimized EasyOCR, and a hybrid natural language processing (NLP) classification module combining regular expressions, keyword matching, and named entity recognition is proposed, confirming that the pipeline preserves quantitative pixel fidelity when applied to institutional and device identifiers embedded in phantom acquisitions.

Abstract

The growing adoption of AI-based healthcare research has increased the need for properly anonymized medical imaging datasets. PHI within DICOM files - particularly burned-in pixel-level text - poses significant privacy and regulatory risks. Existing methods either focus solely on metadata or remove all detected text indiscriminately, sacrificing clinically relevant annotations. This paper proposes a semantic de-identification pipeline integrating YOLOv11n-based text detection, domain-optimized EasyOCR, and a hybrid natural language processing (NLP) classification module combining regular expressions, keyword matching, and named entity recognition. A dual-path architecture processes metadata and pixel-level PHI in parallel, enabling complete DICOM sanitization while preserving non-PHI clinical annotations. The system was evaluated on 1,042 multi-modality DICOM images (CT, MRI, X-ray, ultrasound). As a secondary evaluation, the pipeline was also applied to two tissue-mimicking material (TMM) phantom datasets from TCIA - the RIDER Phantom MRI and Phantom FDA CT (RIDER = Reference Image Database to Evaluate Therapy Response; FDA = Food and Drug Administration) - which served as surrogates for controlled evaluation of metadata and burned-in identifier removal. The system achieves an F1-score of 95.4%, 96.1% recall, a structural similarity index measure (SSIM) of 0.969, a peak signal-to-noise ratio (PSNR) of 28.9 dB, and processes each image in 2.8 s. It achieves SSIM of 0.986 and PSNR of 49.0 dB on RIDER Phantom MRI, and SSIM of 0.974 and PSNR of 31.5 dB on Phantom FDA CT. These results confirm that the pipeline preserves quantitative pixel fidelity when applied to institutional and device identifiers embedded in phantom acquisitions, supporting blinding for domain-generalization studies across institutions. The modular design supports institutional customisation, making it suitable for clinical research workflows and privacy-compliant phantom imaging pipelines.

Read PDF

Similar papers

#generative ai Preprint Aug 2026

Masking Is Not Enough: Generative Restoration for Multimodal De-Identification in Medical AI

ClinX is introduced, an end-to-end multimodal PHI sanitization framework for medical image-text data, and results show that OCR-only masking is not sufficient as a standalone solution, and restoration-based sanitization better preserves clinically relevant visual context while sharply reducing recoverable PHI.

S. Shrestha, Zongxing Xie, Chen Zhao et al. · 0 citations
Jul 2026

Dual U-shaped cross-modal fusion network for lung infection region segmentation

A dual U-shaped network based on convolutional neural network and vision transformer is designed to sufficiently achieve the cross-modal feature fusion of image and text to compensate for the defects of existing datasets.

Shangwang Liu, Mengjiao Zhao · 1 citation
#natural language process... Preprint Sep 2026

SIFTING: A Novel LLM-Based Framework for Structured and Transparent Information Extraction from Clinical Free-Text Reports, with Application to Tumor Staging in Lung Cancer

Background: Large language models (LLMs) show promise for extracting information from clinical free-text documents, but their outputs are often unstructured and lack traceability, complicating validation and adoption in clinical workflows. In this work we introduce SIFTING, an LLM-based framework designed to address th...

Mirco Hess, Gerben van Veenendaal, J. Wakkie et al. · 0 citations
Conference Aug 2026

MedXAIgnosis: a metadata-enhanced graph-fusion pilot for thorax disease classification

Classifying medical images is essential for the diagnosis of thorax diseases, often aided by deep learning (DL) techniques. However, traditional DL approaches typically focus on a single image type input and overlook valuable insights from clinical header data. To address this, MedXAIgnosis introduces a multimodal fram...

Datenji Sherpa, D. Pant, J. Heikkonen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.