A semantic de-identification pipeline integrating YOLOv11n-based text detection, domain-optimized EasyOCR, and a hybrid natural language processing (NLP) classification module combining regular expressions, keyword matching, and named entity recognition is proposed, confirming that the pipeline preserves quantitative pixel fidelity when applied to institutional and device identifiers embedded in phantom acquisitions.
Abstract
The growing adoption of AI-based healthcare research has increased the need for properly anonymized medical imaging datasets. PHI within DICOM files - particularly burned-in pixel-level text - poses significant privacy and regulatory risks. Existing methods either focus solely on metadata or remove all detected text indiscriminately, sacrificing clinically relevant annotations.
This paper proposes a semantic de-identification pipeline integrating YOLOv11n-based text detection, domain-optimized EasyOCR, and a hybrid natural language processing (NLP) classification module combining regular expressions, keyword matching, and named entity recognition. A dual-path architecture processes metadata and pixel-level PHI in parallel, enabling complete DICOM sanitization while preserving non-PHI clinical annotations. The system was evaluated on 1,042 multi-modality DICOM images (CT, MRI, X-ray, ultrasound). As a secondary evaluation, the pipeline was also applied to two tissue-mimicking material (TMM) phantom datasets from TCIA - the RIDER Phantom MRI and Phantom FDA CT (RIDER = Reference Image Database to Evaluate Therapy Response; FDA = Food and Drug Administration) - which served as surrogates for controlled evaluation of metadata and burned-in identifier removal.
The system achieves an F1-score of 95.4%, 96.1% recall, a structural similarity index measure (SSIM) of 0.969, a peak signal-to-noise ratio (PSNR) of 28.9 dB, and processes each image in 2.8 s. It achieves SSIM of 0.986 and PSNR of 49.0 dB on RIDER Phantom MRI, and SSIM of 0.974 and PSNR of 31.5 dB on Phantom FDA CT.
These results confirm that the pipeline preserves quantitative pixel fidelity when applied to institutional and device identifiers embedded in phantom acquisitions, supporting blinding for domain-generalization studies across institutions. The modular design supports institutional customisation, making it suitable for clinical research workflows and privacy-compliant phantom imaging pipelines.
ClinX is introduced, an end-to-end multimodal PHI sanitization framework for medical image-text data, and results show that OCR-only masking is not sufficient as a standalone solution, and restoration-based sanitization better preserves clinically relevant visual context while sharply reducing recoverable PHI.
S. Shrestha, Zongxing Xie, Chen Zhao et al.· 0 citations
A dual U-shaped network based on convolutional neural network and vision transformer is designed to sufficiently achieve the cross-modal feature fusion of image and text to compensate for the defects of existing datasets.
MRD-UNet provides a practical balance between segmentation accuracy and computational efficiency and outperforms baseline CNNs and performs comparably to heavier transformer-based models while using significantly fewer parameters.
Musa Doğan, I. Ozkan· BMC Medical Imaging· 0 citations
Background: Large language models (LLMs) show promise for extracting information from clinical free-text documents, but their outputs are often unstructured and lack traceability, complicating validation and adoption in clinical workflows. In this work we introduce SIFTING, an LLM-based framework designed to address th...
Mirco Hess, Gerben van Veenendaal, J. Wakkie et al.· 0 citations
Classifying medical images is essential for the diagnosis of thorax diseases, often aided by deep learning (DL) techniques. However, traditional DL approaches typically focus on a single image type input and overlook valuable insights from clinical header data. To address this, MedXAIgnosis introduces a multimodal fram...
Datenji Sherpa, D. Pant, J. Heikkonen et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.