Aug 2026· IEEE/ACM International Conference on Connected Health: Applications, Systems and Engineering Technologies· pp. 422-427· 0 citations· 16 references
Abstract
De-identification of clinical notes is critical for protecting patient privacy, yet existing approaches often struggle under real-world variation and provide limited support for auditing and error analysis. By examining the outputs of NER-based systems, we observe three recurring failure modes—false positives, false negatives, and fragmented entity spans—the latter of which can simultaneously introduce both error types under exact-match evaluation.We present a transparent, locally deployable de-identification framework that augments NER-based extraction with two refinement stages: a Verification loop to correct candidate entities and a candidate expansion loop to recover missed protected health information (PHI). Beyond improving extraction quality, the system generates structured artifacts for human review, including error attribution and audit-ready outputs, enabling systematic inspection and iterative improvement.We evaluate the framework on the i2b2 2014 benchmark, a MIMIC-IV radiology subset, and a synthetic dataset simulating distribution shift. Results show substantial robustness gains under variation, increasing entity-level F1 from 61.52% to 87.78% and significantly reducing false negatives, while maintaining competitive performance on benchmark data. These findings highlight the value of combining post-NER refinement with transparency mechanisms for reliable clinical de-identification.
Verified Extraction is introduced, an auditing framework that distinguishes identifiers attributable to fine-tuning data from spurious or prior-driven outputs and quantifies recoverable leakage under explicit query budgets.
Florent Pollet, Tong Wang, Rahul Gupta et al.· 0 citations
Clinical notes contain personally identifiable information (PII), restricting reuse for research and medical AI, especially when data cannot leave an institution. We developed MedDeID, an on-premises framework combining in-house annotation and synthetic-note generation with model training, inference, pseudonymisation a...
Stig Hellemans, T. Stroobants, E. Scheurwegs et al.· 0 citations
Secondary use of electronic health records requires de-identification, yet existing systems miss \emph{institutionally situated} protected health information (PHI) such as hospital abbreviations, building names, and internal codes whose status is locally determined. We ask whether large language models (LLMs) with in-c...
Daniel Palacios, Matthew Neeley, A. A. Otto et al.· 0 citations
Clinical de-identification relies on accurately identifying personally identifiable information (PII). However, manually annotated datasets are costly to construct, while existing synthetic alternatives often provide limited details about their generation process or rely on relatively simple synthesis strategies. We in...
Linh Le, Christian Hoang, Huy Hoang Ha· 0 citations
CAST (Concept-guided Artifact Suppression Tuning), an SAE-based framework for auditable clinical text classification, improves over its corresponding fine-tuned encoder baselines and remains competitive with strong LLM baselines, while producing a feature-level audit trail of the clinical concepts that support each pre...
Jin Mu, Guan-Hua Chen· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.