Skip to content
Conference

Improving Clinical Note De-identification via Post-NER Verification and Candidate Expansion

Aug 2026 · IEEE/ACM International Conference on Connected Health: Applications, Systems and Engineering Technologies · pp. 422-427 · 0 citations · 16 references

Abstract

De-identification of clinical notes is critical for protecting patient privacy, yet existing approaches often struggle under real-world variation and provide limited support for auditing and error analysis. By examining the outputs of NER-based systems, we observe three recurring failure modes—false positives, false negatives, and fragmented entity spans—the latter of which can simultaneously introduce both error types under exact-match evaluation.We present a transparent, locally deployable de-identification framework that augments NER-based extraction with two refinement stages: a Verification loop to correct candidate entities and a candidate expansion loop to recover missed protected health information (PHI). Beyond improving extraction quality, the system generates structured artifacts for human review, including error attribution and audit-ready outputs, enabling systematic inspection and iterative improvement.We evaluate the framework on the i2b2 2014 benchmark, a MIMIC-IV radiology subset, and a synthetic dataset simulating distribution shift. Results show substantial robustness gains under variation, increasing entity-level F1 from 61.52% to 87.78% and significantly reducing false negatives, while maintaining competitive performance on benchmark data. These findings highlight the value of combining post-NER refinement with transparency mechanisms for reliable clinical de-identification.

View source

Similar papers

Privacy Audits for Clinical Large Language Models

Verified Extraction is introduced, an auditing framework that distinguishes identifiers attributable to fine-tuning data from spurious or prior-driven outputs and quantifies recoverable leakage under explicit query budgets.

Florent Pollet, Tong Wang, Rahul Gupta et al. · 0 citations
#machine learning Preprint Sep 2026

MedDeID enables locally governed clinical-text de-identification from real or synthetic training data

Clinical notes contain personally identifiable information (PII), restricting reuse for research and medical AI, especially when data cannot leave an institution. We developed MedDeID, an on-premises framework combining in-house annotation and synthetic-note generation with model training, inference, pseudonymisation a...

Stig Hellemans, T. Stroobants, E. Scheurwegs et al. · 0 citations
Preprint Aug 2026

Institution-Specific LLM Prompting Recovers PHI That De-identification Systems and Their Gold Standards Both Miss

Secondary use of electronic health records requires de-identification, yet existing systems miss \emph{institutionally situated} protected health information (PHI) such as hospital abbreviations, building names, and internal codes whose status is locally determined. We ask whether large language models (LLMs) with in-c...

Daniel Palacios, Matthew Neeley, A. A. Otto et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Meddies-PII: A Multilingual Framework for Personally Identifiable Information Extraction in Clinical De-identification

Clinical de-identification relies on accurately identifying personally identifiable information (PII). However, manually annotated datasets are costly to construct, while existing synthetic alternatives often provide limited details about their generation process or rely on relatively simple synthesis strategies. We in...

Linh Le, Christian Hoang, Huy Hoang Ha · 0 citations
Preprint Aug 2026

Making Clinical Language Models Auditable: Concept-Guided Fine-Tuning for Robust Prediction

CAST (Concept-guided Artifact Suppression Tuning), an SAE-based framework for auditable clinical text classification, improves over its corresponding fine-tuned encoder baselines and remains competitive with strong LLM baselines, while producing a feature-level audit trail of the clinical concepts that support each pre...

Jin Mu, Guan-Hua Chen · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.