Skip to content
Open access

MMAC-Net: A Multi-Modal Multi-Label Attention-Based Deep Learning Approach for Automated ICD-9 Coding of Rare Disease Admissions from Electronic Health Records

Sep 2026 · Applied Sciences · 0 citations · 56 references

Abstract

Automating the identification of International Classification of Diseases (ICD) codes from electronic health records (EHRs) presents a critical challenge, particularly for rare diseases where existing computational methods severely underperform due to extreme long-tail label distributions. To address this, we propose a multi-modal deep learning framework known as MMAC-Net, designed to enhance the retrospective assignment of ICD-9 codes to admissions involving rare pathologies. The model integrates unstructured clinical narratives with structured auxiliary data, specifically pharmacological prescriptions and microbiology events, using a convolutional attention-based architecture. Through a late fusion mechanism, it synthesizes attention-weighted textual representations with dense embeddings of the structured data types. Validation on the MIMIC-III dataset shows consistent improvements over a matched text-only baseline evaluated under an identical protocol. On the full dataset of 8930 ICD codes, the framework achieved a Micro-AUC of 0.997 and Precision@8 of 0.875. On the subset of admissions carrying at least 1 of 568 rare codes, adding the two structured modalities to the text encoder raises Macro-F1 from 0.011 to 0.084 and Micro-F1 from 0.368 to 0.513 relative to the text-only baseline, corresponding to relative increases of 6.69 and 0.39, respectively, while Precision@8 rises from 0.092 to 0.159 and Micro-AUC from 0.966 to 0.985. While extreme class imbalance remains a formidable obstacle, these findings underscore that incorporating structured clinical context partially mitigates the limitations of purely natural language processing approaches. Practically, the framework is intended as a decision-support component that presents a ranked shortlist of candidate codes to a human coder or clinician; by recovering rare codes that text-only systems miss, it targets the under-coding of low-prevalence conditions that degrades registry completeness and downstream epidemiological estimates.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.