Sep 2026· ACM Transactions on Intelligent Systems and Technology· 0 citations· 60 references
Topic Modeling
TL;DR
The proposed AGPNer integrates a heterogeneous dependency fusion encoder, which reconstructs masked entities to enhance token representations and fine-tunes a hybrid dependency modeling block to learn domain-specific patterns in medical texts; an imbalance-adaptive span decoder, which decouples entity and non-entity spans and adaptively assigns them different exponential decay factors to regulate their contributions during training.
Abstract
As a fundamental task in biomedical natural language processing, Medical Named Entity Recognition (MNER) aims to identify and classify medical entities from unstructured medical texts. A major challenge in this task is the prevalence of nested entities, which arise from the syntactic complexity and domain-specific characteristics of medical language. Recently, span-based models have been proposed to handle nested entities by reformulating the task as a multi-label classification problem over all possible spans. However, when directly applied to medical texts, these models encounter two problems. First, existing span-based methods typically rely on general-domain pre-trained language models to learn textual semantics, which hinders their ability to capture contextual dependencies in complex medical texts, thereby limiting their generalization across diverse medical data. Second, these methods treat all candidate spans equally during training, which results in an underestimation of the gradient contributions from entity spans due to their relatively small proportion, thereby degrading recognition performance. To address these issues, we propose AGPNer, a novel method for recognizing nested medical named entities. The proposed AGPNer integrates two key components: (1) a heterogeneous dependency fusion encoder, which reconstructs masked entities to enhance token representations and fine-tunes a hybrid dependency modeling block to learn domain-specific patterns in medical texts; (2) an imbalance-adaptive span decoder, which decouples entity and non-entity spans and adaptively assigns them different exponential decay factors to regulate their contributions during training. Experimental results on four public benchmarks demonstrate that AGPNer achieves absolute improvements in F1-score of 0.90, 1.02, 0.47 and 0.75 percentage points on CMeEE-V1, CMeEE-V2, GENIA, and CLUENER, respectively, showing consistent and competitive performance among the compared baselines.
BELXTR is presented, a novel embedding model based on the multi-vector (a.k.a. late interaction) architecture, which allows to leverage token-level matching information in biomedical entity linking by integrating an existing task-specific training objective and exploring active query expansion.
How should natural language processing models be selected and adapted for global health literature in environments where annotated data and computational resources are limited? This thesis investigates these challenges through experiments on semantic tag discovery, named entity recognition (NER), and multi-label topic...
Genis Skura, A. Geissbühler, Jean-Luc Falcone· 0 citations
This work proposes TdSciNER, a type-driven approach that effectively leverages entity type information to enhance SciNER performance and develops a novel demonstration selection strategy based on sentence similarity and entity type diversity to activate the in-context learning capabilities of LLMs, thereby improving en...
Tong Bao, Yi Zhao, Heng Zhang et al.· Expert systems with applicat...· 0 citations
This analysis shows that, in new sentences, the contextual associations of tokens representing old entity types exhibit a significantly stronger bias towards new entity types compared to their contexts in old sentences, which intensifies the degradation of old knowledge while promoting the overfitting of new knowledge.
Du-Zhen Zhang, Yahan Yu, Xiuyi Chen et al.· IEEE Transactions on Artific...· 0 citations
Instruction tuning has become a common paradigm for applying large language models (LLMs) to biomedical named entity recognition (BioNER). However, existing instruction-tuning approaches still face two key challenges. First, conventional natural-language instructions typically serialize BioNER annotations as flat textu...
Song-Tao Li, Yi-Jia Zhang, Jian-Yuan Yuan et al.· 1 citation
: Biomedical texts naturally contain multiple biological and medical concepts within a document, resulting in a semantically rich and complex structure. Consequently, multi-label text classification (MLTC) has become a suitable framework for comprehensively modeling biomedical texts, including clinical reports, laborat...