Skip to content
Review Open access

The DSMZ Digital Diversity Annotation Hub: a pipeline for database expansion via text mining and human curation

Dec 2025 · Journal of Integrative Bioinformatics · Vol 22 · 1 citation · 28 references
Medicine

Abstract

Abstract Maintaining scientific databases that depend on continuous curation of research literature often requires labor-intensive, slow, and error-prone annotation processes. To address these challenges, we present a pipeline that integrates text mining with expert supervision to support database expansion. Using the BRENDA enzyme database as a case study, we compiled a relation extraction dataset by aligning document-level annotations with literature references through distant supervision. We then developed a neural model that performs entity recognition and relation classification, enabling the extraction of enzyme-strain associations from full-text articles. To close the loop between machine learning and expert curation, we designed a web-based interface that allows annotators to review and refine predicted relations. While preliminary, our initial experiments show the potential of combining weak supervision and human-in-the-loop validation to accelerate the integration of literature-derived information into knowledge bases.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.