Skip to content
#small language model Open access

Evaluating large language models for domain-specific scholarly metadata classification

Sep 2026 · Frontiers in Research Metrics and Analytics · Vol 11 · 0 citations · 26 references
Medicine

TL;DR

The findings indicate that open-source LLMs can support domain-specific scholarly metadata classification without task-specific fine-tuning, however, their moderate and dimension-dependent performance limits their suitability for fully automated fine-grained metadata enrichment.

Abstract

Introduction Large language models (LLMs) have been widely applied to text classification; however, their effectiveness in domain-specific scholarly metadata classification remains insufficiently explored. This study investigates the use of open-source LLMs to classify modeling and simulation research articles across three predefined metadata dimensions: Industry, Application Area, and Modeling Method. Methods Using a curated dataset of research articles from the AnyLogic research repository, we evaluated models from the Qwen2.5, Llama3.1, Mistral, Ministral, and Gemma3 families. We compared Direct, Few-shot, and Chain-of-Thought prompting using either the article title alone or the title and abstract together. Performance was evaluated using Macro-F1 across the three metadata dimensions. Results The best overall configuration, Ministral-3-14B with Direct prompting and the title and abstract as input, achieved a Macro-F1 of 53.8%. Performance varied substantially across the metadata dimensions. Modeling Method was the easiest to classify, reaching a maximum Macro-F1 of 69.1% with Ministral-3–14B. The best Macro-F1 scores for Industry and Application Area were 57.6% and 39.4%, respectively. Including abstracts consistently improved classification performance across the evaluated models, whereas differences among the three prompting strategies were relatively small. Categories that were semantically overlapping, broadly defined, highly imbalanced, or not explicitly mentioned in the article text were particularly difficult to classify. Discussion The findings indicate that open-source LLMs can support domain-specific scholarly metadata classification without task-specific fine-tuning. However, their moderate and dimension-dependent performance limits their suitability for fully automated fine-grained metadata enrichment. Richer contextual information, more clearly defined taxonomies, and human validation are therefore needed to improve the reliability of LLM-assisted scholarly metadata classification.

Read PDF

Similar papers

Automatic Domain Classification of Tabular Datasets Using Large Language Models

It is asserted that the present contribution consists of an interpretable domain palette, a constructed benchmark of diverse tabular datasets, and reproducible code and data to enable further research on domain discovery and domain-aware tooling for tabular data.

Elizaveta Gamper, Irina Deeva · 0 citations
#large language models Open access Sep 2026

Research on Text Information Extraction and Imbalanced Classification Methods for Enterprise Profiling

A comprehensive natural language processing (NLP) pipeline for extracting key information and identifying industries, which outperforms traditional machine learning baselines and single deep learning models, offering more reliable recognition for minority classes.

Xin-Yi Xu · 0 citations
Review Open access Sep 2026

Exploring AI-assisted Classification for Collection Analysis

Academic libraries face ongoing difficulty in aligning their collections with the needs of the communities they serve. Although the literature identified the value of using research, publications, and teaching data to inform collection development, maintaining crosswalks between these activities and library classificat...

Jennifer Moon-Chung, Berenika Maria Webster · 0 citations
Open access Aug 2026

Optimizing sample selection for large language model-based entity matching using AssistEM

AssistEM, a framework for efficient LLM adaptation to EM via principled data selection, demonstrates that selective fine-tuning not only accelerates adaptation but also improves training efficiency (requiring fewer GPU hours), enabling open-source LLMs to rival–and in some cases outperform–closed-source models.

John Bosco Mugeni, Steven J. Lynden, Toshiyuki Amagasa et al. · 1 citation

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.