Skip to content

A two-stage hybrid intelligence framework for subject indexing via semantic embedding and LLM collaborative optimization

Jul 2026 · Journal of information science · 0 citations · 43 references

TL;DR

Fusing semantic embedding with LLM reasoning, the proposed R3 framework can boost accuracy and provide new paths for hybrid intelligence in artificial intelligence–assisted subject indexing, delivering a possible solution for large-scale indexing in practice.

Abstract

Approaches to subject indexing need to consider a trade-off between operational efficiency and the maintenance of indexing quality. Given that existing automated approaches struggle with domain adaptability, this study proposes the recall–rank–rerank (R3) framework—a two-stage hybrid intelligence method for automated subject indexing. By simulating human indexing cognition through case-based reasoning and human–machine collaboration trained on the TIBKAT data set, the results demonstrate that R3 significantly improves efficiency and semantic precision within the standard test collection evaluation framework in information retrieval research. R3 operates without model fine-tuning: its first stage uses cross-lingual semantic embedding to recall and rank candidate subjects from similar documents, while the second employs a large language model (LLM) as a simulated expert for refinement—calibrating semantics, enriching implicit concepts, and reranking relevance. Results show R3 outperforms supervised fine-tuning and semantic recall (R2) methods. On the validation set, it lifts mean average precision (MAP) from 41.19% to 45.24% (a 4% gain) and increases top-5 precision (P@5) by 2.49%. In the LLMs4Subjects task, R3 achieves top recall rates—65.68% for core subjects and 58.56% for all subjects—leading the SemEval’25 shared evaluation. With its lightweight design and strong performance, R3 offers adaptability and scalability, delivering a possible solution for large-scale indexing in practice. Fusing semantic embedding with LLM reasoning, the proposed framework can boost accuracy and provide new paths for hybrid intelligence in artificial intelligence–assisted subject indexing. Implications for the design and evaluation of automated subject indexing systems are discussed.

View source

Similar papers

Open access Sep 2026

An Intelligent Research Paper Assistant: Integrated Retrieval, Explainable Domain Classification, and Research-Gap Ranking

This study presents an end-to-end Intelligent Research Paper Assistant built over 499,999 scholarly abstracts spanning 10 domains and 25 subdomains, integrating six modules: title suggestion, abstract retrieval, methodology generation, dual-branch domain classification, post-hoc explainability, and research-gap ranking...

Hamza Shahbaz · 0 citations
Aug 2026

Automated subject indexing in large-scale library collections: a minimal semantic enhancement strategy at the National Library of Spain

This study addresses a key technical challenge in automated subject indexing: improving prediction accuracy for large-scale library collections. We introduce and validate a hybrid AI framework that combines statistical and semantic embedding approaches to enhance the performance of automated subject heading assignm...

Gökhan Usta · 0 citations
#large language models Open access Sep 2026

SciRep: A Ranking-Aware Representation Model for Scientific Text

Results demonstrate that the proposed ranking-aware distillation mechanism significantly enhances scientific text representation quality while maintaining efficient inference, offering a more effective contrastive learning method for domain-specific retrieval tasks.

Bing-Hao Fu, Jun Wang · 0 citations
Preprint Aug 2026

Factorized Hypothesis Search for Evidence-to-Taxonomy Retrieval

Large-taxonomy retrieval often assumes that the input already expresses the target concept. In many settings, however, the input is indirect evidence, such as a table cell whose meaning depends on its row, column, datatype, and context. We call this mismatch the retrieval readiness gap. Our analysis shows that the curr...

Lin-Hai Ma, Ethan F. Wei, Xue-Qing Peng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.