Skip to content
Open access

An Intelligent Research Paper Assistant: Integrated Retrieval, Explainable Domain Classification, and Research-Gap Ranking

Sep 2026 · Advances in Artificial Intelligence Research · 0 citations · 21 references

TL;DR

This study presents an end-to-end Intelligent Research Paper Assistant built over 499,999 scholarly abstracts spanning 10 domains and 25 subdomains, integrating six modules: title suggestion, abstract retrieval, methodology generation, dual-branch domain classification, post-hoc explainability, and research-gap ranking.

Abstract

Background: The exponential growth of scientific literature creates a practical bottleneck: researchers must manually navigate domain taxonomies, identify research gaps, formulate titles, and bootstrap a methodology before writing begins. Methods: This study presents an end-to-end Intelligent Research Paper Assistant built over 499,999 scholarly abstracts spanning 10 domains and 25 subdomains, integrating six modules: title suggestion, abstract retrieval, methodology generation, dual-branch domain classification, post-hoc explainability, and research-gap ranking. A deterministic six-stage NLP pipeline feeds a TF-IDF encoder (12,000 features); an incremental SGD linear baseline and a fine-tuned DistilBERT transformer (66.96M parameters) are trained in parallel, with SHAP and LIME providing global and local explainability. Results: Domain classification achieves Accuracy = 1.00 (test-set per-domain accuracy 99.78%) for both models. SHAP and LIME correctly recover domain-discriminative vocabulary, and a composite novelty score identifies Social History and Digital Humanities as the most under-researched subdomains. Limitations: Results derive from a generated corpus with probable label co-occurrence inflating domain separability; real deployment requires retraining on verified bibliographic data.

Read PDF

Similar papers

#large language models Open access Sep 2026

SciRep: A Ranking-Aware Representation Model for Scientific Text

Results demonstrate that the proposed ranking-aware distillation mechanism significantly enhances scientific text representation quality while maintaining efficient inference, offering a more effective contrastive learning method for domain-specific retrieval tasks.

Bing-Hao Fu, Jun Wang · 0 citations

Automatic Domain Classification of Tabular Datasets Using Large Language Models

It is asserted that the present contribution consists of an interpretable domain palette, a constructed benchmark of diverse tabular datasets, and reproducible code and data to enable further research on domain discovery and domain-aware tooling for tabular data.

Elizaveta Gamper, Irina Deeva · 0 citations
#natural language process... Preprint Aug 2026

RATIO: A Benchmark for Retrieval Across Typed Ideation Operations in Scientific Literature

RATIO (Retrieval Across Typed Ideation Operations), a large-scale benchmark in which relevance is defined by three operations which are name ideation moves, provides a scalable training and evaluation framework for retrieval components that support literature-grounded ideation, opening up new research avenues on scient...

Ms. S. Sharon, Tom Hope · 0 citations
Open access Aug 2026

Hybrid TF-IDF and Sentence-BERT for Academic Advisor Recommendation Enhancing Topic Matching in Undergraduate Theses

The selection of thesis supervisors is a crucial stage that affects the smooth completion of students’ research. The manual process traditionally used often creates difficulties in finding supervisors whose expertise aligns with the research topic, leading to academic inefficiencies. This study aims to design and devel...

Syahmi Sajid, N. Fadil, Aidizzacky Harizulfadly Mtd et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.