This study presents an end-to-end Intelligent Research Paper Assistant built over 499,999 scholarly abstracts spanning 10 domains and 25 subdomains, integrating six modules: title suggestion, abstract retrieval, methodology generation, dual-branch domain classification, post-hoc explainability, and research-gap ranking.
Abstract
Background: The exponential growth of scientific literature creates a practical bottleneck: researchers must manually navigate domain taxonomies, identify research gaps, formulate titles, and bootstrap a methodology before writing begins. Methods: This study presents an end-to-end Intelligent Research Paper Assistant built over 499,999 scholarly abstracts spanning 10 domains and 25 subdomains, integrating six modules: title suggestion, abstract retrieval, methodology generation, dual-branch domain classification, post-hoc explainability, and research-gap ranking. A deterministic six-stage NLP pipeline feeds a TF-IDF encoder (12,000 features); an incremental SGD linear baseline and a fine-tuned DistilBERT transformer (66.96M parameters) are trained in parallel, with SHAP and LIME providing global and local explainability. Results: Domain classification achieves Accuracy = 1.00 (test-set per-domain accuracy 99.78%) for both models. SHAP and LIME correctly recover domain-discriminative vocabulary, and a composite novelty score identifies Social History and Digital Humanities as the most under-researched subdomains. Limitations: Results derive from a generated corpus with probable label co-occurrence inflating domain separability; real deployment requires retraining on verified bibliographic data.
Results demonstrate that the proposed ranking-aware distillation mechanism significantly enhances scientific text representation quality while maintaining efficient inference, offering a more effective contrastive learning method for domain-specific retrieval tasks.
Bing-Hao Fu, Jun Wang· Applied Sciences· 0 citations
It is asserted that the present contribution consists of an interpretable domain palette, a constructed benchmark of diverse tabular datasets, and reproducible code and data to enable further research on domain discovery and domain-aware tooling for tabular data.
RATIO (Retrieval Across Typed Ideation Operations), a large-scale benchmark in which relevance is defined by three operations which are name ideation moves, provides a scalable training and evaluation framework for retrieval components that support literature-grounded ideation, opening up new research avenues on scient...
The selection of thesis supervisors is a crucial stage that affects the smooth completion of students’ research. The manual process traditionally used often creates difficulties in finding supervisors whose expertise aligns with the research topic, leading to academic inefficiencies. This study aims to design and devel...
Syahmi Sajid, N. Fadil, Aidizzacky Harizulfadly Mtd et al.· Jurnal Edukasi dan Penelitia...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.