This paper compares three families of methods: a TF-IDF baseline, topic-based models (LDA and BERTopic, including clone variants), and embedding-based retrieval using SciBERT with Faiss, and examines explainability through two perspectives: intrinsic topic-based explanations and post-hoc, retrieval-based explanations generated using language models.
Abstract
In this paper, we examine unsupervised, content-based collaboration recommendations using publication text in scholarly settings. We compare three families of methods: a TF-IDF baseline, topic-based models (LDA and BERTopic, including clone variants), and embedding-based retrieval using SciBERT with Faiss. To evaluate model behavior beyond simple lexical matching, we introduce a constrained setting where publication overlap between researchers is partially removed while still using historical co-authorship as proxy ground truth for post-hoc evaluation. Results show clear differences across methods. TF-IDF performs best under full information but drops significantly as overlap is reduced. In contrast, topic-based and embedding-based approaches show more stable performance, suggesting they capture broader distributional similarities, rather than relying only on direct lexical overlap. We also examine explainability through two perspectives: intrinsic topic-based explanations and post-hoc, retrieval-based explanations generated using language models. These provide complementary trade-offs between transparency and human readability.
Impact-DPO integrates temporally informed prompting with direct preference optimization, enabling LLMs to learn comparative influence patterns without explicit graph message passing, and formalizes citation forecasting as pairwise preference learning on temporal text-attributed graphs.
Parham Hamouni, Ebrahim Bagheri· ACM Transactions on Intellig...· 0 citations
It is found that LLMs systematically reinforce the Matthew effect by favoring highly cited papers, with field-specific variation in the rate at which generated references match real papers in bibliometric databases.
Andres Algaba, Vincent Holst, Floriano Tori et al.· Quantitative Science Studies· 2 citations
This study presents a clear and reliable framework for classifying the intent behind scientific citations. It combines multi-model reasoning with concepts from social choice theory. Instead of using a single model, this framework employs three open Large Language Models Gemma, LLaMA, and Mistral. Additionally, we combi...
M. Barchane, Saad Belefqih, El Habib Ben Lahmar et al.· Algorithms· 0 citations
CiteFuncRanker establishes a robust and interpretable ranking-based paradigm for bibliometric research by capturing the nuanced and context-dependent relative preferences between citation roles, and advance citation analysis beyond categorical classification toward a more context-aware and semantically grounded underst...
Yi Wang, Xuan-Min Ruan, Dongqing Lyu et al.· Scientometrics· 0 citations
This study presents one of the first comprehensive evaluations of multiple state-of-the-art (SOTA) large language models (LLMs) for citation function classification, achieving new SOTA results on the ACL-ARC dataset.
Daniel Vodička, Jakub Šmíd, Pavel Král et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.