Aug 2026· International Journal of Intelligent Systems and Data Science· Vol 1· 0 citations· 29 references
TL;DR
The results highlight the role of digital libraries and algorithmic normalization in addressing terminological and data-quality challenges, alongside name disambiguation procedures and network pruning to mitigate citation boundary issues.
Abstract
This paper presents an information technology-driven bibliometric framework for mapping the structural and temporal evolution of interdisciplinary intellectual linkage and apparent conceptual diffusion within the domain of social network analysis. Utilizing large-scale metadata from Web of Science, we construct multiple bibliometric networks representing citation, co-authorship, and keyword relationships. Through computational methods including fractional normalization, Search Path Count (SPC) edge-weighting, and island detection algorithms, we identify latent communities and bibliometric pathways that suggest patterns of intellectual association and citation-mediated influence between social sciences, physics, neuroscience, and behavioral ecology. These computational outputs represent citation-mediated influence patterns and co-occurrence structures, not direct observation of knowledge transmission. Our results highlight the role of digital libraries and algorithmic normalization in addressing terminological and data-quality challenges, alongside name disambiguation procedures and network pruning to mitigate citation boundary issues. This paper contributes to the domain of information technology by demonstrating scalable computational techniques for uncovering hidden intellectual structures and bibliometric evidence of knowledge flows in an increasingly interdisciplinary research landscape.
The academic publishing ecosystem is a vast, heterogeneous network of works, authors, institutions, journals, and topics. Traditional scientometrics reduces it to isolated tabular indicators (h-index, Impact Factor) that ignore topological context and are not designed to capture coordinated illegitimate practices. Building on our companion review, which proposed graph analysis of publishing integrity, this paper implements that approach. We define a heterogeneous multivariate graph model over OpenAlex open data (seven node types, seven edge types) and a methodology based on projections (citation and co-authorship networks), interpretable structural metrics, community detection, and three screening detectors of anomalous publishing patterns. We deliberately avoid binary classification: detectors return ranked candidates with explicit structural evidence for human assessment. On the institutional corpus of VSB - Technical University of Ostrava (2020-2025) with its one-hop citation neighbourhood, community detection recovers real research groups, centralities identify cross-disciplinary bridges, and the screenings flag dense co-authorship cliques, locally closed citation loops, and thematically isolated venues. On a second, venue-centric corpus with external ground truth (journals delisted by Scopus and DOAJ) and size-matched controls, a naive case-control design yields seemingly strong but spurious detectors (a prominence confound), whereas after matching the only robust signal is the breadth of disciplinary scope (AUC 0.70); an open graph-based prestige measure (PageRank over the journal citation network) tracks a JIF proxy while being an order of magnitude more resistant to citation gaming than count-based indicators. We release the method as the open-source library apnet with a reproducible CLI workflow and a web interface; the analysis runs on commodity hardware in minutes.
Systematic literature reviews (SLRs) face challenges from rapid publication growth and low-quality AI-generated content. Simple database queries often retrieve publications that are not thematically coherent, making meaningful clustering difficult. This study aims to develop and evaluate a hybrid method to automate cluster assessment in SLRs, combining statistical measures of semantic similarity (embeddings from nomic-embed-text-v1.5) with bibliometric metrics (shared references, keywords, and Jaccard indices). Three case studies (17, 301, and 3113 publications) from the field of management were analyzed using embedding-based pre-clustering filtration, k-means clustering, t-SNE visualization, and GPT-4 labeling, validated through independent expert assessment (mean rating 4.29/5). Our results show that statistical and bibliometric metrics complement each other: bibliometric metrics uncover intellectual lineage, while statistical metrics assess semantic cohesion. Removing the least relevant articles via pre-clustering filtering consistently enhanced cluster coherence in the analyzed cases and produced more coherent publication sets than the complex Boolean queries used as a baseline. Furthermore, dataset size influences validation: bibliometric metrics can be misleading for small collections due to sparse networks, whereas statistical metrics remain reliable. This hybrid method provides a reproducible, scalable, open-source approach for automated cluster assessment in SLRs and will be implemented in the EmbedSLR open-source software.
Sebastian Matysik, Joanna Wiśniewska, Paweł Karol Frankowski· IEEE Access· 0 citations
How the field has evolved from cognitive and communication-centered perspectives toward a broader socio-technical understanding of digital work is shown and offers a structured agenda for cumulative theory development.
Jan Schmutzer, P. Vrabcová· Journal of Documentation· 0 citations
Research software forms distinct co-usage communities that span traditional disciplinary boundaries, yet the structure of these communities remains largely unexplored. We present a graph-based framework for discovering mathematical software communities and predicting their association with research publications. We construct a software co-usage network from publication-software relationships using a curated swMATH dataset and subsequently apply community detection method, revealing a heterogeneous landscape of mathematical software communities. We formulate publication-to-community mapping as a multi-label classification task and further investigate whether community membership can be predicted from lightweight scholarly metadata. Specifically, we compare two feature representations of scientific publications: Mathematics Subject Classification (MSC) and title-based embeddings. Across a range of models, structured MSC representation consistently provides a stronger precision-recall trade-off, demonstrating that structured domain metadata captures software-community structure more effectively than compressed title-only semantics in this setting. This work highlights the continuing value of structured scholarly metadata for large-scale research software discovery, classification and recommendation.
Maxence Azzouz-Thuderoz, Yuni Susanti, Moritz Schubotz· 0 citations
This study employs computational scientometric modeling to analyze the citation network of the seminal work Usability Evaluation of E-books, addressing the lack of quantitative understanding of how interdisciplinary knowledge evolves in interactive reading design and human–computer interaction. Using bibliometric data from Web of Science and Scopus, 135 citing publications were retrieved and processed through a data-driven citation-tracking algorithm. CiteSpace was utilized to generate dynamic knowledge maps integrating network topology, thematic clustering, and temporal modeling. The analysis reveals a hierarchical and concentric knowledge structure that evolved through three distinct stages: a technology-driven emergence phase (2009–2012), an application expansion phase (2013–2017), and a human-centered convergence phase (2018–present). Bridging nodes such as “visual fatigue” and “human–computer interaction” computationally link previously isolated research clusters, illustrating the mechanisms of interdisciplinary integration. The findings propose a hierarchical computational model of knowledge evolution, demonstrating the effectiveness of algorithmic citation analysis in mapping intellectual structures. This work provides a reproducible methodological framework for analyzing scholarly influence and knowledge development, offering insights into the future evolution of interactive reading design and related computational humanities fields.
Bibliometric indicators - citation counts, h-indexes, co-authorship networks - have long anchored science, technology, and innovation (STI) analytics, yet suffer from temporal lag, semantic shallowness, and an inability to capture the non-linear dynamics of contemporary knowledge ecosystems. Dynamic knowledge graphs and large language models (LLMs) have each been proposed as remedies, but neither is sufficient alone: existing scholarly knowledge graphs remain largely static, while LLM-driven pipelines are prone to hallucination, opacity, and corpus bias without structured grounding. This paper proposes a hybrid, symbolic-first framework integrating all three traditions under explicit methodological constraint. Organized across five layers - an open scholarly data backbone, a dynamic versioned knowledge graph, a constrained LLM-assisted semantic augmentation layer, a multi-layer validation pipeline, and an analytics layer - the framework positions LLMs strictly as generators of provisional candidate enrichments. Candidates become analytically admissible only after passing structural, evidentiary, comparative, and selective expert validation, with full provenance recorded at every stage. The analytics layer supports both established bibliometric indicators and extended graph-based analyses, including trend emergence detection, science-to-technology pathway mapping, and policy-oriented gap analysis. The framework's central theoretical contribution is treating validation as the mediating principle between semantic flexibility and epistemic discipline, enabling STI analytics that is semantically richer and temporally more responsive than static bibliometrics while remaining aligned with the evidentiary standards of science-of-science research. Governance considerations addressing reproducibility, bias, and auditability are also discussed.
Muhsen Hammoud· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.