Hybrid clustering framework for large-scale scientific literature structuring
Abstract
This study addresses the challenge of organizing and interpreting large collections of scientific literature, focusing on a rapidly growing and heterogeneous AI research domain derived from object-detection-with-small-data literature. First, based on the set of keywords using Clarivate Analytics queries, over 18,500 research articles were retrieved, illustrating the problem of manual review infeasibility, leading to machine learning-based approaches. To begin literature filtration, a pipeline combining text preprocessing, TF-IDF feature extraction, and K-means++ clustering was implemented. Clusters were labelled using YAKE keyword extraction and thematically summarized with large language models (ChatGPT-4o and ChatGPT-o1), enabling the interpretation of research domains such as medical imaging, anomaly detection, etc. To compare clustering strategies, we employed neural autoencoder-based embeddings and deep embedded clustering architectures. The autoencoder latent space yielded both convergences and divergences relative to TF-IDF clusters, reflecting higher-level semantic structuring. Direct end-to-end DEC optimization produced unstable cluster assignments, thus we applied a three-phase IDEC training protocol using autoencoder pretraining, centroid initialization with K-means, and joint fine-tuning – this produced substantially more balanced and internally coherent clusters. Results demonstrate that integrating keyword-based and neural methods, supplemented by conversational AI summarization, offers a scalable and interpretable framework for structuring large scientific datasets. This work highlights both methodological strengths and thematic overlaps across AI research, emphasizing the role of hybrid clustering approaches in managing the growing volume of scholarly publications.