A Cognitive-Inspired Scanning Framework for Topic Modeling and Document Classification
Abstract
As digital textual data continues to grow exponentially, it becomes progressively harder to extract information from large collections of documents. Traditional full-text analysis methods often require substantial time and computational resources. This challenge is particularly relevant for massive corpora where important information is dispersed across different documents. This dispersion makes rapid comprehension difficult. Therefore, we propose an innovative method that combines scanning (skimming) reading techniques with Latent Dirichlet Allocation (LDA). This method enables an efficient exploration and an overview analysis of massive text corpora. Our approach leverages scanning as a rapid information extraction strategy in press articles. It identifies key and representative textual segments without the need for exhaustive reading. This preliminary filtering mechanism reduces processing complexity while preserving essential semantic content. To complement this step, we employ LDA to model latent topics and cluster documents based on their thematic structures. Unlike surface-level analysis, LDA exposes hidden topic distributions within the corpus, enabling the identification of both dominant topics and latent subtopics. Through the integration of scanning with LDA, we provide a structured framework to explore extensive collections of press articles. The scanning phase improves efficiency through the refinement of the textual scope. Meanwhile, LDA provides a robust probabilistic topic model and groups articles accordingly. This synergy facilitates a comprehensive understanding of the texts at both macro and micro levels.