COCI, an AI-based framework that automates the extraction of granular, structured metadata from raw CfP text, establishes a foundation for the systematic analysis of grey literature, enabling new research opportunities and shifting the scholarly focus towards non-publisher-based events.
Abstract
Despite its importance, grey literature, including Calls for Papers (CfPs), remains largely overlooked in Metascience and Scientometric analysis due to its unstructured, highly heterogeneous format, which traditional tools struggle to process at scale. However, Large Language Models now offer a pivotal opportunity to devise innovative tools for systematically harvesting and processing such data. In this paper, we introduce COCI, an AI-based framework that automates the extraction of granular, structured metadata from raw CfP text. COCI employs a multi-stage pipeline for entity extraction, followed by author disambiguation against OpenAlex and semantic mapping of topics and conference series. This process identifies key data points, including conference editions, geographic locations, and comprehensive lists of organisers, along with their specific roles and affiliations. By structuring this previously inaccessible information, COCI establishes a foundation for the systematic analysis of grey literature, enabling new research opportunities and shifting the scholarly focus towards non-publisher-based events.
The Conference Organisers and Content Identifier (COCI) is presented, an AI-based framework designed to extract fine-grained, structured metadata from raw CfP texts and bridges the gap between informal scholarly dissemination and structured Semantic Web resources.
Angelo Salatino, Francesco Osborne, Alexis Vizcaino et al.· 0 citations
This hybrid method provides a reproducible, scalable, open-source approach for automated cluster assessment in SLRs and will be implemented in the EmbedSLR open-source software.
Sebastian Matysik, Joanna Wiśniewska, Paweł Karol Frankowski· IEEE Access· 0 citations
An AI-guided framework is developed that aligns AI-assisted metadata extraction with Dublin Core Terms and the FAIR principles for digital libraries, archives, and cultural-heritage repositories and is evaluated as a design-science artefact in which retrieval is not a side feature but a feedback loop.
Wirapong Chansanam, Umawadee Detthamrong, Chunqiu Li et al.· 0 citations
Wikipedia is currently facing multiple crises. LLM-enabled chatbots have normalized the answer-based web in ways that decrease human readership of the encyclopedia while also increasing the burden of the community to adapt to increasing amounts of AI-generated content. The typical answer to emergent issues on Wikiped...
Research teams and organizations often explore unfamiliar free-text collections, from survey comments and reviews to reports and domain documents, before labels, queries or coding schemes exist. At this stage, the first thematic map shapes what users notice, prioritize and carry into downstream analysis, so it should b...
In the 2026 information landscape, information retrieval is increasingly dominated by search engines and generative artificial intelligence (GenAI). While recognizing the genuine value of these technologies, this paper argues that they exhibit documented limitations, such as hallucinated citations, visibility bias, and...
Piero Cavaleri· JLIS.it· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.