The Conference Organisers and Content Identifier (COCI) is presented, an AI-based framework designed to extract fine-grained, structured metadata from raw CfP texts and bridges the gap between informal scholarly dissemination and structured Semantic Web resources.
Abstract
Despite the critical role of grey literature in scholarly communication, artefacts such as Calls for Papers (CfPs) remain largely isolated from modern Scholarly Knowledge Graphs. The unstructured and highly heterogeneous nature of these documents has traditionally hindered their large-scale processing. In this demo paper, we present the Conference Organisers and Content Identifier (COCI), an AI-based framework designed to extract fine-grained, structured metadata from raw CfP texts. COCI employs a multi-stage pipeline that combines Large Language Models (LLMs) with semantic mapping techniques to integrate extracted entities with established knowledge bases, including OpenAlex, DBLP, TIB ConfIDent, and the AIDA Dashboard. By disambiguating authors and semantically aligning topics and conference series, COCI bridges the gap between informal scholarly dissemination and structured Semantic Web resources, laying the foundation for systematic analysis of non-publisher-based academic events.
COCI, an AI-based framework that automates the extraction of granular, structured metadata from raw CfP text, establishes a foundation for the systematic analysis of grey literature, enabling new research opportunities and shifting the scholarly focus towards non-publisher-based events.
Angelo Salatino, Francesco Osborne, Alexis Vizcaino et al.· 0 citations
Wikipedia is currently facing multiple crises. LLM-enabled chatbots have normalized the answer-based web in ways that decrease human readership of the encyclopedia while also increasing the burden of the community to adapt to increasing amounts of AI-generated content. The typical answer to emergent issues on Wikiped...
This paper presents a case study of marking up TEI (Text Encoding Initiative) documents using large language models. We detail our experiments using LLMs for named entity recognition and entity linkage in a corpus of TEI documents. The corpus consists of four journals co-edited by the Swiss-German theologian, Karl Bart...
Jonah Causin, Morgan Fuksa, Arne Käfer et al.· Journal of Humanities and AI· 0 citations
PubLink is presented, a modular toolset designed to bridge gaps in digital publishing by connecting existing systems and standards rather than replacing them, minimizing dependency on any single platform.
E. Bastianello, C. Tomlinson· Proceedings of the 37th ACM...· 0 citations
Large language models augmented with external knowledge produce more factual and context-rich answers. We present a domain-agnostic framework that transforms unstructured textual evidence into a consolidated knowledge graph, in which nodes are concepts with metadata and edges are statements with structured attributes....
Illia Dagil, I. Vergunova· Automation, Control, and Inf...· 0 citations
This work builds four modules (for drug-discovery chemistry, materials science, machine learning, and mineral geochemistry) in the ArticleMiner framework, and evaluates them on 163 papers, including a new geochemistry benchmark with expert-curated ground truth.
Md Abrar Jahin, Craig A. Knoblock, Jay Pujara· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.