Skip to content
Preprint

COCI: Conference Organisers and Content Identifier

Aug 2026 · 0 citations · 15 references
Computer Science

TL;DR

The Conference Organisers and Content Identifier (COCI) is presented, an AI-based framework designed to extract fine-grained, structured metadata from raw CfP texts and bridges the gap between informal scholarly dissemination and structured Semantic Web resources.

Abstract

Despite the critical role of grey literature in scholarly communication, artefacts such as Calls for Papers (CfPs) remain largely isolated from modern Scholarly Knowledge Graphs. The unstructured and highly heterogeneous nature of these documents has traditionally hindered their large-scale processing. In this demo paper, we present the Conference Organisers and Content Identifier (COCI), an AI-based framework designed to extract fine-grained, structured metadata from raw CfP texts. COCI employs a multi-stage pipeline that combines Large Language Models (LLMs) with semantic mapping techniques to integrate extracted entities with established knowledge bases, including OpenAlex, DBLP, TIB ConfIDent, and the AIDA Dashboard. By disambiguating authors and semantically aligning topics and conference series, COCI bridges the gap between informal scholarly dissemination and structured Semantic Web resources, laying the foundation for systematic analysis of non-publisher-based academic events.

View source

Similar papers

Preprint Aug 2026

A Pathway for Assessing Grey Literature: Leveraging AI to Extract Conference Metadata and Organiser Information from Calls for Papers

COCI, an AI-based framework that automates the extraction of granular, structured metadata from raw CfP text, establishes a foundation for the systematic analysis of grey literature, enabling new research opportunities and shifting the scholarly focus towards non-publisher-based events.

Angelo Salatino, Francesco Osborne, Alexis Vizcaino et al. · 0 citations
Open access Sep 2026

Wikimedia research methodologies: a bibliometric history of peer production and AI (2005–2025)

Wikipedia is currently facing multiple crises. LLM-enabled chatbots have normalized the answer-based web in ways that decrease human readership of the encyclopedia while also increasing the burden of the community to adapt to increasing amounts of AI-generated content. The typical answer to emergent issues on Wikiped...

Steve Jankowski · 0 citations
Review Open access Sep 2026

LLM-ASSISTED MARKUP OF ENTITIES IN TEI: A CASE STUDY

This paper presents a case study of marking up TEI (Text Encoding Initiative) documents using large language models. We detail our experiments using LLMs for named entity recognition and entity linkage in a corpus of TEI documents. The corpus consists of four journals co-edited by the Swiss-German theologian, Karl Bart...

Jonah Causin, Morgan Fuksa, Arne Käfer et al. · 0 citations
Conference Sep 2026

Aggregating Duplicate Edges in Text-Derived Knowledge Graphs

Large language models augmented with external knowledge produce more factual and context-rich answers. We present a domain-agnostic framework that transforms unstructured textual evidence into a consolidated knowledge graph, in which nodes are concepts with metadata and edges are statements with structured attributes....

Illia Dagil, I. Vergunova · 0 citations
#artificial intelligence Preprint Sep 2026

ArticleMiner: Ontology-Guided Knowledge Graph Construction from Scientific Publications

This work builds four modules (for drug-discovery chemistry, materials science, machine learning, and mineral geochemistry) in the ArticleMiner framework, and evaluates them on 163 papers, including a new geochemistry benchmark with expert-curated ground truth.

Md Abrar Jahin, Craig A. Knoblock, Jay Pujara · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.