This work proposes GPTKB 2.0, a methodology for constructing disambiguated KBs directly from large language models (LLMs) that incorporates on-the-fly disambiguation of entities, relations and classes, and is meticulously designed to satisfy both scalability and disambiguation accuracy.
Abstract
Automated Knowledge Base Construction (AKBC) is a core NLP task, and recent work proposes generating knowledge bases directly from large language models (LLMs), treating the model itself as the knowledge source. However, LLMs natively possess no representation of entities, leading to duplicate entries as well as conflations. We propose GPTKB 2.0, a methodology for constructing disambiguated KBs directly from LLMs. GPTKB 2.0 incorporates on-the-fly disambiguation of entities, relations and classes, and is meticulously designed to satisfy both scalability and disambiguation accuracy. We analyze the central design decisions and characterize the trade-offs between accuracy, scale, and cost. We execute GPTKB 2.0 at scale, obtaining a materialized KB containing over 1M disambiguated entities and 38.4M triples. This represents the first million-scale LLM-native KB with explicit internal canonicalization of entities, relations, and classes, a significant departure from prior Wikimedia-centric works. GPTKB 2.0 is available at https://gptkb.org/.
GPTKB 2.0 performs context-guided disambiguation during recursive KB construction, separating homonyms and merging synonymous mentions as facts are elicited, and makes this process inspectable: users can browse entities, follow links across the KB, and audit the provenance of individual facts.
Yujia Hu, Tuan-Phong Nguyen, Simon Razniewski· 0 citations
Large language models (LLMs) are increasingly used as knowledge bases (KBs) due to the vast amount of knowledge they acquire during pre-training. While many works focus on extracting single relational triples, most real-world relations are multi-valued and require generating sets of entities. In this paper, we investig...
Timo Pierre Schrader, Annemarie Friedrich, Simon Razniewski et al.· 0 citations
This tutorial presents a unified vision in which structuring serves as the enabling foundation for three pillars of next-generation LLM systems, highlighting how the cooperative interplay between classical KDD techniques and modern LLMs-where KDD defines structural schemas and quality constraints while LLMs execute fle...
Peng-Cheng Jiang, Jiashuo Sun, Wonbin Kweon et al.· Proceedings of the 32nd ACM...· 0 citations
The REAP system combines structured chain-of-thought reasoning, relation-specific query strategies, and a reasoning-based empty-set gate to elicit parametric knowledge, followed by direct extraction into valid JSON arrays for AKBC Shared Task 2026.
BELXTR is presented, a novel embedding model based on the multi-vector (a.k.a. late interaction) architecture, which allows to leverage token-level matching information in biomedical entity linking by integrating an existing task-specific training objective and exploring active query expansion.
The proposed k-Multilingual Concept model allows to uncover novel layers of lexical knowledge in the form of multifaceted conceptual links between naturally disambiguated sets of words.
Francesca Grasso, Vladimiro Lovera, Luigi Di Caro· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.