Skip to content
Preprint

Direct Construction of Disambiguated Knowledge Bases from Large Language Models

Aug 2026 · 0 citations
Computer Science

TL;DR

This work proposes GPTKB 2.0, a methodology for constructing disambiguated KBs directly from large language models (LLMs) that incorporates on-the-fly disambiguation of entities, relations and classes, and is meticulously designed to satisfy both scalability and disambiguation accuracy.

Abstract

Automated Knowledge Base Construction (AKBC) is a core NLP task, and recent work proposes generating knowledge bases directly from large language models (LLMs), treating the model itself as the knowledge source. However, LLMs natively possess no representation of entities, leading to duplicate entries as well as conflations. We propose GPTKB 2.0, a methodology for constructing disambiguated KBs directly from LLMs. GPTKB 2.0 incorporates on-the-fly disambiguation of entities, relations and classes, and is meticulously designed to satisfy both scalability and disambiguation accuracy. We analyze the central design decisions and characterize the trade-offs between accuracy, scale, and cost. We execute GPTKB 2.0 at scale, obtaining a materialized KB containing over 1M disambiguated entities and 38.4M triples. This represents the first million-scale LLM-native KB with explicit internal canonicalization of entities, relations, and classes, a significant departure from prior Wikimedia-centric works. GPTKB 2.0 is available at https://gptkb.org/.

View source

Similar papers

#natural language process... Preprint Aug 2026

GPTKB 2.0: Browsing, Querying, and Auditing a Disambiguated LLM-Derived Knowledge Base

GPTKB 2.0 performs context-guided disambiguation during recursive KB construction, separating homonyms and merging synonymous mentions as facts are elicited, and makes this process inspectable: users can browse entities, follow links across the KB, and audit the provenance of individual facts.

Yujia Hu, Tuan-Phong Nguyen, Simon Razniewski · 0 citations
#natural language process... Preprint Sep 2026

The Canonical Order Problem: When Large Language Models Are Unreliable Knowledge Bases for Multi-Valued Relations

Large language models (LLMs) are increasingly used as knowledge bases (KBs) due to the vast amount of knowledge they acquire during pre-training. While many works focus on extracting single relational triples, most real-world relations are multi-valued and require generating sets of entities. In this paper, we investig...

Timo Pierre Schrader, Annemarie Friedrich, Simon Razniewski et al. · 0 citations
Book Open access Aug 2026

Structure Shapes the Future of DataxLLM Systems: Retrieval, Structuring, and Reasoning

This tutorial presents a unified vision in which structuring serves as the enabling foundation for three pillars of next-generation LLM systems, highlighting how the cooperative interplay between classical KDD techniques and modern LLMs-where KDD defines structural schemas and quality constraints while LLMs execute fle...

Peng-Cheng Jiang, Jiashuo Sun, Wonbin Kweon et al. · 0 citations
#natural language process... Preprint Aug 2026

REAP: Relation-Aware Elicitation and Parsing for Closed-Book Knowledge Base Construction from LLMs

The REAP system combines structured chain-of-thought reasoning, relation-specific query strategies, and a reasoning-based empty-set gate to elicit parametric knowledge, followed by direct extraction into valid JSON arrays for AKBC Shared Task 2026.

T. Bùi, T. Do, Tuan-Phong Nguyen · 0 citations
#natural language process... Preprint Sep 2026

BELXTR: Biomedical Entity Linking via Contextualized Token Retrieval

BELXTR is presented, a novel embedding model based on the multi-vector (a.k.a. late interaction) architecture, which allows to leverage token-level matching information in biomedical entity linking by integrating an existing task-specific training objective and exploring active query expansion.

Samuele Garda, U. Leser · 0 citations

Bridges Between Words and Senses

The proposed k-Multilingual Concept model allows to uncover novel layers of lexical knowledge in the form of multifaceted conceptual links between naturally disambiguated sets of words.

Francesca Grasso, Vladimiro Lovera, Luigi Di Caro · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.