Skip to content

Toward a Unified Statistical Theory of Unsupervised Pretraining and Supervised Neural Knowledge Graph Learning

Jul 2026 · arXiv.org · Vol abs/2607.26346 · 0 citations · 34 references
Mathematics Computer Science

TL;DR

A nonasymptotic risk bound is established that disentangles pretraining representation error from labeled-sample complexity, formally quantifying the benefit of large-scale unlabeled data for downstream knowledge prediction.

Abstract

Knowledge graph learning provides a powerful framework for representing and inferring structured knowledge, with broad practical applications. However, the scarcity of relation-specific labeled triples per entity hinders the training of expressive models, and the ad hoc design of scoring functions limits generalizability and lacks theoretical grounding. We address both issues with a theoretically grounded, end-to-end training framework that extends and subsumes existing methods. Our framework is a two-stage procedure: unsupervised pretraining over heterogeneous corpora followed by supervised learning with multiple relation types. We establish a nonasymptotic risk bound that disentangles pretraining representation error from labeled-sample complexity, formally quantifying the benefit of large-scale unlabeled data for downstream knowledge prediction. Synthetic experiments validate each theoretical component, and real-world experiments confirm the effectiveness of our approach on large-scale knowledge graph benchmarks.

View source

Similar papers

Preprint Aug 2026

Why Does Graph Learning Fail to Fully Benefit from a Text Teacher?

A multimodal model that combines two complementary ideas: a self-supervised method that enables a GNN encoder pretrained on one dataset to operate directly on another dataset with a different node-feature dimensionality, without rebuilding the model or realigning the data is investigated.

Fumiaki Kimino, Ryoma Sato Sokendai, National Institute of Informatics · 0 citations
Conference Open access Sep 2026

LLMs as Parametric Knowledge Sources for Knowledge Graph Completion

This framework performs LLM knowledge elicitation to extract factual knowledge from the model’s internal representations and transforms sentence-level representations into entity-level representations and aligns them within a unified space.

Deyu Chen, Qiyuan Li, Jinguang Gu et al. · 0 citations
Conference 2026

SAGE: Structure-Aware Generative Enhancement for Long-Tail Knowledge Graph Completion

A pre-extractor model based on a hybrid architecture of rules and neural networks is introduced, which is used to identify long tail entities in the dataset and generate several candidate tail entities through relationships to improve the inference performance of the model.

Er-Zhuo Xu · 0 citations
Preprint Aug 2026

SelfGraphRAG: Bridging the Supervision Gap in Graph-Based RAG with Synthetic QA Generation

SelfGraphRAG, a framework that generates question-answer pairs directly from knowledge graph structure and uses them to train a query-conditioned graph retriever, suggests that knowledge graph structure can provide useful supervision for training graph retrievers when labeled data are unavailable.

Ben Lagnese, Manas Gaur · 0 citations
Preprint Aug 2026

Interpretable Unsupervised Community Detection with LLM-Symbolized Structured Processes

LUCID, an LLM-guided, interpretable, training-free, and unsupervised community detection method, designed as a four-stage pipeline that achieves state-of-the-art performance and consistently outperforms leading unsupervised and semi-supervised baselines.

Aoting Zeng, Kai Wang, Jianwei Wang et al. · 0 citations
Conference 2026

Fusing Relation Graph and Mutual Information for Inductive Link Prediction in Knowledge Graphs

Inductive link prediction in knowledge graphs refers to the task of inferring known relations between entities unseen during training. Most existing approach-es are limited to predicting only known relations and struggle to generalize to un-seen relations, which restricts their utility in dynamic settings. To address this challenge, we propose a novel inductive link prediction approach named RGIILP. Specifically, we construct a relation graph from the source knowledge graph and design a neural network model that enables interactive feature propagation be-tween entities and relations. Furthermore, we introduce the mutual information maximization mechanism between global and local representations to capture the global structural information of the graph. Experiments on several benchmark da-tasets demonstrate that RGIILP outperforms existing state-of-the-art methods for inductive link prediction task.

Hong-Bo Liu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.