Skip to content
Conference

Aggregating Duplicate Edges in Text-Derived Knowledge Graphs

Sep 2026 · Automation, Control, and Information Technology · pp. 1346-1349 · 0 citations · 17 references

Abstract

Large language models augmented with external knowledge produce more factual and context-rich answers. We present a domain-agnostic framework that transforms unstructured textual evidence into a consolidated knowledge graph, in which nodes are concepts with metadata and edges are statements with structured attributes. When many documents discuss the same concepts, a small set of entities accumulates many duplicate edges; the framework reconciles these redundant edges while preserving the breadth of supporting evidence. The pipeline has three stages: (i) knowledge extraction, using an artificial intelligence agent or a deterministic LLM pipeline; (ii) edge aggregation, which clusters semantically equivalent edges by combining overlapping numerical, categorical, and textual information while keeping contradictory findings apart; and (iii) a conversational agent that issues structured queries to the aggregated graph and returns answers supported by citations. Building on Retrieval-Augmented Generation (RAG), GraphRAG, and MedGraphRAG, the framework adds a graphlevel consolidation step and evaluates it directly. On a 350-paper nutrition corpus we compare the proposed online LLM claim clustering against trivial bounds and density-based baselines (DBSCAN, HDBSCAN) over 184 high-multiplicity evidence edges. The method compresses these edge sets $3.6 x$ while preserving contradictory evidence rather than collapsing it, keeping opposing-direction findings in separate clusters on every genuinely contested node pair; merge-style aggregation erases such conflicts, and the density baselines either collapse them or barely compress. Against a human-adjudicated subset, our partition matches expert judgment more closely than the LLM reference itself. The resulting graph offers a concise, updatable, provenance-rich source of truth and a foundation for automated knowledge base generation.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.