Jul 2026· IEEE International Conference on Document Analysis and Recognition· pp. 228-244· 0 citations· 37 references
Computer Science
TL;DR
This paper introduces a semi-symbolic framework that integrates word-spotting techniques for post-OCR correction with a knowledge graph representation that enables the agent to access information through synthesized queries that are robust to misinterpretation and hallucination.
Abstract
The emergence of Large Language Models (LLMs) has redefined how users interact with information in digital environments. However, their widespread and often indiscriminate integration has raised significant concerns regarding reliability and trustworthiness issues that are particularly critical when accessing digital libraries and historical archives. How can one leverage the generalization capacity of an LLM without losing the level of accountability required for an archival institution? In this paper, we present an agentic retrieval system designed to deliver more accurate and verifiable access to historical data while preserving much of the flexibility associated with unconstrained LLMs. As a contribution to historical document analysis, we compare traditional Retrieval-Augmented Generation (RAG) with an agentic GraphRAG architecture in their ability to deliver historical information under realistic conditions, including the presence of OCR and transcription errors. We introduce a semi-symbolic framework that integrates word-spotting techniques for post-OCR correction with a knowledge graph representation that enables the agent to access information through synthesized queries. The interleaved collaboration between word spotting and code generation allows the agent to construct strong retrieval queries that are robust to misinterpretation and hallucination, while still leveraging approximate search when noise and uncertainty, common in historical document analysis, would otherwise hinder precise retrieval.
This work introduces a trust based adaptive reranking model- ATM (Adaptive Trust Model) that allocates computational resources according to file level uncertainty, instead of assigning a fixed number of reranker calls per query, which focuses computation only where ranking confidence is low.
Jenny Kalaiarasi.S· Journal of Intelligent Decis...· 0 citations
A literature-based architectural framework for reliable knowledge retrieval systems that separates external knowledge management from LLM-based reasoning and generation is developed and indicates that reliable LLM deployment should be treated as an end-to-end architectural problem rather than solely a model-performance problem.
Bharat Kumar Reddy Karumuri· International Journal of Eng...· 0 citations
Traditional fact-checking methods, while effective, are often too slow to keep up with the speed at which false information circulates online. In recent years, artificial intelligence (AI) has gained popularity as a means of automating the fact-checking process, particularly large language models (LLMs). Although LLMs have demonstrated efficacy in assisting with verification tasks, they are constrained by factors such as the quality of training data and their ability to retrieve pertinent information for verification. They are also prone to ‘hallucinations’, generating plausible but false or misleading information, raising critical concerns about their reliability as independent fact-checkers. This paper explores a hybrid approach that combines retrieval-augmented generation (RAG) with knowledge graphs (KGs) and OpenCTI, a platform for cyber threat intelligence. This approach ensures that the final output is not only informative but also transparent, linking claims to authoritative sources. This makes it a valuable tool for journalists, fact-checkers, or the general public, who increasingly require timely and reliable answers in an information environment shaped by disinformation.
Ashneet Khandpur Singh, Pau Perea Paños, Mario Reyes de los Mozos et al.· Medijske Studije· 0 citations
This work proposes an architecture called K-GRASP (Knowledge Graph-based Retrieval-Augmented Structured Prompting), which combines the representational power of knowledge graphs (KG) with the probabilistic reasoning capabilities of Large Language Models (LLMs) to address the externalisation of tacit knowledge.
Rafael Luna, Gabriel S. Luna, C. E. Barbosa et al.· European Conference on Knowl...· 0 citations
Kontrast is presented, an automatic framework that uses Text-to-SPARQL and LLM reasoning to compare table-based answers with KG evidence and categorize the resulting inconsistencies, and shows that text, tables, and KGs can complement and correct one another through systematic comparison.
Knowledge graphs have become a key resource for integrating heterogeneous data and powering downstream tasks such as question answering, entity linking, and semantic search. They are built and maintained incrementally, either (i) fully automated, e.g., YAGO, (ii) semiautomatically with community oversight, e.g., DBpedia, or (iii) manually through collaborative editing, e.g., Wikidata. Understanding the evolution of knowledge graphs is essential as changes may reflect real-world updates, error corrections, or noise introduced by vandalism, all of which affect the reliability of downstream applications. Among openly available knowledge graphs, Wikidata is the most challenging case to study evolution, with over 120 million entities edited by humans and bots and an edit history spanning more than a decade. Although Wikidata exposes change data in various formats (e.g., periodic dumps and real-time event streams), none support analytical queries over the complete edit history. Therefore, we present WiDiff, a tool that extracts changes from Wikidata's complete edit history and provides a unified interface for large-scale analytical queries over it.
C. Cortes, Lisa Ehrlinger, Lorena Etcheverry et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.