Jun 2026· Theory and Practice of Logic Programming· Vol abs/2607.19365, pp. 1-18· 1 citation· 30 references
Computer Science
TL;DR
A logic-guided data extraction framework combining LLM-based extraction with answer set programming (ASP), which reduces LLM calls and improves extraction quality by mitigating spurious outputs, demonstrating the value of non-monotonic logic programming for controlled semantic extraction.
Abstract
When large language models (LLMs) are used for semantic data extraction from unstructured text, producing candidate relational facts from natural language, they may remain unreliable for tasks requiring complex combinatorial reasoning and global consistency. This paper proposes a logic-guided data extraction framework combining LLM-based extraction with answer set programming (ASP). The LLM produces candidate facts, whereas ASP performs validation, inference, consistency checking, and control. Unlike existing pipelines that query the LLM independently for all target predicates, the proposed approach uses ASP reasoning to identify which predicates are logically admissible at each stage and to guide extraction queries. By interleaving LLM calls with ASP derivation, the framework infers logically implied facts without further extraction and detects inconsistencies early. We formalize the pipeline and prove that, under mild assumptions, it is equivalent to the baseline approach with respect to the final extracted facts, while requiring fewer LLM calls. We also introduce a caching mechanism for logic-based control queries, exploiting monotonicity of conjunctive queries over incrementally constructed fact sets to reduce solver invocations. Experiments on ASP-derived benchmarks show that the framework reduces LLM calls and improves extraction quality by mitigating spurious outputs, demonstrating the value of non-monotonic logic programming for controlled semantic extraction.
HGX is introduced, which natively embeds domain dependencies into an expressive parent-child hierarchy, and achieves up to a 92% reduction in ASP solver invocations compared to flat evaluation strategies, while strictly preserving extraction accuracy.
guided table retrieval is presented, a four-phase pipeline that combines deterministic grounding via hash-based predictors, structural exploration of join-graph reachability, LLM-powered disambiguation of sources and targets, and algorithmic merging into minimal, topologically ordered join trees.
Alekh Jindal, J. Pandey, C. Pavlopoulou et al.· 0 citations
Experiments show that precision improves by filtering invalid candidates, while recall is preserved due to retaining candidates whose constraints are not explicitly violated, and a three-valued constraint semantics that avoids incorrect rejections under open-world assumptions.
E. Kitzelmann· Deutsche Jahrestagung für Kü...· 0 citations
SafeQL is proposed, a search-based refinement paradigm that redefines the role of the DBMS as an active guide in the refinement process, and significantly improves execution accuracy and efficiency compared to regeneration-based methods.
Geonho Lee, Min-Soo Kim· Proceedings of the VLDB Endo...· 0 citations
It is shown that when domain vocabulary and semantics are captured in a well-designed Web Ontology Language (OWL) ontology, Large Language Models (LLMs) can generate accurate structured queries zero-shot, without task-specific fine-tuning, retrieval augmentation, or multi-agent orchestration.
B. Fitch, Cato Elia Kurtz· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.