Skip to content

Distilling Answer Set Programming Theories from Large Language Models

Jul 2026 · arXiv.org · Vol abs/2607.28086 · 0 citations · 46 references
Computer Science

TL;DR

This work takes a neurosymbolic approach to study whether a model can distill complete and correct theories, given a fixed agent harness with the solver in the loop, and releases the code, prompts, and theories distilled.

Abstract

Writing Answer Set Programming (ASP) theories from scratch is a difficult and time-consuming task. We take a neurosymbolic approach to study whether a model can distill complete and correct theories, given a fixed agent harness with the solver in the loop. The protocol is dataset-agnostic: with a single prompt and an empty file as the starting point the model is given a 1-hour time limit to derive a complete theory. We chose VQA as the application domain, three benchmarks (CLEVR, GQA, CLEVRER), as these are publicly available and non-trivial. In order to study the model scale required for solving this task we nine different models: four frontier (Claude Sonnet 4.6, Claude Opus 4.7, GPT-5, DeepSeek V4 Pro), two mid-tier (DeepSeek V4 Flash, gpt-oss-120b), and three open-weights (qwen3.6-27b, gpt-oss-20b, qwen3.5-9b). Three of four frontier models reach 100% on CLEVR and 92.8%-98.8% on GQA; on CLEVRER, Sonnet, Opus, DeepSeek V4 Pro score 92.7%-95.3%. GPT-5 reaches 98.7% on CLEVR but drops to 41.8% on GQA and to 86.7% on CLEVRER. Adding handwritten reference theories from other datasets moves the other three frontier models by at most +/-3.4 pp but reduces GPT-5's accuracy by 3-19 pp. We release the code, prompts, and theories distilled.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

An Empirical Evaluation of Cost-Efficient Large Language Models on Algorithmic Programming Tasks

This study empirically evaluates whether cost-efficient Large Language Models (LLMs) can be trusted to generate enterprise code to a written specification. Three models (Gemini Flash 3, GPT-5.4 mini and Claude Haiku 4.5) were asked to solve 992 algorithmic problems as Java Spring Boot service methods conforming to a ma...

Chandimal Adikari, Nandika Herath · 0 citations
#machine learning Preprint Aug 2026

ClosureBench: A Constructive Benchmark for Compositional Graph Reasoning

ClosureBench is introduced, a constructive benchmark for compositional graph-relational reasoning with programmatically verified ground truth with programmatically verified ground truth: each task's reference answer is computed by executing a program in the Ein tensor-logic language, ensuring machine-verified correctne...

S. Goria · 0 citations
#artificial intelligence Preprint Sep 2026

DatalogBench: Evaluating Large Language Models on Text-to-Datalog Synthesis

Datalog underpins reasoning tasks such as program analysis, but its programs are hard to write. Existing synthesizers automate this task but require users to state their intent as input-output examples. Large language models (LLMs) suggest a more natural route, text-to-Datalog synthesis from a natural-language question...

Yuan Li, Han-Yun Jiang, Guo-Wei Tian et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SAILOR: Solver-Assisted Interactive LLM-based Optimization Recovery

Natural-language descriptions of optimization problems may be incomplete or vague about numerical information that a solver requires, including costs, capacities, demands, bounds, and penalties. A language model can translate the description into code, but when a required value is absent it must either stop or guess. W...

Shaghayegh Sadeghi, Steve Smith, D. C. Del Rey Fernández · 0 citations
#artificial intelligence Preprint Sep 2026

SemVerBench: Benchmarking LLM Comprehension of Version-Constraint Resolution Semantics

Large language model (LLM) coding agents constantly decide whether a version satisfies a constraint such as ^1.2.3 or>=2.0,<3, yet their grasp of version-constraint semantics has never been measured directly. We introduce SemVerBench, the first benchmark of LLM version-constraint resolution semantics across three ecosy...

Qi-Bai Chen, Ze-Ming Liu · 1 citation
#natural language process... Preprint Sep 2026

Large Language Models for Programming: Actually Fixing or Reimplementing Incorrect Code?

Recent studies have shown that Large Language Models can effectively solve problems and fix bugs in diverse programming environments, including competitive programming. Existing approaches primarily evaluate LLM performance in problem solving or bug fixing independently, but do not explore the relationship between thes...

Alexandru Stefan Stoica, Traian Rebedea, M. Mihăescu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.