Skip to content
Preprint

Never the Number: Structural Abstention for AI Systems Whose Answers Are Consumed as Fact

Aug 2026 · 1 citation · 44 references
Computer Science

TL;DR

This work proposes an architectural pattern for text-to-SQL systems, a trusted kernel with a generative shell, resting on one invariant: a component that can fabricate may influence which question the system answers, never which value it returns.

Abstract

Large language models have made natural language interfaces to databases (NLIDB) newly credible, but LLM text-to-SQL systems fail in a way that matters for deployment: a hallucinated column or a mis-aggregated total yields a fluent wrong answer, indistinguishable at the point of use from a right one. Where the consumer cannot inspect the generated query, as in enterprise AI deployments and operational dashboards, and increasingly where the consumer is a tool-using agent rather than a person, accuracy alone is insufficient: nothing marks which answers to distrust. This is a reliability problem before it is an accuracy problem. We propose an architectural pattern for such systems, a trusted kernel with a generative shell, resting on one invariant: a component that can fabricate may influence which question the system answers, never which value it returns. A generative shell interprets underspecified input and phrases replies; a deterministic kernel matches fully specified questions against a bounded set of answerable question shapes and compiles them to queries by deterministic execution. The two meet at a confirmation the user reads before any value is computed, and requests the kernel cannot express are declined rather than approximated. We call this structural abstention, and distinguish it from the statistical abstention of selective prediction and calibrated confidence: refusal here needs no confidence estimate, because unanswerable requests are unrepresentable. We specify the pattern implementation-independently, give a five-decision recipe and work it across three domains, extend the invariant from returned values to the actions of agentic systems, and report a two-year production case study alongside two generative alternatives, a fine-tuned parser and a tool-retrieval agent. We close against enterprise and reliability benchmarks published since.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

Ingest-Time Fact Compilation for Cost-Efficient and Reliable Question Answering over Revised Corpora

Most agentic question answering (QA) systems do an important part of their semantic work at the worst possible time: every time someone asks a question. When a corpus contains revisions, drafts, revocations, deletions, and sources with different levels of authority, the model must reconstruct the governed current state...

Kyle Wild, Yusuke Takahashi, Asako Uraki · 0 citations
#small language model Book Open access Sep 2026

Qualification by Calibration: A Readable Benchmark for Admitting Language Models to Human-Computation Tasks

Evaluating a diverse panel of contemporary models under three administration protocols, it is shown that a naive harness, with a fixed token budget and an unaudited parser, manufactures failing workers out of competent ones, misreading a model that answers essentially every item correctly as badly inaccurate, answer-bi...

Carlos Toxtli, M. Delaflor · 0 citations

Two Surfaces of Ambiguity: Complementary Detection for Text-to-SQL

This work shows that introspection and sampling are complementary but not disjoint on BIRD-Interact-Lite, a Text-to-SQL benchmark with 300 tasks, annotated ambiguities, and an LLM-based user simulator, and proposes a proposed multi-label grammar that substantially narrows it without affecting downstream context.

Leonhard Liu, Patrick K. Erdelt, Two · 0 citations
Preprint Aug 2026

Walking on the DARKSIDE

The evidence supports an architectural claim: when an LLM forward pass is wrapped in an ontology-mediated auditing, structurally inevitable hallucination can be partially recovered.

A. Gangemi, E. Bottazzi · 0 citations
#artificial intelligence Preprint Sep 2026

The Stochastic Shift: A New Evaluation Paradigm for Text-to-SQL with AI Operators

SQL has been augmented with AI operators, enabling modern data analytics platforms to derive insights from both structured and unstructured data. We observe that while current Text-to-SQL systems can successfully generate these AI-augmented queries, reliably evaluating their correctness remains a critical open challeng...

Tarfah Alrashed, Fatma Ozcan, Per Jacobsson et al. · 0 citations
Preprint Aug 2026

RAG Deserves an Index: Why Ingest-Time Compilation Beats Query-Time Interpretation

Nearly every retrieval-augmented question-answering system in production ships with a hidden interpreter: on each query a language model re-derives the meaning of raw corpus text and then throws that work away. Cheaper models do not close the gap: per-token prices have fallen by orders of magnitude while inference spen...

Kyle Wild, Yusuke Takahashi, Asako Uraki · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.