Skip to content

Developing and Benchmarking Verification Algorithms to Improve Text-to-SQL Generation

Jul 2026 · Proceedings of the VLDB Endowment · Vol 19, pp. 3732-3744 · 0 citations · 46 references

TL;DR

This work formalizes Text-to-SQL verification as a standalone task: assessing whether a candidate SQL query semantically satisfies a natural language request, without access to ground-truth labels, and proposes and compares two modular verification strategies.

Abstract

Large Language Models (LLMs) have achieved impressive performance in translating natural language queries into executable SQL. However, these systems remain prone to deceptive failures : generating syntactically valid SQL that executes but fails to capture the user's intent. In this work, we argue that further progress can be made by focusing specifically on verification. In this paper, we formalize Text-to-SQL verification as a standalone task: assessing whether a candidate SQL query semantically satisfies a natural language request, without access to ground-truth labels. We propose and compare two modular verification strategies: Round-Trip Critique , which reverse-translates SQL into natural language to detect semantic drift, and Synthetic Execution Consistency , which uses unit-test-like synthetic inputs to ground verification in execution results. Our evaluation shows that these methods provide a robust signal for identifying incorrect queries, successfully flagging 64% of errors from a state-of-the-art generator and outperforming standard error detectors. We demonstrate two critical applications: (1) auditing foundational benchmarks (Spider, BIRD and KaggleDBQA), revealing that over two-thirds of "generator failures" actually stem from flawed benchmark labels, and (2) enabling Selective Generation, where a system uses verification signals to abstain from answering when confidence is low. Our results show that this paradigm significantly improves the quality of deployed data interfaces by transforming silent failures into explicit abstentions that alert the user to the failure and give them a chance to correct it.

View source

Similar papers

Preprint Aug 2026

TraceSQL: Traceable Answerability Estimation for Reference-Free Text-to-SQL Verification

Results show that SQL verification can be performed with a lightweight learned model while retaining feature-level evidence for inspecting and diagnosing its predictions, and feature attribution shows that the model relies on both semantic grounding and deterministic SQL-structure signals.

N. Shukla, Debasmita Panda, Srutanik Bhaduri et al. · 0 citations
Book Open access Oct 2026

A Transformation-Based Benchmark for Evaluating the Robustness of LLMs in Generating OCL

Large Language Models (LLMs) have shown promising performance in generating Object Constraint Language (OCL) constraints from natural language specifications. However, existing evaluations rely on publicly available UML models, which may overestimate generalization due to potential data leakage and reliance on recurrin...

Hamza Attarwala, Moataz Chouchen, Omar Alam et al. · 0 citations
#artificial intelligence Preprint Sep 2026

The Stochastic Shift: A New Evaluation Paradigm for Text-to-SQL with AI Operators

SQL has been augmented with AI operators, enabling modern data analytics platforms to derive insights from both structured and unstructured data. We observe that while current Text-to-SQL systems can successfully generate these AI-augmented queries, reliably evaluating their correctness remains a critical open challeng...

Tarfah Alrashed, Fatma Ozcan, Per Jacobsson et al. · 0 citations
#artificial intelligence Preprint Sep 2026

DatalogBench: Evaluating Large Language Models on Text-to-Datalog Synthesis

Datalog underpins reasoning tasks such as program analysis, but its programs are hard to write. Existing synthesizers automate this task but require users to state their intent as input-output examples. Large language models (LLMs) suggest a more natural route, text-to-Datalog synthesis from a natural-language question...

Yuan Li, Han-Yun Jiang, Guo-Wei Tian et al. · 0 citations
Preprint Aug 2026

Evaluating Language Models on Cross-Language Code Functional Equivalence

This work investigates whether LLMs can accurately judge functional equivalence across different programming languages in human-written code, a setting that requires deeper reasoning beyond superficial similarity, and identifies a difficulty-dependent breakdown in equivalence judgment.

Hui Sun, Anderson G. Uchôa, Rohit Gheyi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.