Skip to content

Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline

Sep 2026 · 0 citations · 34 references
Computer Science

TL;DR

Applied out-of-domain, the same audit recipe flags 25.5% of BIRD-financial's expert-authored gold SQLs as candidate gold-SQL issues under the authors' annotation protocol under their annotation protocol.

Abstract

Production text-to-SQL pipelines often end with an LLM-as-judge whose agreement with human annotators has never actually been measured. When we checked ours, the deployed gpt-4o-mini judge agreed with two-author gold at only Cohen's kappa = 0.04 on a disagreement-enriched set and 0.42 on a uniform-random spot-check, over-flagging 77.1% of the human-FAITHFUL cases in the enriched set. Most of its over-flags trace back to a single mechanism we call GRADE-HALLUCINATION. A self-hosted Qwen3.6-27B replacement (kappa = 0.72) lands in the same range as Claude Opus 4.7 (kappa = 0.71); the head-to-head is underpowered at n = 96, but for the deployment decision that hardly matters, since Qwen costs roughly 1/300 as much per call. Ensembling does not help for free. Pairing the weak judge with a stronger one degrades agreement, whereas three strong judges under unanimity routing reach kappa = 0.79 at 89.7% auto-coverage. Applied out-of-domain, the same audit recipe flags 25.5% of BIRD-financial's expert-authored gold SQLs as candidate gold-SQL issues under our annotation protocol. Code and pre-registration are at https://github.com/JamesL404/synca-audit.

View source

Similar papers

#machine learning Preprint Sep 2026

Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation

This work investigates LLM-based evaluators of natural language generation quality mechanistically through an eight-attack perturbation taxonomy across the Readability and Adequacy dimensions of NLG quality, a generation pipeline that produces paired clean and corrupt summaries with controlled error intensity and expli...

Himil Vasava, Ming-Zhou Jiang · 0 citations

QuanReview: Offline, Auditable Reconciliation of Human and LLM Span Annotations

Structured span annotations, such as quantities with their units, uncertainty modifiers, and event classes, are expensive to create and hard to keep trustworthy once language models enter the loop. We present QuanReview, an open-source system for auditing and correcting such annotation layers. QuanReview aligns two ann...

Matteo Musacchio, Juan Cruz Giner Pulero, Isabel Castañeda et al. · 0 citations
Preprint Aug 2026

Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation Independence

It is shown that prior scores, even when included only as context metadata, anchor judgments and systematically shift ratings toward their values, and effective mitigation must be validated for the intended model and task or domain.

A. Kapetanović, Kemal Altwlkany, Andro Merćep et al. · 0 citations
Review

Judging the LLM Judges: A Human-Centric Validation of LLM-Generated Training Data for Software Retrieval

A lightweight quality-assessment protocol is presented for LLM-generated synthetic training data and applied to 13,579 synthetic user reviews generated from GitHub issues across four open-source Android applications, high-lighting the need for hybrid human-AI verification when synthetic data is used in security-critica...

Ogtay Hasanov, Saad Ezzini · 0 citations
#machine learning Preprint Oct 2026

SALUS: Automated Auditing of NL-to-SQL Benchmarks through Weak Supervision of Multi-Agent Output

Natural language to SQL (NL-to-SQL) benchmarks are foundational to progress in data analysis research, yet recent work has shown that widely-used benchmarks contain significant annotation errors. These errors silently corrupt evaluation metrics, penalize correct model output, and distort the field's understanding of st...

Shi-Yuan Zhou, Ashwin Gerard Colaco, Sainyam Galhotra et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.