Skip to content

NL-PAC: Specification Ambiguity and Certified Minimax Risk Floors in LLM-Mediated Supervision

Jul 2026 · arXiv.org · Vol abs/2607.08961 · 0 citations · 65 references
Computer Science Mathematics

TL;DR

Natural Language PAC is introduced, a framework that uses a fixed model's thresholded decoding law to define admissible labels and candidate targets and extending it to human interpretations requires external validation.

Abstract

Large language models increasingly provide labels, evaluations, and feedback for tasks specified in natural language. When a specification admits multiple readings but the supervision channel does not reveal which is operative, additional labels reduce sampling error without resolving the resulting identification problem. We introduce Natural Language PAC (NL-PAC), a framework that uses a fixed model's thresholded decoding law to define admissible labels and candidate targets. The probability that multiple labels are admissible equals the diameter of the pointwise-admissible target class, and under target-blind supervision every learner incurs worst-case risk of at least half this diameter, at every sample size; the exact randomized minimax risk over this class is attained by a data-independent strategy. Finite-sample confidence bounds make these quantities certifiable from held-out unlabeled inputs. In a frozen Qwen~2.5--3B audit, one prespecified prompt yields a positive model-relative certificate, whereas a paraphrase and exact-rule controls yield zero. A held-out bridge audit finds that supplied candidate reading clauses fail the admissibility condition needed to transfer the certificate to coherent readings. The guarantee is specific to the audited model, prompt, threshold, and input distribution; extending it to human interpretations requires external validation.

View source

Similar papers

Preprint Aug 2026

Asymptotic Risk Calibration for Selective Question Answering

A-CRC-QA is a post-hoc calibration framework for uncertainty-aware selective question answering that reformulates selection-conditioned error control as a linear expectation constraint and applies a monotonized empirical-risk calibration procedure inspired by conformal risk control.

Shufan Lin, Sijin Dong · 0 citations
Book Open access Aug 2026

CALM-TS: Risk-Controlled LLM Labeling for Time-Series via Calibrated Selective Gating

CALM-TS is presented, a risk-controlled weak-supervision pipeline that bounds the labeling error rate while maximizing coverage, and yields auditable evidence chains and a coverage-based, label-free drift indicator.

Yi-Fan Xiao, Shi-Jie Li, Yu Huang · 0 citations
#natural language process... Preprint Sep 2026

Calibration as a First-Class Criterion in LLM Evaluation

Calibration of language models -- the alignment between expressed or implicit confidence and empirical correctness -- is a well-studied subfield within NLP. Methods to measure it already exist. The problem is adoption: outside this subfield, NLP research regularly introduces new models, datasets, and benchmarks without...

Mario Sanz-Guerrero, Katharina von der Wense · 0 citations
Preprint Sep 2026

When is Test-Time Adaptation Identifiable From Unlabeled Evidence?

Test-time adaptation (TTA) offers many ways to update a deployed model without labels, but choosing the wrong update can make a strong source model worse. Recent methods therefore try to predict which adaptation will work from unlabeled test data. We ask a prior question: does the evidence given to the selector contain...

Kartik Jhawar, Li-Po Wang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.