Natural Language PAC is introduced, a framework that uses a fixed model's thresholded decoding law to define admissible labels and candidate targets and extending it to human interpretations requires external validation.
Abstract
Large language models increasingly provide labels, evaluations, and feedback for tasks specified in natural language. When a specification admits multiple readings but the supervision channel does not reveal which is operative, additional labels reduce sampling error without resolving the resulting identification problem. We introduce Natural Language PAC (NL-PAC), a framework that uses a fixed model's thresholded decoding law to define admissible labels and candidate targets. The probability that multiple labels are admissible equals the diameter of the pointwise-admissible target class, and under target-blind supervision every learner incurs worst-case risk of at least half this diameter, at every sample size; the exact randomized minimax risk over this class is attained by a data-independent strategy. Finite-sample confidence bounds make these quantities certifiable from held-out unlabeled inputs. In a frozen Qwen~2.5--3B audit, one prespecified prompt yields a positive model-relative certificate, whereas a paraphrase and exact-rule controls yield zero. A held-out bridge audit finds that supplied candidate reading clauses fail the admissibility condition needed to transfer the certificate to coherent readings. The guarantee is specific to the audited model, prompt, threshold, and input distribution; extending it to human interpretations requires external validation.
A-CRC-QA is a post-hoc calibration framework for uncertainty-aware selective question answering that reformulates selection-conditioned error control as a linear expectation constraint and applies a monotonized empirical-risk calibration procedure inspired by conformal risk control.
CALM-TS is presented, a risk-controlled weak-supervision pipeline that bounds the labeling error rate while maximizing coverage, and yields auditable evidence chains and a coverage-based, label-free drift indicator.
Yi-Fan Xiao, Shi-Jie Li, Yu Huang· Proceedings of the 32nd ACM...· 0 citations
This work studies whether verbalized confidence can support risk-controlled deferral in small open-weight language models, evaluating eleven instruction-tuned models from three families on ARC-Challenge and TruthfulQA with 25,168 local predictions.
Calibration of language models -- the alignment between expressed or implicit confidence and empirical correctness -- is a well-studied subfield within NLP. Methods to measure it already exist. The problem is adoption: outside this subfield, NLP research regularly introduces new models, datasets, and benchmarks without...
Mario Sanz-Guerrero, Katharina von der Wense· 0 citations
Test-time adaptation (TTA) offers many ways to update a deployed model without labels, but choosing the wrong update can make a strong source model worse. Recent methods therefore try to predict which adaptation will work from unlabeled test data. We ask a prior question: does the evidence given to the selector contain...
An ensemble of LoRA adapters induces a credal set whose lower and upper probabilities expose the spread of plausible predictive distributions rather than collapsing to a single softmax output, and derives two complementary commitment scores.
S. K. Manchingal, S. Nikolenko, Fabio Cuzzolin· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.