Skip to content
Preprint

Conditional Evaluation of Language Models with Cheap Auxiliary Signals

Aug 2026 · 0 citations · 49 references
Computer Science Mathematics

TL;DR

This work empirically evaluates the primary performance-profile estimator on MATH-500, ScienceQA, MMLU, WinoGrande, HellaSwag, TruthfulQA, GSM8K, and ARC and proves calibration-free identification, unbiasedness for grouped profiles, local oracle optimality within centered linear augmentations, and first-order adaptivity to the estimated coefficient.

Abstract

Aggregate accuracy hides where models succeed and fail. Estimating conditional performance profiles from gold labels alone is expensive, while cheap auxiliary signals such as LLM-judge scores, pairwise comparisons, confidence scores, and judge-disagreement features can be collected for every benchmark item but are often biased or miscalibrated. We propose LACE (Local Augmented Control-Variate Evaluation), a semi-supervised estimator for conditional LLM evaluation. The key step is local centering: after subtracting the conditional mean of a cheap signal within the target profile region, any linear augmentation has zero conditional mean and therefore cannot change the estimand. The augmentation coefficient is used only for efficiency, and a local ridge control variate combines a gold-label residual mean from the labeled subset with a cheap-signal mean from the full item pool. We prove calibration-free identification, unbiasedness for grouped profiles, local oracle optimality within centered linear augmentations, and first-order adaptivity to the estimated coefficient. The resulting gain formula is governed by a population local $R^2$, which characterizes how the efficiency attainable from the cheap signals varies across profile values. We also derive corresponding estimators for direct paired model gaps and deployment-weighted scores. We empirically evaluate the primary performance-profile estimator on MATH-500, ScienceQA, MMLU, WinoGrande, HellaSwag, TruthfulQA, GSM8K, and ARC.

View source

Similar papers

#machine learning Review Sep 2026

CORDIAL: Calibrating Ordinal LLM Outputs from Few Labels

A large language model (LLM) can turn a text into a distribution over an ordered scale, but that distribution is a noisy measurement: saturated, compressed or exaggerated, and biased in a consistent direction. We propose CORDIAL, which treats the model's output as a noisy reading of the true label and corrects it with...

Xiang-Wei Wang, Peng Wang, Saman K. Halgamuge · 0 citations
Preprint Aug 2026

No Universal Signal Predicts Sample-Level LLM Regression under Version Updates

It is found that signal effectiveness is task-dependent: confidence is strongest on MCQ and simpler math, while likelihood/KL signals give the most frequent gains on harder math and code; no signal is universally best across model updates either, and some cross-version signals stay informative even when confidence fail...

Jiang-li Sheng, Yiwei Lu · 0 citations
Preprint Aug 2026

Accuracy and Order Sensitivity Diverge Under Label-Free Strategies

This paper test whether preventing a model from seeing option labels while committing to an answer removes positional influence and, in turn, improves performance, and evaluates two different strategies for mitigating bias.

Karleen Hanna, Feng Chen · 1 citation
Preprint Aug 2026

Blind to the Pivotal Vote: Aggregate Independence Metrics Miss Where Verification Actually Helps

LLM judge panels are a standard evaluation tool, but prior work reports highly correlated panel errors: nine judges provide roughly the effective information of two independent ones, and aggregation closes only a small fraction of the gap. A natural remedy--a signal from a different evidence source, e.g., executing a t...

Shu Yang · 0 citations
#artificial intelligence Preprint Sep 2026

Competence-Gated Pooling of Language Models and Priors for Event Forecasting

A competence gate is introduced that estimates domain-level source weights from resolved outcomes, shrinks uncertain estimates toward a global weight, and recalibrates the pooled forecast and provides a practical approach for selective model use based on measured marginal value.

Aditi Tiwari, Aashrith Bandaru, Heng Ji · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.