Skip to content
#small language model Book Open access

Qualification by Calibration: A Readable Benchmark for Admitting Language Models to Human-Computation Tasks

Sep 2026 · Proceedings of the 2026 ACM Conference on Human-AI Complementarity and Alignment · 0 citations · 15 references

TL;DR

Evaluating a diverse panel of contemporary models under three administration protocols, it is shown that a naive harness, with a fixed token budget and an unaudited parser, manufactures failing workers out of competent ones, misreading a model that answers essentially every item correctly as badly inaccurate, answer-biased, and overconfident.

Abstract

When several language models are wired together into a human-computation pipeline, the system routes, arbitrates, and escalates work according to how confident each model says it is, so a team builder needs a way to screen candidate models the way crowdsourcing has long screened human contributors: with a small, inspectable qualification test. We present such an instrument, a compact battery of yes-or-no questions across everyday domains on which a model reports an answer and a confidence, with every gold label independently audited. Our central contribution, however, is the audited harness and the interface properties it measures, not accuracy discrimination: qualification verdicts for model workers are only as valid as the harness that administers them. Evaluating a diverse panel of contemporary models under three administration protocols, we show that a naive harness, with a fixed token budget and an unaudited parser, manufactures failing workers out of competent ones, misreading a model that answers essentially every item correctly as badly inaccurate, answer-biased, and overconfident. Properly administered, every model in the panel proves admissible on accuracy for these common-knowledge items, and the differentiators that remain, and that transfer to disjoint downstream tasks, are interface properties: calibration of stated confidence, format discipline, and availability under budget. A controlled manipulation further shows the calibration axis is dissociable from accuracy. We release the items, the audited protocol, all per-trial responses under every protocol, and the scoring code, so the check, and the audit of the check, are each a single command.

Read PDF

Similar papers

#artificial intelligence Preprint Sep 2026

Evaluating and Benchmarking the System One Model Jev

Jev is a commercial System One model from TypeSafe AI that does not generate text: given a state and typed questions, it returns a choice from fixed options, a position on a rubric, or the probability that a statement is true, with probabilities the vendor describes as calibrated. Such models target small decisions in...

Tobias Deußer, L. Sparrenberg, R. Sifa · 0 citations
Preprint Aug 2026

What Proves You Wrong: Benchmarking Language Models on Falsifiable Research Ideation

The Lit2Test benchmark centers on a six-field contract organized around a falsifying outcome, so that every proposal precommits the observation that would prove it wrong, making its quality decidable in the first place rather than merely arguable.

Zi-Yue Wang, Aomufei Yuan, Yi-Ran Yao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

CruxBench: A Benchmark of Information Discovery

Benchmarks for large language models (LLMs) typically evaluate the accuracy of answers against fixed reference labels. But a central step in many complex real-world tasks is identifying which questions are worth asking in the first place: decomposing a difficult problem into subquestions -- which we call cruxes -- whos...

Hui Dai, Li-Na Piao, Nick Merrill et al. · 0 citations
#small language model Preprint Aug 2026

Grading Needs a Rubric, Not Intelligence

Small language models can grade open-ended examination answers as reliably as substantially more expensive models when they grade against an explicit rubric. We test this claim as the design principle behind any-to-bench: a frontier model reads source documents once, at ingestion, to extract each question and its rubri...

Jhen-Ke Lin · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 30, 2026

This game-playing AI is the new champ at Stratego

Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.