Skip to content

Measuring the Checker: Mutation Analysis for GPU-Kernel Benchmark Oracles

Sep 2026 · 0 citations · 29 references
Computer Science

TL;DR

Mutation analysis as an adequacy metric for kernel-benchmark oracles is introduced: deterministic rules inject 10 compilable faults into verified CUDA implementations of 188 KernelBench problems, 7{,}384 of them with an independent kill witness; any test protocol is scored by the fraction it detects.

Abstract

Benchmarks for LLM-generated GPU kernels decide correctness with a few random inputs and a loose floating-point tolerance, and their verdicts now feed leaderboards and reinforcement-learning rewards. Recent work agrees these checkers are weak and patches them by hand---extra input distributions, fuzzing recipes, tighter tolerances---with no way to \emph{measure} whether any patch suffices. We introduce mutation analysis as an adequacy metric for kernel-benchmark oracles: deterministic rules inject 10{,}303 compilable faults into verified CUDA implementations of 188 KernelBench problems, 7{,}384 of them with an independent kill witness; any test protocol is scored by the fraction it detects. The official check misses \textbf{one in six} witnessed faults (16.9%), deterministically, and the misses are skewed by family: 8.7% of arithmetic faults escape, but 78.6% of precision faults do. The metric explains why (a tolerance blind band growing with reduction size; a measured ceiling on input aggressiveness set by legitimate floating-point variance), audits the strongest existing patch (KernelBench-Verified's gain splits into $+4.0$ points from hidden inputs and $+4.5$ from tighter tolerance, a split its authors could not compute), and exposes a published fuzzing recipe that rejects \emph{correct} kernels 107 times. Optimizing suites over the kill matrix reaches 98.0% detection with two inputs per problem (94.8% held-out), and the measurement's fault taxonomy teaches a test generator more than the raw faults themselves. Across 48 whole architectures, the blindness grows with scale, concentrating in deep homogeneous pipelines, and two problems prove unrefereeable: their official references violate the benchmark's own tolerance against fp64. We release everything as \href{https://huggingface.co/datasets/Elfsong/KernelBench-M}{KernelBench-M}.

View source

Similar papers

Preprint Sep 2026

Evaluating Shaker for Flaky Test Detection in Python Projects

The first empirical evaluation of Shaker for Python is presented, comparing non-order-dependent flaky tests from the ground-truth dataset of Gruber et al., and finding that Shaker provides no statistically significant detection advantage over plain re-execution.

G. Leal, Denini Silva, Leopoldo Teixeira · 0 citations
Preprint Aug 2026

Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure

Design guidance for measurement under strategic optimization is distill design guidance for measurement under strategic optimization: held-out probes retain validity only on non-enumerable axes; gates must measure held-out performance, not just correctness; and a transfer rate is interpretable only with per-failure mec...

Víctor Gallego · 0 citations
#machine learning Preprint Aug 2026

Deterministic LLM Inference Across GPU Kernels: Power-of-Two INT8 Quantization Scales and the Limits of Tolerance-Based Conformance

It is found that every one of five epilogue faults -- scale precision, double rounding, multiplication order, output truncation, fused ordering -- moves the output by at most a single bfloat16 spacing, and by exactly one whenever it moves it at all, across 5,880 cells.

Teng-Ruei Chen · 1 citation
Preprint Sep 2026

Robustness-Aware Evaluation and Enhancement of Mutation-Based Fuzzing for Bug Discovery

Splitting, a black-box wrapper that copies a fuzzer's queue state after a bug trigger and continues from that state in multiple branches, directing more effort toward the discovered region, provides a practical way to measure and improve fuzzers.

Zi-Rui Liu, Meng-Fan Xu, Juan Zhai et al. · 0 citations
Preprint Sep 2026

Efficient Branch-and-Bound Testing and Verification of zkVMs

ZEBRA is a fully automated verification and bug-detection framework that reduces zkVM verification to a solution-set cardinality problem over a canonical trace space, where redundancies such as null-row padding and non-deterministic permutations are eliminated prior to counting.

Hideaki Takahashi, Suman Jana, Junfeng Yang · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.