Skip to content
Book Open access

Trustworthy A/B Patterns and the Winner's Curse: Lessons from Eight Large-Scale Replications

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 7533-7544 · 0 citations · 62 references

TL;DR

Trustworthy A/B Patterns is a community replication effort to evaluate selected patterns at high statistical power, finding that even at this scale, there was not enough power for key business metrics, such as revenue per user and purchase conversion rate, within practical time horizons.

Abstract

Claims of large lifts in A/B tests are widespread, yet many are only supported by small online experiments that are likely underpowered. Trustworthy A/B Patterns is a community replication effort to evaluate selected patterns at high statistical power. We report on eight A/B tests across four patterns (rounded buttons, page performance, coupon-code field, and sticky call-to-action), with a median of 2.4M users per experiment and 80% power at our pre-selected minimum detectable effects (MDEs) of 0.3% to 2.2%. We find that: (1) even at this scale, we did not have enough power for key business metrics, such as revenue per user and purchase conversion rate, within practical time horizons, and had to resort to surrogate metrics, such as click-through rate, add-to-cart rate, and capped add-to-carts (count); (2) previously reported effects for these patterns are highly exaggerated: across all eight replications, estimated effects were substantially smaller than previously claimed; only two showed statistically significant effects in the expected direction at α=0.05, and one was statistically significant in the opposite direction. While it is possible that some patterns have larger effects in certain conditions, we believe it is more likely that many of the prior estimates came from underpowered experiments (power below 50%), which exaggerate treatment effects. We conclude with lessons from running the community project for over one and a half years.

Read PDF

Similar papers

#artificial intelligence Preprint Sep 2026

Who Judges Matters: Measuring Family-Conditioned Preference in LLM-as-Judge Panels

Who the judge is can affect an LLM-as-judge result, but measuring that effect without confusing it with candidate quality is difficult. We study four open-weight families (Llama 3.1, Qwen 2.5, Gemma 2, and Yi 1.5) in a fully crossed pairwise design with 9,312 judgments. A common per-family statistic is strongly confoun...

David Ababio Awuni, Luke Achenie, Benjamin Tei Partey et al. · 1 citation
#machine learning Review Aug 2026

The Exceedance Design Effect: Effective Sample Size for Thresholds under Clustering

Suppose we want a cutoff that 90% of a population falls below. We estimate it from a sample, and another sample would give a different cutoff and a different fraction below it. We ask how much that fraction varies when observations come in independent groups, such as pupils in classrooms or sentences in news articles....

A. Noonan · 0 citations
Preprint Aug 2026

When Can You Trust Offline Evaluation of Equal-Cost Top-k Allocation? A Controlled, Reproducible Benchmark and Practitioner's Guide

This work benchmarks six estimators across five datasets and two known-effect sweeps, and validate the mechanisms against a non-simulated paired reference, finding that weak overlap is governed by logger-target action alignment, not by logging sharpness alone.

Bin-Shuang Li · 0 citations
#machine learning Preprint Sep 2026

GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents

Comparing and selecting task-oriented LLM agents increasingly relies on a low-cost offline evaluation gate: persona-driven LLM user-simulators converse with each candidate, an LLM-as-a-judge scores the transcripts, and the higher-scoring agent is promoted. We introduce GAUGE, a reusable offline protocol that measures w...

Umesh Bodhwani, Thanh Tran, Kai-Lin Wei · 3 citations
Aug 2026

Stars Are Where You Draw the Line: Reassessing Methods for Identifying Star Entrepreneurs

Entrepreneurship outcomes tend to be right-skewed and heavy-tailed, with a small fraction of “star” firms often accounting for a disproportionate share of value creation. Yet, how to define the “star” entrepreneurs driving these outcomes remains highly contested. In this paper, we argue that star identification is be...

Boris Nikolaev · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.