When Fixed-Candidate Offline Evaluation Changes Model Selection in Two-Stage Recommenders
Abstract
Two-stage recommenders first retrieve a small candidate set from a large catalog and then apply a more detailed model to rank only those items. A common offline test instead gives every system the same candidates so that their scoring rules can be compared fairly, but this fixed-candidate protocol may select a different winner when the complete pipelines use different retrievers. We contrast it with own-candidate evaluation, which scores each pipeline on the items it retrieves, and decompose end-to-end utility into candidate survival and conditional ranking quality. Across ten benchmark instances from eight public datasets, the protocols select different top pipelines in eight cases; on OTTO, Dressipi, and Diginetica, disagreement persists under both tested K = 20 anchors through strict reversal or near-suppression. These controlled experiments establish the existence and mechanism of the mismatch—not its prevalence in production—and motivate reporting own-candidate utility whenever complete pipelines are compared.