Conference
Open access
2026
Your Reasoning Model is Secretly a Reward Model - Optimization-Free Verification from Experience
This paper introduces C LUE (Clustering and Experience-based Verification) , a training-free, non-parametric verifier that improves selection and reranking in Large Language Model outputs and finds that correct and incorrect solutions exhibit measurable geometric differences in their hidden-state trajectories.
Zhenwen Liang, Ruosen Li, Yujun Zhou et al.
· Annual Meeting of the Associ... · 0 citations