Metrics That Write Themselves: Evolving an Evaluator from Its Own Blind Spots
EvalCEGAR is a pool of small Python operators that each flag a candidate for one named defect, or abstain, and vote, and borrows counterexample-guided abstraction refinement from program verification to score agents against a reliable automatic metric.
Xing Zhang, Ya Cui, Guanghui Wang et al.
· 0 citations