When Agent Metrics Measure Different Things: An Evidence-Grounded Audit of the Praxa AI Pipeline
Agent evaluations can be numerically correct while measuring a different construct from the one implied by their labels. We present a retrospective measurement audit of selected Praxa AI implementation files, historical evaluation artifacts, and operational records. A 139-case offline routing report contains 112 passes...