Backtest Lie Detector: Benchmarking Large Language Models as Point-in-Time Auditors for Financial Research Workflows
Abstract
Large language models (LLMs) are increasingly used to assist with financial research workflows, including strategy development, code review, and methodology auditing. However, financial backtests and event studies are susceptible to subtle point-in-time validity errors (such as ticker time-travel, filing clock leakage, accounting availability violations, and sur-vivorship bias) that can silently invalidate research conclusions. We introduce the Backtest Lie Detector, a 141-case benchmark for evaluating whether LLMs can identify and repair point-in-time workflow errors in financial research. The benchmark spans four modules covering identifier validity, filing timestamps, accounting data availability, and universe construction. We evaluate GPT-4o under generic and specialized prompts; GPT-4o-mini, Claude Sonnet 4.5, Claude Sonnet 4.6, and Claude Haiku 4.5 under the generic prompt; and rule-based baselines. On this benchmark we observe a safety-utility tradeoff: the GPT-4o generic prompt achieves the highest accuracy (83.0%) but produced a 4.7% false-valid rate, while the specialized prompt produced no false-valid approvals at lower accuracy. Claude Sonnet 4.6 combined high accuracy (81.6%) with a 0% false-valid rate and a much lower false-invalid rate (12.2%) than Sonnet 4.5. GPT-4o-mini reached a 0% false-valid rate only by rejecting nearly all valid workflows (97.6% false-invalid rate), showing that low false-valid rates must be interpreted jointly with false-invalid rates. The evaluated models act as useful first-pass auditors but require calibration, deterministic checks, and human oversight.