Skip to content
Preprint

Janus: An Algorithm-Evaluator Co-Evolution Framework for LLM-Driven Discovery under Expensive Evaluation Budgets

Aug 2026 · 0 citations · 47 references
Computer Science

TL;DR

Janus is introduced, a framework that uses LLMs to co-evolve target programs and executable proxy evaluators to address label scarcity and extend evaluator-guided LLM discovery from tasks with cheap, scalable feedback to scientific domains where trustworthy evaluation is scarce and expensive.

Abstract

LLM-driven program discovery relies on rapid evaluator feedback, but many scientific and engineering tasks require high-fidelity simulations, hardware execution, or physical experiments, making each evaluation expensive. Cheap surrogate evaluators can reduce this cost, yet fixed surrogates are vulnerable to search-induced distribution shift and are difficult to fit reliably from sparse, search-biased labels. We introduce Janus, a framework that uses LLMs to co-evolve target programs and executable proxy evaluators. To address label scarcity, Janus leverages domain knowledge encoded in LLMs to generate task-specific evaluator programs and calibrates them using real outcomes. To mitigate distribution shift, Janus evolves evaluators alongside target programs, selects them using a promotion-aligned objective, and maintains region-conditioned portfolios with online credit updates. Because proxy predictions remain fallible, Janus uses them only to prioritize candidates and requires real validation before candidates can enter the target-program population or update the incumbent. Across five scientific and engineering design tasks, Janus achieves a larger area under the best-so-far improvement curve over the real-evaluation budget and higher final performance than a matched baseline that evolves only target programs. On average, Janus reaches 99/% of the baseline's final improvement with 59.1/% fewer real evaluations. Evolved proxy evaluators also rank promising candidates more accurately than their seed versions. Together, these results extend evaluator-guided LLM discovery from tasks with cheap, scalable feedback to scientific domains where trustworthy evaluation is scarce and expensive.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation

PTA-IRT is proposed, a Privileged Trajectory-Aware Item Response Theory framework that fuses process and outcome signals and consistently outperforms prior IRT baselines on score and ranking recovery across four SWE benchmarks.

Kefeng Duan, De-Wu Zheng, Yanlin Wang et al. · 0 citations
Preprint Sep 2026

AlphaOpsBench: Benchmarking End-to-End Alpha Strategy Operationalization in Prediction Markets

Large language models increasingly generate quantitative trading strategies, yet existing benchmarks assume standardized assets, numerical features, or directly compilable strategy representations---assumptions that prediction-market strategies violate, since a coarse idea may leave the traded outcome, causal informati...

Huai-Yu Jia, Ming-Xuan Zhao, Jin-Cheng Gao et al. · 0 citations
Preprint Aug 2026

Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection

We present a novel approach to efficient LLM harness optimization through adaptive validation task selection. Harness optimization iteratively rewrites the harness code based on validation performance, enabling substantial performance gains without updating the underlying model weights. Existing approaches, however, ev...

Atsuyuki Miyai, Kiyoharu Aizawa, Toshihiko Yamasaki · 2 citations
Preprint Aug 2026

Loreley: Repository-Scale Program Evolution with Quality-Diversity Search

This work compares configured Loreley QD, sequential champion editing, and independent root proposals in a matched Zstandard experiment and finds that Sequential had the highest observed 48-job mean and median and established a QD advantage.

Mo Chen · 0 citations
#natural language process... Preprint Sep 2026

OptiArena: Can LLMs Improve Executable Algorithms under Fixed Resource Budgets?

Static QA and code-generation benchmarks only partially capture the role that large language models (LLMs) now play as coding agents and research tools. We introduce OptiArena, a budget-controlled testbed for studying whether LLMs can improve executable game-playing algorithms through five rounds of code edits within a...

Wen-Jun Peng, Xin-Yu Wang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.