This work contributes a cost-controlled characterization of when cheap-tier search substitutes for target-tier search, and where it fails.
Abstract
Evolutionary optimization of LLM prompts and agentic programs (e.g., GEPA) is dominated by fitness evaluation: scoring each candidate runs an answering LLM over a validation set, so the evaluator's price tier dictates total search cost. We restructure that search by decoupling the three roles an LLM plays, running the high-volume answering role on the cheapest tier, reserving a strong model for the rare reflection/variation operator, then exploiting upward cross-tier transfer to deploy the cheaply evolved prompt on a stronger target. We contribute a cost-controlled characterization of when cheap-tier search substitutes for target-tier search, and where it fails. Across four tasks (HotpotQA, IFBench, LiveBench-Math, HoVer) and eleven models in four model families, the resulting prompt matches or exceeds same-tier optimization while placing over 96% of search tokens on the cheapest tier, at 5.6-14x lower search cost, rising to 25-54x where reasoning tiers emit long chains of thought on every fitness call.
It is shown that model selection is a recommendation problem, and AgentRec, a multi-stage retrieval-and-ranking framework that progressively narrows hundreds of candidate models using increasingly expensive but more faithful evaluation signals, is introduced.
Chen Luo, Yu-Lin Liu, Yi Liu et al.· Proceedings of the 20th ACM...· 0 citations
Janus is introduced, a framework that uses LLMs to co-evolve target programs and executable proxy evaluators to address label scarcity and extend evaluator-guided LLM discovery from tasks with cheap, scalable feedback to scientific domains where trustworthy evaluation is scarce and expensive.
Xi-Meng Liu, Qianlong Wang, Ying-Ming Mao et al.· 0 citations
Budget-constrained agentic search arises when an LLM agent must refine candidates under a small evaluation budget, because validation is expensive, generation requires multiple model calls, or both. In this regime, standard MCTS allocates budget poorly: exploration bonuses dominate at low visit counts, unpromising sibl...
EcoAgent-Bench is introduced, in which every task specifies priced actions and an explicit budget, and results show that completion under a budget and economical action selection are distinct properties.
Jie Wu, Ming Gong, Feixiang Cheng et al.· 1 citation
CRM+RCCR, an architecture-agnostic cost-aware objective that encodes cost preference into continuous relevance targets through per-pair independent scoring, eliminating multi-positive dilution while regularizing queries with similar routing preferences to be closer in the routing space.
Tao Yu, Yi-Fei Qu, Zhi-Qing Cui et al.· 1 citation
Can an agent learn a numerical search strategy through executable practice and then transfer that strategy as text? We study low-budget black-box optimization, where unaided language models remain well below strong classical optimizers. During development, an agent repeatedly writes and evaluates optimizer programs. It...
Yi Wu, Zheng Ren, Zhi-Yu Hu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.