Which LLM to Fine-Tune? Agent-Driven Model Selection at Scale
Abstract
Open-source model hubs now host over two million public AI models, yet teams building customer-facing AI systems must still determine which model to fine-tune for production deployment—a decision that shapes the quality, latency, and cost experienced by hundreds of millions of users. At Amazon, we spent over years of iterating on this process across multiple production use cases, where model selection remained manual, slow, and heavily biased toward a small set of familiar model families despite the rapidly expanding open-source ecosystem. We show that model selection is a recommendation problem, and introduce AgentRec, a multi-stage retrieval-and-ranking framework that progressively narrows hundreds of candidate models using increasingly expensive but more faithful evaluation signals. Across public benchmarks and Amazon production systems, AgentRec reduces model selection from multi scientist-weeks to couple unattended GPU-hours while matching or exceeding the quality of exhaustive manual exploration. Our results suggest that, for industrial teams deploying fine-tuned LLMs at scale, model selection can evolve from an ad-hoc bottleneck into a repeatable and continuously automated system for discovering high-quality models under real-world deployment constraints.