SkillRet: A Large-Scale Benchmark for Skill Retrieval in LLM Agents
Task-specific fine-tuning on SkillRet improves NDCG@10 by 12.9 points over the strongest prior retriever and by 16.2 points over the strongest off-the-shelf retriever, establishing SkillRet as a strong benchmark and foundation for future research on retrieval in large-scale agent systems.