Conference
Open access
2026
AutoTaskEval: Towards Domain-Specific and Fine-Grained Evaluation for LLMs
An automated framework that constructs domain-specific benchmarks directly from unstructured corpora and systematically discovers tasks, enriches contextual grounding via iterative Socratic prompting, and generates diverse, progressively challenging evaluation instances that preserve established model-level evaluation trends are proposed.
Qingqing Lyu, Linjuan Wu, Yongliang Shen et al.
· Annual Meeting of the Associ... · 0 citations