Jul 2026
LLM-as-a-Judge for Evaluating System Responses in Conversational Music Recommendation
This paper presents the first user study to empirically assess the reliability of LLM-as-a-judge for evaluating CRS responses, and finds that LLM-based judges exhibit moderate positive alignment with human assessments and outperform all reference-based baselines.
Seungheon Doh, B. Sguerra, Sergio Oramas et al.
· arXiv.org · 0 citations