Discovering and Repairing Blind Spots in LLM-as-a-Judge Evaluation of Multi-Turn AI Systems
The widespread use of Large Language Models (LLMs) as automated evaluators for multi-turn conversational AI systems is due to their scalability and adaptability. Nonetheless, LLM-as-a-judge systems often have systemic blind spots, leading to the neglect of certain flaws because to insufficient, inflexible, or biased as...