Evaluating the Evaluator: Summarization Metrics and LLM-Judges beyond English
This paper generates a new multilingual summary meta-evaluation dataset (BASSE) by generating a new multilingual summary meta-evaluation dataset (BASSE), which comprises human judgments on 2,040 abstractive summaries generated either manually or by five Large Language Models (LLMs) with four different prompts.