Beyond Literal Meaning: How LLMs Interpret Yemeni Proverbs
Results show that instruction-tuned models like GPT-4o and Gemini 1.5 Pro outperform smaller models in both automatic and human evaluations, and LLM-as-a-Judge evaluation correlates strongly with human assessment.