Skip to content

Beyond Literal Meaning: How LLMs Interpret Yemeni Proverbs

2026 · International Conference on Language Resources and Evaluation · pp. 1071-1080 · 0 citations · 37 references
Computer Science

TL;DR

Results show that instruction-tuned models like GPT-4o and Gemini 1.5 Pro outperform smaller models in both automatic and human evaluations, and LLM-as-a-Judge evaluation correlates strongly with human assessment.

View source

Similar papers

Open access 2026

Semantic Parsing for Evaluating Large Language Models: Separating Linguistic Abilities with YARN

A layer-wise analysis indicates that surface-level features such as temporality and negation are captured more reliably than deeper semantic phenomena like quantification in large language models, highlighting the limited capacity of current LLMs to generate fully formal meaning representations.

Rémi De Vergnette, Maxime Amblard · 0 citations
#machine learning Preprint Sep 2026

Evaluation of Contextual Understanding in Large Language Models

Large Language Models (LLMs) demonstrate impressive performance across diverse NLP tasks, yet their ability to exhibit genuine contextual understanding remains uncertain. Traditional evaluation metrics such as perplexity, BiLingual Evaluation Understudy (BLEU), or surface-level accuracy fail to reveal how well LLMs ext...

Subavarshana Arumugam, Mamta Nallaretnam, K. Wickramasinghe et al. · 0 citations
Open access 2026

Of Words and Meaning: A Grammatical and Semantic Benchmark for Faroese LLM Understanding

Evaluating language technology for low-resource languages faces a fundamental challenge: the scarcity of native benchmarks suitable for systematic assessment. For Faroese, no such evaluation frameworks exist. We address this gap by presenting the first benchmark suite for Faroese semantic understanding and grammatical...

Iben Nyholm Debess, Barbara Scalvini, Bolette S. Pedersen · 0 citations
Preprint Aug 2026

Interpretable Cross-Lingual Alignment in Small Language Models: Probing Cultural and Pragmatic Reasoning in Japanese-English Bilingual LLMs

J-PragEval-v0 is introduced, a minimal-pair benchmark isolating four such phenomena from surface fluency, and Pragmatic Representation Steering is specified, a parameter-free inference-time method that edits residual-stream activations along the class-mean-difference directions probing identifies.

F. Braun · 1 citation
#artificial intelligence Preprint Sep 2026

Multilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu

Multilingual large language models (LLMs) are increasingly used for open-ended text generation, yet their behaviour in low-resource languages remains poorly understood. In this work, we question how correct and reliable is the generation of multilingual LLMs when used for the task of story generation. We consider Urdu...

Farah Adeeba, A. Khan, Rajesh Bhatt et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.