Skip to content

NormasTCU — A Brazilian Portuguese IR dataset and an evaluation of LLM-as-a-judge for relevance assessment

Aug 2026 · Information Processing & Management · 0 citations · 44 references
Computer Science

TL;DR

The results suggest that LLMs could effectively support scalable relevance assessment in specialized Portuguese corpora when evaluated using nDCG or MRR (rank-aware metrics), but they should be avoided when relying on precision or recall.

Abstract

Portuguese Information Retrieval (IR) lacks public datasets, and relevance assessment for specialized collections remains costly. While Large Language Models (LLMs) increasingly support relevance assessment, their reliability in non-English specialized domains remains unclear. We introduce NormasTCU (https://huggingface.co/datasets/LeandroRibeiro/NormasTCU), a Brazilian Portuguese IR dataset with 14,469 legal documents, 46 queries, and 3,048 human judgments over 812 query-document pairs. Using NormasTCU, we evaluated LLM-as-a-judge for relevance assessment by prompting three models with two prompt techniques to grade these pairs. We then compared the rankings of 15 IR systems derived from LLM-generated and human reference qrels. LLMs consistently showed a positive scoring bias (mean absolute error: 0.46--0.66 on a 0-2 scale). Furthermore, pair-level agreement with human judgments achieved only fair to moderate levels, with Cohen's kappa ranging from 0.32 to 0.53. Despite this bias, LLM-generated judgments often yielded highly similar system rankings for nDCG@10 and MRR (observed Kendall's tau greater than or equal 0.90, although the bootstrap confidence intervals did not always remain above this threshold), but were less reliable for P@10 and R@10. Notably, LLM-based rankings were sometimes more strongly correlated with the reference ranking than individual human annotations were. As a practical implication, our results suggest that LLMs could effectively support scalable relevance assessment in specialized Portuguese corpora when evaluated using nDCG or MRR (rank-aware metrics), but they should be avoided when relying on precision or recall.

Read PDF

Similar papers

Jul 2026

Benchmarking and AI-assisted human-like evaluation of retrieval-augmented generation for Arabic and English documents

An important practical reproducible framework for multilingual RAG benchmarking and insights for optimizing performance on resource-constrained devices are contributed and support the latent language hypothesis by suggesting an internal model bias toward high-resource languages.

B. J. Mohd, Khalil M. Ahmad Yousef, Salah G. Abu Ghalyon · 0 citations
#natural language process... Preprint Sep 2026

When Consistency Does Not Mean Reliability: Evaluating Local LLM Judges Against Human Ratings

The results show that high self-consistency does not necessarily indicate high agreement with human judgments when using local LLMs as automatic judges, and highlight the need to evaluate both consistency and human alignment when using local LLMs as automatic judges.

Aakash Kumar Tiwari · 1 citation
#natural language process... Preprint Sep 2026

V\={a}kQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering

V\={a}kQA, a Telugu SQA benchmark of 2,001 factoid question-answer pairs across six domains, is introduced and it is observed that Telugu phrasing retains cultural specificity that is lost in translation, speech input introduces phonetic confusions that alter question meaning, and cascaded ASR-MT errors compound progre...

Bhavana Akkiraju, Ravi Sastry Kolluru, D. SriCharan et al. · 0 citations

Optimising Retrieval for Linguistic Question-Answering in European Portuguese: A Benchmark on Ciberdúvidas Da Língua Portuguesa

A novel benchmark framework for linguistic QA retrieval, empirical evidence supporting monolingual IR-specialised models, and insights into retrieval robustness under paraphrastic variation are included, enabling improved QA systems for specialised and low-resource environments.

Pedro Moura, Inês Gama, F. Batista et al. · 0 citations
Conference Jul 2026

Cross-Lingual Information Retrieval for Indonesian–Javanese Documents Using LLM-Based Reranking

Cross-lingual information retrieval (CLIR) for low-resource regional languages remains challenging due to limited annotated training data, morphological complexity, and linguistic diversity. This paper systematically investigates Indonesian–Javanese CLIR by integrating traditional sparse lexical retrieval, multilingual...

Raden Mohamad Adrian Ramadhan Hendar Wibawa, Ika Alfina, Evi Yulianti · 0 citations
Preprint Aug 2026

L3Cube-IndicQuest v2: A Large-Scale Multilingual Benchmark for Evaluating Factual Knowledge of Large Language Models Across Indic Languages

We present L3Cube-IndicQuest v2, a large-scale gold-standard multilingual question-answering benchmark for evaluating the India-specific factual knowledge of Large Language Models (LLMs). The benchmark comprises 3,471 curriculum-grounded English question--answer pairs spanning nine domains, curated from educational cur...

Rinit Jain, Tirthraj Mahajan, Advait Joshi et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.