Skip to content

Author

M.J.F. Valdez

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Open access 2026

Evaluating Retrieval-Augmented Generation on Social Bias Benchmarks across Small Language Models

: As small-scale, open-source Large Language Models (LLMs) proliferate for on-device and privacy-centric applications, understanding the trade-offs between their utility and behavioural reliability becomes critical. This study evaluates a suite of instruction-tuned LLMs, Gemma, Llama, and Qwen ( ≤4 B parameters), treating the model family, and scale as the primary units of analysis. Retrieval-Augmented Generation (RAG) is employed as a controlled experimental condition to assess utility gains on the Natural Questions (NQ) benchmark, while utilizing native, non-augmented configurations to establish a fairness baseline via the Bias Benchmark for QA (BBQ). The findings reveal that while RAG significantly enhances utility, often doubling Exact Match (EM) scores, these gains are non-uniform and architecture-dependent, with certain families exhibiting greater "retrieval-readiness" than others. Paradoxically, the fairness analysis shows that providing explicit context in disambiguated settings can increase stereotype engagement rather than suppressing it. These results suggest a fundamental disconnect between a model's capacity for factual accuracy and its ability to maintain social fairness, highlighting the need for multi-dimensional evaluation frameworks for small-scale systems.

M.J.F. Valdez, Arghir-Nicolae Moldovan · 0 citations