Large language models are increasingly used as judges to measure social bias in text, yet the passages they judge are often noisy, containing typos, informal spelling, and broken punctuation. The consequences of such surface noise for social bias measurement remain unclear. To investigate this question, we apply five realistic noise conditions at multiple intensity levels to 3,822 stereotype-related responses and compare the resulting bias judgments with those on the original text. We find that such surface noise does not degrade bias measurement symmetrically: it is far more likely to turn neutral judgments into biased ones than biased judgments into neutral ones, by up to a 120x margin. We further observe two non-obvious effects across four LLM judges: in the most fragile judge the distortion is at its purest at mild, realistic noise levels, where erasure is scarcest, and as judges grow robust it attenuates toward parity rather than reversing. Bias measured on noisy text is therefore systematically overestimated, most in the categories that matter most for fairness.
DongHyun Ryu, Jaehyeok Lee, Yeongjun Hwang et al.· 0 citations
Camellia is introduced, a benchmark for evaluating entity-centric cultural biases in nine Asian languages, spanning six Asian cultures, and it is found that LLMs can struggle with context understanding in some Asian languages, creating performance gaps between cultures in entity extraction.
Tarek Naous, Anagha Savit, Carlos Rafael Catalan et al.· arXiv.org· 2 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.