Aug 2026· Innovations in Systems and Software Engineering· Vol 22· 0 citations· 20 references
Computer Science
TL;DR
It is observed that bias vulnerability increases in Hindi and Bengali compared to English, particularly under manipulative prompts, which highlights the importance of multilingual bias evaluation and provides practical guidance for selecting commercial language models in bias-sensitive applications.
The results show that LLMs systematically adapt their responses to align with prompt framing, even in factual contexts, which suggests that prompt framing can outweigh factual consistency in model responses.
Mudar Adas, Polina Tsvilodub, Michael Franke et al.· 0 citations
It is found that answer format does substantially alter measured outcomes, including reversals in order rankings, and the importance of treating answer format as a substantive component of LLM evaluation and motivate multi-format designs for more robust model assessment is highlighted.
K. Merzlyakova, Sebastian Padó, Franziska Weeber· 1 citation
Benevolence bias is identified and measure, a small but consistent tendency for aligned LLMs to lean toward the kinder, safer, more socially approved answer on value-laden survey questions, and is easy to diagnose and straightforward to fix.
Yuanzi Li, Jun-Hao Wang, Minghui Liu et al.· 0 citations
A multilingual German-English benchmark dataset that combines community-sourced stereotypes from German-speaking queer individuals with a German translation of WinoQueer is introduced, showing that language models reproduce anti-queer stereotypes, with variation across identities and models.
Large language models are increasingly used to read résumés and judge who advances in hiring, a task once reserved for people and now handed to systems whose reasoning is hard to inspect. Whether these models carry the demographic biases that have long shaped human hiring is therefore an urgent question, and the publis...
While English language models have been widely examined for anti-queer bias, Dutch models remain understudied. To address this gap, we developed a culturally and linguistically adapted Dutch dataset based on the English WinoQueer benchmark, containing pairs of stereotypical and counter-stereotypical sentences. To valid...
Jiska Beuk, Gerasimos Spanakis· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.