It is found that answer format does substantially alter measured outcomes, including reversals in order rankings, and the importance of treating answer format as a substantive component of LLM evaluation and motivate multi-format designs for more robust model assessment is highlighted.
Abstract
Gender bias or other social biases in large language models (LLMs) are frequently evaluated with question answering or survey benchmarks where the LLM needs to give a response in a predefined answer format. It is well known in survey science that the answer format has a substantial impact on answers, just as LLMs are sensitive to the prompt wording. However, to our knowledge it has not been studied yet how changes in answer format impact the measurement of gender bias in LLMs and their alignment with human response distributions. We evaluate three instruction-tuned models on the BBQ benchmark and OpinionQA survey data across closed-ended, Likert-scaled and open-ended formats, comparing bias measurement and distributional alignment under otherwise identical conditions. We find that answer format does substantially alter measured outcomes, including reversals in order rankings. These differences arise because each format elicits distinct response behaviours, such as forced-choice selection, scale-based distributions and refusal in free-text generation. Our findings highlight the importance of treating answer format as a substantive component of LLM evaluation and motivate multi-format designs for more robust model assessment.
It is observed that bias vulnerability increases in Hindi and Bengali compared to English, particularly under manipulative prompts, which highlights the importance of multilingual bias evaluation and provides practical guidance for selecting commercial language models in bias-sensitive applications.
Koushik Deb, Imon Mukherjee, Debarshi Kumar Sanyal· Innovations in Systems and S...· 0 citations
The results show that LLMs systematically adapt their responses to align with prompt framing, even in factual contexts, which suggests that prompt framing can outweigh factual consistency in model responses.
Mudar Adas, Polina Tsvilodub, Michael Franke et al.· 0 citations
Benevolence bias is identified and measure, a small but consistent tendency for aligned LLMs to lean toward the kinder, safer, more socially approved answer on value-laden survey questions, and is easy to diagnose and straightforward to fix.
Yuanzi Li, Jun-Hao Wang, Minghui Liu et al.· 0 citations
Professional communication is increasingly mediated by LLMs - but do these models serve all users equally? We show that when prompts contain linguistic features more commonly used by women (hedges, tag questions, collective reference), they systematically elicit shorter, less sophisticated, and less formal responses ac...
It is demonstrated that gender bias undermines both reliability and fairness in LLM-based fake news detection, highlighting the need for bias-aware evaluation and mitigation strategies.
Razieh Chalehchaleh, R. Farahbakhsh, N. Crespi· 0 citations
Large language models (LLMs) are increasingly used to generate media, but whether their content perpetuates gender stereotypes is unknown: standard benchmarks rely on selection-based formats rather than long-form generation, and surveyed baselines for local gender associations are scarce outside the West. We collect ge...
Sharif Kazemi, Tanya Popli, Neil K. R. Sehgal et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.