A gradient-based head attribution strategy where the Token-level Max-Margin loss is backpropagated to the attention maps enables a large-scale causal analysis of attention heads, making it suitable for LLMs.
Paweł Mąka, Yusuf Can Semerci, Jan Scholtes et al.· 0 citations
An AI audit of popular chatbots using real commercial-advice queries shows that neither isolated responses nor API observations can be assumed to represent the commercial advice consumers encounter, and independent audits of AI-mediated commercial advice should account for repeated responses, consumer-facing conditions...
Lucas G. Uberti-Bona Marin, Thales Bertaglia, Giovanni Astante et al.· 0 citations
Automatic speech recognition (ASR) systems exhibit unequal error rates across speaker groups, motivating interventions on their internal representations. We ask whether speaker-linked attributes that are linearly readable from pretrained ASR encoders yield useful directions for reducing group word-error-rate (WER) gaps...
Nicolas Bourrel, Abderrahmane Issam, Gerasimos Spanakis· 0 citations
While English language models have been widely examined for anti-queer bias, Dutch models remain understudied. To address this gap, we developed a culturally and linguistically adapted Dutch dataset based on the English WinoQueer benchmark, containing pairs of stereotypical and counter-stereotypical sentences. To valid...
Jiska Beuk, Gerasimos Spanakis· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.