Tracing Stereotypes from Representation to Output in Multilingual LLMs
To investigate internal mechanisms of linear probing, attribution patching, sparse autoencoders and feature ablation in Llama-3.1-8B, Qwen3-8B, and Gemma-2-9B, linear probing performance peaks substantially earlier than attribution in all three models.
Ariun-Erdene Tumurchuluun, Yusser Al Ghussin, Pinzhen Chen et al.
· 0 citations