Deep and shallow biases in language models
A bias depth score is introduced that measures both how strongly a model prefers its top answer under direct prompting and whether that answer survives scenario reframing and shows that Deep biases are consistently harder to remove than Shallow biases.
A. Vo, V. Dang, Khai-Nguyen Nguyen et al.
· 0 citations