Preprint
Jul 2026
Minionese: Comprehensive Benchmark and Mechanistic Study of Multilingual LLM Safety
It is shown that English-only safety evaluations are insufficient; they require accounting for script family, perturbation type, and per-language alignment coverage, and a geometric mechanistic analysis of refusal failure across language tiers.
Chigozirim Ifebi, Brent Kong, Ayushi Mehrotra
· 0 citations