Sep 2026· International Journal for Research in Applied Science and Engineering Technology· 0 citations
TL;DR
This survey provides the first comprehensive review of research on code-switched LLM safety and robustness evaluation, and aims to guide researchers and practitioners in developing more robust evaluation frameworks and alignment techniques for real-world multilingual deployments.
Abstract
Large language models (LLMs) have demonstrated remarkable capabilities across languages, yet their safety
alignment remains predominantly evaluated in monolingual, especially English, settings. Code-switching (also known as codemixing)—the alternation between two or more languages within a single utterance or conversation—is a pervasive phenomenon
in multilingual societies and digital communication. Recent evidence reveals that code-switched inputs can substantially degrade
the safety robustness of LLMs, enabling jailbreaks and harmful outputs that monolingual safety mechanisms fail to intercept.
This survey provides the first comprehensive review of research on code-switched LLM safety and robustness evaluation. We
systematize existing attack methodologies (including Code-Switching Red-Teaming, Multilingual Blending, and attributional
analyses), evaluation datasets and metrics, empirical findings on attack success rates and linguistic factors, and emerging
mitigation strategies. We further situate these works within the broader multilingual safety literature, highlight critical gaps in
linguistic coverage, cultural contextualization, and mechanistic understanding, and outline a research agenda toward equitable,
linguistically inclusive LLM safety. Our synthesis aims to guide researchers and practitioners in developing more robust
evaluation frameworks and alignment techniques for real-world multilingual deployments.
This work presents a systematic mechanistic analysis of multilingual safety using sparse autoencoder features, sparse interpretable directions in the residual stream associated with harmful and harmless model behavior across three instruction-tuned LLMs, eight languages, and all model layers to qualify the language-uni...
As large language models are deployed across multilingual environments, the benchmarks used to evaluate their safety remain largely designed for high-resource, English-dominant contexts. Trust & Safety systems increasingly rely on automated classifiers and generative models to moderate harmful content across dozens of...
This work investigates cross-lingual safety transfer in four African languages, Twi, Hausa, Amharic, and Swahili, using LoDNA, a new safety dataset that pairs literal translations with culturally localized prompts to demonstrate superficial safety alignment.
Abigail Oppong, P SAM SAHIL, Tadesse Destaw Belay et al.· 1 citation
Multi-SALLM, a benchmarking framework designed to systematically evaluate Large Language Models’ ability to generate secure code, reveals three key findings: functional correctness and security are closely related but not equivalent, and sampling strategy is a critical risk factor.
Mohammed Latif Siddiq, Noshin Ulfat, Nishat Raihan et al.· International Conference on...· 0 citations
Reproduction test generation translates a natural-language issue description into executable tests that fail on the original code and pass after the issue is resolved, providing executable evidence for verifying candidate patches. Existing benchmarks are constructed for individual programming languages, preventing a un...
Kazuki Kusama, Sota Nakashima, Haruka Tokumasu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.