Skip to content
Review Open access

A Comprehensive Survey on Code-Switched LLM Safety and Robustness Evaluation

Sep 2026 · International Journal for Research in Applied Science and Engineering Technology · 0 citations

TL;DR

This survey provides the first comprehensive review of research on code-switched LLM safety and robustness evaluation, and aims to guide researchers and practitioners in developing more robust evaluation frameworks and alignment techniques for real-world multilingual deployments.

Abstract

Large language models (LLMs) have demonstrated remarkable capabilities across languages, yet their safety alignment remains predominantly evaluated in monolingual, especially English, settings. Code-switching (also known as codemixing)—the alternation between two or more languages within a single utterance or conversation—is a pervasive phenomenon in multilingual societies and digital communication. Recent evidence reveals that code-switched inputs can substantially degrade the safety robustness of LLMs, enabling jailbreaks and harmful outputs that monolingual safety mechanisms fail to intercept. This survey provides the first comprehensive review of research on code-switched LLM safety and robustness evaluation. We systematize existing attack methodologies (including Code-Switching Red-Teaming, Multilingual Blending, and attributional analyses), evaluation datasets and metrics, empirical findings on attack success rates and linguistic factors, and emerging mitigation strategies. We further situate these works within the broader multilingual safety literature, highlight critical gaps in linguistic coverage, cultural contextualization, and mechanistic understanding, and outline a research agenda toward equitable, linguistically inclusive LLM safety. Our synthesis aims to guide researchers and practitioners in developing more robust evaluation frameworks and alignment techniques for real-world multilingual deployments.

Read PDF

Similar papers

#machine learning Preprint Aug 2026

When Safety Speaks a Language: A Mechanistic Analysis of Safety-Language Identity Entanglement in LLMs

This work presents a systematic mechanistic analysis of multilingual safety using sparse autoencoder features, sparse interpretable directions in the residual stream associated with harmful and harmless model behavior across three instruction-tuned LLMs, eight languages, and all model layers to qualify the language-uni...

Apoorva Upadhyaya, Sandipan Sikdar · 0 citations
Review Open access Sep 2026

Evaluating Multilingual Safety Benchmarks for Low-Resource Languages in the Majority World

As large language models are deployed across multilingual environments, the benchmarks used to evaluate their safety remain largely designed for high-resource, English-dominant contexts. Trust & Safety systems increasingly rely on automated classifiers and generative models to moderate harmful content across dozens of...

Alisar Mustafa, Cherry Wu · 0 citations
Preprint Aug 2026

The Illusion of Cross-Lingual Safety in Low-Resource Languages

This work investigates cross-lingual safety transfer in four African languages, Twi, Hausa, Amharic, and Swahili, using LoDNA, a new safety dataset that pairs literal translations with culturally localized prompts to demonstrate superficial safety alignment.

Abigail Oppong, P SAM SAHIL, Tadesse Destaw Belay et al. · 1 citation
Open access Aug 2026

Multi-SALLM: a multilingual security assessment of generated code

Multi-SALLM, a benchmarking framework designed to systematically evaluate Large Language Models’ ability to generate secure code, reveals three key findings: functional correctness and security are closely related but not equivalent, and sampling strategy is a critical risk factor.

Mohammed Latif Siddiq, Noshin Ulfat, Nishat Raihan et al. · 0 citations
Preprint Sep 2026

Multi-SWT-Bench: A Multilingual Benchmark for Reproduction Test Generation

Reproduction test generation translates a natural-language issue description into executable tests that fail on the original code and pass after the issue is resolved, providing executable evidence for verifying candidate patches. Existing benchmarks are constructed for individual programming languages, preventing a un...

Kazuki Kusama, Sota Nakashima, Haruka Tokumasu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.