Skip to content
Preprint

Redteaming Leading Arabic LLMs with ASAS

Aug 2026 · 0 citations · 16 references
Computer Science

TL;DR

The Arabic Safety Index (ASAS) is introduced, the first fully human-curated Arabic benchmark for redteaming LLMs and provides a culturally grounded benchmark and redteaming protocol to drive progress in Arabic LLM safety.

Abstract

As the adoption of large language models (LLMs) grows in Arabic-speaking regions, ensuring their safety and cultural alignment is increasingly critical. However, Arabic LLM safety remains underexplored, especially in adversarial evaluation settings. We introduce the Arabic Safety Index (ASAS), the first fully human-curated Arabic benchmark for redteaming LLMs. ASAS contains 801 prompts spanning 8 safety categories and 8 attack strategies, with ideal responses in Modern Standard Arabic (MSA). We conduct a redteaming evaluation across seven leading models with Arabic capabilities, including GPT-4o, Claude 3.7 Sonnet, and regional models such as ALLaM and FANAR. Human annotators rate responses using a structured 4-point safety scale, revealing that most models fail to defend against 50% of unsafe prompts. Our findings highlight major safety gaps in high-harm categories such as weapons and illicit substances, with direct and obfuscation-based attacks proving most effective. The results also show that language alignment does not readily transfer across languages, and that automated safety judges (e.g., GPT-4o) perform poorly compared to human annotators. ASAS provides a culturally grounded benchmark and redteaming protocol to drive progress in Arabic LLM safety.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks

Large language models (LLMs) are increasingly used in multilingual settings, yet their safety is still evaluated primarily in English. This limits our understanding of how alignment failures manifest in low-resource and culturally diverse languages. We introduce IndicSafeEval, a persuasion-based jailbreak evaluation fr...

Saikat Mondal, Mamta Mamta, Deeksha Varshney et al. · 0 citations
#natural language process... Preprint Sep 2026

AraDynFact: Dynamic Evaluation of Factual Knowledge in Arabic

As Large Language Models (LLMs) continue to scale both in size and capabilities, their proficiency in the Arabic Language has seen significant advancement. However, a critical gap remains: the extent of their factual knowledge and cultural sensitivity to the diverse Arabic-speaking world remains largely underexplored....

Ignacio Iacobacci, Faroq Al-Tam, Zhao-Zhi Qian et al. · 0 citations
Preprint Aug 2026

SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs

SurakshaEval is introduced, a novel safety benchmark composed of human-written prompts spanning real-world scenarios, explicitly designed for ten major Indian languages - Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Punjabi, Tamil, and Telugu - along with English.

Debopriyo Banerjee, K. R. Kavitha, Angana Borah et al. · 0 citations
#natural language process... Preprint Aug 2026

ArabicDialectSafety: A Dialect-Aware Benchmark for Arabic Content Safety Classification

This paper introduces a dual-task evaluation framework for binary safe/unsafe detection and granular harm classification across dialects, and evaluates seven frontier LLMs as response generators on harmful dialectal Arabic prompts and observes unsafe generation rates below 5 percent across models.

Wajdi Zaghouani, Mukut Biswas, K. Aldous et al. · 0 citations
Preprint Aug 2026

The Illusion of Cross-Lingual Safety in Low-Resource Languages

This work investigates cross-lingual safety transfer in four African languages, Twi, Hausa, Amharic, and Swahili, using LoDNA, a new safety dataset that pairs literal translations with culturally localized prompts to demonstrate superficial safety alignment.

Abigail Oppong, P SAM SAHIL, Tadesse Destaw Belay et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.