The findings indicate that Turkish LLM safety cannot be inferred from general model capability alone and should be assessed through language-specific, culturally aware, and continuously updated adversarial benchmarks.
Abstract
Large language models (LLMs) are increasingly deployed in multilingual settings, yet their safety behavior under Turkish harmful prompts and prompt injection attempts remains insufficiently characterized. This study evaluates the adversarial robustness of 55 open- and closed-source LLMs under paired Turkish and English harmful prompt conditions. We constructed a benchmark of 790 Turkish adversarial prompts, translated the prompts into English for cross-lingual comparison, and applied both prompt sets to the model pool. Model responses were labeled as harmful, harmless, or hallucinatory, and safety was analyzed using safety scores, Turkish–English ranking differences, and inter-rater reliability based on Fleiss’ kappa. The results reveal substantial variation across models. Closed-source systems generally achieved higher safety scores and stronger filtering behavior, whereas open-source and Turkish-oriented models showed a wider performance distribution. GPT-5.4 ranked first in the Turkish tests with a 99.37% safety score but decreased to 96.71% in the English tests, while Qwen3.5:27B ranked first in English with 97.47%. These differences suggest that safety mechanisms are not fully language-invariant. Hallucination also emerged as a distinct safety risk, particularly in Turkish evaluations. The findings indicate that Turkish LLM safety cannot be inferred from general model capability alone and should be assessed through language-specific, culturally aware, and continuously updated adversarial benchmarks.
IndicSafeEval, a persuasion-based jailbreak evaluation framework for Indian languages, is introduced and it is observed that different risk categories exhibit different levels of vulnerability, with some types of harmful content being significantly more susceptible to persuasion-based jailbreaks than others.
Saikat Mondal, Mamta Mamta, Deeksha Varshney et al.· 0 citations
This work examines emoji-augmented prompts as a test case for robustness, evaluating 50 prompts across four open-source LLMs, showing substantial variation in robustness.
TanglishGuard is introduced, a novel benchmark designed to systematically assess the robustness of LLM safety mechanisms against code-mixed inputs and contributes to the development of more robust and equitable AI systems by ensuring safety mechanisms are tested against authentic communication patterns rather than sani...
Vivina Vijesh, Dr. E. Gothai, M. Student· International Journal of Inn...· 0 citations
The Arabic Safety Index (ASAS) is introduced, the first fully human-curated Arabic benchmark for redteaming LLMs and provides a culturally grounded benchmark and redteaming protocol to drive progress in Arabic LLM safety.
F. Abed, Haidar Khan, M Saiful Bari et al.· 0 citations
This survey provides the first comprehensive review of research on code-switched LLM safety and robustness evaluation, and aims to guide researchers and practitioners in developing more robust evaluation frameworks and alignment techniques for real-world multilingual deployments.
Pulagam Naveen Kumar, Sowjanya Bojja, K. Anoosha et al.· International Journal for Re...· 0 citations
The threat posed by adversarial prompts to large language models is becoming harder to ignore. Problems including prompt injection, jailbreaking, phishing, and Unicode-based attacks are now widespread. Most existing solutions protect against only one threat type, operate in English only, and provide no explanation for...
Abdullah M. Abughallous, Somia Abufakher· IEEE Jordan Conference on Ap...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.