Towards trustworthy foundation models: a systematic review of safety, evaluation, and defence mechanisms
Abstract
Large Language Models (LLMs) and foundation models are increasingly deployed in security-critical and high-impact settings, including healthcare, cybersecurity, software engineering, and intelligent infrastructure. Their open-ended interfaces and multimodal capabilities create new attack surfaces, where prompt injection, jailbreaking, hallucination, adversarial fine-tuning, and cross-modal manipulation can compromise reliability, integrity, privacy, and user trust. This paper presents a systematic literature review of 94 peer-reviewed studies identified through database searching, backward reference snowballing, venue screening, and quality assessment within a January 2021–July 2026 search window. The review synthesises evidence on safety and robustness vulnerabilities, red-teaming practices, mitigation mechanisms, evaluation metrics, and unresolved research gaps for LLMs and related foundation-model systems. The findings show that LLM vulnerabilities are rarely isolated: prompt-level, model-level, data-level, and deployment-level risks often interact, particularly in high-stakes domains. Red-teaming has become more structured, but existing practices remain inconsistent across threat models, benchmarks, attack settings, and reporting standards. The reviewed defences include prompt hardening, input/output filtering, adversarial training, retrieval-augmented grounding, cross-model verification, gatekeeper models, and domain-specific safety layers; however, their effectiveness is difficult to compare because studies use heterogeneous metrics and evaluation protocols. Major gaps include fragmented benchmarking, limited multilingual and low-resource assessment, insufficient adaptive-adversary and long-term deployment studies, and weak integration between technical defences and governance mechanisms. The review indicates that trustworthy foundation-model deployment requires standardised threat models, reproducible red-teaming protocols, transparent safety metrics, and defence-in-depth strategies tailored to domain-specific security risks.