RAG-Safety-Bench: Reliable Evaluation of Retrieval-Augmented LLM Safety
Results are reported, showing an inverse relationship between benign and unsafe capability, strong evidence that baseline safety guardrails do not lead to downstream safety guarantees in the RAG case, and model-specific support for previous findings that even benign documents can lead to unsafe generation in retrieval-...