Jul 2026
Sound Probabilistic Safety Bounds for Large Language Models
A novel framework for computing rigorous bounds on the probability that a large language model (LLM) generates harmful output to a given prompt is proposed and a new application of the Clopper-Pearson confidence intervals is studied to obtain probably approximately correct bounds.
Mahdi Nazeri, Anne-Kathrin Schmuck, S. Soudjani et al.
· arXiv.org · 0 citations