Skip to content

Sound Probabilistic Safety Bounds for Large Language Models

Jul 2026 · arXiv.org · Vol abs/2607.20286 · 0 citations · 35 references
Computer Science

TL;DR

A novel framework for computing rigorous bounds on the probability that a large language model (LLM) generates harmful output to a given prompt is proposed and a new application of the Clopper-Pearson confidence intervals is studied to obtain probably approximately correct bounds.

Abstract

We propose a novel framework for computing rigorous bounds on the probability that a large language model (LLM) generates harmful output to a given prompt. We study a new application of the Clopper-Pearson confidence intervals to obtain probably approximately correct (PAC) bounds for this problem. As our main technical contribution, we propose an algorithm that leverages features in the latent space to prioritize exploring branches in the auto-regressive generation tree that are more likely to produce harmful outputs. Our approach in particular enables the efficient computation of useful lower bounds, even in scenarios where the true harm probability is extremely small, and crucially, the obtained lower bounds are sound, i.e., formally proven to be less than the actual harmfulness probability: our experimental results demonstrate the effectiveness of our method by computing non-trivial lower bounds on state-of-the-art LLMs. This study newly enables the evaluation and statistical certification of LLMs.

View source

Similar papers

#machine learning Preprint Sep 2026

Quantifying Behavioral Tails in Black-Box Language Models

We introduce RareTrap, a framework for estimating the probability of severe behaviors in black box large language models (LLMs). A key challenge for probability estimation is defining a tractable distribution over the input space. To accomplish that, RareTrap uses a surrogate LLM and constructs a geometry-aware mapping...

Elsayed Eshra, Ali Al-Lawati, Dongwon Lee et al. · 0 citations
#artificial intelligence Preprint Aug 2026

How to Verify Probabilistic Consistency of Predictive Models

The interactive PCP is constructed, which shows a protocol in which a polynomial-time verifier can verify the approximate consistency of (P,Q), and places l_2-approximate probabilistic consistency of explicit claims in NP, with certificates of length O(mn + log B) in the input bit-precision B.

Orr Paradise, Oliver E. Richardson, Y. Bengio et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond Average Safety: Chance-Constrained LLM Fine-tuning

This work proposes a chance-constrained formulation for safety-preserving fine-tuning that limits the fraction of safety examples whose degradation relative to a reference model exceeds a prescribed threshold and develops a constraint-aware gradient descent method that treats the majorized constraint as a safe set in p...

Taha Entesari, Mahyar Fazlyab · 0 citations
#artificial intelligence Preprint Sep 2026

Direct Optimization of Generators for Search in Automated Theorem Proving

This work extends Compute-Aligned Training to this setting through an abstraction of policy-guided search, deriving tractable, trace-supported losses and introduces a search-agnostic uniform-allocation (UA) loss that accounts for the budget without specifying the specific search.

Adam Ousherovitch, A. Tewari · 0 citations
#artificial intelligence Preprint Sep 2026

Learning to Discover Interesting Mathematics

Recently, Large Language Models (LLMs) have been increasingly able to solve advanced mathematical problems, including many that have been open for decades. This opens the door to expansion of mathematical knowledge at unprecedented scale. Yet, while LLMs may be able to conjecture and prove more and more theorems, it re...

Niket Patel, A. Rammal, Amaury Hayat et al. · 0 citations
#artificial intelligence Preprint Sep 2026

TCSAlgBench: Benchmarking Automated Proving for Research-Level Theoretical Computer Science

Large language models perform strongly on competition mathematics, but their research-level reasoning remains difficult to evaluate systematically. Theoretical computer science (TCS) connects algorithm design to explicit guarantees and fundamental limits, providing a setting for evaluating whether models can justify co...

Chu-Tong Yang, Xi-Yuan Zhang, Yu Huang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.