A novel framework for computing rigorous bounds on the probability that a large language model (LLM) generates harmful output to a given prompt is proposed and a new application of the Clopper-Pearson confidence intervals is studied to obtain probably approximately correct bounds.
Abstract
We propose a novel framework for computing rigorous bounds on the probability that a large language model (LLM) generates harmful output to a given prompt. We study a new application of the Clopper-Pearson confidence intervals to obtain probably approximately correct (PAC) bounds for this problem. As our main technical contribution, we propose an algorithm that leverages features in the latent space to prioritize exploring branches in the auto-regressive generation tree that are more likely to produce harmful outputs. Our approach in particular enables the efficient computation of useful lower bounds, even in scenarios where the true harm probability is extremely small, and crucially, the obtained lower bounds are sound, i.e., formally proven to be less than the actual harmfulness probability: our experimental results demonstrate the effectiveness of our method by computing non-trivial lower bounds on state-of-the-art LLMs. This study newly enables the evaluation and statistical certification of LLMs.
We introduce RareTrap, a framework for estimating the probability of severe behaviors in black box large language models (LLMs). A key challenge for probability estimation is defining a tractable distribution over the input space. To accomplish that, RareTrap uses a surrogate LLM and constructs a geometry-aware mapping...
Elsayed Eshra, Ali Al-Lawati, Dongwon Lee et al.· 0 citations
The interactive PCP is constructed, which shows a protocol in which a polynomial-time verifier can verify the approximate consistency of (P,Q), and places l_2-approximate probabilistic consistency of explicit claims in NP, with certificates of length O(mn + log B) in the input bit-precision B.
Orr Paradise, Oliver E. Richardson, Y. Bengio et al.· 0 citations
This work proposes a chance-constrained formulation for safety-preserving fine-tuning that limits the fraction of safety examples whose degradation relative to a reference model exceeds a prescribed threshold and develops a constraint-aware gradient descent method that treats the majorized constraint as a safe set in p...
This work extends Compute-Aligned Training to this setting through an abstraction of policy-guided search, deriving tractable, trace-supported losses and introduces a search-agnostic uniform-allocation (UA) loss that accounts for the budget without specifying the specific search.
Recently, Large Language Models (LLMs) have been increasingly able to solve advanced mathematical problems, including many that have been open for decades. This opens the door to expansion of mathematical knowledge at unprecedented scale. Yet, while LLMs may be able to conjecture and prove more and more theorems, it re...
Niket Patel, A. Rammal, Amaury Hayat et al.· 0 citations
Large language models perform strongly on competition mathematics, but their research-level reasoning remains difficult to evaluate systematically. Theoretical computer science (TCS) connects algorithm design to explicit guarantees and fundamental limits, providing a setting for evaluating whether models can justify co...
Chu-Tong Yang, Xi-Yuan Zhang, Yu Huang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.