Skip to content
Review

A Threshold Exceedance Framework for CBRN Uplift Evaluation in Frontier Language Models

Jul 2026 · arXiv.org · Vol abs/2607.12200 · 0 citations · 26 references
Computer Science

TL;DR

A Threshold Exceedance Criteria (TEC) framework is introduced that decomposes an uplift study into independently executable components: determining non-expert participant eligibility, defining the CBRN threat scope for the study, and statistically estimating material uplift.

Abstract

As frontier language models advance, policymakers and model developers need methods for assessing whether model access materially increases a non-expert actor's ability to plan high-consequence Chemical, Biological, Radiological, or Nuclear (CBRN) misuse relative to public tools alone. Existing CBRN evaluations differ in non-expert definitions, threat scope, baselines, scoring rubrics, and decision rules, making results difficult to compare across studies. We introduce a Threshold Exceedance Criteria (TEC) framework that decomposes an uplift study into independently executable components: determining non-expert participant eligibility, defining the CBRN threat scope for the study, and statistically estimating material uplift. We then operationalize the TEC framework in a large-scale empirical study using a design that determines two forms of uplift: generative (where a model assists plan creation from scratch) and revisionist (where a model assists refinement of an existing plan). The study produced attack plans across the CBRN domains, which we evaluated through subject-matter-expert review to estimate generative and revisionist uplift. Applying the framework, our empirical study revealed domain heterogeneity: under this controlled pre-release evaluation, model-assisted plans sometimes received expert-equivalent instructional ratings, but confirmed material uplift was limited to the radiological domain. These findings informed mitigation and deployment-governance decisions rather than characterizing deployed model behavior. We conclude with methodological lessons for future CBRN uplift evaluations, emphasizing prespecified criteria, explicit baselines, separation of generative and revisionist estimates, and careful distinction between preliminary screening signals and confirmed risk determinations.

View source

Similar papers

Review Aug 2026

FinRiskAtlas: Decision-Aligned Evaluation of Large Language Models for Financial Risk Review

FinRiskAtlas is introduced, a Chinese-language benchmark that evaluates financial LLMs along two complementary dimensions: operation execution under fixed evidence states and evidence-state control under evolving review conditions, and shows that broad financial capability scores do not fully capture where models are r...

Su-Yang Zhong, Jingzhe Zhu, Qi Xu et al. · 0 citations
Preprint Jul 2026

WuYuEval: A Multi-Level Benchmark for Large Language Models in Solid Waste Management

WuYuEval is introduced, a multi-level benchmark for evaluating LLMs in SWM across foundational knowledge, domain reasoning, and expert decision-making and suggests that visible deliberation helps only when it remains anchored to units, assumptions, and engineering constraints; otherwise, it may drift from decisive answ...

Yi Zhang, Hongyang Wang, Zheng Hao Leong et al. · 0 citations
Jul 2026

INS-ActBench: A Comprehensive Benchmark for Assessing Professional Actuarial Capability of Large Language Models

Large Language Models (LLMs) have shown strong potential in financial reasoning, but existing benchmarks often evaluate domain knowledge, numerical reasoning, long-context understanding, and tool use in separate settings. This limits their ability to assess realistic professional workflows that require auditable, conte...

Changyu Chen, Chenwei Lin, X. Xu · 0 citations
#artificial intelligence Preprint Sep 2026

FUSE: An Evaluating Framework for Dangerous Capabilities of LLMs

A modular framework that evaluates each model through three orthogonal pipelines under a unified protocol, aggregating results into a standardized dangerous-capability profile $\phi$, showing that scaling and alignment progress do not uniformly translate into safety.

Zheng-Yi Jin, Ru Zhang, Xiao Chen et al. · 0 citations
Preprint Aug 2026

HugSelect: An Explainable Multi-Criteria Decision-Support Framework for foundation-model selection

It is argued that foundation-model selection should be treated as an explicit, auditable software-component selection task rather than as keyword search, popularity ranking, or opaque conversational advice, and HugSelect, an explainable decision-support framework for foundation-model selection is proposed.

Alireza Joonbakhsh, Arda Canser Adalı, Slinger Jansen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.