It is argued that AI safety evaluations should systematically account for modality, multi-run consistency, search conditions, and response-level behaviors to better reflect how deployed AI systems behave in practice.
Abstract
Large language model (LLM) benchmark evaluations are routinely used to support claims about model safety, reliability, and deployment readiness. Yet most evaluations rely on a single access modality (model APIs), perform a single run per prompt, and report accuracy as the primary outcome metric, without accounting for conditions such as web search that may have effects on model behavior in deployment. We audit these assumptions for one of the most widely-used LLMs, comparing two modalities, ChatGPT's chat UI and OpenAI's API, with and without web search enabled. We use a stratified total sample of 401 prompts from two popular benchmarks, BBQ and SafetyBench, collecting 4,812 total responses across three repeated runs per prompt. Beyond standard performance measures, we evaluate model output dimensions including response consistency, response text similarity, citation grounding, and abstention behavior. For instance, chat UI responses were less accurate than API responses on both benchmarks with search disabled. Enabling web search reduced accuracy by up to 8 percentage points, and even reversed the direction of modality performance trends for one benchmark. Repeated runs of the same prompt produced inconsistent responses in up to 21\% of prompts. The two modalities also grounded answers in different citations, and abstention behavior was also inconsistent across both modalities. These results illustrate that, even within a model family, reporting only simple accuracy metrics can obscure important forms of model behavioral variation relevant to AI safety assessments. We argue that AI safety evaluations should systematically account for modality, multi-run consistency, search conditions, and response-level behaviors to better reflect how deployed AI systems behave in practice.
Power analysis is critical for assuring rigor and validity of quantitative research yet remains underutilized due to technical challenges associated with specialized software. At the same time, large language models (LLMs) are being rapidly integrated into research practice, raising interest in their potential to assist statistical and research design tasks. However, despite their widespread adoption, the reliability of LLMs in supporting statistically rigorous procedures has not been systematically evaluated, posing risks for unexamined or overly optimistic use. To address this gap, we evaluated four widely used LLMs—ChatGPT (GPT-3.5, GPT-4, GPT-4o) and Llama 3.2—across two experiments. Experiment 1 examined whether LLMs could calculate required sample sizes for common statistical tests (two-sample t-test, one-way ANOVA, and χ² goodness-of-fit test) under different prompting strategies, including direct calculation versus R/Python code generation. Experiment 2 assessed models’ ability to identify missing input parameters necessary for power analysis, which is a task that requires methodological understanding. Results revealed that GPT-4 and GPT-4o performed well when generating R code for sample size estimation, but struggled with direct numerical calculation. Furthermore, while LLMs were able to detect missing information, their reliability varied by statistical context. Findings suggest that while LLMs may offer support in structuring and initiating power analysis, they cannot substitute for expert judgment. Overall, the study underscores the importance of critically evaluating LLM performance in statistically demanding tasks. Responsible integration of LLM requires critical oversight, cross-verification, and methodological evaluation.
Hajung Kim, Jia Qi, Zhe Feng et al.· Journal of Behavioral Data S...· 0 citations
A fundamental disconnect is suggested between a model's capacity for factual accuracy and its ability to maintain social fairness, highlighting the need for multi-dimensional evaluation frameworks for small-scale systems.
M.J.F. Valdez, Arghir-Nicolae Moldovan· International Conference on...· 0 citations
This work presents the most comprehensive evaluation of LLM safety capabilities to date, systematically testing models across datasets that are organized into four distinct categories, and uncovers critical blind spots.
A large-scale assessment of the effectiveness and robustness of these automated pipelines is conducted by evaluating five widely used benchmark suites across 26 open-source SLMs under a unified judging rubric, which reveals a capability-safety confound that mixes model capability with apparent safety.
A probe corpus of 42 retracted, fraudulent, and pseudoscientific papers is paired with a methodology for eliciting and scoring single-shot model engagement with each paper's framing, indicating an urgent need for guardrail infrastructure for scientific deployment of language models.
V. Rodionov, Shamil Assylbekov· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.