SafetyJudge-LLM shows that local open-weight LLMs can support semantic safety judging, but their reliability must be evaluated across multiple dimensions.
Abstract
Background: LLM-as-a-judge workflows are increasingly used to evaluate open-ended model outputs, but the judge model can itself become a source of error in safety assessment. SafetyJudge-LLM audits local open-weight LLMs as semantic safety judges. Methods: This study reused a fixed set of previously reviewed safety-boundary responses and their hidden reference labels. Two independent human evaluations (R1 and R2) quantified reference-layer ambiguity. Seven local open-weight judge models were evaluated under a common Ollama inference protocol. A paired C6 sensitivity analysis reran llama3.2:3b and qwen3:8b through Hugging Face Transformers. Results: The final judge-output matrix contained 10,612 retained outputs. R1–R2 agreement was 95.45% (Cohen’s κ = 0.612) overall but 47.80% (κ = 0.341) in secondary cases. Several judge models detected more than 90% of confirmed safety-boundary failures, but high detection was not always accompanied by low false-unsafe behavior on control cases. Output-format reliability also varied across models: overall label parseability was 98.11%, while strict JSON schema compliance was 92.55%. The llama3.2:3b schema-failure rate persisted across engines (52.06% under Ollama; 59.60% under Transformers), whereas qwen3:8b maintained complete compliance. Conclusions: SafetyJudge-LLM shows that local open-weight LLMs can support semantic safety judging, but their reliability must be evaluated across multiple dimensions.
Large language model (LLM) safety judges from different model families can disagree by an order of magnitude on the same responses, with reported attack success rate (ASR) ranging from 0.04 to 0.42 on a single benign-only fine-tuned model. This is not ordinary classifier noise: the gap is driven by judges whose unsafe recall is too low to detect harmful outputs. Existing pipelines select judges by aggregate metrics such as balanced accuracy, which can look acceptable while false negatives silently suppress ASR. We propose a confidence-interval (CI) filtered two-stage judge-screening framework. Stage 1 retains only judges whose Clopper–Pearson 95% lower bound on unsafe recall exceeds a predefined floor; Stage 2 selects the highest balanced-accuracy judge among the survivors. When no judge has sufficient calibration evidence, the framework emits an explicit low-confidence or abort verdict, and downstream ASR can be accompanied by a Rogan–Gladen first-order prevalence correction. We evaluate the framework on seven labeled prompt sources and nine candidate judges spanning five base model families in two judge categories. In small-calibration regimes, it reduces the rate of selecting judges that fail the unsafe-recall screen by an order of magnitude. Across five English sources it produces a pass and reject map that is stable under the unweighted bound, with one of 45 judge-source cells reversing under design weighting; on a Chinese multilingual jailbreak source, every English-trained judge fails Stage 1 and the framework abstains. A controlled supervised fine-tuning (SFT) case study with varying benign-to-safety data ratios shows that on the same benign-only (1:0) responses the reported ASR spans a factor of 12.5 across judges, and that the lowest-recall judge underestimates the human-labeled ASR by a factor of 6.0; CI-filtered screening prevents the corresponding ASR underestimation.
Ji-Xiang Yang, Jun-Fei Yi, Jin-Han Li et al.· IEEE Access· 0 citations
This work proposes Sequential Epistemic and Action-Level Validation (SEAV), a verification-centric jailbreak evaluation framework that decomposes responses into ordered steps and evaluates both validity and correctness, and shows that enforcing correctness substantially reshapes measured robustness.
Qilong Wu, Sahil Wadhwa, Pranab Mohanty et al.· 0 citations
Generative AI systems are increasingly producing real-world artifacts, however their efficacy and validity are often evaluated via context-free LLM-scoring. These judges can be miscalibrated by irrelevant in-context reference examples, creating false confidence and allowing low-quality or harmful outputs to pass evaluation. We study this failure mode as context-induced miscalibration and introduce DA-RAC, a distance-aware reference-anchored calibration method for LLM judges. DA-RAC retrieves semantically and structurally similar labeled anchors for each judgement scenario, weights them by distance, and exposes neighborhood difficulty as a calibration and triage signal. On multi-run LLM-judge evaluation benchmarks, it improves calibration and reduces false-pass risk relative to zero-shot, chain-of-thought evaluation, and static-anchor baselines. Mechanistic analysis shows that judge scores vary systematically with anchor distance, while static references can induce misleading decision boundaries. Thus LLM-judgement requires not only better models, but also calibrated, auditable reference selection, especially when automated evaluation is used to support high-impact AI generated artifacts. Judgments should be grounded in relevant, inspectable, and contestable interpretive artifacts.
Chengyu Wu, Vishal Anand, J. Mandivarapu et al.· 0 citations
The results show that high judge agreement can coexist with weak sensitivity to changes in the construct being evaluated, motivating joint reporting of invariance and sensitivity and auditing the validation set itself.
Jianlin Chen, Wen-Hui Chen, Ziyao Lin et al.· 0 citations
ConsensusGrade provides a diagnostic framework for interpreting automated scores relative to observed human variability; inside-envelope rates should not be interpreted as stand-alone measures of grading accuracy.
Cătălin Anghel, A. Anghel, Mihai Vlase et al.· Applied System Innovation· 0 citations
A two-stage framework for medical hypothesis verification in multiple-choice settings that manages this tradeoff through targeted ontology grounding, applied only when the model abstains, and shows that abstention is not random but reflects genuine uncertainty, with abstained predictions associated with lower confidence.
Uma Ranjan, Kunal Tilaganji, Aditya Koul et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.