Skip to content
Preprint

BeTaL-GBI: Admission-Aware Benchmark Tuning and Full-Stack Verification of Geometric Belief Interfaces

Aug 2026 · 0 citations · 21 references
Computer Science

TL;DR

Whether an enterprise verification architecture can isolate interface failure, task competence, policy admissibility, policy admissibility, and control integrity while keeping claims auditable is highlighted.

Abstract

A verification substrate is more credible when exposing errors in its own claims, not just model outputs. GBI-DCSE v3 falsified an architectural claim: the reported Fisher value epsilon ~ 0.066 satisfies the kappa^2<= 10^4 budget only on the slice [epsilon, 3, 4, 5], while the full box [epsilon, 20]^4 requires epsilon ~ 0.326472. This erratum highlights whether an enterprise verification architecture can isolate interface failure, task competence, policy admissibility, and control integrity while keeping claims auditable. BoundaryBench v0.1 established the baseline: Qwen3-4B-Instruct-2507 completed 768 frozen executions, but 0% cleared the contract (369 failed parsing, 399 failed validation), limiting downstream selectivity metrics. This companion study evaluates three successive improvements. First, BeTaL-GBI v0.2 applies Benchmark Tuning with an LLM-in-the-loop over 2,218,750,380 grid points, separating format admission from conditional performance (rho_adm = N_admitted/N; rho_task = N_verified/N_admitted). Following schema repair, a model-free feedback search achieves a 2.87% mean held-out target gap, outperforming non-feedback baselines (13.61%, 11.46%). Second, GBI v2 swaps static keys for a reference-independent witness state W and policy P. Across 512 synthetic tasks, a 16-gate policy detects all 116 injected severe contradictions and accepts all 99 clean records (broad denominator: 4.27%). Hallucinator and evidence-forger surrogates are blocked with zero silent promotions. Third, GBI-DCSE v3 maps 99 claims to machine-readable evidence: 95 of 96 testable claims pass, with 148 standalone checks executed without failure. The harness exercises signed ledgers, PBFT quorums, and enclave forgery across 62 configurations. Under synthetic conditions, GBI-DCSE is a selective, policy-versioned, self-auditing test and routing substrate.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

SemVerBench: Benchmarking LLM Comprehension of Version-Constraint Resolution Semantics

SemVerBench is introduced, the first benchmark of LLM version-constraint resolution semantics across three ecosystems (npm, PEP 440, Cargo), and six frontier models are evaluated: Opus significantly outperforms all other models, and Sonnet outperforms the OpenAI models (McNemar).

Qi-Bai Chen, Ze-Ming Liu · 1 citation
#artificial intelligence Preprint Aug 2026

Beyond the Answer Key: Robustness Evaluation of Large Language Models for Step-Level Mathematical Verification

Evaluating state-of-the-art open LLMs reveals a significant robustness gap, and shows that reliable process-level verification remains challenging, and evaluator robustness should be measured separately from solver accuracy, even in a simple algebraic domain with exact ground truth.

Fatemeh Mazdarani, Carlos Toxtli · 2 citations
Preprint Aug 2026

EnterpriseRAG: Benchmarking LLM Instruction Adherence and Robustness under Non-Ideal Enterprise Retrieval

Enterprise RAG deployments face a critical reliability gap: while LLMs satisfy 80% of individual constraints, only 26.8% of responses meet all requirements simultaneously, revealing a 57-point orchestration gap, which provides a reproducible foundation for measuring and closing these gaps.

Hui Miao, Xin Sun, Bo Wang et al. · 0 citations
Open access Sep 2026

Raw CVE/CWE Retrieval Does Not Improve LLM-Based Vulnerability Detection in Python: A Pre-Specified Null Result and a Retrieval-Dose Audit

Retrieval-Augmented Generation (RAG) over vulnerability databases is widely expected to improve LLM-based vulnerability detection. We report a pre-specified evaluation in which it does not, and derive from it an evaluation protocol. On a near-balanced benchmark of 100 Python 3.12 snippets, raw CVE/CWE retrieval never m...

Patrick Deininger, S. Rappl, Helmut Lindner et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.