Skip to content
Open access

WikiConstraintCalibration: Content-Aware Cost Preference Estimation for Constraint Threshold Selection in Wiki-Grounded LLM Agents

Sep 2026 · Next-Generation Computing Systems and Technologies · 0 citations · 10 references

Abstract

Wiki-grounded LLM agents enforce output constraints through safety rules and scope boundaries derived from structured knowledge bases, yet selecting an appropriate constraint scorer threshold $\theta$ remains challenging. Since outputs are blocked when $s(o) \geq \theta$, a low threshold may over-restrict safe responses, whereas a high threshold may admit dangerous outputs. Existing approaches largely rely on manual, domain-agnostic tuning without systematically incorporating wiki content. This paper proposes WikiConstraintCalibration, a framework that extracts eight wiki-derived features spanning hazard indicators (dangerous assertion density, immediate-action fraction, domain risk prior, and high-risk constraint fraction) and reliability indicators (citation rate, completeness, staleness mean, and stale fraction). These features determine a domain-specific cost ratio $\rho=C_{fn}/C_{fp}$, representing the relative cost of missing dangerous outputs versus blocking safe ones. The ratio determines a safety weight $w_s$ fornavigating an empirical Safety--Function Pareto frontier, while constrained optimisation identifies a recommended operating region $[\theta_l,\theta_h]$. A dynamic module further incorporates staleness signals to adapt class-conditional response policies. Experiments across four deployed wiki-agent domains show that the cost ratio captures domain risk differences, ranging from 5.76 for GP medical to 1.89 for AI course. The GP domain selects the lowest feasible threshold, $\theta^\ast=0.32$, driven by a score-distribution cliff near $\theta\approx0.30$. Ablation analysis identifies domain risk prior as the dominant calibration factor, while content-level features provide additional discrimination. Overall, WikiConstraintCalibration provides an auditable, content-grounded approach to threshold calibration by linking wiki characteristics and domain risk to empiricalsafety--function trade-offs. Future work will address limited test-set sizes and domain-generic safety judging through human annotation and domain-specific scoring.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.