IH-Benchmark is a conflict-centered benchmark for instruction-hierarchy robustness across direct system-user conflicts (S>U) and tool-mediated user-tool (U>T) conflicts and suggests that instruction-hierarchy robustness is not a single capability, but a set of behaviors that must be evaluated across conflict surfaces, constraint types, and attack presentations.
Abstract
When a language model receives conflicting instructions from different priority levels, which one does it actually follow? This question lies at the heart of reliable LLM deployment. Existing benchmarks answer this only partially, often focusing on a single hierarchy edge or adapting public datasets with limited tool-use coverage. We present IH-Benchmark, a conflict-centered benchmark for instruction-hierarchy robustness across direct system-user conflicts (S>U) and tool-mediated user-tool (U>T) conflicts. IH-Benchmark is built from a human-authored taxonomy of 44 constraint families across generic, health, finance, retail, and coding settings, and evaluates scenarios with a uniform binary pass/fail protocol combining a predicate DSL with category-scoped LLM judges. Across 37 evaluated models, hierarchy compliance ranges from 98.2% to 20.5%. We find that strong S>U compliance is not a reliable proxy for U>T robustness: several models preserve system constraints under direct user conflict but degrade sharply when conflicting instructions appear in tool outputs. Constraint hardening also reveals a split between models: some failures are largely fixed by stronger warnings, while others persist across all strictness levels. Finally, the most revealing failures are often subtle rather than overtly dangerous; models resist unauthorized purchases or bulk ticket closure more reliably than injected disclaimers or small factual distortions. These results suggest that instruction-hierarchy robustness is not a single capability, but a set of behaviors that must be evaluated across conflict surfaces, constraint types, and attack presentations.
It is shown that user-preferring conflict resolution can coexist with a readable internal arbitration signal, and successful intervention depends on the geometry of the readout rather than probe accuracy alone, while directions selected mainly for pooled separability steer poorly.
Enrique Balp-Straffon, Chih-Hao Hsu, Rushiraj Gadhvi et al.· 0 citations
SkillSafe-Bench is introduced, a controlled benchmark that scores skill-merged models on static refusal, adaptive jailbreak robustness, and capability retention under a conservative two-judge AND rule, and the static effect of merging is base-conditional.
KC-Bench is introduced, a controlled multi-turn benchmark for measuring model-level behavior across world-knowledge conflicts, input inconsistencies, and multi-source temporal conflicts, and provides a reproducible diagnostic for developing conflict-aware reasoning and execution safeguards.
Yaxing Lyu, Sheng-Jie Zhou, B. Toh et al.· 0 citations
Per-server prompt engineering is therefore a workaround rather than a fix; it is argued that MCP host applications should provide an explicit mechanism that places server instructions ahead of tool selection in the client LLM's deliberation.
Minhan Cho, Soyoung Park, Kihyeon Jeong et al.· 0 citations
Accurately evaluating the instruction-following ability of Large Language Models (LLMs) is crucial for their practical deployment. Existing evaluation methods mainly rely on static assessment of single-pass generations, making it difficult to comprehensively measure instruction-following performance or models’ capabili...
Ya-Chao Fu, Qu-Fei Zhang, Meng-Na Zhu et al.· 2026 12th International Conf...· 0 citations
SemVerBench is introduced, the first benchmark of LLM version-constraint resolution semantics across three ecosystems (npm, PEP 440, Cargo), and six frontier models are evaluated: Opus significantly outperforms all other models, and Sonnet outperforms the OpenAI models (McNemar).
Qi-Bai Chen, Ze-Ming Liu· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.