Skip to content

IH-Benchmark: A Conflict-Centered Benchmark for Instruction-Hierarchy Robustness in LLM Applications

Jul 2026 · arXiv.org · Vol abs/2607.25987 · 2 citations · ⚡ 1 influential · 22 references
Computer Science

TL;DR

IH-Benchmark is a conflict-centered benchmark for instruction-hierarchy robustness across direct system-user conflicts (S>U) and tool-mediated user-tool (U>T) conflicts and suggests that instruction-hierarchy robustness is not a single capability, but a set of behaviors that must be evaluated across conflict surfaces, constraint types, and attack presentations.

Abstract

When a language model receives conflicting instructions from different priority levels, which one does it actually follow? This question lies at the heart of reliable LLM deployment. Existing benchmarks answer this only partially, often focusing on a single hierarchy edge or adapting public datasets with limited tool-use coverage. We present IH-Benchmark, a conflict-centered benchmark for instruction-hierarchy robustness across direct system-user conflicts (S>U) and tool-mediated user-tool (U>T) conflicts. IH-Benchmark is built from a human-authored taxonomy of 44 constraint families across generic, health, finance, retail, and coding settings, and evaluates scenarios with a uniform binary pass/fail protocol combining a predicate DSL with category-scoped LLM judges. Across 37 evaluated models, hierarchy compliance ranges from 98.2% to 20.5%. We find that strong S>U compliance is not a reliable proxy for U>T robustness: several models preserve system constraints under direct user conflict but degrade sharply when conflicting instructions appear in tool outputs. Constraint hardening also reveals a split between models: some failures are largely fixed by stronger warnings, while others persist across all strictness levels. Finally, the most revealing failures are often subtle rather than overtly dangerous; models resist unauthorized purchases or bulk ticket closure more reliably than injected disclaimers or small factual distortions. These results suggest that instruction-hierarchy robustness is not a single capability, but a set of behaviors that must be evaluated across conflict surfaces, constraint types, and attack presentations.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

How Language Models Choose Sides: Internal Representations of Instruction Hierarchy

It is shown that user-preferring conflict resolution can coexist with a readable internal arbitration signal, and successful intervention depends on the geometry of the readout rather than probe accuracy alone, while directions selected mainly for pooled separability steer poorly.

Enrique Balp-Straffon, Chih-Hao Hsu, Rushiraj Gadhvi et al. · 0 citations
Preprint Aug 2026

When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs

SkillSafe-Bench is introduced, a controlled benchmark that scores skill-merged models on static refusal, adaptive jailbreak robustness, and capability retention under a conservative two-judge AND rule, and the static effect of merging is base-conditional.

Yu Ma, Hongli Shi, Jing Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents

KC-Bench is introduced, a controlled multi-turn benchmark for measuring model-level behavior across world-knowledge conflicts, input inconsistencies, and multi-source temporal conflicts, and provides a reproducible diagnostic for developing conflict-aware reasoning and execution safeguards.

Yaxing Lyu, Sheng-Jie Zhou, B. Toh et al. · 0 citations
Preprint Aug 2026

LLM within MCP Matters: Measuring Inefficient Resource Utilization Driven by LLMs

Per-server prompt engineering is therefore a workaround rather than a fix; it is argued that MCP host applications should provide an explicit mechanism that places server instructions ahead of tool selection in the client LLM's deliberation.

Minhan Cho, Soyoung Park, Kihyeon Jeong et al. · 0 citations
Conference Aug 2026

A Dynamic Evaluation Framework for LLM Instruction Following: Multi-dimensional Verification and Iterative Feedback

Accurately evaluating the instruction-following ability of Large Language Models (LLMs) is crucial for their practical deployment. Existing evaluation methods mainly rely on static assessment of single-pass generations, making it difficult to comprehensively measure instruction-following performance or models’ capabili...

Ya-Chao Fu, Qu-Fei Zhang, Meng-Na Zhu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SemVerBench: Benchmarking LLM Comprehension of Version-Constraint Resolution Semantics

SemVerBench is introduced, the first benchmark of LLM version-constraint resolution semantics across three ecosystems (npm, PEP 440, Cargo), and six frontier models are evaluated: Opus significantly outperforms all other models, and Sonnet outperforms the OpenAI models (McNemar).

Qi-Bai Chen, Ze-Ming Liu · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.