Prompt Sensitivity or Evaluation Artifact? A Task-Aware Analysis for Large Language Models
Abstract
Prompt sensitivity is widely treated as a model robustness deficiency. Yet the extent to which measured sensitivity reflects genuine model instability, rather than artifacts of the evaluation method used to measure it, remains largely underexplored. We introduce Evaluation-Attributable Sensitivity (EAS), a per-instance metric that quantifies the disagreement between heuristic-based and LLM-judge-based sensitivity scores computed across a controlled prompt design, together with Signed EAS for directional analysis. Using the framework we evaluate nine models spanning five families, Llama-3, Qwen-2.5, Mistral, Gemma and OpenAI across ARC-Challenge, BoolQ, and SQuAD with 200-items per dataset, yielding 43,200 prompt-response-judge triples. We introduce a four-zone taxonomy that classifies instances as Artifact, Genuine, Underdetected or Stable. Across the evaluated models and datasets, the direction of EAS was more strongly associated with task format than with model family or parameter scale: token-F1 overestimates open-ended sensitivity on SQuAD, whereas exact-match heuristics underestimate sensitivity on ARC-Challenge and BoolQ. On SQuAD, token-F1 accuracy falls up to 49% below judge accuracy. Format directives displayed the only prompt factor with consistent measurable effects, though their direction varies across model families. These findings position prompt sensitivity as a joint property of model behavior, task format, and evaluation regime, motivating EAS as a reusable diagnostic for metric-induced variability in LLM evaluation. The official implementation and experimental results are available at https://github.com/sayumimuthu/llm-prompt-sensitivity-evaluation-artifact.git.