Skip to content
Conference

Prompt Sensitivity or Evaluation Artifact? A Task-Aware Analysis for Large Language Models

Aug 2026 · Moratuwa Engineering Research Conference · pp. 1082-1087 · 0 citations · 22 references

Abstract

Prompt sensitivity is widely treated as a model robustness deficiency. Yet the extent to which measured sensitivity reflects genuine model instability, rather than artifacts of the evaluation method used to measure it, remains largely underexplored. We introduce Evaluation-Attributable Sensitivity (EAS), a per-instance metric that quantifies the disagreement between heuristic-based and LLM-judge-based sensitivity scores computed across a controlled prompt design, together with Signed EAS for directional analysis. Using the framework we evaluate nine models spanning five families, Llama-3, Qwen-2.5, Mistral, Gemma and OpenAI across ARC-Challenge, BoolQ, and SQuAD with 200-items per dataset, yielding 43,200 prompt-response-judge triples. We introduce a four-zone taxonomy that classifies instances as Artifact, Genuine, Underdetected or Stable. Across the evaluated models and datasets, the direction of EAS was more strongly associated with task format than with model family or parameter scale: token-F1 overestimates open-ended sensitivity on SQuAD, whereas exact-match heuristics underestimate sensitivity on ARC-Challenge and BoolQ. On SQuAD, token-F1 accuracy falls up to 49% below judge accuracy. Format directives displayed the only prompt factor with consistent measurable effects, though their direction varies across model families. These findings position prompt sensitivity as a joint property of model behavior, task format, and evaluation regime, motivating EAS as a reusable diagnostic for metric-induced variability in LLM evaluation. The official implementation and experimental results are available at https://github.com/sayumimuthu/llm-prompt-sensitivity-evaluation-artifact.git.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.