Skip to content
Open access

Response Stability of Large Language Models Under Meaning-Preserving Prompt Variation

2026 · International journal of research and innovation in social science · 0 citations

Abstract

Single-prompt assessments provide limited evidence about whether task-relevant response properties remain stable when the same communicative intention is reformulated. This study examined response stability under controlled, meaning-preserving English prompt variation using 648 responses from 12 researcher-selected argumentative policy and education topics, six prompt variants, three accessed consumer systems, and three repeated runs. The analysis combined response-length screening, exact-duplication and cross-topic reuse detection, latent semantic analysis (LSA), structural instruction checks, bootstrap confidence intervals, and matched non-parametric comparisons. Under a fixed 100-dimensional LSA configuration, aggregate semantic stability was .975 for ChatGPT (95% bootstrap CI [.965, .985]), .968 for Claude [.948, .985], and .860 for Gemini [.791, .924]. ChatGPT and Claude showed comparably high semantic stability, whereas the accessed Gemini configuration was lower and more variable. All three run-specific Friedman tests yielded χ²(2) = 24.000, p = 6.14 × 10⁻⁶, Kendall’s W = 1.000; this value indicates consistent within-run rank ordering across the 12 topics, not model superiority, and includes ties produced by exact repetition. ChatGPT and Claude met the experimental 180–250-word instruction in 100% of responses, compared with 33.3% for Gemini. Repetition complicated interpretation: Claude returned byte-identical six-variant blocks in 66.7% of run-topic blocks and Gemini in 33.3%, while Gemini Run 2 produced only 25 distinct texts across 72 responses without any complete block. LSA estimates also varied with representation settings, so model ordering based on small semantic-score differences was not robust. Only three of nine runs retained version and tier information, those documented configurations used unequal access tiers, and other session and inference conditions could not be reconstructed. The findings therefore describe the accessed consumer configurations under this task design, not controlled vendor-level or model-family differences. The study supports a multidimensional operational profile that reports semantic preservation, instruction following, topical relevance, structural adaptation, and repetition as related but distinct dimensions.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.