Skip to content
Open access

Prompt Escalation for Lightweight Large Language Models: An Empirical Evaluation of Cost–Performance Trade-Offs

Sep 2026 · Applied Sciences · 0 citations · 18 references

Abstract

Prompt escalation can increase the resource requirements of lightweight large language models (LLMs) without improving predictive performance. We evaluated four instruction-tuned 2–4B models on Grade School Math 8K (GSM8K), CommonsenseQA (CSQA), Recognizing Textual Entailment (RTE), and the binary Stanford Sentiment Treebank (SST-2). Zero-shot (ZS), few-shot (FS), chain-of-thought (CoT), and few-shot CoT (FS+CoT) were compared across 64 conditions using predictive performance, tokens, latency, and stochastic response consistency. Demonstrations came from training splits; FS+CoT used worked rationales with automatic screening and a partial manual audit. Latency was measured separately with synchronization, warm-up exclusion, and counterbalanced prompt order. After Holm adjustment, 12.5% of predictive-performance contrasts were significant, and ZS was significantly outperformed in 1 of 48 contrasts, compared with significant differences in 100% of total-token and 75.0% of latency contrasts. Strategy rankings varied by model and task. A single auxiliary 7B model showed no significant accuracy gain over ZS but did not establish a general scale effect. Scenario-weight sensitivity frequently favored ZS, with model–task exceptions. These results support ZS as a low-overhead reference within the four evaluated benchmarks and the stated model, quantization, prompt, and decoding settings. The rankings and guidelines have not been validated for summarization, code generation, or multi-turn dialogue.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.