Skip to content

Influence of Prompt Engineering on Small Language Models for Guarded Query Routing

Jul 2026 · arXiv.org · Vol abs/2607.24801 · 0 citations · 42 references
Computer Science

TL;DR

The results show that prompt optimization techniques enable SLMs to handle out-of-distribution queries gracefully, without changing the models'weights, while weaker models may still need weight-level adaptation or schema-aware training.

Abstract

We study the problem of guarded query routing, where we assume that a user query first meets a router that either determines the ideal endpoint for in-distribution queries or rejects out-of-distribution queries that are potentially unsafe or out of the system's scope. We investigate whether compact open-weight Small Language Models (SLMs) can jointly handle both tasks under latency constraints. We evaluate 22 models on GQR-Bench and score them with the harmonic mean of in-distribution and out-of-distribution accuracy. We find that mid-scale SLMs come close to frontier model routing quality at much lower latency. Still, many compact models fail because they do not reliably follow the required output format. However, our results show that prompt optimization techniques enable SLMs to handle such cases gracefully, without changing the models'weights. Moreover, few-shot prompt optimization raises Mistral 7B from 81.79 to 90.87 GQR-Score and lifts Qwen3.5 9B to 95.74, the best optimized score in our study and within 0.3 points of the strongest unoptimized larger model: Gemma 3 27B at 96.01. The bare DSPy signature, without in-context exemplars, is the most effective strategy for Granite 4 Tiny, raising its score from 54.29 to 83.05. These results show that prompt optimization is a useful first step for guarded query routing, while weaker models may still need weight-level adaptation or schema-aware training

View source

Similar papers

Preprint Aug 2026

EXPLAIN Yourself! Finding Query Planner Stalls Across DBMSes

It is found that although the queries triggering slow planning are largely DBMS-specific, recurring pathologies involving correlated subqueries, CTE expansion, repeated subquery expressions, disjunctive joins, and constant folding affect multiple systems.

Geoffrey X. Yu, Ryan Marcus, Tim Kraska · 0 citations
#small language model Preprint Aug 2026

Most of the LLM Routing Gap Is Task Type

This paper argues that a small win does not show that routing did anything, the authors' or anyone else's, and argues that a small win does not show that routing did anything, theirs or anyone else's.

Janghoon Lee · 1 citation

SEFRQO-Plus: A Self-Evolving Query Optimizer via Large Language Models and Retrieval-Augmented Generation

This paper designs a feedback-oriented vector database with query structure-aware embed-dings to support effective similarity search, and incorporates multi-layered histories as references to enrich the feedback, and enhances the prompt optimization work-flow by utilizing multi-dimensional historical feedback to drive...

Han-Wen Liu, Qihan Zhang, Ryan Marcus et al. · 0 citations
Preprint Aug 2026

EXCISE: Query-Side Exclusion for Late-Interaction Retrieval

Late-interaction retrievers handle exclusion queries poorly. When a user asks for X but not Z, the additive MaxSim score promotes documents covering Z, a problem we call exclusion inversion. We show that no readout of the frozen vectors recovers the constraint, because the difficulty lies in identifying the excluded to...

Mohammed Ali, Abdelrahman Abdallah, Adam Jatowt · 0 citations
Conference Open access 2026

A Survey on Context Injection Strategies for Long-Context Language Models: Three Perspectives

This survey argues that context injection strategy, rather than context capacity, is the defining research challenge for long-context LLM deployment, and proposes a three-axis analytical framework revealing that injection performance is jointly governed by selection, representation, and scheduling.

Aicha Dakir, M. El Hajji, Tarek Ait Baha et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.