Skip to content
Book Open access

Beyond Top-e: Simulation-Based Interactive Evaluation for Query Suggestions

Jul 2026 · Annual International ACM SIGIR Conference on Research and Development in Information Retrieval · pp. 3167-3174 · 0 citations · 32 references
Computer Science

TL;DR

This work develops and validate an LLM-based selection model, systematically analysing how varying levels of intent information and selection strategies affect its ability to approximate human selection behaviour, and develops and validate an LLM-based selector that is used to benchmark multiple query suggestion systems across diverse datasets under interactive conditions.

Abstract

Evaluating query suggestion systems in a manner that reflects real-world query formulation remains a persistent challenge. Most offline methodologies adopt static assumptions, such as users accepting all or the top-e suggestions, ignoring the inherently selective and intent-driven nature of interactive search. While online experiments provide realistic behavioural signals, they are costly, difficult to scale, and often irreproducible. To bridge this gap, we introduce SIQSE (Simulation-based Interactive Query Suggestion Evaluation), a framework that models query reformulation as an interactive selection task performed by a simulated user. In SIQSE, a Large Language Model (LLM) acts as a surrogate user that progressively selects suggestions according to contextual relevance and explicit search intent. Unlike static offline protocols, this simulation captures the iterative and selective dynamics of real query formulation. Our contributions are twofold. First, we develop and validate an LLM-based selection model, systematically analysing how varying levels of intent information and selection strategies affect its ability to approximate human selection behaviour. Second, we employ this selector to benchmark multiple query suggestion systems across diverse datasets under interactive conditions. Importantly, while the selector is LLM-based, the final evaluation is computed exclusively through ranking-based effectiveness metrics over the rankings produced by selected expansions, ensuring that system performance reflects retrieval quality rather than alignment with the surrogate user model. By modelling round-based interaction while maintaining metric independence, SIQSE offers a scalable, reproducible evaluation paradigm that brings offline assessment closer to the complexity of real-world search behaviour. To facilitate adoption and reproducibility, we release SIQSE as an open-source Python library.

Read PDF

Similar papers

Open access Aug 2026

An Analytical Study of Behavior-Aware Retrieval-Augmented Generation Frameworks in Enterprise Software Ecosystems for Optimizing User Navigation and Decision Support

It is demonstrated that incorporating real-time user behavioral context is critical for transforming generative AI utilities into proactive workflow accelerators within complex, data-dense corporate environments.

Sri Charan Chowdary Konidina · 0 citations
Preprint Sep 2026

Can Agents Win the Video Browser Showdown?

The results show that modern agents can autonomously operate interactive video retrieval systems to solve many search tasks from an initial intent description, achieving performance competitive with strong historical expert-operated systems in several settings.

Bastian Jäckl, Zuzana Vopálková, Daniel A. Keim et al. · 0 citations
#artificial intelligence Preprint Open access Aug 2026

Query Implied Generative Engine Optimization

Query Implied Generative Engine Optimization (QI-GEO) is proposed, which approximates document's intent space and identifies content that may be missing yet relevant to answer potential user queries and suggests that document-derived approximations of user intents can improve visibility without relying on explicit quer...

S. Ramakrishna, William B. Andreopoulos · 0 citations
Aug 2026

FilterPilot: An Interactive Assistant for Adapting Filtering Predicate to Table Content

FilterPilot, an LLM-powered interactive assistant designed to adapt filtering predicates to table content, employs a novel iterative recall-then-verify paradigm, combining LLM-based query reformulation with table-value feedback to dynamically expand search terms.

Shi-Wen Wu, Yu-Qi Wang, Yue Pang et al. · 0 citations
Book Aug 2026

SPEAR: Selection-aware Personalized End-to-end Adaptive Rewriting and Retrieval for Community Search

Query reformulation bridges user intent and retrieval in e-commerce search, yet production systems optimize rewrite quality and retrieval effectiveness separately, leaving the two stages structurally misaligned. Path-based architectures unify them end-to-end but were designed for personalization, where relevance is not...

Wen-Bin Wu, Yu-Zhong Wu, Yu-Fan Xu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.