The results show that modern agents can autonomously operate interactive video retrieval systems to solve many search tasks from an initial intent description, achieving performance competitive with strong historical expert-operated systems in several settings.
Abstract
Searching large video collections is typically an interactive process in which users play two roles. First, they hold the search intent: the underlying goal that determines what content they seek and why. Second, users must operationalize this intent through an iterative search loop. Users translate their intent into queries, browse the retrieved candidates, and refine their queries based on the results. In this paper, we investigate the capabilities of modern Vision Language Models (VLM) and agentic approaches to reach search goals interactively and fully autonomously. Specifically, we study whether a provided initial specification of a search goal might be sufficient to solve traditionally interactive search tasks with an agentic system. Provided that the involved VLMs are not aware of the whole large video dataset in advance, the key challenge lies in the effective combination of an existing interactive video search system and a smart VLM agent controlling the system. While the search system provides indexing and efficient querying, the VLM-based agents analyze top-ranked items and make decisions about next actions. Our results show that modern agents can autonomously operate interactive video retrieval systems to solve many search tasks from an initial intent description, achieving performance competitive with strong historical expert-operated systems in several settings.
Many emerging AI systems can search and reason over web content, but most are designed for task-bounded objectives such as answering a question or producing a report. In contrast, many real-world workflows require thematic web-data collection: identifying and gathering large numbers of topically relevant documents dist...
Michael West, Eduard Dragut· Proceedings of the VLDB Endo...· 0 citations
ITER, an agent interaction-aware dense retriever trained using agent trajectory learning signals, is introduced andlations show that structured interaction history and pre-search reasoning provide complementary retrieval context, while previously visited and useful documents provide the strongest trajectory-relative su...
Hao-Dong Chen, Shuai Wang, Yu Yin et al.· 0 citations
Long video understanding often behaves like a visual needle-in-a-haystack problem: query-relevant evidence is sparsely distributed across long temporal spans, while packing dense frames into a single VLM context incurs \textit{context rot} and high cost. Existing video agents often rely on query-agnostic offline prepro...
Sen Yang, Bo-Qiang Duan, Jing Yang et al.· 1 citation
The overall objectives of the CLEF 2017 Dynamic Search Lab are described, the resources created for the pilot task, the resources created for the pilot task and the evaluation methodology adopted.
CAP is introduced, a scalable benchmark for evaluating browser agents on cross-site, human-like web tasks that require non-trivial UI interactions and visual understanding and a decomposition-and-recomposition pipeline that first abstracts each website into a structured site card capturing user-facing functions, comple...
Existing computer-use agent benchmarks do not fully evaluate agents acting as assistants. A useful assistant retrieves information across complex, multi-step workflows, synthesizes it into artifacts (documents, presentations, spreadsheets), and navigates program interfaces to produce a coherent final product. Such work...
Alexander Gill, Md Farhan Ishmam, X. Nguyen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.