Skip to content
Preprint

Can Agents Win the Video Browser Showdown?

Sep 2026 · 0 citations · 40 references
Computer Science

TL;DR

The results show that modern agents can autonomously operate interactive video retrieval systems to solve many search tasks from an initial intent description, achieving performance competitive with strong historical expert-operated systems in several settings.

Abstract

Searching large video collections is typically an interactive process in which users play two roles. First, they hold the search intent: the underlying goal that determines what content they seek and why. Second, users must operationalize this intent through an iterative search loop. Users translate their intent into queries, browse the retrieved candidates, and refine their queries based on the results. In this paper, we investigate the capabilities of modern Vision Language Models (VLM) and agentic approaches to reach search goals interactively and fully autonomously. Specifically, we study whether a provided initial specification of a search goal might be sufficient to solve traditionally interactive search tasks with an agentic system. Provided that the involved VLMs are not aware of the whole large video dataset in advance, the key challenge lies in the effective combination of an existing interactive video search system and a smart VLM agent controlling the system. While the search system provides indexing and efficient querying, the VLM-based agents analyze top-ranked items and make decisions about next actions. Our results show that modern agents can autonomously operate interactive video retrieval systems to solve many search tasks from an initial intent description, achieving performance competitive with strong historical expert-operated systems in several settings.

View source

Similar papers

Aug 2026

A Demo of Interactive Thematic Data Collection on the Live Web

Many emerging AI systems can search and reason over web content, but most are designed for task-bounded objectives such as answering a question or producing a report. In contrast, many real-world workflows require thematic web-data collection: identifying and gathering large numbers of topically relevant documents dist...

Michael West, Eduard Dragut · 0 citations
Preprint Aug 2026

ITER: Interaction-Aware Retrieval for Agentic Search

ITER, an agent interaction-aware dense retriever trained using agent trajectory learning signals, is introduced andlations show that structured interaction history and pre-search reasoning provide complementary retrieval context, while previously visited and useful documents provide the strongest trajectory-relative su...

Hao-Dong Chen, Shuai Wang, Yu Yin et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Online Video Agent Harness for Long Video Understanding

Long video understanding often behaves like a visual needle-in-a-haystack problem: query-relevant evidence is sparsely distributed across long temporal spans, while packing dense frames into a single VLM context incurs \textit{context rot} and high cost. Existing video agents often rely on query-agnostic offline prepro...

Sen Yang, Bo-Qiang Duan, Jing Yang et al. · 1 citation
Review

Lab Overview

The overall objectives of the CLEF 2017 Dynamic Search Lab are described, the resources created for the pilot task, the resources created for the pilot task and the evaluation methodology adopted.

E. Kanoulas, Leif Azzopardi · 1 citation
Preprint Aug 2026

CAP: A Scalable Benchmark for Evaluating Cross-Site Browser Agents with Complex Actions and Perception

CAP is introduced, a scalable benchmark for evaluating browser agents on cross-site, human-like web tasks that require non-trivial UI interactions and visual understanding and a decomposition-and-recomposition pipeline that first abstracts each website into a structured site card capturing user-facing functions, comple...

Zejun Xu, Taiyi Chen, Jin Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge

Existing computer-use agent benchmarks do not fully evaluate agents acting as assistants. A useful assistant retrieves information across complex, multi-step workflows, synthesizes it into artifacts (documents, presentations, spreadsheets), and navigates program interfaces to produce a coherent final product. Such work...

Alexander Gill, Md Farhan Ishmam, X. Nguyen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.