Skip to content

Overreliance on AI in Information-seeking from Video Content

Mar 2026 · arXiv.org · Vol abs/2603.19843 · 2 citations · 99 references
Computer Science

TL;DR

It is found that AI assistance increases accuracy by 3-7% when participants viewed the relevant video segment, and by 27-35% when they did not, while efficiency increases by 10% for short videos and 25% for longer ones, and self-reported confidence in answers remains stable across all three conditions.

Abstract

The ubiquity of multimedia content is reshaping online information spaces, particularly in social media environments. At the same time, search is being rapidly transformed by generative AI, with large language models (LLMs) routinely deployed as intermediaries between users and multimedia content to retrieve and summarize information. Despite their growing influence, the impact of LLM inaccuracies and potential vulnerabilities on multimedia information-seeking tasks remains largely unexplored. We investigate how generative AI affects accuracy, efficiency, and confidence in information retrieval from videos. We conduct an experiment with around 900 participants on 8,000+ video-based information-seeking tasks, comparing behavior across three conditions: (1) access to videos only, (2) access to videos with LLM-based AI assistance, and (3) access to videos with a deceiving AI assistant designed to provide false answers. We find that AI assistance increases accuracy by 3-7% when participants viewed the relevant video segment, and by 27-35% when they did not. Efficiency increases by 10% for short videos and 25% for longer ones. However, participants tend to over-rely on AI outputs, resulting in accuracy drops of up to 32% when interacting with the deceiving AI. Alarmingly, self-reported confidence in answers remains stable across all three conditions. Our findings expose fundamental safety risks in AI-mediated video information retrieval.

View source

Similar papers

Preprint Sep 2026

MultiVENT-Raw: A Benchmark for Retrieval and Reasoning over Raw Videos

MultiVENT-Raw is released, a multilingual collection of nearly 120,000 primarily raw videos paired with 130 events and 222 event-centric queries, along with human-annotated video relevance judgments and human-extracted key facts for relevant videos.

Reno Kriz, David Etter, Alexander Martin et al. · 0 citations
Preprint Aug 2026

Auditing Exposure to Harmful Content on TikTok using Multimodal Language Models: A Cross-National, Age-Stratified Study

Online video platforms can expose young users to harmful content, but independent audits remain difficult because video annotation is costly and moderation judgments vary across languages. We audit TikTok in France, Italy, and Sweden with sockpuppet accounts representing four age personas (13, 16, 19, 40), collecting 3...

Hamidreza Saffari, F. Pierri · 0 citations
Review Open access Aug 2026

ChatGPT on TikTok: Multimodal Representations of Artificial Intelligence and Engagement Patterns

This research examined how public communication about emerging technologies is shaped through platform-specific multimodal conventions and engagement dynamics on TikTok to provide insights for stakeholders who seek to promote responsible AI communication on video-sharing platforms where algorithms shape which content i...

Xi-Ran Liu, M. S. Schäfer · 1 citation

From Web(logs) to Web(AI): Questions, Platforms, and Methods across Twenty Editions of ICWSM

Over twenty editions, the ICWSM community has examined social life online as platforms, interactions, and research methods have changed. What can this body of research tell us at this critical juncture, as AI increasingly reshapes how people communicate online? We analyzed 2,139 indexed contributions from 2007 to 2026,...

Koustuv Saha, Eshwar Chandrasekharan · 0 citations
Open access Sep 2026

The Readability of Generative AI Policies: An Empirical Analysis of OpenAI’s Informational Documents

In recent years, media attention has focused on artificial intelligence, particularly on chatbot services and generative intelligence. ChatGPT, created by OpenAI, was one of the earliest online tools and rapidly gained popularity. Users are indeed exposed to a service with privacy notifications and conditions of use th...

Jacopo Bassetta, D. Perpetuini, Maria Teresa Giusti et al. · 0 citations
#natural language process... Preprint Aug 2026

WildSEEK: Evaluating Language Models for Information-Seeking

This work introduces WildSEEK, a manually annotated dataset of 3k information-seeking queries from real user interactions, and an evaluation framework for LLM-generated responses, and finds that over a third of information-seeking queries are high-risk and more often analytical.

Tanise Ceron, Joachim Baumann, Elisa Bassignana et al. · 0 citations

Related blog posts

Microsoft Research Blog Jul 8, 2026

Flint: A visualization language for the AI era

Short chart specifications are easy to write, but often produce uninspiring results. Flint is an open-source visualization language that offers a middle path, letting AI agents create expressive charts from compact, human-editable specifications. The post Flint: A visualization language for the AI era appeared first on Microsoft Research.

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.