This work develops an alternative, retrieval-based account of human oversight and posit that error detection is more effective when oversight-relevant information is accessible to users at the moment of review, and shows that self-generated explanations improve error detection and strengthen recall of verification-relevant reasoning.
Abstract
Large language models (LLMs) are increasingly embedded in organizational work, yet their errors often pass human review. Prior research locates such failures in users'capability to review LLM output or their engagement in doing so. We develop an alternative, retrieval-based account of human oversight and posit that error detection is more effective when oversight-relevant information is accessible to users at the moment of review. Across two randomized lab-in-the-field experiments with 640 customer-facing employees, we show that self-generated explanations improve error detection and strengthen recall of verification-relevant reasoning, while cues that reactivate such reasoning help sustain detection under repeated LLM use. Theoretically, we identify information retrievability as a distinct precondition for effective oversight and specify generative encoding and cue-supported reactivation as mechanisms that build and sustain it. Practically, lightweight onboarding self-explanations and daily retrieval cues can make human oversight more resilient as LLM use becomes routine.
Organizations increasingly use oversight loops where one large language model (LLM) audits another's outputs alongside procedural traces of claimed steps. A common concern about such LLM-as-a-judge pipelines is that detailed traces make overseers gullible. Using signal detection theory, we audit five LLM overseers on 1...
It is shown how fact-checking, a generally desirable behavior, can interfere with belief tracking in LLMs and how suppressing this attention at decoding time recovers accuracy only partially and only in some models, calling for future work on intervention methods.
Quang Minh Nguyen, Luis Frentzen Salim· 0 citations
Improvements in large language model (LLM) capability can increase accuracy, reasoning, and access to relevant information while leaving unassessed whether a particular output is appropriately presented for epistemic uptake. This creates an evaluative gap: first-order capability metrics can register genuine gains while...
Qian Wu· Ethics and Information Techn...· 0 citations
Can LLMs reason through new information like humans, or do they merely retrieve cached opinions? This is critical for silicon sampling, where LLM personas simulate public opinion at scale. Current evaluations test only whether personas hold the right opinions -- a static snapshot. But opinion research increasingly depe...
This work proposes Sequential Epistemic and Action-Level Validation (SEAV), a verification-centric jailbreak evaluation framework that decomposes responses into ordered steps and evaluates both validity and correctness, and shows that enforcing correctness substantially reshapes measured robustness.
Qilong Wu, Sahil Wadhwa, Pranab Mohanty et al.· 0 citations
We investigate how LLM-mediated explanations can preserve evidence trails, support orientation, and reduce interpretation effort under cognitive and temporal constraints. While LLMs can make XAI artifacts accessible, fluent summaries may obscure provenance and hinder verification. We introduce the XAI-Seeking Principle...
Valentin Grimm, J. Rubart, Eelco Herder et al.· Proceedings of the 37th ACM...· 0 citations
Known for his clear and elegant writing style, Bertsekas shaped fields from control and optimization to large-scale computation and artificial intelligence.
MIT News · Artificial Intelligence· news.mit.eduJul 7, 2026
The professor of physics and inaugural director of the NSF AI Institute for Artificial Intelligence and Fundamental Interactions will lead LNS and continue his research in particle physics.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.