This work introduces EgoCITE (Egocentric Context-augmented Indexing and Time-aware Evidence retrieval), a long-horizon agentic memory framework for egocentric QA that improves accuracy over agentic memory baselines by at least 4.4--14.2% while achieving 36$\times$ lower cost than long-context LLM agents.
Abstract
Long-horizon egocentric memory transforms continuous first-person video and audio into a searchable record of past experiences. We demonstrate two bottlenecks in existing systems: indices built from context-poor captions are unreliable for agentic search, while retrieval ignores a question's temporal intent. To address both bottlenecks, we introduce EgoCITE (Egocentric Context-augmented Indexing and Time-aware Evidence retrieval), a long-horizon agentic memory framework for egocentric QA. EgoCITE comprises three components. EgoScheme uses local multimodal context to turn fragmentary video captions and speech transcripts into self-contained atomic memory indices. EgoIndex organizes complementary action, activity, utterance, and conversation representations into searchable multi-view memory indices at multiple granularities. EgoRetrv combines semantic search with question-conditioned temporal relevance scoring and curation of retrieved evidence. We evaluate EgoCITE on EgoLifeQA, EgoMem, and EgoR1-Bench in terms of answer accuracy and target-event retrieval alignment. EgoCITE improves accuracy over agentic memory baselines by at least 4.4--14.2% while achieving 36$\times$ lower cost than long-context LLM agents.
Imprint is proposed, an interaction-centric memory framework that formulates long-horizon egocentric memory as an online memory compression problem rather than summarization, and demonstrates that memory compression provides a scalable and retrieval-effective foundation for long-horizon egocentric question answering.
Evaluation of state-of-the-art open-source and closed-source MLLMs reveals that current MLLMs function as lossy summarizers rather than faithful memorizers, highlighting the need for architectures with genuine long-term spatiotemporal memory.
Weitao Chen, Jiaxing Hu, Xie Tianyidan et al.· 0 citations
TrajWiki is proposed, a trajectory-based memory framework for long-horizon conversational agents that improves long-horizon dialogue performance across both open-source and closed-source LLM backbones, while providing greater interpretability and diagnostic visibility into memory evolution, retrieval failures, and answer generation.
Jingyu Sun, Yuyang Xue, Mingyang Li et al.· 0 citations
This paper presents AdaMM, a framework that jointly supports retrieval and analytic memory that extracts provenance-linked attribute-value observations from dialogue, images, and contextual metadata, discovers recurring field structures, and materializes them for analytical access.
Zhoujin Tian, Hao Zhang, Yao Tian et al.· 0 citations
This work proposes ViSAGE, a multimodal agentic memory framework that constructs self-correcting, entity-centric memories and applies bidirectional memory refinement to propagate delayed identity evidence, retroactively unifying historical records and improving future reasoning.
Xinkui Zhao, Enbo Chen, Yifan Zhang et al.· 0 citations
LongEgoRefer is introduced, a novel and challenging benchmark constructed from long-form videos in the Ego4D dataset that defines a demanding spatio-temporal grounding problem that requires models to identify both when an event occurs and where the referred object appears within extended video sequences.
Shunya Kato, Taiki Miyanishi, Shuhei Kurita et al.· 0 citations
Known for his clear and elegant writing style, Bertsekas shaped fields from control and optimization to large-scale computation and artificial intelligence.