Skip to content
Preprint

Remember-R1: Mitigating Long-Context Visual Forgetting through Reinforcement Learning

Aug 2026 · 0 citations · 30 references
Computer Science

TL;DR

Remember-R1 is proposed, a reinforcement learning framework that mitigates long-context visual forgetting by applying process-level supervision directly on the original reasoning trajectory, demonstrating its effectiveness in mitigating long-context visual forgetting.

Abstract

Multimodal large language models (MLLMs) increasingly rely on long chain-of-thought reasoning for complex tasks. However, as reasoning sequences lengthen, models may gradually rely less on visual evidence and more on accumulated textual context, leading to visual forgetting. Existing approaches do not directly constrain how visual evidence is used and maintained along the original reasoning trajectory, leaving long-context visual forgetting insufficiently addressed. To address this issue, we propose Remember-R1, a reinforcement learning framework that mitigates long-context visual forgetting by applying process-level supervision directly on the original reasoning trajectory. Specifically, Remember-R1 introduces rewards that encourage broader coverage of matched visual keywords, stronger persistence of visual dependence in later reasoning steps, and greater focus on question-relevant image regions. Experiments across multiple model scales and diverse multimodal benchmarks demonstrate that Remember-R1 consistently improves reasoning performance. Additional analyses further show that it slows the decline of visual attention during generation, supporting its effectiveness in mitigating long-context visual forgetting.

View source

Similar papers

Preprint Aug 2026

TRAM: Enhancing Multimodal Reasoning with Trajectory-Derived Auxiliary Memory

TRAM (TRajectory-derived Auxiliary Memory), a training-free method that augments standard decoding with an auxiliary memory pathway derived from the model's own reasoning trajectory, shows that TRAM improves performance over vanilla decoding on mathematical, scientific, and general visual reasoning tasks without additional training.

Kang Liu, Zijing Wang, Yongkang Liu et al. · 0 citations
Preprint Aug 2026

Balancing Efficiency and Efficacy: Training-Free Attention-Guided Switching Between Explicit and Latent Thoughts for MLLMs

A novel training-free inference strategy for MLLMs that explicitly decouples perception and reasoning is presented, and a novel metric, the vision-to-text attention ratio, is proposed, to dynamically gauge the model's cognitive focus.

Haoqiang Kang, Liupeng Li, Kuofeng Gao et al. · 0 citations
Book Open access Aug 2026

Learning to Forget: Emotional Salience as a Compression Mechanism for Long-Term AI Memory

Recent advancements in Large Language Model (LLM) agents have largely focused on extending context windows or implementing massive Retrieval-Augmented Generation (RAG) systems to retain long-term history. However, this store-everything approach causes high computational costs and digital hoarding, paradoxically leading to digital amnesia where key emotional contexts are buried under trivial data. To challenge this paradigm, we introduce the Affective Memory Architecture, drawing from the amygdala's role in memory modulation to equip AI agents with the essential capacity to actively forget. Unlike static summarization, our framework structures multimodal inputs into an Affective Scene Graph (ASG) and dynamically adjusts the memory resolution based on emotional salience. High-arousal core memories are preserved in rich, high-resolution episodic detail; low-salience routines are aggressively downsampled using novel Optical Context Compression to minimal vision tokens; and frequently reactivated patterns are consolidated into crystallized semantic insights. Through quantitative proof-of-concept modeling, we demonstrate that systematically managing the trivial not only resolves the digital hoarding problem but actively reduces proactive interference, enhancing overall recall clarity. Ultimately, this work offers a scalable, privacy-friendly blueprint for resource-efficient AI capable of evolving with users over time, fundamentally shifting the goal of AI memory from total recall to meaningful retention.

SoYeop Yoo, Sunghoon Im · 0 citations
Preprint Aug 2026

Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning

Multimodal large language models increasingly use visual chain-of-thought (Visual CoT) to reason about spatial, temporal, and embodied environments. By generating intermediate reasoning images, Visual CoT provides an intuitive mechanism for visual foresight but introduces substantial inference overhead, which is particularly problematic for proactive video reasoning. We ask whether models can learn to think visually during training while reasoning directly at inference. We introduce Internalized Visual Thinking (IVT), a post-training framework that jointly optimizes textual prediction and next-embedding prediction over unlabeled videos. Given a partially observed video, IVT predicts latent representations of future frames together with the target textual answer, encouraging the model to capture motion, object transitions, interactions, and latent intent. At inference, IVT generates the answer directly without synthesizing or re-encoding future frames. We conduct controlled studies across target representations, decoder designs, prediction horizons, data mixtures, training curricula, and predictive objectives. IVT improves over direct-answer fine-tuning on all six evaluation settings while retaining the same inference pathway. Compared with explicit Visual CoT, IVT achieves comparable or better performance and reduces average end-to-end latency by more than 5x. Together, our findings suggest that explicit pixel-space generation at inference time, as used in visual chain-of-thought, may not be necessary for effective proactive video reasoning. Predictive world modeling can be internalized during training to produce multimodal reasoners that are both more accurate and substantially more efficient.

Xiaoyu Zhu, Xinke Deng, Suresh Taddewadikar et al. · 0 citations
Preprint Aug 2026

ChronoVision: Temporal Reasoning via Latent State Reconstruction

This work proposes ChronoVision, a multimodal framework designed to align visual logic with latent imagery, and introduces Vbvr-VQA, a novel dataset that evaluates temporal tracking by reformulating video reasoning into a strict image-ordering task.

Yifan Shen, Jian Xu, Boyi Li et al. · 1 citation
Preprint Aug 2026

GLaQ: Grounding Latent Queries in Visual Evidence for Multimodal Reasoning

GLaQ, a grounded latent-query framework that replaces sequential latent rollout with a fixed set of context-conditioned queries grounded in the original visual tokens, is introduced, suggesting that direct query-to-image grounding can recover localized evidence from the full image without external visual operations or autoregressive latent rollouts.

Ze-Sheng Yang, Lingling Zhang, Xinyu Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.