Skip to content

Anchoring What Matters: A Dual-Level Learning Framework for Visually-Grounded Multimodal Reasoning

Sep 2026 · 0 citations · 40 references
Computer Science

TL;DR

PIVOT introduces a self-calibrated experience replay mechanism, which selectively collects and replays visually-grounded historical experiences as stable reference anchors for policy optimization, and designs a vision-guided advantage allocation mechanism to allocate additional vision-aware advantages to tokens based on their local visual support and impact on downstream reasoning.

Abstract

Reinforcement learning with verifiable rewards (RLVR) has significantly improved the reasoning capabilities of large vision-language models (LVLMs). However, standard on-policy RLVR algorithms face a critical optimization bottleneck in preserving and reinforcing visually grounded reasoning behaviors: valuable visually-grounded reasoning trajectories are discarded after a single update, while uniform token advantage allocation prevents the model from reinforcing critical perception or reasoning steps. To bridge this gap, we propose PIVOT, a dual-level learning framework that anchors policy optimization around informative visual reasoning signals. Specifically, PIVOT introduces a self-calibrated experience replay mechanism, which selectively collects and replays visually-grounded historical experiences as stable reference anchors for policy optimization. Building upon this, we further design a vision-guided advantage allocation mechanism to allocate additional vision-aware advantages to tokens based on their local visual support and impact on downstream reasoning. Extensive experiments across diverse benchmarks demonstrate that PIVOT achieves highly competitive performance in enhancing the multimodal reasoning capabilities of LVLMs.

View source

Similar papers

Review Open access Aug 2026

From Multi-Agent Reinforcement Learning to Agentic AI: A Comprehensive Literature Review of Algorithmic Advances and Decision-Analytic Implications (2020-2025)

This literature review synthesizes 57 peer-reviewed and openly archived contributions published since 2019 into a thematic taxonomy spanning value-decomposition algorithms, trust-region and sequence-model policy methods, and LLM-based agentic frameworks, and discusses implications for applied decision analytics.

B. Rai, Milena Popović · 0 citations
Book Open access Aug 2026

Salient-Q: A Saliency-Guided Vision-Language Framework for Medical Image De-Identification

Salient-Q, a saliency-guided vision-language framework that addresses challenges of PHI identification and localization through three architectural innovations, outperforms other baseline models in terms of PHI identification and localization.

Sicheng Zhou, Zai-Fu Zhan, Lei Wu et al. · 0 citations
Conference Aug 2026

AI-Driven Human Behavior Analysis for Smart Surveillance Using Computer Vision and Explainable AI

The fast development of intelligent surveillance systems has enhanced the need to have intelligent and dependable analysis of human behaviour through computer vision methods. To overcome these obstacles, this paper presents an AI-enabled platform for analyzing human behavior in intelligent surveillance settings with th...

Anshu Vashisth, Gagandeep Kaur, Neha et al. · 0 citations
Open access Aug 2026

İnsan Resursları Sistemlərində İnformasiya Təhlükəsizliyi

Məqalədə iqtisadiyyatın qlobal rəqəmsal transformasiyası kontekstində kadr məlumatlarının təhlükəsizliyinin və məxfiliyinin təmin edilməsində müasir informasiya texnologiyalarının kritik rolu tədqiq olunur. Təşkilatların rəqəmsal insan resursları platformalarına keçidi və toplanan şəxsi məlumatların həcminin artması in...

Könül Tahirova · 0 citations
Open access Aug 2026

Strateji Qərar Qəbulu Sistemlərində Neyron Şəbəkə Əsaslı Data Fusion Modellərinin Multi-Laylı Arxitekturası və Təhlükəsizlik Aspektləri

Bu tədqiqat işində müasir dövlət və regional idarəetmə sistemlərində heterogen məlumatların sintezi üçün hazırlanmış “DataFusion Strateji Əməliyyat Sistemi” platformasının elmipraktiki əsasları təqdim olunur. Tədqiqatın mərkəzində dayanan Model Architecture v4.0 müxtəlif formatlı (SQL, NoSQL, coğrafi və mətn) verilənlə...

Rüstəm Şəfaqətov · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.