Skip to content

On Asymmetric Optimization of Reasoning and Perception in Vision-Language Model Post-Training

May 2026 · arXiv.org · Vol abs/2605.29496 · 0 citations · 48 references
Computer Science

TL;DR

A controlled diagnostic framework with two synthetic tasks that disentangle perception from reasoning reveals a consistent perception-reasoning asymmetry: post-training improves reasoning more substantially than perception, though the underlying mechanism differs across training paradigms.

Abstract

Post-training has greatly improved reasoning in frontier vision-language models, yet its gains for perception remain comparatively limited, creating a bottleneck for end-to-end visual reasoning. To investigate this gap, we introduce a controlled diagnostic framework with two synthetic tasks that disentangle perception from reasoning. Our analysis reveals a consistent perception-reasoning asymmetry: post-training improves reasoning more substantially than perception, though the underlying mechanism differs across training paradigms. For supervised fine-tuning (SFT), this asymmetry stems from token imbalance, with perception occupying a smaller fraction of tokens in chain-of-thought supervision. Reweighting the loss boosts end-to-end performance by up to 18.2 points. For reinforcement learning (RL), the asymmetry instead arises from reward coupling, as outcome rewards correlate more strongly with reasoning than perception. Adding a perception-aware reward improves end-to-end accuracy by up to 6.0 points; when ground-truth perception rewards are unavailable, a reliable surrogate provides useful signal, yielding gains of 2.2 points. Beyond the controlled setting, these strategies also improve real-world visual reasoning, with gains of up to 3.3 points across three benchmarks. Overall, we diagnose the causes of asymmetric optimization and provide actionable guidance that benefits both synthetic and realistic settings.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

UniCAR-RL: Seeing Better before Thinking Deeper in Visual Mathematics

Multimodal Large Language Models (MLLMs) often struggle with complex mathematical visual reasoning primarily due to a lack of fine-grained perception, causing initial visual hallucinations to directly trigger cascading reasoning failures. In traditional end-to-end reinforcement learning (RL), sparse rewards fail to dec...

Yu-Zhe Li, Hao Yan, Hao Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MM-OPD: Towards One More Bottleneck Between Perception and Reasoning

Recent multimodal large language models (MLLMs) advance visual reasoning by strengthening both perception and reasoning, implicitly assuming a process that transitions seamlessly from perception to reasoning. However, we observe a counterintuitive phenomenon that challenges this assumption: holding the model, question,...

Jin-Tao Tong, Yujing Lou, Zhan-Ming Shen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

TTRSD: Test-Time Reinforcement Learning with Self-Distillation for Vision-Language Models

TTRSD separates update direction from update position, determined by group-relative advantages, from update position, determined by visual sensitivity, without requiring ground-truth labels, external verifiers, or a separate teacher, demonstrating cross-dataset generalization while preserving inherent reasoning integri...

Shu-Ning Wang, Zhi-Heng Wu, Xun-Lan Zhou et al. · 0 citations
#machine learning Preprint Sep 2026

From Perception to Integration: Revisiting the Internal Dynamics of Reasoning in Vision-Language Models

Vision-language models (VLMs) can answer simple visual questions, but often struggle when one question requires several visual judgments. We study this gap with controlled tasks for feature binding, numerosity, spatial relations, and amodal completion, together with a Composite task that combines them. Matched counterf...

Rong-Yu Xu, Prayag Tiwari, Shao-Lei Zhang · 0 citations

Delayed Convergence and Emergent CoT Reliance in RL-Tuned Language Models

This analysis reveals that intermediate log-probability is an unreliable indicator for reasoning capability; instead, reasoning performance results from a shift of internal confidence allocation where RL fine-tuning delays internal convergence, exhibiting prolonged mid-layer exploration before converging sharply at the...

Pablo Pérez-Lázaro, Rocío Aznar-Gimeno, F. J. Lacueva-Pérez et al. · 0 citations
Preprint Sep 2026

PRIME: Perception Feedback with Situational Memory Embeddings in VLA Models

Current Vision-Language-Action (VLA) models for autonomous driving operate primarily through feedforward inference across the perception--reasoning--planning hierarchy. While modern architectures maintain temporal recurrence within the perceptual module, early perception remains blind to downstream reasoning and naviga...

Erik Deinzer, Naya Baslan, L. Paparusso et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Jun 3, 2026

MIT researchers teach AI models to interpret charts

The new ChartNet training dataset could improve the accuracy of vision-language models that help analyze business trends or interpret scientific figures.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.