Skip to content

SD-MAR: Multi-image Analytical Reasoning via Synthetic Data and Reinforcement Learning

Jul 2026 · arXiv.org · Vol abs/2607.14333 · 0 citations · 35 references
Computer Science

TL;DR

SD-MAR (Synthetic Data for Multi-image Analytical Reasoning) is introduced, a framework for training and evaluating VLMs on multi-image analytical reasoning that constructs paired visual scenarios through controlled perturbations and generates reasoning tasks spanning semantic change attribution and quantitative comparison.

Abstract

Vision Language Models (VLMs) demonstrate strong perceptual abilities but remain limited in tasks requiring analytical reasoning across multiple visual states, such as multi-image comparison, change detection, and multi-step visual inference. These capabilities are critical for real-world multimodal applications where reasoning must be grounded in systematic differences between visual contexts. However, existing benchmarks rarely require both explicit visual comparison and analytical reasoning, leaving this capability underexplored. To address this gap, we introduce SD-MAR (Synthetic Data for Multi-image Analytical Reasoning), a framework for training and evaluating VLMs on multi-image analytical reasoning. SD-MAR constructs paired visual scenarios through controlled perturbations and generates reasoning tasks spanning semantic change attribution and quantitative comparison. We further train VLMs using GRPO-lite with Backward Discounted Allocation (BDA), a reinforcement learning approach that removes KL regularization to encourage stronger policy optimization while allocating greater credit to the later reasoning steps where analytical conclusions are formed. Experiments on Qwen2.5-VL-7B and InternVL3-8B show that GRPO-lite fine-tuning on SD-MAR improves in-domain accuracy by up to 36.95%, with Qwen2.5-VL-7B outperforming GPT-4.1 on the SD-MAR benchmark. Importantly, out-of-domain generalization is preserved or improved: performance remains within 1% on MME, MMMU-Pro, and MathVista, while improving by up to 4% on MMBench. LLM-as-judge evaluation further demonstrates consistent improvements in logical coherence and explanation quality across both models.

View source

Similar papers

Jul 2026

Be Consistent! Enhancing Robust Visual Reasoning in LVLMs with Consistency Constraints

ConVLM is presented, which improves LVLM reasoning through Group Relative Policy Optimization-based reinforcement learning with a novel consistency reward, and ConVBench, a complex vision-centric reasoning benchmark in which each image is paired with two logically equivalent questions across six categories.

Li-Qiang Jing, Xiong Zhou, Siddharth Varia et al. · 1 citation · ⚡1
Jul 2026

MIRROR: Learning from the Other View for Multi-Modal Reasoning

Modality-Informed Reciprocal Reasoning Optimization (MIRROR), a reinforcement learning approach for improving multimodal reasoning via self supervision, is developed and improves over standard RL and yields more accurate and consistent behavior across modalities.

Wen Ye, Yuxiao Qu, Aviral Kumar et al. · 0 citations
Preprint Aug 2026

TRAM: Enhancing Multimodal Reasoning with Trajectory-Derived Auxiliary Memory

TRAM (TRajectory-derived Auxiliary Memory), a training-free method that augments standard decoding with an auxiliary memory pathway derived from the model's own reasoning trajectory, shows that TRAM improves performance over vanilla decoding on mathematical, scientific, and general visual reasoning tasks without additi...

Kang Liu, Zi-Jing Wang, Yongkang Liu et al. · 0 citations
Preprint Aug 2026

SCOUT: Unlocking Enhanced Spatial Reasoning via Structured Chain-of-Thought and Multi-Objective Process Reward

This work designs a structured Chain-of-Thought (CoT) framework that explicitly models 3D environmental perception to ensure robust spatial understanding and reasoning and introduces a novel RL algorithm featuring multi-objective process rewards and a tailored advantage estimation method, facilitating fine-grained cred...

Zile Zhou, Huining Yuan, Weichen Zhang et al. · 0 citations
Preprint Aug 2026

Video-FLAIR: Not Whether to Reason, But How

Video-FLAIR is introduced, a training framework that learns to select the appropriate reasoning mode for each query using reinforcement learning, and yields a supervision signal for learning adaptive reasoning without per-query annotations.

Yogesh Kulkarni, Pooyan Fazli · 0 citations
Preprint Aug 2026

ChronoVision: Temporal Reasoning via Latent State Reconstruction

This work proposes ChronoVision, a multimodal framework designed to align visual logic with latent imagery, and introduces Vbvr-VQA, a novel dataset that evaluates temporal tracking by reformulating video reasoning into a strict image-ordering task.

Yi-Fan Shen, Jian Xu, Boyi Li et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.