Skip to content

Filling Before Advancing: Capability-Gap-Driven Post-Training for Scenario-Specialized Remote Sensing MLLMs

Jul 2026 · arXiv.org · Vol abs/2607.22205 · 0 citations · 45 references
Computer Science

TL;DR

This work instantiates FBA for coastal harbor understanding, a representative multi-source scenario, by constructing CPRS (Coastal-Port Remote Sensing), a three-layer supervision dataset coupled with three ordered stages: domain-bridge convergence for shared RS priors across target and bridging scenarios under different modalities, and evidence-grounded scenario tuning for downstream performance.

Abstract

Remote sensing multimodal large language models (RS-MLLMs) have improved general aerial-image understanding. However, Earth observation applications require fine-grained scenario specialization, constrained by scarce high-quality scenario data and incomplete capability coverage. We formulate this adaptation as a capability-gap-driven post-training problem and propose filling before advancing (FBA). Rather than relying on single-stage supervised fine-tuning (SFT) over target-domain samples, FBA first fills prerequisite capability gaps before advancing toward scenario specialization. We instantiate FBA for coastal harbor understanding, a representative multi-source scenario, by constructing CPRS (Coastal-Port Remote Sensing), a three-layer supervision dataset coupled with three ordered stages: (1) RS semantic anchoring for overhead-view visual-language alignment; (2) domain-bridge convergence for shared RS priors across target and bridging scenarios under different modalities; and (3) evidence-grounded scenario tuning for downstream performance. We construct HarborEval, an eight-track diagnostic benchmark covering perception, spatial understanding, robustness, and generation. Under comparable training budgets, HarborEval increases from 57.95 with Direct-SFT to 70.29 with FBA on LLaVA-v1.5, and from 81.09 to 83.37 on Qwen3-VL. FBA also outperforms Collapsed-SFT and leads on harbor-related VRSBench/RSVQA subsets and OpenEval. Stage-wise and role-replacement analyses validate progressive gap filling and stage-specific roles. Public examples and release updates for CPRS, HarborEval, code, and trained weights are available at https://github.com/Z0ngL1ng/filling-before-advancing.

View source

Similar papers

Aug 2026

RS-CoT: enhancing remote sensing vision-language models with self-consistency sampling

RS-CoT is proposed, a unified framework that enhances multimodal reasoning through CoT-based supervised fine-tuning (SFT) and self-consistency sampling (SCS) and self-consistency sampling (SCS)guided preference optimization for reliable multimodal reasoning in remote sensing.

Ke-Hua Feng, Yong Feng, Wen-Gang Zhang et al. · 0 citations
Preprint Sep 2026

VPRef: A Cross-Domain Benchmark for Referring Remote Sensing Image Segmentation

Rapid advancements in vision-language models have propelled Referring Remote Sensing Image Segmentation (RRSIS) to the forefront of Earth observation. However, practical deployments suffer severe performance degradation under a coupled dual-drift paradigm: visual domain drift from cross-spatial-resolution mismatches an...

Quan-Wei Liu, Tao Huang, Jia-Qi Yang et al. · 0 citations
Sep 2026

PromptRefine: source-free unsupervised domain adaptation for remote sensing image classification via CLIP filtering

This work proposes PromptRefine, a white-box SFUDA framework designed to anchor target adaptation through cross-modal intelligence, and leverages the zero-shot semantic priors of large-scale vision–language models to rectify source-biased predictions via a dynamic prompt fine-tuning mechanism.

Unknown authors · 0 citations
Review Sep 2026

Aperture: Training-Free Multiscale Concept Bottlenecks for Remote Sensing

While earth observation models have advanced substantially, they still lack interpretability. While concept-bottleneck models provide interpretability and expert interaction, they are either too expensive to train for the remote sensing domain or perform poorly without annotation. We posit that in expert domains like r...

Rishabh Mondal, Nipun Batra, Utkarsh Mall · 0 citations
Conference Open access Sep 2026

Unified Sequence Modeling for Remote Sensing: A Parameter-Efficient Foundation Model via Prompt-Driven Granularity Alignment

RS-Florence is proposed, a compact unified model that addresses remote sensing perception systems through a Prompt-Driven Sequence-to-Sequence framework, which maps images and task-specific prompts into a unified sequence of natural language and discrete geometric tokens.

Yang Liu, Wei-Xing Luo, Huai-Zhou Qi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SatOV: Restoring Spatial Priors for Training-Free Open-Vocabulary Segmentation in Remote Sensing Imagery

Open-vocabulary semantic segmentation (OVS) of remote sensing imagery is a challenging pixel-level task requiring strong generalization and adaptation to the spatial characteristics of remote sensing data. Although existing vision-language foundation models perform well in general domains, their image-level classificat...

Chang-Hao Zhao, Ling-Lin Zeng, Hai Liu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.