Skip to content
Preprint

Summarize First, Download Later: Onboard VLMs for Bandwidth-Efficient Earth Observation

Aug 2026 · 0 citations · 22 references
Computer Science

TL;DR

This work proposes a "Summarize First, Download Later" paradigm that exploits recent advances in onboard edge computing and Vision-Language Models (VLMs) and substantially reduces bandwidth consumption while accelerating time-to-insight for time-sensitive missions.

Abstract

Modern Earth observation (EO) satellites carry increasingly advanced sensors that produce vast volumes of high-resolution, multispectral data, yet downlink capacity remains a critical bottleneck -- often causing significant latency or the loss of valuable observations within limited contact windows. We propose a"Summarize First, Download Later"paradigm that exploits recent advances in onboard edge computing and Vision-Language Models (VLMs). Rather than indiscriminately downlinking raw imagery, the system follows a three-phase interaction protocol: the satellite first transmits concise natural language summaries generated by a quantized onboard VLM; ground operators then issue targeted Visual Question Answering (VQA) queries to verify scene relevance (e.g., wildfires or maritime anomalies); and full-resolution images are downloaded only when critical information is confirmed. This transforms the downlink from passive bulk transfer into an active, semantics-aware dialogue. We implement and evaluate the system on a resource-constrained NVIDIA Jetson platform, and experiments on diverse remote sensing scenes show that the proposed strategy substantially reduces bandwidth consumption while accelerating time-to-insight for time-sensitive missions.

View source

Similar papers

Book Open access Aug 2026

Achieving Efficient Storage and Communication via Collaboration

CoOrbit is presented, a collaborative EO system that conserves onboard storage by retaining only changed and non-overlapping regions and achieves over 108.6× reduction in storage cost and 41.6× reduction in communication size compared to existing EO pipelines.

Rui-Chen Li, Yufan Wu, Zhengyi Hu et al. · 0 citations
Jul 2026

IR275K: A Benchmark for Infrared Multi-Frame Super-Resolution Toward Efficient Remote Sensing

The introduction of IR275K provides a reproducible foundation for accuracy--efficiency evaluation of infrared MFSR methods, while the architectural analysis offers a concrete starting point for spatially aware SSM design under resource-constrained infrared sensing.

Jie Deng, Heyang Wang, Changxin Wang et al. · 0 citations
Open access 2026

Real-Time Onboard AI for Remote Sensing: Cloud Detection

An unsupervised domain adaptation (UDA) scheme based on diffusion-driven style transfer is adopted, which produces sensor-stylized training data reflecting sensor-specific characteristics without any labeled onboard imagery, enabling robust cloud detection under real operational conditions.

Changmin Lee, Minsu Kim, Chanhee Jung et al. · 0 citations
Preprint Sep 2026

Lightweight Adaptation of General-Purpose VLMs for Multispectral and SAR Image Understanding

General-purpose vision-language models (VLMs) now support strong visual recognition, instruction following, and generation. However, most pretrained visual encoders are built around three-channel natural images and do not directly accommodate observations such as native multispectral measurements or synthetic aperture...

Shan-Ji Liu, Ke-Lu Yao, Jun-Xiao Xue et al. · 0 citations
Open access Sep 2026

On-board efficiency: comparing compressed deep learning and classical models for Earth observation

A unified experimental framework is introduced that jointly evaluates compressed deep learning models and CML ensembles on two representative EO benchmarks and provides empirical guidance for selecting model families and compression strategies when designing future on-board EO systems.

G. Di Palma, Alessio Pardini, Lan-Pei Li et al. · 0 citations
Preprint Aug 2026

RSVideo: Are Your Vision-Language Models Ready for Remote Sensing Videos?

Remote-sensing videos enable real-time observation of changes in target attributes, short-term activities, and scene evolution. They record motion, actions, interactions, and scene changes that cannot be captured by isolated images. Existing models primarily target single images or discrete temporal observations spanni...

Hongjie Zhou, Shiqin Wang, Haoyang Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.