This work proposes a "Summarize First, Download Later" paradigm that exploits recent advances in onboard edge computing and Vision-Language Models (VLMs) and substantially reduces bandwidth consumption while accelerating time-to-insight for time-sensitive missions.
Abstract
Modern Earth observation (EO) satellites carry increasingly advanced sensors that produce vast volumes of high-resolution, multispectral data, yet downlink capacity remains a critical bottleneck -- often causing significant latency or the loss of valuable observations within limited contact windows. We propose a"Summarize First, Download Later"paradigm that exploits recent advances in onboard edge computing and Vision-Language Models (VLMs). Rather than indiscriminately downlinking raw imagery, the system follows a three-phase interaction protocol: the satellite first transmits concise natural language summaries generated by a quantized onboard VLM; ground operators then issue targeted Visual Question Answering (VQA) queries to verify scene relevance (e.g., wildfires or maritime anomalies); and full-resolution images are downloaded only when critical information is confirmed. This transforms the downlink from passive bulk transfer into an active, semantics-aware dialogue. We implement and evaluate the system on a resource-constrained NVIDIA Jetson platform, and experiments on diverse remote sensing scenes show that the proposed strategy substantially reduces bandwidth consumption while accelerating time-to-insight for time-sensitive missions.
CoOrbit is presented, a collaborative EO system that conserves onboard storage by retaining only changed and non-overlapping regions and achieves over 108.6× reduction in storage cost and 41.6× reduction in communication size compared to existing EO pipelines.
Rui-Chen Li, Yufan Wu, Zhengyi Hu et al.· Conference on Applications,...· 0 citations
The introduction of IR275K provides a reproducible foundation for accuracy--efficiency evaluation of infrared MFSR methods, while the architectural analysis offers a concrete starting point for spatially aware SSM design under resource-constrained infrared sensing.
Jie Deng, Heyang Wang, Changxin Wang et al.· arXiv.org· 0 citations
An unsupervised domain adaptation (UDA) scheme based on diffusion-driven style transfer is adopted, which produces sensor-stylized training data reflecting sensor-specific characteristics without any labeled onboard imagery, enabling robust cloud detection under real operational conditions.
Changmin Lee, Minsu Kim, Chanhee Jung et al.· IEEE Geoscience and Remote S...· 0 citations
General-purpose vision-language models (VLMs) now support strong visual recognition, instruction following, and generation. However, most pretrained visual encoders are built around three-channel natural images and do not directly accommodate observations such as native multispectral measurements or synthetic aperture...
Shan-Ji Liu, Ke-Lu Yao, Jun-Xiao Xue et al.· 0 citations
A unified experimental framework is introduced that jointly evaluates compressed deep learning models and CML ensembles on two representative EO benchmarks and provides empirical guidance for selecting model families and compression strategies when designing future on-board EO systems.
G. Di Palma, Alessio Pardini, Lan-Pei Li et al.· Pattern Analysis and Applica...· 0 citations
Remote-sensing videos enable real-time observation of changes in target attributes, short-term activities, and scene evolution. They record motion, actions, interactions, and scene changes that cannot be captured by isolated images. Existing models primarily target single images or discrete temporal observations spanni...
Hongjie Zhou, Shiqin Wang, Haoyang Chen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.