Skip to content

Author

Shi-Long Jing

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

2026

FSM: Frequency-Enhanced Spatiotemporal Model for Streaming Remote Sensing Object Detection

Remote sensing object detection is a fundamental task in ground scene observation and analysis. Despite the currently discrete-frame detectors achieves remarkable performance, they still suffer from three critical limitations: 1) mainstream architectures regress spatial locations independently, making it difficult to exploit cross-frame temporal cues for disambiguating occlusions and motion blur; 2) the high-frequency details are prone to degradation during long-term temporal modeling in the complex remote sensing backgrounds where the foreground and background are highly similar, leading to memory disturbance; and 3) traditional feature pyramid networks (FPNs) merely adopt simple concatenation to fuse multiscale feature maps, neglecting the synergy between shallow details and deep semantics. To address these issues, we propose the Frequency-enhanced spatiotemporal model (FSM). Specifically, we utilize a streaming inference pipeline to continuously process image sequences and build upon the state space model (SSM) to develop the spatiotemporal Mamba module (SMM) that dynamically memorizes and updates target location states across consecutive frames. Moreover, we construct a frequency enhancement module (FEM) to counteract feature degradation in long temporal sequences under challenging remote sensing backgrounds, generating more discriminative temporal representations for the SSM. In addition, we design an adaptive bidirectional FPN (ABiFPN) to selectively fuse shallow and deep features through a learnable scalar mechanism, restoring the small target information forgotten in deep layers and achieving fine-grained feature synergy. Experimental results demonstrate that the proposed FSM achieves 76.6% and 38.7% mAP@0.5:0.95 on the EMRS-Frame and SAT-MTB datasets, outperforming all current state-of-the-art detection methods.

Shi-Long Jing, Heng-Yi Lv, Yu-Chen Zhao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.