Rotation Invariant and Edge Aware Mamba for Oriented Remote Sensing Object Detection
Abstract
Oriented remote sensing object detection (oriented RSOD) extends general remote sensing object detection (RSOD) by localizing aerial objects with rotated bounding boxes, which is essential for dense and arbitrarily oriented targets in overhead imagery. High-resolution remote sensing scenes require broad contextual reasoning, yet CNN backbones are limited by local receptive fields and Transformer backbones incur quadratic cost in global token interactions. Mamba offers an efficient alternative because its state-space sequence modeling captures long-range dependencies with linear complexity. However, vanilla visual Mamba still lacks the geometric, boundary, and scale-aware priors required by oriented RSOD. To address this gap, we propose RIEMamba, a Mamba-based backbone that preserves efficient long-sequence modeling while embedding remote-sensing-specific priors. First, Rotation-Invariant Edge-Aware Pixel Difference Convolution (RIEPDC) integrates gradient-based operators into pixel-difference convolution with SO(2)-equivariant structural modeling to enhance rotation-consistent boundary representation. Second, Variable Multi-Scale Scanning (VMSS) fuses scanning information from two scale windows to balance local detail capture with contextual modeling. Experiments show that RIEMamba achieves state-of-the-art results on DOTA-v1.0 (80.11%/82.49% mean Average Precision (mAP) for single/multi scale), HRSC2016 (98.62% mAP under VOC 2012), and DIOR-R (68.55% mAP), with only slight computational overhead.