Event-Guided Spatiotemporal Cube Learning for Infrared Video Small-Object Detection
Abstract
Infrared video small-object detection remains a challenging problem due to the extremely weak target appearance, low signal-to-noise ratio, cluttered thermal backgrounds, and frequent temporal inconsistency across frames. Most existing detectors follow a frame-wise paradigm, where each frame is processed independently, or temporal cues are introduced through additional aggregation modules. Such designs often suffer from insufficient motion perception, background-induced false alarms, and increased computational complexity. In this article, we propose Cube-IRVSOD, a compact yet effective cube-to-trajectory detection framework for infrared video small-object detection. The key idea is to reformulate consecutive infrared frames as a spatiotemporal cube, enabling short-term motion cues to be encoded as intrinsic input structures rather than external temporal dependencies. Specifically, a cube data stream generator converts sampled consecutive frames into grayscale intensity maps and stacks them into a cube representation, allowing the detector to jointly perceive spatial appearance and temporal evolution. To further suppress background-dominated pseudo-motion, we introduce an event flow encoder that derives sparse event maps from interframe brightness changes and employs SNN-based temporal encoding to generate motion-aware priors. These priors guide sparse feature enhancement over multilevel spatial representations, emphasizing moving target responses while reducing stationary clutter interference. In addition, a trajectory-to-frame inference scheme with confidence voting aggregates predictions from overlapping cube streams, improving temporal stability and localization reliability. Extensive experiments demonstrate that Cube-IRVSOD achieves superior detection accuracy and a favorable accuracy–efficiency tradeoff under challenging scenarios.