Skip to content

HSMTrack: Heterogeneous-State Motion Tracking for Vision-Sensor Pipelines

Aug 2026 · IEEE Sensors Journal · Vol 26, pp. 24708-24716 · 0 citations

Abstract

Motion-only multiobject tracking (MOT) suffers from ID switches in uniform-appearance and deformation-heavy scenes. In these settings, appearance cues become less reliable, so stable identities depend mainly on motion information. Existing methods often process all bounding-box variables together, which can weaken cues needed for prediction and matching. We address this problem by treating each trajectory as a heterogeneous multivariate time series (MTS) and redesigning the motion-only pipeline for embedding, encoding, and matching. HSMTrack separates box-state variables before modeling their temporal and cross-variable relationships, then uses deformation-aware matching for identity association. The method requires no appearance branch and can serve as a post-detection motion module in vision-sensor tracking pipelines. Its SSM-based encoder has linear complexity with respect to trajectory length, reducing modeling cost compared with attention-based alternatives. HSMTrack achieves 59.6 IDF1 and 42.9 AssA on DanceTrack, and 77.9 IDF1 and 67.2 AssA on SportsMOT. Under a unified end-to-end protocol, it reaches 34.1 frames/s on RTX 4090 and 10.5 frames/s on Jetson Orin NX.

View source

Similar papers

Jul 2026

VidMap: Exploiting Temporal Structure for Video-Based Structure-from-Motion

This work introduces a system that combines the strong sequential constraints of SLAM with the flexibility and global optimization of offline SfM, enabling the metric reconstruction of arbitrary, long, uncalibrated videos.

Zador Pataki, Paul-Edouard Sarlin, Marc Pollefeys · 1 citation · ⚡1
Open access 2026

HOMA-ST: High-Order Motion-Guided Cross-Attention for UAV Single-Object Tracking

Single-object tracking from unmanned aerial vehicles (UAVs) is complicated by small target size, rapid camera ego-motion, and frequent occlusion, all of which degrade the appearance cues that Transformer trackers rely on. We present HOMA-ST, which recovers a complementary motion signal by decomposing the short-term dyn...

Peng Gu, Zhanlin Qiu, L. Cherikbayeva et al. · 0 citations
Conference Sep 2026

GSC-ByteTrack: a motion-aware and scale-adaptive multi-object tracking framework

In the UAV traffic monitoring scene, vehicle targets usually have the characteristics of small scale, dense distribution and complex camera motion. These factors will reduce the detection reliability and interfere with the data association process, which can easily lead to trajectory breakage and identity switching pro...

Lin-Zhao Cui, Zhao-Yu Liu, Yu Chen · 0 citations
Open access 2026

Motion-Appearance Synergistic Dual-Branch Joint Decision for Airborne Infrared Multi-Object Tracking

Airborne infrared small-object tracking is crucial for applications such as autonomous reconnaissance and border surveillance. Unlike visible-light imagery, infrared data provides stable imaging in low-light and hazy conditions. However, tracking in this domain is exceptionally challenging due to the diminutive size of...

Ming-Yu Hong, Xue Jin, Yuan Liu et al. · 1 citation
Open access Aug 2026

TraTeTrack: historical trajectory-guided temporal modeling for visual object tracking

Currently, prevalent object-tracking methods are mostly trained using image pairs (i.e. a template image and a search image), and this training paradigm makes it difficult to capture temporal correlations in the object’s continuous motion. Meanwhile, methods that depend on consecutive video frame training incur a drast...

Yong Tao, Haibin Wang, Jing-Lin Ma et al. · 0 citations
Open access Aug 2026

Event-based Optical Flow Using Spatio-temporal Registration.

This paper introduces a spatio-temporal registration framework to increase accuracy of current state-of-the-art event-by-event flow estimation, while also introducing a twofold algorithm acceleration approach and a real-time implementation strategy to mitigate the impact of computation scaling with event rate.

Zhi-Chao Li, Arren J. Glover, Lorenzo Natale et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.