Skip to content

Leveraging Scene-Invariant Priors for Autoregressive Human Trajectory Prediction

Nov 2026 · IEEE Robotics and Automation Letters · Vol 11, pp. 12511-12518 · 0 citations · 28 references

Abstract

Accurate human trajectory prediction is essential for autonomous driving and robot navigation. Despite substantial progress, deep learning-based approaches often suffer from the train-inference gap caused by distributional discrepancies between training and test environments. The goal-guided framework mitigates this issue by exploiting the scene-invariant prior that human motion is typically goal-driven. However, existing goal-guided approaches do not fully leverage additional scene-invariant priors and therefore still face two key limitations: insufficient diversity in predicted goals and the lack of explicit modeling of human kinematic consistency. In this paper, we propose SIPTraj, a goal-guided trajectory prediction framework that integrates scene-invariant priors at both the long-term intention and short-term motion levels. For goal estimation, we introduce goal candidates derived from offline clustering of trajectory endpoints as a data-driven approximation of scene-invariant long-term intention modalities, and develop a weighted Farthest Point Sampling (FPS) strategy that balances spatial coverage with predicted likelihoods to improve goal diversity while preserving semantic plausibility. For trajectory completion, we develop a kinematic-aware dual-stream autoregressive decoder that jointly predicts aligned position and velocity. The decoder adopts a kinematic initialization with residual refinement (KIRR) strategy: future states are initialized using a constant-velocity kinematic model and refined via reciprocal cross-attention, which injects kinematic consistency into the decoding process and leads to more natural and accurate trajectory prediction. Extensive experiments demonstrate that SIPTraj achieves state-of-the-art performance on the ETH/UCY and SDD benchmarks.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

DiffWAM: A Fast and Efficient Navigation World Action Model

Pretrained video foundation models encode rich semantic and spatiotemporal priors for embodied navigation, yet converting these priors into UAV motion typically requires expensive future-video synthesis and geometric reconstruction. We investigate whether the motion implicit in future visual prediction can instead be r...

Morui Zhu, Yu-Ze Wu, Xi-Jie Huang et al. · 0 citations
Preprint Sep 2026

READ: Learning Risk-Informed Fields for End-to-End Autonomous Driving

Autonomous driving requires more than recognizing what is present in a scene: a planner must determine how road structure, surrounding agents, and their motion states should influence a future maneuver. Existing learning-based planners can capture these influences through latent scene features and trajectory decoders,...

Zhi-Yuan Liu, Yuan-Xin Tian, Ze-Hong Ke et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Planning-Aligned Pretraining of BEV Representations with Sparse Action-Conditioned Targets for End-to-End Autonomous Driving

PAVER, Planning-Aligned BEV Encoder Pretraining is introduced, where from a single LiDAR sweep, PAVER constructs sparse risk and unknown targets describing occupied and unobserved evidence along rule-based ego motions, preserving the downstream architecture and camera-only inference.

Jaeha Song, Soonmin Hwang · 0 citations
Oct 2026

MaTF: Maneuver-Aware Temporal Fusion for Trajectory Prediction Under Arbitrary Observation Length

Trajectory prediction is essential for many robotic applications, yet most existing models rely on fixed-length observations and struggle with temporally irregular inputs. In real-world settings, prediction difficulty further increases when agents exhibit strong maneuverability, as their future motions depend on distin...

Shuobo Wang, Wen-Yuan Qin, Yong-Zhao Hua et al. · 0 citations
#machine learning Preprint Sep 2026

EgoNeMo: Transferable Map of Pedestrian Dynamics via Egocentric LiDAR Scan

This paper proposes a transferable Map of Dynamics (MoD) framework that generalizes to unknown environments using only egocentric 3D LiDAR point clouds to overcome the long-standing limitation of traditional MoD methods. While MoDs are essential for encoding human motion characteristics to enable accurate pedestrian tr...

Azusa Sawada, Allan Wang, Hideo Saito et al. · 0 citations
Preprint Aug 2026

Gaussian-Mixture Latent Flow for Stochastic 3D Human Motion Prediction

This work proposes a latent flow-based model equipped with a data-driven Gaussian mixture prior that more effectively disentangles diverse human behaviors than conventional single-modal priors and enables natural uncertainty quantification through tractable likelihood computation.

Yue Ma, Frederick W. B. Li, Xiaohui Liang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.