Skip to content

TSCA-Net: Temporal-Spatial Clique Attention for Interpretable Multimodal Pedestrian Trajectory Prediction

Jul 2026 · arXiv.org · Vol abs/2607.11939 · 0 citations · 51 references
Computer Science

TL;DR

TSCA-Net achieves state-of-the-art performance, with average ADE/FDE of 0.13/0.20 m on ETH/UCY and 6.95/10.43 pixels on SDD.

Abstract

Accurate pedestrian trajectory prediction in crowded environments remains challenging due to the multimodal uncertainty of human motion and the variable complexity of motion dynamics across different scene contexts. Existing goal-conditioned models rely on static displacement structures that assign equal weight to all historical time steps, standard graph attention mechanisms, and fixed-capacity motion decoders that cannot adapt to local prediction complexity. To address these limitations, we propose TSCA-Net, a trajectory prediction framework built upon three complementary modules. The Temporal-Spatial Clique Attention (TSCA) module introduces learnable temporal gating into clique-based goal-history interaction, enabling time-aware modulation of historical observations relative to each candidate goal. The Cross-Pedestrian Clique Potential (CPCP) module models asymmetric pairwise agent relationships through a dynamic clique potential framework with a time-varying social graph. The Adaptive KAN Grid Refinement (AKGR) mechanism dynamically adjusts the B-spline grid resolution of a Kolmogorov-Arnold Network-augmented LSTM decoder based on per-agent goal distribution entropy, balancing model expressiveness against overfitting across varying motion complexities. Extensive experiments on the ETH/UCY and Stanford Drone Dataset benchmarks demonstrate that TSCA-Net achieves state-of-the-art performance, with average ADE/FDE of 0.13/0.20 m on ETH/UCY and 6.95/10.43 pixels on SDD. Comprehensive ablation studies confirm the complementary contributions of all three proposed modules.

View source

Similar papers

Preprint Aug 2026

Social Graph Mamba: Forecasting Pedestrian Movements Based on Social Context

Forecasting pedestrian motion has always been fundamental for autonomous navigation in crowded environments. While attention-based methods achieve strong performance, they suffer from quadratic computational complexity in modeling social interactions, limiting scalability. Additionally, the existing methods often achie...

H. Nguyen, Yen-Chen Liu · 0 citations
Open access Aug 2026

Stochastic human trajectory prediction via interaction-aware diffusion model

The Interaction-Aware Diffusion Model (IADM) is proposed, a novel diffusion-based framework considering both human motions and surrounding scene layout by treating the social and scene interactions as conditions in the parameterized reverse Markov chain.

Zhong Zhang, Nuoran Wang, Song Gao et al. · 0 citations

G-VTM: A Multimodal Vision-Trajectory Model for Generalized Vehicle Trajectory Prediction

G-VTM, a generalized vision-trajectory model, is proposed, which captures global map semantics while modeling scenario-and direction-aware interaction based on intuitive visual perception and achieves strong generalized performance under heterogeneous traffic conditions.

Xinyue Zhang, Letian Gong, Yan Lin et al. · 0 citations
Open access Jul 2026

An Interpretable and Edge Deployable Spatio-Temporal Trajectory Prediction for Autonomous Driving

A comprehensive Explainable AI (XAI) evaluation framework is introduced, including temporal sensitivity analysis, interaction-aware perturbation studies, spatial influence analysis, and gradient-based feature attribution methods that provide insights into how the model captures temporal motion dependencies, neighboring...

R. Megalingam, Naveen Prasaad Selvarajan, Pritty Vijay · 0 citations
Sep 2026

A Spatio-Temporal Graph Network with Informer-TCN Fusion for Vehicle Trajectory Prediction on Highways

Accurate vehicle trajectory prediction is essential for the driving safety and efficiency of autonomous vehicles. However, this task remains challenging due to the complex spatial interactions among traffic participants and the wide range of temporal dependencies in motion sequences. To address these issues, this paper...

Zi-Yan Liang, Rui Yuan, Peng-Ying Zhou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.