Skip to content
Open access

Enhanced multi-agent interaction modeling with differential transformers for trajectory prediction.

Aug 2026 · Neural Networks · Vol 205 Pt B, pp. 109491 · 0 citations · 51 references
Medicine

TL;DR

A general framework that explicitly models the discrepancy between each position and its context to enhance multi-agent trajectory prediction and is a general architecture that consistently achieves improved performance compared to the base models.

Abstract

Predicting the future trajectories of multiple agents is intrinsically challenging due to complex dynamic agent-scene interactions, diverse motion patterns, and the inherent uncertainty of real-world behaviors. Transformer-based architectures have recently become a prevailing paradigm for multi-agent trajectory prediction, as they leverage the multi-head self-attention mechanism to flexibly aggregate information across agents and time steps into contextual representations, where each position is updated by taking a weighted sum over all input tokens. However, in the self-attention approach, it is unknown if the current input or the rest of the inputs has a stronger influence on the updated representation. Therefore, there exists information lost at each model layer and the accumulated information lost could have unpredictable consequences for trajectory prediction. To address this issue of information lost, this paper introduces a framework based on the differential Transformer architecture. It is a general framework that explicitly models the discrepancy between each position and its context to enhance multi-agent trajectory prediction. The core idea of our framework is to introduce a differential encoding layer that encodes the difference between a token's self-representation and the contextual representation aggregated from all other positions in the sequence. The learned differential encoding is combined with the original representation as the updated feature. By integrating differential encodings, the representational capability of the generated trajectory embeddings is improved. The implementation of our differential encoding module is highly parallelizable and can be efficiently integrated into the Transformer-based models. Extensive evaluations are conducted on two popular benchmark datasets. The experimental results show that our differential Transformer is a general architecture that consistently achieves improved performance compared to the base models.

Read PDF

Similar papers

2026

SG-SRC: Semantics-Guided State Residual Coupling for Dynamic Multi-Agent Trajectory Prediction

In autonomous driving, heterogeneous agents interact asymmetrically and change dynamically. This often causes coordination inconsistency, policy mismatch and long-horizon prediction drift. The problem becomes more challenging when communication availability, interaction connectivity, and information freshness change dy...

Bin Wang, Wei Liu, Wei She et al. · 0 citations
Oct 2026

MaTF: Maneuver-Aware Temporal Fusion for Trajectory Prediction Under Arbitrary Observation Length

Trajectory prediction is essential for many robotic applications, yet most existing models rely on fixed-length observations and struggle with temporally irregular inputs. In real-world settings, prediction difficulty further increases when agents exhibit strong maneuverability, as their future motions depend on distin...

Shuobo Wang, Wen-Yuan Qin, Yong-Zhao Hua et al. · 0 citations
Open access Sep 2026

A non-local multi-head spatiotemporal attention LSTM for vehicle trajectory prediction

A non-local multi-head spatiotemporal attention based long short-term memory model (NL-MHA-LSTM) is introduced which employs an attention mechanism to assign context weights to relevant neighbor vehicles and extends beyond pairwise effects to model long-range dependencies.

S. Rashid, M. A. Khan, Usman Akram et al. · 0 citations
Preprint Sep 2026

V2X-WAM: A Cooperative World Action Model for End-to-End Autonomous Driving

Vehicle-infrastructure cooperation can complement onboard sensing with broader and more informative observations of the traffic environment, providing valuable support for end-to-end autonomous driving. However, existing cooperative driving methods mainly exploit roadside information to enhance the representation of th...

Jun-Wei You, Wei-Zhe Tang, Can Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

CoDrive: Cross-Vehicle World-Consistent Video Generation with Precise Trajectory Control for Cooperative Driving

Real-world driving is inherently multi-agent, yet most existing driving world models generate observations from a single ego vehicle. Independently extending them to multiple vehicles does not ensure that different agents observe a consistent shared world. We present CoDrive, a cross-vehicle, multi-view driving video g...

Yu Meng, Bai-Ning Zhao, Jun-Tao Wu et al. · 0 citations
Preprint Sep 2026

MM-Future: Multi-Mode Joint World-Action Modeling for Autonomous Driving

MM-Future, a world-action model that generates multiple paired scene-action hypotheses and models bidirectional interaction within each pair, shows consistent improvements over both single-mode and action-only variants, validating the benefit of multi-mode joint world-action modeling.

Shuai Liu, Hechangle Gong, Hao Jiang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.