Skip to content

M3-Former: Multimodal Transformer with Mixture-of-Experts for Long-Term Vessel Trajectory Prediction

Aug 2026 · 0 citations · 34 references
Computer Science

TL;DR

The proposed framework establishes a semantic-guided hierarchical prediction paradigm, in which high-level navigational intent and local motion dynamics are jointly modeled for robust long-term vessel trajectory forecasting.

Abstract

To address the challenges of behavioral multimodality, limited semantic utilization, and long-term error accumulation in vessel trajectory prediction, this paper proposes M3-Former, a multimodal trajectory prediction framework enhanced by large language models (LLMs). The proposed framework incorporates vessel static attributes and navigational intent as semantic priors for long-term trajectory modeling. Specifically, a unified multimodal representation space is constructed, in which static semantic information is encoded by a pre-trained LLM and aligned with dynamic trajectory features through self-attention. To jointly capture global route planning and local motion variations, a dual-granularity Mixture-of-Experts (MoE) architecture is introduced, where sequence-level experts model global navigation trends and token-level experts refine fine-grained maneuvering behaviors. In addition, a Steering-Weighted Cross-Entropy loss is designed to alleviate the long-tail distribution of sparse turning samples and improve prediction accuracy in critical maneuvering scenarios. Experiments on a real-world Danish AIS dataset demonstrate that M\textsuperscript{3}-Former consistently outperforms state-of-the-art baselines across prediction horizons from 1 to 4 hours. In the 4-hour prediction task, the proposed method reduces Average Displacement Error (ADE) and Final Displacement Error (FDE) by 4.4\% and 5.1\%, respectively, compared with the strongest baseline. Qualitative and ablation analyses further verify that semantic fusion effectively reduces long-term trajectory drift, while the dual-granularity MoE improves robustness in complex waterways and route-branching scenarios. The proposed framework establishes a semantic-guided hierarchical prediction paradigm, in which high-level navigational intent and local motion dynamics are jointly modeled for robust long-term vessel trajectory forecasting.

View source

Similar papers

Oct 2026

MaTF: Maneuver-Aware Temporal Fusion for Trajectory Prediction Under Arbitrary Observation Length

Trajectory prediction is essential for many robotic applications, yet most existing models rely on fixed-length observations and struggle with temporally irregular inputs. In real-world settings, prediction difficulty further increases when agents exhibit strong maneuverability, as their future motions depend on distin...

Shuobo Wang, Wen-Yuan Qin, Yong-Zhao Hua et al. · 0 citations
Open access Sep 2026

A non-local multi-head spatiotemporal attention LSTM for vehicle trajectory prediction

A non-local multi-head spatiotemporal attention based long short-term memory model (NL-MHA-LSTM) is introduced which employs an attention mechanism to assign context weights to relevant neighbor vehicles and extends beyond pairwise effects to model long-range dependencies.

S. Rashid, M. A. Khan, Usman Akram et al. · 0 citations
Nov 2026

Leveraging Scene-Invariant Priors for Autoregressive Human Trajectory Prediction

Accurate human trajectory prediction is essential for autonomous driving and robot navigation. Despite substantial progress, deep learning-based approaches often suffer from the train-inference gap caused by distributional discrepancies between training and test environments. The goal-guided framework mitigates this is...

Ge Sun, Jun Ma · 0 citations
Conference Aug 2026

MTME-transformer for AIS-based vessel trajectory prediction

It is suggested that multi-scale local motion modelling can stably improve the accuracy of AIS-based vessel trajectory prediction and is evaluated under a unified data preprocessing, resampling and multi-step autoregressive prediction framework.

Qi Xu, Hua-Sheng Nong, Tian-Wei Ma et al. · 0 citations
Open access Aug 2026

HCAD-Net: end-to-end parking network with historical context and attention-based dual-decoder

A vision-based end-to-end autonomous parking framework trained through imitation learning that introduces a historical context fusion encoder to capture temporal dependencies from past vehicle motions, a dual-stream attention decoder to enhance interaction between scene features and trajectory representations, and kine...

Da Zheng, Bing-Li Zhang, Xinyu Wang et al. · 0 citations

Related blog posts

Microsoft Research Blog Aug 11, 2026

Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

Radiology AI is evolving beyond report generation. CARE-X explores a unified approach that combines flexible reasoning, calibrated predictions, and measurement-based tools for chest X-ray interpretation. The post Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement appeared first on Microsoft Research.

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.