Skip to content
Preprint

Multi-Person Human Motion Forecasting in Complex Scenes

Aug 2026 · 0 citations · 59 references
Computer Science

TL;DR

Object-Conditioned Social Diffusion is proposed, a conditional diffusion model that integrates motion history, multi-person interactions, and object cues into a single framework that reduces the two-second path error, produces more realistic long-term forecasts, and supports sampling multiple plausible futures.

Abstract

Accurately forecasting the movement of people in complex scenes requires reasoning over the past and present state of the entire environment. In this context, effectively incorporating object information and social interactions into a unified framework remains particularly challenging. To address this, we propose Object-Conditioned Social Diffusion (OCSD), a conditional diffusion model that integrates motion history, multi-person interactions, and object cues into a single framework. OCSD uses an object-conditioning mechanism that modulates denoising at every timestep, enabling fine-grained human-object reasoning, and a social encoder that models the interactions between all humans in the scene. As a result, our model naturally handles varying group sizes, complex social interactions, and supports sampling multiple plausible futures. Extensive experiments show that OCSD achieves state-of-the-art results on the Humans in Kitchens (HiK) and HOI-M3 benchmarks. It reduces the two-second path error by 121.5 mm (31.3%) on HiK and 130.5 mm (33.2%) on HOI-M3 compared to prior work, and produces more realistic long-term forecasts.

View source

Similar papers

Open access Aug 2026

Stochastic human trajectory prediction via interaction-aware diffusion model

The Interaction-Aware Diffusion Model (IADM) is proposed, a novel diffusion-based framework considering both human motions and surrounding scene layout by treating the social and scene interactions as conditions in the parameterized reverse Markov chain.

Zhong Zhang, Nuoran Wang, Song Gao et al. · 0 citations
Aug 2026

Anticipating Object Interactions Via Aggregation and Distillation of Spatio-Temporal Knowledge From Vision Language Models.

ST-KAD sets a new state of the art, demonstrating accurate what-when-where prediction of future interactions, and confirms that the prior-informed aggregation and teacher-student distillation generalize beyond anticipation to spatial localization, validating the generality of the design.

Yang Liu, Dejie Yang, Minghang Zheng et al. · 1 citation
Preprint Aug 2026

Gaussian-Mixture Latent Flow for Stochastic 3D Human Motion Prediction

Stochastic human motion prediction aims to forecast future motion distributions. Although recent studies have achieved strong performance in terms of accuracy and diversity, they often overlook plausibility (e.g., resulting in physically unrealistic predictions) and uncertainty quantification, both of which are essenti...

Yue Ma, Frederick W. B. Li, Xiaohui Liang · 0 citations
Preprint Sep 2026

MC-DeTra: Motion-Consistent Joint Object Detection and Socially-Aware Trajectory Forecasting in Bird's-Eye-View Images

Unified models for object detection and trajectory forecasting aim to merge perception and prediction for autonomous driving, refining actor trajectories directly over shared bird's-eye-view (BEV) images rasterized from LiDAR and high-definition maps. Their accuracy on dynamic, moving actors, however, remains the harde...

Vladislav Diuzhev, Dmitry Yudin · 0 citations
Preprint Sep 2026

Modality-Autoregressive World-Action Models

World-action models (WAMs) jointly model future observations and actions, typically predicting the future as RGB images. Other visual modalities such as depth, pretrained visual features, and point tracks can more efficiently capture geometric, semantic, and motion features. However, how best to combine these modalitie...

Adam Hung, B. Duisterhof, D. Ramanan et al. · 0 citations
Open access 2026

UnifiedPoseNet-A Lightweight Shared Pose Estimation Model for Human and Vehicles

Pose estimation is a key aspect of action recognition, study of behaviors and spatial relationships from video data. Traditional pose estimation methods are designed exclusively for keypoint detection of either humans or vehicles and incur high inference time and complexity. In this paper, we present UnifiedPoseNet, a...

Reenie Tanya, Balika J. Chelliah · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.