Skip to content

Anticipating Object Interactions Via Aggregation and Distillation of Spatio-Temporal Knowledge From Vision Language Models.

Aug 2026 · IEEE Transactions on Pattern Analysis and Machine Intelligence · Vol PP · 1 citation
Medicine

TL;DR

ST-KAD sets a new state of the art, demonstrating accurate what-when-where prediction of future interactions, and confirms that the prior-informed aggregation and teacher-student distillation generalize beyond anticipation to spatial localization, validating the generality of the design.

Abstract

The increasing prevalence of wearable cameras has driven the development of egocentric (first-person) systems that assist human activities proactively by anticipating imminent interactions. A central challenge in this domain is active object interaction anticipation from first-person video-predicting what interaction will occur, when it will happen, and where it will take place. This involves forecasting (1) what interaction category (verb-noun pair), (2) when (time-to-interaction), and (3) where (the active object's location, bounding box in the last observed frame). However, existing approaches rely on limited prior knowledge about active objects and their state changes, and they struggle to (1) predict diverse state-change interactions, (2) handle temporal uncertainty of changes, and (3) localize the active object accurately in the presence of spatial distractors. To address these problems, we propose ST-KAD. It consists of a Spatial-Temporal Knowledge Aggregator that integrates rich commonsense priors to enhance what-when-where interaction anticipation and guides the model's attention toward informative cues, and a Teacher-Student Distillation framework that enables efficient inference without access to oracle inputs by transferring knowledge from an oracle-informed teacher model to a query-based student decoder. On two egocentric anticipation benchmarks (Ego4D-STA, EPIC-Kitchens-STA), ST-KAD sets a new state of the art, demonstrating accurate what-when-where prediction of future interactions. Moreover, results on four active object detection benchmarks (Ego4D-AOD, EPIC-Kitchens-AOD, MECCANO, 100DOH) further confirm that our prior-informed aggregation and teacher-student distillation generalize beyond anticipation to spatial localization, validating the generality of the design.

View source

Similar papers

#small language model Review Aug 2026

Vision-Language Models for Egocentric Video: From Hand-Object Interaction to Embodied AI

This survey presents a critical review of VLMs for egocentric video understanding, tracing the progression from conventional recognition architectures to multimodal foundation models and embodied systems, and examines how first-person perception and multimodal foundation models support wearable assistance, robot skill...

M. Zamani, Fatemeh Ziaeetabar · 0 citations
Preprint Aug 2026

Multi-Person Human Motion Forecasting in Complex Scenes

Object-Conditioned Social Diffusion is proposed, a conditional diffusion model that integrates motion history, multi-person interactions, and object cues into a single framework that reduces the two-second path error, produces more realistic long-term forecasts, and supports sampling multiple plausible futures.

Serdar Ozsoy, Lars Doorenbos, Juergen Gall · 0 citations
#artificial intelligence Preprint Sep 2026

From Where to How: Continuous 4D Interaction Forecasting from Egocentric Video

Egocentric 4D interaction forecasting aims to anticipate both where future interactions will occur in 3D and how the human body will move to realize them, providing an important capability for assistive robotics and human-computer interaction. Existing methods struggle to translate semantic understanding into precise c...

Qiao-Hui Chu, Haoyu Zhang, Meng Liu et al. · 0 citations
Preprint Aug 2026

GESTO: Human-Centric Spatio-Temporal Memory for Reasoning in Dynamic Scenes

Robots operating in human environments need memories that capture not only what objects exist and where, but also how people use them over time and how individual interactions compose into goal-directed activities. Existing 4D scene graphs preserve object and place histories but omit activity structure, whereas activit...

Ermanno Bartoli, Buwei He, Dennis Rotondi et al. · 0 citations
Jul 2026

Egocentric Online Action Segmentation via Parametric Context Memory Learning

To facilitate smart wearable devices or human-like robotics with real-time first-person perspective perception ability, recent researchers proposed the Egocentric Online Action Segmentation (EOAS) task. It requires models to recognize what is happening in egocentric streaming videos and discriminate the starting and en...

Xun Jiang, Xing Xu, Chong Liu et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Ambient @ EgoProactive 2026 : Proactive Egocentric Assistance with Visually Grounded Supervision

This submission to the EgoProactive track of the ECCV 2026 Wearable AI Challenge is presented, which ranked first in the large-model division and second in the<=2B division, suggesting that visual grounding is more important than annotation volume for this task.

Logesh Kumar Umapathi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.