Aug 2026
Anticipating Object Interactions Via Aggregation and Distillation of Spatio-Temporal Knowledge From Vision Language Models.
ST-KAD sets a new state of the art, demonstrating accurate what-when-where prediction of future interactions, and confirms that the prior-informed aggregation and teacher-student distillation generalize beyond anticipation to spatial localization, validating the generality of the design.
Yang Liu, Dejie Yang, Minghang Zheng et al.
· IEEE Transactions on Pattern... · 1 citation