Evaluating semantic similarity between videos is a fundamental challenge in computer vision, essential for tasks ranging from out-of-distribution (OOD) detection to video retrieval. However, defining and labeling video similarity is notoriously difficult and expensive due to the complex spatio-temporal nature. In this...
Enrico Pallotta, Sina Raoufi, Lars Doorenbos et al.· 0 citations
Object-Conditioned Social Diffusion is proposed, a conditional diffusion model that integrates motion history, multi-person interactions, and object cues into a single framework that reduces the two-second path error, produces more realistic long-term forecasts, and supports sampling multiple plausible futures.
Serdar Ozsoy, Lars Doorenbos, Juergen Gall· 0 citations
This work proposes the first video-language-model post-training technique for mistake detection, which uses a tailored reward function to encourage the model to identify discrepancies between an instruction and the corresponding video, and generalizes especially well to unseen procedures.
Federico Spurio, Olga Zatsarynna, Lars Doorenbos et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.