Preprint
Jul 2026
LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension
LongEgoRefer is introduced, a novel and challenging benchmark constructed from long-form videos in the Ego4D dataset that defines a demanding spatio-temporal grounding problem that requires models to identify both when an event occurs and where the referred object appears within extended video sequences.
Shunya Kato, Taiki Miyanishi, Shuhei Kurita et al.
· 0 citations