Skip to content
Preprint

Event ActivityNet: A Large-Scale Simulated-Event Benchmark for Untrimmed Action Understanding

Aug 2026 · 0 citations · 30 references
Computer Science

TL;DR

Event ActivityNet is introduced, a large-scale simulated-event benchmark derived from human-annotated, untrimmed ActivityNet videos that provides a scalable benchmark for long-horizon event modeling, although native-camera evaluation remains essential for deployment-oriented conclusions.

Abstract

Long-horizon event-based action understanding remains underexplored because existing datasets largely comprise short, trimmed clips, while collecting native event streams with dense temporal annotations is costly. We introduce Event ActivityNet, a large-scale simulated-event benchmark derived from human-annotated, untrimmed ActivityNet videos. It comprises 3,263 videos, 200 action classes, and 106.94 hours, with matched 5-bin and 9-bin event-voxel representations, temporal action annotations, and timestamped captions. The benchmark supports annotated-segment action recognition, auxiliary event-language alignment, and causal online temporal action localization. We generate event voxels directly from non-interpolated source videos in decoded frame order, retain per-video rational nominal or average frame-rate metadata for approximate time mapping, and use action-center reconstruction LPIPS as a soft diagnostic of retained reconstructable content. We establish baselines for adaptive event framing, prompt-caption alignment, and event-only, RGB-only, and RGB-event localization. Under a progressive nested-scale training protocol, recognition Top-1 accuracy increases from 52.25 to 66.42, while online temporal localization average mAP improves from 21.7 to 29.0. Moreover, staged Event ActivityNet pretraining followed by native-event fine-tuning consistently outperforms target-only and joint-from-scratch training across multiple supervision budgets. Event ActivityNet provides a scalable benchmark for long-horizon event modeling, although native-camera evaluation remains essential for deployment-oriented conclusions.

View source

Similar papers

#computer vision Preprint Sep 2026

Interpretable Temporal Video Reasoning with EventGraph and EventField

The results indicate that structured temporal representations can support both performance and inspectability by preserving symbolic structure, capturing temporal continuity, and providing human-readable diagnostics for video reasoning.

Durgendra Narayan Singh · 0 citations
Jul 2026

EgoPlay: Event-Triggered Video Editing for Egocentric Streams

We introduce EgoPlay, an event-triggered video-to-video editor for egocentric streams, obtained by fine-tuning a pretrained V2V diffusion transformer on event-conditioned data built primarily from Ego4D. Given a monocular video and an event-triggered prompt of the form"when X happens, do Y,"EgoPlay infers whether and w...

Jinjie Mai, G. Qian, W. Menapace et al. · 0 citations
Conference Aug 2026

ActionLMM: captioning long-video actions with memory-augmented VLMs

This work introduces ActionLMM, a memory-augmented vision-language model for long-video action summarization that aligns visual and motion modalities through joint representation learning and leverages a novel dual-memory mechanism to retain both local motion details and global temporal structure.

Rui-Rui Li, Dari Abdullah Alrwoaily, Turgut Sofuyev et al. · 0 citations
Preprint Aug 2026

StreamSoccer: Event-Driven Memory for Streaming Soccer Commentary

Results indicate that event memory supports streaming soccer commentary across temporal scopes while controlling long-history computation.

Chen-Xi Shao, Bo-Zhong Wang, Jiaxin Huang et al. · 0 citations
Preprint Aug 2026

NBA_Streaming: A Large-Scale Benchmark for Fine-Grained Basketball Commentary Generation in Continuous Streams

Live basketball commentary generation requires determining when an event is sufficiently observable and describing it before subsequent events unfold. However, existing methods are primarily designed for pre-segmented clips or complete videos, making them unsuitable for continuous streams. Existing datasets also provid...

Lifang Wu, Yuyang Wu, Yangdong Gao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

EM^2Mem: Event-Centric Multimodal Memory for Large Language Models

EM^2Mem is proposed, an event-centric multimodal memory framework that binds heterogeneous evidence to event anchors during memory construction, enabling compact evidence readout over grounded multimodal events rather than modality-specific fragments.

Yijun Chen, Yangfan Zheng, Yanyang Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.