Skip to content

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Aug 2026

ActionLMM: captioning long-video actions with memory-augmented VLMs

The success of large language models (LLMs) has inspired the development of foundation-level multimodal systems that integrate vision and language. However, current video-language models—such as Video-LLaMA and VideoChat—struggle with fine-grained human motion understanding and fail to summarize long videos effectively. Meanwhile, motion-focused models are limited to short clips and lack mechanisms to capture long-range spatiotemporal context. We introduce ActionLMM, a memory-augmented vision-language model for long-video action summarization. It aligns visual and motion modalities through joint representation learning and leverages a novel dual-memory mechanism to retain both local motion details and global temporal structure. To support evaluation, we propose a large-scale benchmark dataset with 33,887 longform action videos and 169,435 caption annotations across 1920 action categories. Experiments show that ActionLMM significantly outperforms prior methods, offering a robust and scalable solution for fine-grained human action understanding.

Ruirui Li, Dari Abdullah Alrwoaily, Turgut Sofuyev et al. · 0 citations