Skip to content
Conference

AVT-PAC: A Pipeline for Multimodal Action Prediction and Captioning

Jul 2026 · International Conference on Signal Processing and Communications · pp. 1-5 · 0 citations · 26 references

Abstract

Action prediction from frames and videos is a well-studied problem. Models trained with a single modality, mostly vision, will fail in low-light conditions. Recent works have attempted to predict action categories using vision-language and audio-visual models. A challenge, however, is that some dataset annotations lack temporal ground truth and include only the vision modality. Relying on transformers and an intelligent Vision Language Model (VLM) is a viable solution, but deploying them on edge devices could lead to reduced performance and hallucinations. This work presents an Audio-Visual-Text (AVT)-based multi-step Pipeline for Action Prediction and Captioning (AVT-PAC) to address this problem. First, for an input video, we identify the area to focus on using the Region-of-Interest (ROI) Extraction module. CLIP and CLAP encoders are used for ROI prediction. However, the ROI extracted region may vary in duration, resulting in a large number of frames to be processed. To avoid learning from redundant frames, we uniformly sample key frames within the ROI extracted region using a keyframe extraction module. These keyframes are then used to train an Audio-Visual Action and Text-Aware Representation (AVATAR) model to predict actions and captions. Through systematic experiments, we demonstrated that the proposed AVATAR-TCN model beats the present state-of-the-art (SOTA) baselines on the AVE dataset. Code is available in https://github.com/Ifovia/AVT-PAC

View source

Similar papers

Conference Aug 2026

ActionLMM: captioning long-video actions with memory-augmented VLMs

This work introduces ActionLMM, a memory-augmented vision-language model for long-video action summarization that aligns visual and motion modalities through joint representation learning and leverages a novel dual-memory mechanism to retain both local motion details and global temporal structure.

Rui-Rui Li, Dari Abdullah Alrwoaily, Turgut Sofuyev et al. · 0 citations
Open access Jul 2026

Deep Features Evaluation Method of Human Action Recognition Based on Convolutional Neural Network

A compact and deployable Convolutional Neural Network-Long Short Term Memory (CNN-LSTM) framework that combines a 2D convolutional backbone (AlexNet) for frame-level descriptors with an LSTM head for sequence modeling is proposed, indicating a robust, real-time-capable solution for video understanding in both offline a...

H. Khan, Altaf Hussain · 0 citations
Preprint Sep 2026

Modality-Autoregressive World-Action Models

World-action models (WAMs) jointly model future observations and actions, typically predicting the future as RGB images. Other visual modalities such as depth, pretrained visual features, and point tracks can more efficiently capture geometric, semantic, and motion features. However, how best to combine these modalitie...

Adam Hung, B. Duisterhof, D. Ramanan et al. · 0 citations
Open access 2026

Cross-Modal Temporal Alignment for Action Grounding in Videos

A cross-modal temporal alignment framework that combines a multi-scale temporal convolutional encoder with capsule-based dynamic routing, jointly optimizing temporal boundary prediction, cross-modal semantic alignment, and capsule diversity is proposed.

Gengtian Shi, Chen-Hao Wu, Shao-Fei Wang et al. · 0 citations
Open access Sep 2026

ActNet: focus-aware multi-scale CNN for human activity recognition from images

Recognizing human actions from still images is a challenging task due to the absence of temporal information and the need to infer actions from subtle pose and contextual cues. In this article, we propose ActNet, a novel deep convolutional neural network (CNN) architecture that combines multi-scale feature learning wit...

Şafak Kılıç · 0 citations
Preprint Sep 2026

SV-WAM: An Efficient Surround-View World-Action Model for End-to-End Autonomous Driving

World models (WMs) have demonstrated strong potential for end-to-end autonomous driving by learning predictive representations of future scene dynamics. However, generating future videos during inference introduces substantial computational overhead, leading many recent driving WMs to adopt a single front camera as inp...

Jin-Yang Wang, Shi-Wei Li, Jun-Jian Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.