Skip to content
Open access

Frame Importance Voting for Video Scene Classification

Jul 2026 · JOIV: International Journal on Informatics Visualization · Vol 10, pp. 1635 · 0 citations

TL;DR

This paper proposes a framework called Frame Importance Voting (FIV), where frame importance weighting and voting are merged as part of a shared inference process to enhance the temporal classification of video data without additional computational burden.

Abstract

Classification of video data is challenging due to the temporally repeated measures and the differential importance of different frames in a video clip. This paper proposes a framework called Frame Importance Voting (FIV), where frame importance weighting and voting are merged as part of a shared inference process to enhance the temporal classification of video data without additional computational burden. Spatial video features were extracted from the video clips using a ResNet-50 architecture, and the temporal relationships were modeled using a two-layer transform encoder. Frame significance was derived adaptively by summing the transformer's attention and confidence scores for each frame, and predictions for categories were made by summing the frame predictions using weighted voting. Results on Kinetics-400 (50 categories, 10,000 clips, 16-32 frames per video) confirmed that FIV achieved 77.4% top-1, 92.5% top-5, and 76.8% top-1, outperforming the aggregators by up to 5.6% (using about 37 million parameters).

Read PDF

Similar papers

Jul 2026

VideoSEMA: a scalable and efficient Mamba-like attention for video understanding

It is proved that the computationally cheaper split space-time attention is equivalent to full space-time attention and is promising to extend VideoSEMA to longer videos with a dilated/sparse temporal attention.

N. Tran, Fanghui Xue, Shuai Zhang et al. · 0 citations
Conference Jul 2026

Towards an Efficient and Unified Strategy for Video Understanding Applications

Understanding long-range videos remains a key challenge in computer vision due to high temporal redundancy and computational burden. Despite strong performance of recent models, they are constrained in terms of scalability and generalization when applied to longer video sequences. In this work, we present Keyframe-base...

Rahul Kumar, S. Channappayya · 0 citations
Open access 2026

CNN-ViT Hybrid Architecture With Time-Color Coding for Classifying Behavioral Events in Video Sequences

Classifying behavioral events in video sequences requires modeling morphology and motion under severe data constraints. This paper presents a hybrid two-stream CNN-ViT architecture combining EfficientNetB4 for local spatial features with a Vision Transformer branch for global context. Time-color coding (TCC) transforms...

M. Barulina, I. Kovalenko, E. Ahremenko et al. · 0 citations
Open access Jul 2026

Development of a method for redundant frame removal in video sequences based on assessment of the black-and-white frames similarity

Video sequences redundant frame removal currently plays an important role in various video preprocessing applications related to video classification, change detection, and object recognition (e.g., defects, faults, human actions, etc.). This paper proposes a video sequence subsampling method based on estimating the si...

A. Kolodenkova, M. O. Bochkarev · 0 citations
Preprint Aug 2026

Frame-Level Evaluation in Weakly Supervised Video Anomaly Detection Mostly Measures Video-Level Ranking

Weakly supervised video anomaly detectors are trained with video-level labels but are commonly evaluated as temporal localizers using Micro-AUROC or AP over pooled test frames. Because these metrics compare frames from different videos, a detector can score well by separating videos without accurately ordering moments...

Inpyo Song, Jangwon Lee · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.