Jul 2026· JOIV: International Journal on Informatics Visualization· Vol 10, pp. 1635· 0 citations
TL;DR
This paper proposes a framework called Frame Importance Voting (FIV), where frame importance weighting and voting are merged as part of a shared inference process to enhance the temporal classification of video data without additional computational burden.
Abstract
Classification of video data is challenging due to the temporally repeated measures and the differential importance of different frames in a video clip. This paper proposes a framework called Frame Importance Voting (FIV), where frame importance weighting and voting are merged as part of a shared inference process to enhance the temporal classification of video data without additional computational burden. Spatial video features were extracted from the video clips using a ResNet-50 architecture, and the temporal relationships were modeled using a two-layer transform encoder. Frame significance was derived adaptively by summing the transformer's attention and confidence scores for each frame, and predictions for categories were made by summing the frame predictions using weighted voting. Results on Kinetics-400 (50 categories, 10,000 clips, 16-32 frames per video) confirmed that FIV achieved 77.4% top-1, 92.5% top-5, and 76.8% top-1, outperforming the aggregators by up to 5.6% (using about 37 million parameters).
It is proved that the computationally cheaper split space-time attention is equivalent to full space-time attention and is promising to extend VideoSEMA to longer videos with a dilated/sparse temporal attention.
N. Tran, Fanghui Xue, Shuai Zhang et al.· arXiv.org· 0 citations
Understanding long-range videos remains a key challenge in computer vision due to high temporal redundancy and computational burden. Despite strong performance of recent models, they are constrained in terms of scalability and generalization when applied to longer video sequences. In this work, we present Keyframe-base...
Rahul Kumar, S. Channappayya· International Conference on...· 0 citations
Classifying behavioral events in video sequences requires modeling morphology and motion under severe data constraints. This paper presents a hybrid two-stream CNN-ViT architecture combining EfficientNetB4 for local spatial features with a Vision Transformer branch for global context. Time-color coding (TCC) transforms...
M. Barulina, I. Kovalenko, E. Ahremenko et al.· IEEE Access· 0 citations
Video sequences redundant frame removal currently plays an important role in various video preprocessing applications related to video classification, change detection, and object recognition (e.g., defects, faults, human actions, etc.). This paper proposes a video sequence subsampling method based on estimating the si...
A. Kolodenkova, M. O. Bochkarev· Вестник Ростовского государс...· 0 citations
Weakly supervised video anomaly detectors are trained with video-level labels but are commonly evaluated as temporal localizers using Micro-AUROC or AP over pooled test frames. Because these metrics compare frames from different videos, a detector can score well by separating videos without accurately ordering moments...
Inpyo Song, Jangwon Lee· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.