Aug 2026· 2026 IEEE International Conference on Cybernetics and Intelligent Systems (CIS) and IEEE International Conference on Robotics, Automation and Mechatronics (RAM)· pp. 97-102· 0 citations· 25 references
Abstract
Event cameras can capture human motion with low latency and high dynamic range, providing an emerging sensing paradigm for action recognition. A large body of recent work converts event data into a sequence of dense frames and feeds them into a video classification model for prediction. Although the frame-based methods take advantage of pre-trained backbones from the image domain to achieve high accuracy, most existing works sacrifice the sparsity of events due to dense processing, thereby increasing model complexity and latency. To boost the efficiency of frame-based methods while fitting the sparse nature of events, we propose the Sparsity-Aware Vision Transformer Network (SVTNet) with two novel event-oriented designs. Instead of direct event-to-frame conversion, we introduce an event representation method named Progressive Cumulative Event Representation (PCER) that adaptively integrates spatiotemporal cues across multiple temporal slices to better preserve fine-grained information. To leverage the spatial sparsity of events to accelerate model inference, we present the Motion Prior-Guided Token Sparsification Module (MPTS) that prunes redundant visual tokens hierarchically based on the joint motion-semantic importance. Comprehensive experiments show that SVTNet achieves state-of-the-art accuracy on multiple benchmark datasets and maintains much lower model complexity than existing frame-based methods.
This work proposes FLEET (Feature Learning from Events via Efficient Tokenization), a feature extractor that processes event sequences directly and decouples inference cost of the feature extractor's backbone from the sensor's resolution, enabling end-to-end learning without auxiliary losses.
T. Gottwald, Maximilian Schier, Melanie Schaller et al.· 0 citations
This work investigates how temporal information can be encoded directly within the event representation, proposing a confidence-normalized continuous multi-timescale representation based on logarithmic B-spline temporal encoding together with a geometry-aware local confidence mechanism that exploits the spatial structu...
Fredrik Lundell, Per-Erik Forssén, Mårten Wadenbäck et al.· 1 citation
This article proposes ActNet, a novel deep convolutional neural network architecture that combines multi-scale feature learning with a focus-aware attention mechanism to address the problem of recognizing human actions from still images.
Autonomous navigation requires precise and efficient semantic segmentation, yet existing frame-based approaches remain limited by motion blur, glare, latency, and the low temporal resolution (20-30 FPS) of conventional cameras, which leads to information loss between frames. Event cameras have emerged as an alternative...
Dalia Hareb, Jean Martinet, Benoit Miramond et al.· 0 citations
LiFR v2 is presented, a unified propagation-completion-memory framework for causal anytime and streaming dense prediction from an RGB keyframe and events and introduces an Event-Guided Completion Module (EGCM) and a History Retrieval Module (HRM) to reuse completed representations across successive queries.
Tao Wan, Xiao-Shan Wu, Yi-Fei Yu et al.· 0 citations
Infrared video small-object detection remains a challenging problem due to the extremely weak target appearance, low signal-to-noise ratio, cluttered thermal backgrounds, and frequent temporal inconsistency across frames. Most existing detectors follow a frame-wise paradigm, where each frame is processed independently,...
Feng-Xiang Xu, Nan Zhang, Ting-Fa Xu et al.· IEEE Transactions on Geoscie...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.