Aug 2026· IEEE Internet of Things Journal· 0 citations· 70 references
Computer Science
TL;DR
CoDAT is proposed, a Collaborative Dual-Attention Transformer that replaces conventional multi-head attention with a lightweight dual-branch module: Spatial Convolutional Attention (SCA) for local aggregation and Strided Single-Head Attention (SSHA) for global context.
Abstract
Real-time human action recognition on Internet-of-Things (IoT) edge devices requires models that capture rich spatio-temporal cues within strict latency, memory, and power envelopes. Current 3D CNNs, video transformers, and shift-based ViT deliver high accuracy but come at computational costs that preclude edge IoT deployment. This paper proposes CoDAT, a Collaborative Dual-Attention Transformer that replaces conventional multi-head attention with a lightweight dual-branch module: Spatial Convolutional Attention (SCA) for local aggregation and Strided Single-Head Attention (SSHA) for global context. SSHA jointly compresses the spatial resolution and channel dimensions of the query, key, and value tensors via stride-based sparse projection, then fuses the resulting global and local features at a markedly reduced cost. To enable temporal communication across frames, a parameter-free TShift module is embedded in each block. Extensive experiments on Jetson AGX Orin and Raspberry Pi 5 demonstrate that CoDAT achieves an energy-accuracy balance in both image and action recognition. On ImageNet-1K, CoDAT-M runs 2x faster than EfficientViT384 and FastViT-S12 at comparable accuracy, and CoDAT-L matches ViT-S with 3x fewer parameters at 2x higher throughput. On Kinetics-400 and MA-52, CoDAT achieves competitive Top-1 accuracy against state-of-the-art CNN, transformer, and hybrid baselines while running up to 2.9x faster than VSwin-T, 2x faster than ViT-Temporal-Shift variants, and 5x faster than UniFormer-B. On UCF-101, CoDAT-S384 matches TokShift and LAPS while being 6x faster and requiring up to 13x fewer FLOPs, establishing an efficiency-accuracy balance for real-time action recognition in edge IoT perception systems. Code is available at https://github.com/novendrastywn/CoDAT .
Mobile badminton video analysis faces three coupled challenges: fine-grained strokes differ mainly in brief local body motions, the high-speed shuttlecock is easily degraded by motion blur and occlusion, and cascaded task-specific models impose excessive latency and energy overhead on mobile devices. To address these i...
Yong Wang· International journal of pat...· 0 citations
A novel action recognition method, named MICA-Net, which combines data from multiple sensors to improve the efficiency of the HAR model, and a new compact version of a wrist-worn sensor device with Wi-Fi connectivity to an edge device, enhancing usability in human-machine interaction applications.
Trung-Hieu Le, Thai-Khanh Nguyen, T. Tran et al.· ACM Transactions on Multimed...· 0 citations
The Equipment-Primed Network (EP-Net) is proposed, a heterogeneous dual-stream architecture that treats sports equipment as a primary semantic cue for action discrimination and a Cross-Modal Channel Attention (CMCA) module that projects equipment features into the behavior-feature space and performs directional channel...
Xiaocui Sang, Chang-Wei Gu, Lei Zhao· Applied Sciences· 0 citations
This work proposes LITEWAY, a modality-agnostic, fully convolutional framework for multichannel sensor time series that replaces recurrent temporal modeling with structured convolutional decomposition and achieves competitive macro F1 while reducing model size.
Dominique Nshimyimana, Vitor Fortes Rey, Meng-Xi Liu et al.· 0 citations
Recognizing human actions from still images is a challenging task due to the absence of temporal information and the need to infer actions from subtle pose and contextual cues. In this article, we propose ActNet, a novel deep convolutional neural network (CNN) architecture that combines multi-scale feature learning wit...
A hybrid CNN-SSM-Attention backbone for 12-lead ECG classification is introduced, which provides strong supervised baselines under a compact parameter budget and improves transfer, particularly in reduced-label settings and under both full fine-tuning and LoRA-based adaptation.
Y. Bazi, Sarah Aljuhani, M. M. Al Rahhal et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.