Skip to content

CAT-Net: a coordinate attention transformer network for workplace activity recognition

Aug 2026 · Neural computing & applications (Print) · Vol 38 · 0 citations · 22 references

TL;DR

A hybrid deep learning (DL) model that integrates Coordinate Attention (CA), Convolutional Neural Networks (CNN), and Transformer encoders for better HAR achieves superior performance and proved its effectiveness for workplace safety monitoring applications.

View source

Similar papers

Open access 2026

A Deep Spatio-Temporal Framework for Multi-Class Traffic Prediction and Accident Detection in Surveillance Video

Experimental results demonstrate that Z-score standardization improves classification performance, and the feasibility and robustness of the proposed framework in real-world traffic environments are indicated.

Dhartee Patel, Jinal Ahir, Namrata Shroff et al. · 0 citations
Open access Sep 2026

ActNet: focus-aware multi-scale CNN for human activity recognition from images

Recognizing human actions from still images is a challenging task due to the absence of temporal information and the need to infer actions from subtle pose and contextual cues. In this article, we propose ActNet, a novel deep convolutional neural network (CNN) architecture that combines multi-scale feature learning wit...

Şafak Kılıç · 0 citations
Open access Aug 2026

CoDAT: Collaborative Dual-Attention Transformer with Low-Cost Temporal Modeling for Efficient Edge Action Recognition

CoDAT is proposed, a Collaborative Dual-Attention Transformer that replaces conventional multi-head attention with a lightweight dual-branch module: Spatial Convolutional Attention (SCA) for local aggregation and Strided Single-Head Attention (SSHA) for global context.

Novendra Setyawan, Chi-Chia Sun, Mao-Hsiu Hsu et al. · 0 citations
Conference Aug 2026

ViTBN: A Vision Transformer with Batch Normalization for Crowd Behaviour and Anomaly Detection in Surveillance Videos

Intelligent crowd behaviour analysis is critical for modern video surveillance systems to enhance public safety and enable timely anomaly detection. This study presents ViTBN, a vision transformer-based framework augmented with batch normalization to improve feature robustness and the stability of training. The propose...

Ayushi Tiwari, A. S. Kushwaha · 0 citations
#edge computing Open access Sep 2026

Deep Learning-Based Forensic Detection of Suspicious Activities in CCTV Systems

A deep learning-based forensic framework for real-time detection of suspicious human activity in CCTV videos, trained without relying on any external sensors is proposed, and incorporates anonymization of personal identities and local edge-based processing to prevent raw data exposure.

Qazi Mazhar Ul Haq, Muhammad Imran, M. Waqas et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.