Skip to content
Conference

RCF-Net: Degradation-Aware Hybrid CNN–Transformer for Child Face Identification in Surveillance

Jul 2026 · 2026 4th International Conference on Sustainable Computing and Smart Systems (ICSCSS) · pp. 1209-1218 · 0 citations · 21 references

Abstract

Child face identification from surveillance video remains difficult because facial crops are frequently low-resolution, blurred, partially occluded, and captured under unstable illumination. Age-related facial variation further increases the difficulty of maintaining discriminative identity embeddings for children. This paper presents RCF-Net, a degradation-aware hybrid CNN–Transformer architecture that combines surveillance-oriented image degradation, dual-branch local/global feature extraction, and learnable cross-attention fusion. MTCNN is used for face detection and alignment, ArcFace supervision is used for discriminative embedding learning, and DeepSORT can optionally be integrated to improve temporal identity consistency in video streams. To address deployment concerns raised by surveillance use, the revised framework also specifies age-progression handling, latency-aware scheduling for live video, multi-camera scaling, and adversarial/spoof-risk safeguards. Experiments are conducted using public face datasets, namely VGGFace2, CASIA-WebFace, CelebA, and IMDB-WIKI, with child-oriented filtering and synthetic surveillance degradations. Compared with representative CNN, transformer, and hybrid baselines, RCF-Net achieves the best overall accuracy of 91.4% and yields the strongest robustness under low-resolution, blur, and occlusion stress tests. The results indicate that explicit degradation modeling and local-global feature fusion are complementary for surveillance-oriented child face identification.

View source

Similar papers

Open access Aug 2026

An Attention-Enhanced Lightweight CNN Framework with MTCNN Detection and TripletEmbedding Recognition for Occlusion-Robust Automated Attendance from Surveillance Video

Manual and card/barcode-based attendance recording remains slow, error-prone, and vulnerable to proxy marking, motivating fully automated, camera-based alternatives for schools and organizations. This paper proposes an Attention-Enhanced Lightweight CNN framework that couples MTCNN multi-scale face detection with a CBA...

K. B, L. C. · 0 citations
Conference Jul 2026

A Hybrid CNN–Transformer Network for Robust Masked and Occluded Face Recognition in Smart Surveillance Systems

Face recognition systems applied to smart surveillance settings often experience poor performance when the faces are partially occluded by a mask or other objects. Occlusions eliminate critical facial information, which makes face identification much more difficult for traditional deep learning models. To solve this is...

R. R, Anbalagan E · 0 citations
Review Open access Aug 2026

PERSON RE-IDENTIFICATION BASED ON DEEP LEARNING NETWORKS: A SURVEY

This survey offers an updated and focused review of deep learning-based ReID methods, encompassing research from 2020 to 2025, and investigates in-depth the engineering aspects, including system integration, real-time performance, and sensor constraints, which are often overlooked in reviews of earlier work.

Zahraa A. Faisal, N. E. El Abbadi · 0 citations
Conference Jul 2026

Vision Transformer with Attention Rollout for Deepfake Face Image Detection and Localization

Generative AI and synthetic media generation tools have enabled widespread media manipulation tools and raised important privacy concerns with misinformation, identity fraud and the verification of authenticity of media. Most of the current convolution-based deepfake detection methods are hard to be deployed in real sc...

K. V. Sai Phani, Shaik Mahaboob Jailan, F. Mahammad et al. · 0 citations
Open access Aug 2026

A Hybrid Vision Transformer and EfficientNet-B3 Framework for Facial Expression Recognition

A hybrid architecture that combines Vision Transformers (ViTs) to capture global context with EfficientNet-B3 for multi-scale feature extraction and highlights the promise of hybrid deep learning architectures in tackling real-world facial expression recognition challenges.

Sasan Karamizadeh, Saman Shojae Chaeikar, Mazdak Zamani · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.