Mar 2026· 2026 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops (WACVW)· pp. 603-612· 1 citation· 29 references
Computer Science
Abstract
Illegal waste dumping poses significant environmental and public health challenges worldwide, requiring automated surveillance systems for detection and prevention. This paper presents our solution for the IWDD 2026 Contest, addressing the dual challenge of detecting illegal dumping events in surveillance videos and localizing the exact moment of occurrence. We employ X3D-M, an efficient 3D convolutional network pretrained on Kinetics-400, combined with a sliding window inference strategy for temporal localization. Through systematic hyperparameter optimization across 96 configurations and ablation studies examining nine combinations of fine-tuning strategies and loss functions, we identify key design choices for this application domain. Our experiments reveal that differential learning rates-applying lower rates to the pretrained backbone while training the classifier more aggressively-outperform both frozen backbones and uniform fine-tuning. The optimal system achieves an F1-score of 0.8387 and a temporal F1score of 0.7742 on our test set, with 92.3% of correct detections within the temporal tolerance window. Operating at over 8 times real-time speed with only 2.97M parameters, our approach demonstrates that efficient video classification architectures can be effectively adapted for specialized surveillance applications through careful transfer learning and inference design.
Detecting oil spills at sea is a difficult vision problem. Slicks typically have soft edges, scatter into irregular patches, and resemble several harmless features of the sea surface. This paper presents a controlled benchmark of six YOLO-based object detectors applied to this task, namely YOLOv8n, YOLOv8s, YOLO11n, YO...
Mohamed Mahmoud Ain Dhib, M. Lachgar, Mohamedou Cheikh Tourad et al.· EPJ Web of Conferences· 0 citations
A deep learning-based forensic framework for real-time detection of suspicious human activity in CCTV videos, trained without relying on any external sensors is proposed, and incorporates anonymization of personal identities and local edge-based processing to prevent raw data exposure.
Qazi Mazhar Ul Haq, Muhammad Imran, M. Waqas et al.· Arab Journal of Forensic Sci...· 0 citations
This work proposes a detection framework called distillation alignment YOLO (DA-YOLO) for PPE detection and introduces a teacher–student distillation framework with consistency constraints across predictions and high-order features extracted from baseline that enables the student model to achieve strong generalization...
Chonghua Zhou, Ruixuan Zhang, Yi-Xin Fu et al.· Multimedia Systems· 0 citations
Conventional video surveillance based on pixel-level deep-learning models is resource hungry, processes gigabytes of video material, and retains biometric identifying data. This paper describes a lightweight, privacy-sensitive alternative that uses skeletal pose estimation to replace pixel-based processing. We only pro...
S. L. Jany Shabu, P. Asha, P.Asmitha Priyaa et al.· 2026 7th International Confe...· 0 citations
Background: Human-elephant conflict poses a significant threat to both wildlife conservation and rural livelihoods, particularly in regions bordering forest reserves. Traditional observation methods are often time-consuming, error-prone and limited under challenging environmental conditions, highlighting the need for a...
Seng-phil Hong· Indian Journal of Agricultur...· 0 citations
The YOLOv12 network is adopted as the baseline model and the ADown module is introduced to improve downsampling efficiency while maintaining lightweight performance, and the BN-CGLU is incorporated into the A2C2f module to enhance the model’s nonlinear representation capability.
Ao-Bo Yue, Puchun Chen, Yan Yang· Journal of Real-Time Image P...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.