Occlusion-Aware Image and Video Perception for Vulnerable Road User Detection: Methods, Benchmarks, and Deployment Challenges
Abstract
Occlusion remains one of the main failure points in traffic-scene perception, and the errors it causes do not follow a single, predictable pattern. In a still frame, a camera may capture only a pedestrian’s head or upper torso. In video, a tracker can lose that person for several frames and assign a different identity when the person reappears. Cyclists and riders create a related problem because the body parts that remain visible may overlap visually with cars, trucks, or nearby road users. This review examines how image- and video-based methods handle such incomplete observations. The review covers 96 studies published between 2009 and 2025. These studies investigate visible-to-full-body localization, interactions among neighboring instances, feature completion, transformer-based detection, RGB-thermal fusion, tracking, and the less frequently studied problem of inferring fully hidden pedestrians. Because these tasks produce different outputs, we did not combine their results. Pedestrian and crowd detection, multispectral detection, hidden-target prediction, and multi-object tracking are discussed under their original evaluation protocols. CityPersons miss rate measures frame-level detection; MOT17 HOTA includes association across frames; hidden-pedestrian F1 measures prediction quality. We also recorded runtime, hardware constraints, model size, and sensor requirements when the source papers provided those details. The published evidence is strongest for pedestrians who remain partly visible. Much less evidence is available for cyclists, riders, and fully hidden targets. Many deployment claims are also difficult to assess because the papers do not report results for specific occlusion levels or provide enough detail about computational cost.