An end-to-end multi-modal pipeline for person search in unconstrained CCTV environments
In real-world surveillance, situations arise in which two people wear nearly identical clothing or in which identifying features are obscured by heavy occlusions and shifting poses. In these environments, traditional uni-modal systems that rely on static appearance do not perform well and often produce false matche...