A generative AI integrated multimodal ReID framework designed explicitly for robustness under missing cues and low latency deployment is proposed, with a cost aware early-exit cascade that prioritizes inexpensive, high confidence evidence and only triggers expensive modalities for ambiguous cases.
Abstract
Person re-identification (ReID) is essential for multi-camera surveillance and tracking, yet remains difficult due to viewpoint and illumination changes, occlusion, background clutter, and low resolution imagery. We propose a generative AI integrated multimodal ReID framework designed explicitly for robustness under missing cues and low latency deployment. The key idea is a cost aware early-exit cascade that prioritizes inexpensive, high confidence evidence and only triggers expensive modalities for ambiguous cases. Our system integrates (i) global visual embeddings from segmented person regions, (ii) automatically generated fine grained semantic attribute descriptions generated by vision-language models (VLMs), and (iii) optional facial embeddings when face observations are reliable. To optimize the balance between accuracy and latency, we use a cost aware early-exit cascade instead of fusing all modalities. Specifically, we first inspect the top-k retrieval results to determine whether the query is unambiguous. If the best match is clearly separated from the remaining candidates, we stop early and return the result to minimize latency; in ambiguous cases, we keep multiple hypotheses and invoke additional modalities (face/semantic) with adaptive reliability weighting to refine the decision. We report person re-identification performance using mAP and Rank-1 accuracy on the Market-1501 and DukeMTMC-reID benchmarks. The proposed adaptive early-exit cascade resolves 60.7% of DukeMTMC-reID queries and 68.4% of Market-1501 queries without invoking semantic reasoning, reducing computational overhead while maintaining competitive retrieval performance.
A novel end-to-end Diffusion-Driven Dual-stream Framework (D 3 F), which seamlessly integrates generative structural priors from Diffusion Transformers (DiT) into vision-language ReID, achieving state-of-the-art (SOTA) performance on both occluded and holistic ReID benchmark datasets.
A unified empirical evaluation on the Market-1501, Occluded-DukeMTMC, and MSMT17 datasets is presented, revealing that latent feature-space refinement and semantic cross-modal alignment offer superior stability, scalability, and robustness compared to explicit pixel-level generation.
Multimodal 3D object detection is fundamental to robust perception in autonomous driving because it integrates complementary information from LiDAR and camera sensors. However, existing methods often fail to maintain robustness under out-of-distribution (OOD) corruptions caused by sensor noise, adverse weather, and env...
Zi-Ying Song, Lin Liu, Hong-Yu Pan et al.· 0 citations
Passage-based top-view person re-identification aims to match individuals across short overhead walking videos. This privacy-preserving setting is challenging because overhead cameras suppress facial and body-part cues, compress pedestrian appearance, and make models vulnerable to scene shortcuts from floors, ramps, an...
Hien Pham Duy, Bao Tran, Tien Do et al.· International Conference on...· 0 citations
Person search is challenging due to limitations in identity representation. Existing methods rely on one-hot encoding, ignoring semantic relationships among pedestrians. This leads to a fragmented feature space and reduces generalization ability, especially in large-scale scenarios with a significant proportion of unla...
Xi Yang, He-Xun Zhou, Hai-Yang Zhu et al.· Neural Networks· 0 citations
Known for his clear and elegant writing style, Bertsekas shaped fields from control and optimization to large-scale computation and artificial intelligence.