Skip to content

A Generative AI Integrated Multimodal Framework for Low-Latency Multi-Camera Person Re-Identification

Sep 2026 · 0 citations · 21 references
Computer Science Engineering

TL;DR

A generative AI integrated multimodal ReID framework designed explicitly for robustness under missing cues and low latency deployment is proposed, with a cost aware early-exit cascade that prioritizes inexpensive, high confidence evidence and only triggers expensive modalities for ambiguous cases.

Abstract

Person re-identification (ReID) is essential for multi-camera surveillance and tracking, yet remains difficult due to viewpoint and illumination changes, occlusion, background clutter, and low resolution imagery. We propose a generative AI integrated multimodal ReID framework designed explicitly for robustness under missing cues and low latency deployment. The key idea is a cost aware early-exit cascade that prioritizes inexpensive, high confidence evidence and only triggers expensive modalities for ambiguous cases. Our system integrates (i) global visual embeddings from segmented person regions, (ii) automatically generated fine grained semantic attribute descriptions generated by vision-language models (VLMs), and (iii) optional facial embeddings when face observations are reliable. To optimize the balance between accuracy and latency, we use a cost aware early-exit cascade instead of fusing all modalities. Specifically, we first inspect the top-k retrieval results to determine whether the query is unambiguous. If the best match is clearly separated from the remaining candidates, we stop early and return the result to minimize latency; in ambiguous cases, we keep multiple hypotheses and invoke additional modalities (face/semantic) with adaptive reliability weighting to refine the decision. We report person re-identification performance using mAP and Rank-1 accuracy on the Market-1501 and DukeMTMC-reID benchmarks. The proposed adaptive early-exit cascade resolves 60.7% of DukeMTMC-reID queries and 68.4% of Market-1501 queries without invoking semantic reasoning, reducing computational overhead while maintaining competitive retrieval performance.

View source

Similar papers

D 3 F: Diffusion-Driven Dual-Stream Framework for Occluded Person Re-Identification

A novel end-to-end Diffusion-Driven Dual-stream Framework (D 3 F), which seamlessly integrates generative structural priors from Diffusion Transformers (DiT) into vision-language ReID, achieving state-of-the-art (SOTA) performance on both occluded and holistic ReID benchmark datasets.

Xiaohao Xie, Weihao Meng, Wen-Hua Jiao · 0 citations
Review Aug 2026

Occluded person re-identification: a taxonomic survey and reproducible empirical benchmark

A unified empirical evaluation on the Market-1501, Occluded-DukeMTMC, and MSMT17 datasets is presented, revealing that latent feature-space refinement and semantic cross-modal alignment offer superior stability, scalability, and robustness compared to explicit pixel-level generation.

Ishani Sharma, Puneet Kapoor, Pankaj Vaidya · 0 citations
Preprint Sep 2026

Towards robust multimodal 3D object detection via visual foundation models

Multimodal 3D object detection is fundamental to robust perception in autonomous driving because it integrates complementary information from LiDAR and camera sensors. However, existing methods often fail to maintain robustness under out-of-distribution (OOD) corruptions caused by sensor noise, adverse weather, and env...

Zi-Ying Song, Lin Liu, Hong-Yu Pan et al. · 0 citations
Conference Aug 2026

Segmentation-Guided Scene Perturbation for Passage-Based Top-View Person Re-Identification

Passage-based top-view person re-identification aims to match individuals across short overhead walking videos. This privacy-preserving setting is challenging because overhead cameras suppress facial and body-part cues, compress pedestrian appearance, and make models vulnerable to scene shortcuts from floors, ramps, an...

Hien Pham Duy, Bao Tran, Tien Do et al. · 0 citations
Sep 2026

Multimodal-guided self-distillation for unified person search.

Person search is challenging due to limitations in identity representation. Existing methods rely on one-hot encoding, ignoring semantic relationships among pedestrians. This leads to a fragmented feature space and reduces generalization ability, especially in large-scale scenarios with a significant proportion of unla...

Xi Yang, He-Xun Zhou, Hai-Yang Zhu et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.