A novel end-to-end Diffusion-Driven Dual-stream Framework (D 3 F), which seamlessly integrates generative structural priors from Diffusion Transformers (DiT) into vision-language ReID, achieving state-of-the-art (SOTA) performance on both occluded and holistic ReID benchmark datasets.
A generative AI integrated multimodal ReID framework designed explicitly for robustness under missing cues and low latency deployment is proposed, with a cost aware early-exit cascade that prioritizes inexpensive, high confidence evidence and only triggers expensive modalities for ambiguous cases.
León Fernando, C. Dombawala, P. Hettigoda et al.· 0 citations
Existing latent diffusion models (LDMs) for pose-guided person image synthesis (PGPIS) face an inherent trade-off between semantic consistency and texture fidelity. This dilemma stems from their reliance on a single conditioning pathway: semantic guidance is robust but prone to over-smoothing high-frequency details, wh...
Person search is challenging due to limitations in identity representation. Existing methods rely on one-hot encoding, ignoring semantic relationships among pedestrians. This leads to a fragmented feature space and reduces generalization ability, especially in large-scale scenarios with a significant proportion of unla...
Xi Yang, He-Xun Zhou, Hai-Yang Zhu et al.· Neural Networks· 0 citations
Transformer-based detectors model long-range context effectively, yet their representations remain dominated by RGB appearance and can become unreliable in cluttered, occluded, or crowded scenes. We present BMF-DETR, a pseudo-depth-guided detector that introduces RGB-derived geometric structure without requiring a dept...
Hai Wang, Jun-Hao Wen, Chun-Lai Yang et al.· Applied Informatics· 0 citations
Multi-modal object Re-Identification (ReID) benefits from complementary information across heterogeneous imaging modalities. To further enrich semantic representation, text descriptions have recently been incorporated as an additional modality. However, recent vision-language approaches often treat text descriptions as...
Wei-Xiang Zhou, Yu-Hao Wang, Xing-Guo Xu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.