Skip to content

D 3 F: Diffusion-Driven Dual-Stream Framework for Occluded Person Re-Identification

· 0 citations · 25 references

TL;DR

A novel end-to-end Diffusion-Driven Dual-stream Framework (D 3 F), which seamlessly integrates generative structural priors from Diffusion Transformers (DiT) into vision-language ReID, achieving state-of-the-art (SOTA) performance on both occluded and holistic ReID benchmark datasets.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

A Generative AI Integrated Multimodal Framework for Low-Latency Multi-Camera Person Re-Identification

A generative AI integrated multimodal ReID framework designed explicitly for robustness under missing cues and low latency deployment is proposed, with a cost aware early-exit cascade that prioritizes inexpensive, high confidence evidence and only triggers expensive modalities for ambiguous cases.

León Fernando, C. Dombawala, P. Hettigoda et al. · 0 citations
Open access 2026

SSGLD: Spatial-Semantic Guided Latent Diffusion for Person Image Synthesis

Existing latent diffusion models (LDMs) for pose-guided person image synthesis (PGPIS) face an inherent trade-off between semantic consistency and texture fidelity. This dilemma stems from their reliance on a single conditioning pathway: semantic guidance is robust but prone to over-smoothing high-frequency details, wh...

Yi-Heng Xu, Zhen-Yin Zhang, Geng-Shen Chen · 0 citations
Sep 2026

Multimodal-guided self-distillation for unified person search.

Person search is challenging due to limitations in identity representation. Existing methods rely on one-hot encoding, ignoring semantic relationships among pedestrians. This leads to a fragmented feature space and reduces generalization ability, especially in large-scale scenarios with a significant proportion of unla...

Xi Yang, He-Xun Zhou, Hai-Yang Zhu et al. · 0 citations
Open access Sep 2026

BMF-DETR: Pseudo-Depth-Guided Bidirectional Multi-Strategy Fusion for End-to-End Object Detection

Transformer-based detectors model long-range context effectively, yet their representations remain dominated by RGB appearance and can become unreliable in cluttered, occluded, or crowded scenes. We present BMF-DETR, a pseudo-depth-guided detector that introduces RGB-derived geometric structure without requiring a dept...

Hai Wang, Jun-Hao Wen, Chun-Lai Yang et al. · 0 citations
Preprint Sep 2026

Incentive Noise and Structural Prior Infusion for Multi-modal Object Re-Identification

Multi-modal object Re-Identification (ReID) benefits from complementary information across heterogeneous imaging modalities. To further enrich semantic representation, text descriptions have recently been incorporated as an additional modality. However, recent vision-language approaches often treat text descriptions as...

Wei-Xiang Zhou, Yu-Hao Wang, Xing-Guo Xu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.