Diff-MM: Exploring Pre-Trained Text-to-Image Generation Models for Unified Multi-Modal Object Tracking.
A unified multi-modal tracker Diff-MM is proposed by exploiting the multi-modal understanding capability of the pre-trained text-to-image generation model by harnessing the extensive prior knowledge in the generation model to achieve a unified tracker with uniform parameters for RGB-N/D/T/E tracking.