TCTracker: a RGB-T tracking via text-guided and contrast-learning enhancement
Abstract
In complex environments, RGB-T object tracking relies on the complementary features of RGB and thermal infrared (TIR) modalities to enhance accuracy and robustness. However, traditional approaches often overlook the guiding role of semantic features, and boundary box-based initialization can lead to ambiguity and drift when the target undergoes appearance changes or occlusions. To address these challenges, we propose TCTracker, a CLIP-based RGB-T tracking algorithm driven by large language models. Unlike traditional methods, TCTracker utilizes large language models to generate textual descriptions of the target. It then uses cross-modal contrastive learning to guide the backbone network in learning target representations based on these descriptions (e.g., target color and type). This approach effectively leverages the rich semantic information in image-text pairs. During tracking, textual information guides the interaction of cross-modal features, fully exploiting the complementary advantages of multimodal data. Additionally, TCTracker enhances the response of target feature regions by leveraging the correlation between text and image, thereby improving target regression accuracy and refining scale estimation. Our extensive evaluation on three leading RGB-T tracking benchmarks demonstrates that TCTracker achieves competitive performance compared to state-of-the-art methods.