A unified decision-driven framework: Long-term tracking via visual-language models with motion estimation.
Current vision-language multimodal long-term tracking methods are highly dependent on large-scale data training and complex cross-modal models. This not only leads to substantial computational overhead but also restricts their deployment and application in resource-constrained scenarios. To resolve this contradiction,...