VLG-RMOT: Calibration-before-control for end-to-end referring multi-object tracking.
Referring multi-object tracking (RMOT) aims to detect and track all objects that satisfy a natural-language expression in video. End-to-end RMOT relies on object queries for detection, association, and temporal updating, so language must function as both an alignment cue and a control signal for evolving queries. Exist...