Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Sep 2026

VLG-RMOT: Calibration-before-control for end-to-end referring multi-object tracking.

Referring multi-object tracking (RMOT) aims to detect and track all objects that satisfy a natural-language expression in video. End-to-end RMOT relies on object queries for detection, association, and temporal updating, so language must function as both an alignment cue and a control signal for evolving queries. Exist...

Hong-Shen Zhao, Ming Dai, Fei Xie et al. · 0 citations
Preprint Jul 2026

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding

Spatio-Temporal Video Grounding (STVG) aims to retrieve the visual trajectory of a specific object from a video stream as described by a natural language expression. However, most advanced methods struggle to balance global context modeling with precise boundary localization. Due to the prohibitive computational costs...

Kai Chen, Ming Dai, Wenxuan Cheng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.