Vision-language models face a fundamental scaling bottleneck: the number of visual tokens grows with both temporal duration and spatial resolution, making long-video understanding expensive for the vision encoder and the language model. Existing methods often compress visual tokens after dense encoding, creating a mism...
Yulong Liu, Xiao-Tian Han, Jun-Yuan Shang et al.· 0 citations
This work proposes Test-Time Curriculum (TTC), a simple and model-agnostic framework that adapts a detector on unlabeled test data through curriculum-based self-training and substantially improves overall detection performance under diverse unseen-generator shifts, establishing a practical and effective test-time adapt...
Yiqian Zhang, Zheyuan Gu, Xiangzhao Hao et al.· 0 citations
Memory-Augmented Compression is proposed, a training-free framework that constructs reusable reasoning memories from historical traces and retrieves them as prefill-side scaffolds to compensate for information lost during compression.
Si-Meng Zhang, Yi-Long Chen, Wenyuan Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.