Embodied systems must make knowledge acquired during one encounter usable in another despite changes in viewpoint, motion, and illumination. Yet aggregate cross-video accuracy conflates failures of local perception with failures to preserve observation identity, establish correspondence, and compose evidence, obscuring...
Yue-Dong Tan, Lei-Tao Qi, Yu Liu et al.· 0 citations
The Region Token Interface (\method{}) adapts a diffusion model to these tokens, with the region count drawn at random during fine-tuning so one checkpoint serves every budget.
Eduard Zamfir, C. Reisswig, Zongwei Wu et al.· 0 citations
Extensive experiments demonstrate that SiConMo achieves a state-of-the-art accuracy-efficiency trade-off among lightweight semantic segmentation models, highlighting simplicity as a powerful design principle.
Mian Muhammad Naeem Abid, Nancy Mehta, Zong-Wei Wu et al.· International Conference on...· 0 citations
TRaM-VSR, a Token Routing and Merging framework for adaptive token allocation, leveraging both context-aware video priors and network-level priors accelerates inference significantly while preserving state-of-the-art reconstruction quality and robust temporal consistency is proposed.
Sicheng Gao, Zhuyun Zhou, Yixuan Liu et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.