Training-free token reduction accelerates vision transformers by removing redundant tokens across layers, recovering most of the original accuracy at a fraction of the compute. These methods, however, are designed and evaluated primarily on clean data, and under real-world distribution shift their accuracy gap to the u...
Temporal Foundation Models (TFMs) aim to generalize across domains, datasets, and tasks. Yet, their development remains constrained by fragmented, task-specific data formats, annotations, and processing pipelines. We introduce TimeNet, an open-source data standard and scalable infrastructure that decouples temporal dat...
Martin Maritsch, Timo Stoffregen, Thomas Kaar et al.· 0 citations
A simple branch-regenerate-distill algorithm, Trajectory-Intervention Self-Distillation (TISD), which forces a teacher-selected branch action, returns suffix generation to the student, and distills the full trajectory under the privileged-context-conditioned teacher, support teacher-guided branching as a way to expose...
Taeckyung Lee, Rinat Amankos, Jeonghye Kim et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.