Preprint
Jul 2026
UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models
Extensive experiments demonstrate that the on-device latency-informed design combined with the tailored training strategy establishes a new state-of-the-art for efficient LVLM encoding, significantly outperforming existing encoder-centric baselines while operating on-device at nearly 1.7xthe speed.
Ioannis Maniadis Metaxas, Adrian Bulat, Alberto Baldrati et al.
· 0 citations