Preprint
Jul 2026
Look Less, Think Faster: Joint Token-Compute Adaptation for Multimodal LLMs
SmartVL is proposed, a unified adaptive inference framework that jointly controls vision token number and model compute capability in response to varying input contents and compute budgets and consistently outperforms prior adaptive methods and achieves superior accuracy-efficiency Pareto frontiers.
Pengcheng Wang, Zhiquan Wang, Jayoung Lee et al.
· 0 citations