Multimodal large language models (MLLMs) incur substantial computational overhead due to the reliance on hundreds of visual tokens to represent images. While token pruning has emerged as a promising approach to reduce the inference cost of MLLMs, existing methods typically reassign position embeddings to the retained t...
Ya-Hong Wang, Zhang-Kai Ni, Jun-Cheng Wu et al.· 0 citations
LoMeVQA, a comprehensive benchmark consisting of 206K longitudinal visual question answering (VQA) pairs for temporal medical image analysis, is proposed and MedLong-8B, which achieves state-of-the-art performance across all tasks is introduced.
Zhilin Wu, Zhangkai Ni, Cheng Yang et al.· arXiv.org· 0 citations
Vorch-Streamer, a post-training framework that addresses real-time long-form Text-to-Audio-Video (T2AV) streaming and enables real-time long-form avatar audio-video streaming, and bounded causal context and four-step denoising are presented.
Meng-Lin Han, Yang Ding, Yu-Lei Lu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.