MMSep is proposed, a training-free multimodal separator localization and compression framework that improves efficiency in both prefilling and decoding and consistently reduces latency while maintaining competitive generation quality and reasoning accuracy.
Ming-Jie Ma, Yichao Ma, Jiannan Cao et al.· Proceedings of the 32nd ACM...· 0 citations
Existing research on Efficient Multimodal Large Language Models (EMLLMs) primarily focuses on reducing the number of visual tokens in the prefilling stage, which is tailored to short-answer inference scenario. However, in more complex multimodal reasoning tasks, models are often required to generate lengthy intermediat...
Mingjie Ma, Yichao Ma, Jiannan Cao et al.· Proceedings of the 32nd ACM...· 0 citations
SKILLFORGE is introduced, a framework that decomposes formal code synthesis into a library of atomic, reusable skills, each targeting a specific subtask such as specification inference, body synthesis, invariant generation, error diagnosis, or targeted repair, and defined by a prompt template, tool binding, and decidab...
Yan-Ming Liu, Xinyue Peng, Jiannan Cao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.