In real-world applications, current multimodal large models are often overestimated in their ability to understand scientific charts. To assess their true capabilities and identify key performance bottlenecks, we conducted an in-depth study on scientific chart understanding. Charts in scientific literature often featur...
Ling-Dong Shen, Qigqi, Kun Ding et al.· IEEE Transactions on Image P...· 0 citations
Most modern multimodal large language models (MLLMs) build on a pretrained visual encoder that provides a strong visual prior. Encoder-free MLLMs instead learn visual representations directly from raw pixels, offering a simple and unified architecture, but their scaling behavior has not been systematically characterize...
A novel Continual Audio-Visual Segmentation (CAVS) task, aiming to continuously segment new classes guided by audio, and a Collision-based Multi-modal Rehearsal (CMR) framework, designed to address challenges of multi-modal semantic drift and co-occurrence confusion.
Yuyang Hong, Qi Yang, Tao Zhang et al.· arXiv.org· 3 citations· ⚡1
This work proposes DiffPDE, a framework leveraging discrete diffusion language models for targeted code repair by introducing a localized re-masking and infilling strategy, which achieves competitive accuracy, outperforms same-scale AR models, and significantly accelerates repair.
Wen-Xuan Guo, Yuyang Hong, Lu-Bin Fan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.