Most modern multimodal large language models (MLLMs) build on a pretrained visual encoder that provides a strong visual prior. Encoder-free MLLMs instead learn visual representations directly from raw pixels, offering a simple and unified architecture, but their scaling behavior has not been systematically characterize...
A novel Continual Audio-Visual Segmentation (CAVS) task, aiming to continuously segment new classes guided by audio, and a Collision-based Multi-modal Rehearsal (CMR) framework, designed to address challenges of multi-modal semantic drift and co-occurrence confusion.
Yuyang Hong, Qi Yang, Tao Zhang et al.· arXiv.org· 3 citations· ⚡1
This work proposes DiffPDE, a framework leveraging discrete diffusion language models for targeted code repair by introducing a localized re-masking and infilling strategy, which achieves competitive accuracy, outperforms same-scale AR models, and significantly accelerates repair.
Wen-Xuan Guo, Yuyang Hong, Lu-Bin Fan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.