Skip to content

Author

Shiming Xiang

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining

Most modern multimodal large language models (MLLMs) build on a pretrained visual encoder that provides a strong visual prior. Encoder-free MLLMs instead learn visual representations directly from raw pixels, offering a simple and unified architecture, but their scaling behavior has not been systematically characterize...

Lin Chen, Bo-Lin Ni, Qi Yang et al. · 0 citations

Taming Modality Entanglement in Continual Audio-Visual Segmentation

A novel Continual Audio-Visual Segmentation (CAVS) task, aiming to continuously segment new classes guided by audio, and a Collision-based Multi-modal Rehearsal (CMR) framework, designed to address challenges of multi-modal semantic drift and co-occurrence confusion.

Yuyang Hong, Qi Yang, Tao Zhang et al. · 3 citations · ⚡1
#artificial intelligence Preprint Aug 2026

DiffPDE: Masked Diffusion Language Models as PDE Solver

This work proposes DiffPDE, a framework leveraging discrete diffusion language models for targeted code repair by introducing a localized re-masking and infilling strategy, which achieves competitive accuracy, outperforms same-scale AR models, and significantly accelerates repair.

Wen-Xuan Guo, Yuyang Hong, Lu-Bin Fan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.