Recent medical multimodal models have benefited from larger corpora, broader modality coverage, and stronger reasoning-oriented training, yet effective data design across continued pretraining (CPT) and post-training remains challenging. Medical sources vary substantially in structure, granularity, and information dens...
Guang-Hao Zhu, Ze-Yu Liu, Zhitian Hou et al.· 0 citations
Model merging efficiently combines specialized large language models (LLMs) without joint retraining, but can substantially alter expert routing in Mixture-of-Experts (MoE) models. Such \emph{routing drift} is often interpreted as routing failure, raising a fundamental question that remains unclear: \emph{does routing...
Yuan-Yi Wang, Yang-Gan Gu, Su Lu et al.· 0 citations
Accurate physical simulation is fundamental to science and engineering, yet conventional numerical solvers incur high costs when handling complex geometries, varying boundary and initial conditions, and diverse physical parameters. Recent deep-learning-based methods offer faster solutions, while limited flexibility and...
Peng-Wei Liu, Xingyu Ren, Peng-Kai Wang et al.· Communications AI & Computin...· 0 citations
MedPIC-Bench makes conditional rule application measurable and highlights the limitations of static medication-safety accuracy for assessing patient-specific reliability among medical-specific LLMs, whose average CF performance trails that of general LLMs.
Zhitian Hou, Yuhang Liu, Peng-Kai Wang et al.· 0 citations
DART (Decoded Attention over Recurrent sTates), which retains the chunk state contributions produced by the Mamba-2 chunked scan as chunk state memories, decodes token-conditioned keys and values from these memories, and performs state-memory attention (SMA) over the resulting KV pairs.
Yixiao Qian, Song Chen, Pengkai Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.