Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Sep 2026

Lingshu: Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning.

Multimodal Large Language Models (MLLMs) excel at understanding generic visual content, such as landscapes, objects, and events, thanks to extensive datasets and advanced training regimes. However, their effectiveness in medical applications remains limited due to the inherent discrepancies between data and tasks in medical scenarios and those in the general domain. Existing medical MLLMs face the following critical deficiencies: 1) inadequate coverage of medical knowledge beyond imaging; 2) elevated propensity for hallucinations due to suboptimal data curation; and 3) limited reasoning capacity tailored to complex medical tasks. To address these challenges, we first propose a comprehensive data-curation procedure that 1) efficiently acquires rich medical knowledge data not only from medical imaging but also from extensive medical texts and general domain data; and 2) synthesizes high-quality medical captions, visual question answering, and reasoning samples. Leveraging the curated data, we build a multimodal dataset imbued with extensive medical knowledge and develop our medical-specialized MLLM, Lingshu-Med, which undergoes multi-stage training to embed the medical expertise and enhance task-solving capabilities progressively. We also investigate reinforcement learning with verifiable rewards to further refine Lingshu-Med's medical reasoning abilities. For rigorous assessment, we introduce MedEvalKit, a unified evaluation framework that consolidates the leading multimodal and textual medical benchmarks for standardized, fair, and efficient model assessment. On three core medical tasks-multimodal QA, textual QA, and radiology report generation, Lingshu-Med consistently outperforms existing multimodal baselines in most tasks. Moreover, we conduct five case studies drawn from real-world clinical scenarios that illustrate its practical utility in medical contexts.

Wei-Wen Xu, H. Chan, Long Li et al. · 0 citations
#computer vision Preprint Sep 2026

On the Design Fundamentals of Pixel Text Representation Learning

This work investigates the fundamental design principles required for robust visual text representation learning and trains Pixel Linguist II, a native-resolution vision encoder trained with on-the-fly rendering, unified contrastive grounding, and a multilingual curriculum over 280M training examples.

Chaohao Yuan, Rui-Feng Yuan, Zhuoxu Huang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.