Skip to content

Author

Xiaotang Gai

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

A multimodal vision-language model for comprehensive dental diagnosis and enhanced clinical practice

Oral diseases affect billions of people, yet specialist dental expertise remains unevenly distributed, and diagnosis often requires synthesis across diverse imaging modalities. Existing artificial intelligence systems mostly address isolated tasks, limiting their applicability in comprehensive dental assessment. Here we introduce DentVLM, a dental vision-language model that jointly interprets images and text, supports expert-level oral disease diagnosis across seven dental imaging modalities and 36 tasks. Developed using 110,447 images and 2.46 million bilingual visual question-answer pairs, DentVLM outperforms leading proprietary, open-source and domain-specific medical models on internal and external tests. In a study of 32 participants, DentVLM surpasses junior readers, matches intermediate general practitioners and approaches senior specialists. In collaborative workflows, it raises junior and intermediate readers toward specialist-level performance and reduces diagnostic time for all readers by 15.0-37.0%. These results establish DentVLM as a clinical decision support tool for reducing specialist care gaps and broadening access to high-quality dental expertise. DentVLM is a dental vision-language model developed to support dental diagnosis across seven oral imaging modalities and 36 tasks. It matches intermediate general practitioners, approaches senior specialists, and reduces diagnostic time by 15.0-37.0% in collaborative clinical workflows.

Zijie Meng, Jinxiang Hao, Xi-Wei Dai et al. · 1 citation
#artificial intelligence Preprint Aug 2026

MedUAG: Unified Understanding and Generation for Medical Multimodal Models

This work develops MedUAG, an end-to-end trained unified medical model that achieves strong performance across a wide array of understanding and generation tasks, establishing a competitive baseline and paving the way for next-generation medical multimodal systems.

Zijie Meng, Yuncheng Zhang, Hualiang Wang et al. · 0 citations