Skip to content

Author

Zijie Meng

We have 4 of 28 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

A multimodal vision-language model for comprehensive dental diagnosis and enhanced clinical practice

Oral diseases affect billions of people, yet specialist dental expertise remains unevenly distributed, and diagnosis often requires synthesis across diverse imaging modalities. Existing artificial intelligence systems mostly address isolated tasks, limiting their applicability in comprehensive dental assessment. Here we introduce DentVLM, a dental vision-language model that jointly interprets images and text, supports expert-level oral disease diagnosis across seven dental imaging modalities and 36 tasks. Developed using 110,447 images and 2.46 million bilingual visual question-answer pairs, DentVLM outperforms leading proprietary, open-source and domain-specific medical models on internal and external tests. In a study of 32 participants, DentVLM surpasses junior readers, matches intermediate general practitioners and approaches senior specialists. In collaborative workflows, it raises junior and intermediate readers toward specialist-level performance and reduces diagnostic time for all readers by 15.0-37.0%. These results establish DentVLM as a clinical decision support tool for reducing specialist care gaps and broadening access to high-quality dental expertise. DentVLM is a dental vision-language model developed to support dental diagnosis across seven oral imaging modalities and 36 tasks. It matches intermediate general practitioners, approaches senior specialists, and reduces diagnostic time by 15.0-37.0% in collaborative clinical workflows.

Zijie Meng, Jinxiang Hao, Xi-Wei Dai et al. · 1 citation
Preprint Aug 2026

Motion Beyond Morphology: Bootstrapping Cross-Category Motion Transfer from Abstract Motion Representations

Video motion transfer aims to animate a target object using dynamics from a reference video. Existing formulations largely rely on fixed structural correspondence, which becomes ill-defined when reference and target objects differ substantially in morphology, articulation, or deformation mechanisms. We introduce Motion Beyond Morphology, a perspective that seeks to transfer motion beyond fixed structural correspondence, by preserving dynamics that remain meaningful across different target morphologies. To realize this, we propose a two-stage framework. Stage~I learns complementary multi-granularity abstract motion views and uses them to bootstrap cross-category video pairs that preserve transferable dynamics across diverse morphologies. Stage~II internalizes this supervision into direct reference-video-conditioned generation, removing the need for explicit motion extraction at inference. We further introduce OpenVMT-Dataset and OpenVMT-Bench for training and evaluating image- and text-conditioned motion transfer across Same, Near, and Far category gaps. Extensive experiments demonstrate state-of-the-art motion fidelity and target preservation. Project page: https://miniz233.github.io/MotionBeyondMorphology/

Zhixue Fang, Zhimin Zhang, Bi'an Du et al. · 0 citations
#artificial intelligence Preprint Aug 2026

MedUAG: Unified Understanding and Generation for Medical Multimodal Models

This work develops MedUAG, an end-to-end trained unified medical model that achieves strong performance across a wide array of understanding and generation tasks, establishing a competitive baseline and paving the way for next-generation medical multimodal systems.

Zijie Meng, Yuncheng Zhang, Hualiang Wang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

DentAgent: Evidence-Centric Multi-Agent Coordination for Multimodal Dental Reasoning

DentAgent is introduced, an evidence-centric multi-agent framework, in which the Orchestrator coordinate five specialized agents spanning various modalities, which supports its value for broadly applicable and traceable multimodal dental reasoning, and highlights its potential as a technical foundation for population oral health assessment and management.

Zijie Meng, Xi-Wei Dai, Yixuan Tang et al. · 0 citations