Skip to content
Review

[Expert consensus on evaluating large language models for aided diabetes diagnosis and treatment (2026 edition)].

Unknown authors
Sep 2026 · Zhonghua yi xue za zhi · Vol 106 33, pp. 3478-3491 · 0 citations
Medicine

TL;DR

This consensus is primarily intended for clinical application in China, emphasizes that physicians retain ultimate responsibility for diagnosis and treatment, and proposes a three-tier evaluation framework covering dimensions, methods, and indicators across five domains: accuracy and reliability, safety, clinical utility and value, user experience and interactivity, ethics and compliance.

View source

Similar papers

Open access Aug 2026

Evaluation methods for Traditional Chinese Medicine large language models: A clinical practice study

To address the critical gap that existing medical large language model evaluation systems are predominantly based on Western medical paradigms and lack specialized assessment standards for the field of traditional Chinese medicine (TCM), this study constructs a comprehensive evaluation framework specifically design...

Nanxing Xian, Wen Zhu, Lei Zhang et al. · 0 citations
Review Open access Mar 2026

Artificial Intelligence in Healthcare Practice: Validation, Fairness, and Regulatory Challenges: A Systematic Review

AI demonstrates strong potential to improve the effectiveness, safety, and quality of healthcare, however, broader clinical adoption remains constrained by regulatory requirements, interpretability gaps, data quality issues, and workflow integration challenges, underscoring the need for stronger validation practices an...

Ghulam Hussain Noori, Shaista Bibi, Seung Won Lee · 0 citations
Review Open access Aug 2026

Large Language Models and Medical AI Systems for Healthcare Diagnosis: A Systematic Review

Although promising, LLM-based systems are not yet reliable enough for autonomous medical diagnosis, and multiple recommendations for future research are contained to ensure a high level of safety, transparency, and clinical applicability for LLMs and other AI/ML-related technologies and devices.

M. U. K. Gunawardhna, Pirunthavi Wijikumar, D. Weerasinghe · 0 citations
Open access Aug 2026

Stepwise Diagnostic Evaluation of Chinese Large Language Models: Comparative Study of Common and Rare Diseases

Evaluated large language models for common diseases and rare diseases using clinical vignettes within a hypothetico-deductive framework demonstrated relatively strong diagnostic performance for common diseases such as COPD, but lower and less stable performance for rare diseases such as RP.

Jiayi Wang, Jiao Yang, Rui Guo · 0 citations
#artificial intelligence Preprint Sep 2026

Medical Knowledge Simplification for Patients in the Era of LLMs: A Case Study on Diabetes

This case study demonstrates the potential of LLMs to improve the accessibility of medical knowledge for patient education through the implementation and evaluation of MediClear, an LLM-based medical knowledge simplification system enhanced with Retrieval-Augmented Generation (RAG).

Pallika Kafle, Yi-Peng Zhou, Guan-Feng Liu et al. · 0 citations
Open access Aug 2026

The Potential of Large Language Models in Family Medicine Practice: Measuring the Artificial Intelligence Anxiety Among Family Medicine Residents

Large language models (LLMs) are increasingly being proposed as decision-support and administrative tools in primary healthcare. However, their integration into family medicine, a discipline grounded in continuity, trust, and holistic decision-making, may be shaped by clinicians’ emotional, ethical, and professional co...

G. Kolcu, Nebahat Bilge Balım, Funda Yıldırım Baş et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.