Skip to content
Review Open access

Generating guideline-concordant and safe recommendations for diabetic kidney disease management via a hierarchical retrieval-augmented large language model.

Jul 2026 · npj Digital Medicine · 0 citations
Medicine

TL;DR

This study demonstrates that anchoring LLMs with authoritative knowledge effectively mitigates hallucination risks and enhances clinical reliability, and reveals that the RAG-enhanced system significantly outperforms unaugmented models in providing accurate, guideline-compliant recommendations.

Abstract

Managing diabetic kidney disease (DKD) is inherently complex, requiring clinicians to synthesize patient history, fluctuating biomarkers, and evolving treatment guidelines. While large language models (LLMs) show promise in medical decision support, their clinical adoption is hindered by factual inaccuracies and a lack of specific reasoning required for individualized patient management. To address this, we developed a hierarchical multi-agent system that integrates a locally deployed retrieval-augmented generation (RAG) framework with a cloud-based advanced reasoning engine, grounding responses in a curated corpus of clinical guidelines. We conducted a multi-center retrospective validation using 267 patient cases. The system's performance was evaluated against baseline models through a blinded review by twelve independent physicians across clinical dimensions including accuracy, safety, and factuality. Our evaluation reveals that the RAG-enhanced system significantly outperforms unaugmented models in providing accurate, guideline-compliant recommendations. Notably, it substantially reduced safety-critical errors, particularly in identifying medication contraindications related to renal function stages, while achieving high inter-rater reliability. This study demonstrates that anchoring LLMs with authoritative knowledge effectively mitigates hallucination risks and enhances clinical reliability. The proposed framework functions as a reliable on-demand assistant for DKD management, providing guideline-grounded decision support for primary care providers.

Read PDF

Similar papers

#artificial intelligence Preprint Sep 2026

Medical Knowledge Simplification for Patients in the Era of LLMs: A Case Study on Diabetes

This case study demonstrates the potential of LLMs to improve the accessibility of medical knowledge for patient education through the implementation and evaluation of MediClear, an LLM-based medical knowledge simplification system enhanced with Retrieval-Augmented Generation (RAG).

Pallika Kafle, Yi-Peng Zhou, Guan-Feng Liu et al. · 0 citations
Review Sep 2026

[Expert consensus on evaluating large language models for aided diabetes diagnosis and treatment (2026 edition)].

This consensus is primarily intended for clinical application in China, emphasizes that physicians retain ultimate responsibility for diagnosis and treatment, and proposes a three-tier evaluation framework covering dimensions, methods, and indicators across five domains: accuracy and reliability, safety, clinical utili...

Unknown authors · 0 citations
Review Aug 2026

Enhancing Clinical Decision-Making Using Generative AI-Powered Knowledge Retrieval Systems: A Review of Emerging Approaches and Challenges

The study concludes that evidence-based, auditable, locally adaptable, locally adaptable, and supervised by licensed clinician retrieval systems with generative AI can support safer, faster, and more relevant decision-making processes in clinical settings.

Sonam Kumari · 0 citations
Aug 2026

Evaluation of Large Language Model-Generated Recommendations in Glaucoma Surgical Decision-Making.

LLM-based chatbots can provide acceptable surgical guidance for straightforward primary surgical cases, but their utility is limited in high-risk or complex clinical settings, and their rationale quality was significantly lower in complex scenarios than in primary ones.

Müge Toprak, Büşra Yılmaz Tuğan, N. Yüksel · 0 citations
#artificial intelligence Preprint Sep 2026

A Structured Debate-Mixture-of-Agents Framework for Complex Clinical Diagnostic Decision Support

Large language models (LLMs) show potential for medical tasks, but their single-turn question-answer format does not reflect how clinical diagnosis is performed in practice. As a result, they remain limited in complex diagnostic settings. We developed Debate-Mixture-of-Agents (DMoA), a novel multi-agent framework that...

Changda Xia, L. Ouyang, Hui-Min Wang et al. · 0 citations
Review Open access Aug 2026

Large Language Models and Medical AI Systems for Healthcare Diagnosis: A Systematic Review

Although promising, LLM-based systems are not yet reliable enough for autonomous medical diagnosis, and multiple recommendations for future research are contained to ensure a high level of safety, transparency, and clinical applicability for LLMs and other AI/ML-related technologies and devices.

M. U. K. Gunawardhna, Pirunthavi Wijikumar, D. Weerasinghe · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.