Skip to content
Review

Can artificial intelligence accurately assess systematic review quality? Benchmarking large language models for AMSTAR 2 appraisal in dental evidence synthesis.

Aug 2026 · Evidence-Based Dentistry · 0 citations · 23 references
Medicine

TL;DR

Perplexity demonstrated the highest accuracy and agreement with expert assessments of the methodological quality of systematic reviews, suggesting its potential as a supportive AI tool for AMSTAR-2-based appraisal in dental evidence synthesis.

View source

Similar papers

Review Open access Jul 2026

A comparative analysis of readability, quality, and reliability in large language model outputs pertaining to knee osteoarthritis queries

This study aims to comparatively examine the readability, accuracy, and quality of responses provided by artificial intelligence (AI)-based chatbots such as Perplexity, ChatGPT-5, and Gemini to questions about knee osteoarthritis (KOA), which accounts for approximately four-fifths of the global osteoarthritis (OA) burd...

Erdem Maraşlı, E. Ozduran, Volkan Hancı · 1 citation
Review Open access Sep 2026

Benchmarking Large Language Model Performance in Generating and Assessing Radiology Objective Structured Clinical Examination

High-quality radiology assessment questions are essential for education competency evaluation but labor-intensive to create. To compare four large language models (LLMs) in generating and evaluating radiology objective structured clinical examination (OSCE)–style questions and responses. Fifty Radio...

Ankush Ankush, Samriddhi Burman, Sydney Smith et al. · 0 citations
#small language model Review Open access Aug 2026

Toward Automating the Selection of Articles Reporting EQ-5D Data for Systematic Literature Reviews Using Large Language Models: Algorithm Development and Evaluation Study

The models reproduced human screening tendencies despite the small dataset size, demonstrating the technical feasibility of LLM-assisted article selection and providing the first demonstration of LLM-assisted identification of EQ-5D data in biomedical literature.

Gábor Kertész, J. Czere, Z. Zrubka et al. · 0 citations
Review Sep 2026

Guiding LLM Peer Reviewers: The Impact of Score Anchors on Review Evidence and Accuracy

Large language models are increasingly used for research quality evaluation, with prior work exploring their scoring accuracy and the plausibility of review rationales exploring their scoring accuracy and the plausibility of review rationales.

Judita Preiss, YunHong Yang · 0 citations
Review Open access Sep 2026

Large Language Model versus Clinician Written Summaries of Research Papers.

An enterprise LLM, prompted in POEM style, produced accurate, low-error clinical summaries that matched or exceeded expert-edited POEMs and were generally preferred by reviewers, though further research is needed to assess broader applicability and impact.

Richard Guthmann, Robert Martin, Erin Lee et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.