Can artificial intelligence accurately assess systematic review quality? Benchmarking large language models for AMSTAR 2 appraisal in dental evidence synthesis.
Aug 2026· Evidence-Based Dentistry· 0 citations· 23 references
Medicine
TL;DR
Perplexity demonstrated the highest accuracy and agreement with expert assessments of the methodological quality of systematic reviews, suggesting its potential as a supportive AI tool for AMSTAR-2-based appraisal in dental evidence synthesis.
This study aims to comparatively examine the readability, accuracy, and quality of responses provided by artificial intelligence (AI)-based chatbots such as Perplexity, ChatGPT-5, and Gemini to questions about knee osteoarthritis (KOA), which accounts for approximately four-fifths of the global osteoarthritis (OA) burd...
Erdem Maraşlı, E. Ozduran, Volkan Hancı· PLoS ONE· 1 citation
High-quality radiology assessment questions are essential for education competency evaluation but labor-intensive to create.
To compare four large language models (LLMs) in generating and evaluating radiology objective structured clinical examination (OSCE)–style questions and responses.
Fifty Radio...
Ankush Ankush, Samriddhi Burman, Sydney Smith et al.· Radiology Advances· 0 citations
The models reproduced human screening tendencies despite the small dataset size, demonstrating the technical feasibility of LLM-assisted article selection and providing the first demonstration of LLM-assisted identification of EQ-5D data in biomedical literature.
Gábor Kertész, J. Czere, Z. Zrubka et al.· JMIR Formative Research· 0 citations
Large language models are increasingly used for research quality evaluation, with prior work exploring their scoring accuracy and the plausibility of review rationales exploring their scoring accuracy and the plausibility of review rationales.
An enterprise LLM, prompted in POEM style, produced accurate, low-error clinical summaries that matched or exceeded expert-edited POEMs and were generally preferred by reviewers, though further research is needed to assess broader applicability and impact.
Richard Guthmann, Robert Martin, Erin Lee et al.· Journal of the American Boar...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.