Skip to content

Right for the Wrong Reasons: A Benchmark for Hallucination and Clinical Safety in AI Health Triage

· 0 citations · 9 references

TL;DR

An empirical benchmark for evaluating clinical triage systems that assesses explanation quality alongside decision outcomes, and provides a reproducible, checkpoint-based evaluation pipeline and outline a roadmap for bias stress-testing, hallucination mitigation, and open benchmark release.

View source

Similar papers

Review Open access Sep 2026

Retrieve-Then-Verify for Evaluating Evidence Support and Hallucination in Large Language Model-Generated Medical Information: Empirical Study.

Retrieval-based evidence verification provides a reproducible and transparent approach for evaluating the reliability of AI-generated medical information, with direct relevance to digital health practice, evidence-based medicine, and medical informatics.

Zhao-Hui Liang, Cynthia Sheffield, Gisela Butera et al. · 0 citations
Conference Open access Sep 2026

Factual Hallucination in Medical Large Language Models: Typology, Evaluation, and Systematic Governance

Large language models (LLMs) have demonstrated transformative potential in clinical documentation generation, diagnostic assistance, and patient consultation. However, their tendency toward “hallucination”— generating semantically fluent but factually inconsistent content with established medical knowledge or input con...

Zi-Jiao Liu · 0 citations
Review Open access 2026

AI Hallucination And Fabricated References: A Growing Crisis For Medical Researchers A Narrative Review

Large language models (LLMs) have introduced a new research integrity threat into the biomedical literature i.e. fabricated references that appear authentic but correspond to no existing publication. This narrative review distinguishes fabrication, the invention of an entirely non-existent source, from unfaithfulness,...

M. Rana, M. Alkhlewi, Turki Abdulaziz Alsohaibani et al. · 0 citations
2026

Research protocol: Evaluating Chain-of-Thought Reasoning in LLMs for Complex Clinical Psychiatric Cases

Background Psychiatric medicine presents unique diagnostic and therapeutic challenges, often involving multimorbidity, polypharmacy, and atypical presentations requiring complex reasoning. Artificial Intelligence, particularly Large Language Models (LLMs), is emerging as a support tool in these settings. However, the c...

Vittorio De Vita, Bianca Destro Castaniti, Antonio Cristiano et al. · 0 citations
Preprint Aug 2026

The Judgment-Consequence Gap: LLM Moral Reasoning in Healthcare Decisions

As large language models (LLMs) enter high-stakes domains such as healthcare, understanding their moral reasoning becomes essential. Decisions about scarce medical resources often hinge on judgments of responsibility, particularly when patients'own actions contribute to illness. We investigate how LLMs reason about res...

Hadi Hosseini, Samarth Khanna, Leona Pierce · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.