Skip to content
Open access

Can large language models serve as consultants for forensic cause of death analysis? A multidimensional evaluation

Jul 2026 · Frontiers in Artificial Intelligence · Vol 9 · 0 citations · 53 references
Medicine

TL;DR

LLMs can provide limited auxiliary value in cause of death analysis but should not replace the final judgment of forensic experts, and open source LLMs can further mitigate data privacy concerns and provide practical support for cause of death analysis.

Abstract

Introduction Large language models (LLMs) have been proposed as decision support tools in medicine, yet their role in forensic cause of death analysis remains unexplored. Methods In this study, we used 118 real-world cases spanning diverse categories of death to systematically evaluate the performance of four representative LLMs (GPT-4o, OpenAI o3, Gemini-2.5pro, and DeepSeek-R1) in forensic cause of death analysis. Two senior forensic pathologists independently evaluated each model’s decision-making capabilities regarding inference quality and conclusion accuracy. These metrics were assessed using an expert scoring system with a 5-point Likert scale, with original analytical statements and legally valid expert opinions serving as objective gold standards. In a sub-study, we examined the application potential of the locally deployed open-source model DeepSeek-R1:32b. Additionally, a targeted retrospective analysis was conducted to quantify the incidence and typologies of AI hallucinations. Results DeepSeek-R1 demonstrated a statistically significant advantage in inference quality scores over GPT-4o (p = 0.015, rrb = 0.28) and Gemini-2.5pro (p = 0.000003, rrb = 0.46), while no statistically significant differences were observed among the four models in terms of conclusion accuracy scores. The locally deployed DeepSeek-R1:32b model also showed no statistically significant difference from GPT-4o in conclusion accuracy scores. However, hallucinations persistently appear in the response reports of all LLMs. Discussion LLMs can provide limited auxiliary value in cause of death analysis but should not replace the final judgment of forensic experts. LLMs still require expert oversight to ensure evidence integrity and mitigate risks such as hallucination. Open source LLMs can further mitigate data privacy concerns and provide practical support for cause of death analysis.

Read PDF

Similar papers

Conference Open access Sep 2026

Factual Hallucination in Medical Large Language Models: Typology, Evaluation, and Systematic Governance

Large language models (LLMs) have demonstrated transformative potential in clinical documentation generation, diagnostic assistance, and patient consultation. However, their tendency toward “hallucination”— generating semantically fluent but factually inconsistent content with established medical knowledge or input con...

Zi-Jiao Liu · 0 citations
Open access Jul 2026

Explainable AI-Based Clinical Decision Support System for Early Prediction of Pathological Findings in Forensic Practice

Background: The number of cases, subjectivity of the examiner and difficulty of integrating the different modes of evidence are all increasing challenges in forensic pathology. Whilst identification of the cause of death is strongly provided for, and characterization of trauma is a critical unmet need, neither can be a...

Nedha Abdulla P · 0 citations
Open access Aug 2026

The use of large language models in automated depression detection.

Current performance estimates of LLMs with respect to depression screening are most likely optimistic, but when restricted to smaller models that could be locally deployed (for privacy protection) in a clinical setting, LLMs do not detect depression with sufficient accuracy, sensitivity, or specificity to be used in a...

Sing-Hui Ling, W. Chorney · 0 citations
Open access Aug 2026

Evaluation of Diagnostic Accuracy of Open-Source and Proprietary Large Language Models Across Multi-System Clinical Cases

A reproducible estimate of diagnostic retrieval accuracy across four widely used model configurations is provided to establish a baseline for further clinical validation and establish a baseline for further clinical validation.

Lalwani Saurabh, Bodetti Dr.Vishala, Gor Kishan et al. · 0 citations
Jul 2026

IyawoBench v2.0: Extended Diagnostic Evaluation of Large Language Model Clinical Triage in Nigerian Primary Care

IyawoBench v2.0 provides both a rigorous benchmark and a diagnostic framework transferable to any triage-style clinical AI evaluation, and proposes the Escalation Bias Index and Expected Deployment Cost as novel metrics that expose failure modes hidden by conventional accuracy and sensitivity scores.

Anthonio Oladimeji Gabriel, Dimeji AbdulSobur Olawuyi · 0 citations
Open access Jul 2026

Evaluating Large Language Models for AI-Assisted Decision Support in Legal Capacity Assessment: A Comparative Study Using Interdisciplinary Medical Board Recommendations as the Expert Medical Reference Standard

LLMs demonstrated agreement with IMBD recommendations on standardized medico-legal case vignettes, supporting further investigation of their potential role as AI-assisted decision-support tools under expert supervision.

H. Aydoğan, Muhammet Sevindik, Zeynep Unat Öztürk et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.