LLMs can provide limited auxiliary value in cause of death analysis but should not replace the final judgment of forensic experts, and open source LLMs can further mitigate data privacy concerns and provide practical support for cause of death analysis.
Abstract
Introduction Large language models (LLMs) have been proposed as decision support tools in medicine, yet their role in forensic cause of death analysis remains unexplored. Methods In this study, we used 118 real-world cases spanning diverse categories of death to systematically evaluate the performance of four representative LLMs (GPT-4o, OpenAI o3, Gemini-2.5pro, and DeepSeek-R1) in forensic cause of death analysis. Two senior forensic pathologists independently evaluated each model’s decision-making capabilities regarding inference quality and conclusion accuracy. These metrics were assessed using an expert scoring system with a 5-point Likert scale, with original analytical statements and legally valid expert opinions serving as objective gold standards. In a sub-study, we examined the application potential of the locally deployed open-source model DeepSeek-R1:32b. Additionally, a targeted retrospective analysis was conducted to quantify the incidence and typologies of AI hallucinations. Results DeepSeek-R1 demonstrated a statistically significant advantage in inference quality scores over GPT-4o (p = 0.015, rrb = 0.28) and Gemini-2.5pro (p = 0.000003, rrb = 0.46), while no statistically significant differences were observed among the four models in terms of conclusion accuracy scores. The locally deployed DeepSeek-R1:32b model also showed no statistically significant difference from GPT-4o in conclusion accuracy scores. However, hallucinations persistently appear in the response reports of all LLMs. Discussion LLMs can provide limited auxiliary value in cause of death analysis but should not replace the final judgment of forensic experts. LLMs still require expert oversight to ensure evidence integrity and mitigate risks such as hallucination. Open source LLMs can further mitigate data privacy concerns and provide practical support for cause of death analysis.
Large language models (LLMs) have demonstrated transformative potential in clinical documentation generation, diagnostic assistance, and patient consultation. However, their tendency toward “hallucination”— generating semantically fluent but factually inconsistent content with established medical knowledge or input con...
Background: The number of cases, subjectivity of the examiner and difficulty of integrating the different modes of evidence are all increasing challenges in forensic pathology. Whilst identification of the cause of death is strongly provided for, and characterization of trauma is a critical unmet need, neither can be a...
Nedha Abdulla P· International Journal of Adv...· 0 citations
Current performance estimates of LLMs with respect to depression screening are most likely optimistic, but when restricted to smaller models that could be locally deployed (for privacy protection) in a clinical setting, LLMs do not detect depression with sufficient accuracy, sensitivity, or specificity to be used in a...
Sing-Hui Ling, W. Chorney· Acta Psychologica· 0 citations
A reproducible estimate of diagnostic retrieval accuracy across four widely used model configurations is provided to establish a baseline for further clinical validation and establish a baseline for further clinical validation.
Lalwani Saurabh, Bodetti Dr.Vishala, Gor Kishan et al.· Indian Journal of Computer S...· 0 citations
IyawoBench v2.0 provides both a rigorous benchmark and a diagnostic framework transferable to any triage-style clinical AI evaluation, and proposes the Escalation Bias Index and Expected Deployment Cost as novel metrics that expose failure modes hidden by conventional accuracy and sensitivity scores.
LLMs demonstrated agreement with IMBD recommendations on standardized medico-legal case vignettes, supporting further investigation of their potential role as AI-assisted decision-support tools under expert supervision.
H. Aydoğan, Muhammet Sevindik, Zeynep Unat Öztürk et al.· Healthcare· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.