Skip to content

Author

Lori Israelian

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

Comparing large language model performance in assessing fever in the returning traveller using clinician and AI-generated cases

Background: Fever in the returning traveller is a common but challenging presentation with a broad, geography-dependent differential diagnosis. Timely assessment can be difficult for front-line clinicians, especially outside tropical medicine settings. Large language models (LLMs) may support clinical decision-making by generating differential diagnoses, but comparative evaluation workflows remain underdeveloped. Objective: To develop and evaluate an LLM workflow to assess the reliability and efficacy of LLM-generated differential diagnoses for written fever-in-returning-traveller cases, compare performance across four LLMs, and secondarily evaluate LLM case-generation capability. Methods: We studied 21 travel-related diagnoses using four LLMs (ChatGPT 5.2 Thinking, GPT-o3, Llama 3.2, and Mistral Instruct). Infectious diseases fellows created cases across the diagnoses (n=84; 4/diagnosis), which were used to calibrate AI case-generation prompting. Each LLM then generated matched cases (n=84; 21/model). Clinician and AI-generated cases were pooled, assigned unique IDs, and randomized so each analysis model received equal numbers of clinician and AI cases (n=168 analyses; 42/model). Feasibility outcomes were completion, retries/system errors, and logged response time. Expert qualitative scoring is underway. Results: We completed dataset generation and model-output acquisition for 252 outputs across 21 diagnoses. Case-generation retries occurred in 17/84 outputs (20.2%), all first-pass refusals were from Llama 3.2, all resolved with repeat prompting. Case-analysis retries were rare (1/168, 0.6%; single no-response, resolved), with no persistent failures. Balanced allocation was achieved. Median case-analysis response times for cloud-based models were 11s (ChatGPT 5.2 Thinking) and 4s (GPT-o3). Conclusions: Preliminary findings demonstrate a feasible, scalable, and reproducible workflow for comparative LLM evaluation in fever in the returning traveller assessment. Model-specific refusal behaviour is an important implementation consideration. Ongoing expert qualitative scoring will determine comparative reliability and efficacy. Current conclusions are limited to operational feasibility and dataset generation reliability.

Bhavya Gandhi, Leo Morjaria, Lori Israelian et al. · 0 citations