Comparing large language model performance in assessing fever in the returning traveller using clinician and AI-generated cases
Background: Fever in the returning traveller is a common but challenging presentation with a broad, geography-dependent differential diagnosis. Timely assessment can be difficult for front-line clinicians, especially outside tropical medicine settings. Large language models (LLMs) may support clinical decision-making by generating differential diagnoses, but comparative evaluation workflows remain underdeveloped. Objective: To develop and evaluate an LLM workflow to assess the reliability and efficacy of LLM-generated differential diagnoses for written fever-in-returning-traveller cases, compare performance across four LLMs, and secondarily evaluate LLM case-generation capability. Methods: We studied 21 travel-related diagnoses using four LLMs (ChatGPT 5.2 Thinking, GPT-o3, Llama 3.2, and Mistral Instruct). Infectious diseases fellows created cases across the diagnoses (n=84; 4/diagnosis), which were used to calibrate AI case-generation prompting. Each LLM then generated matched cases (n=84; 21/model). Clinician and AI-generated cases were pooled, assigned unique IDs, and randomized so each analysis model received equal numbers of clinician and AI cases (n=168 analyses; 42/model). Feasibility outcomes were completion, retries/system errors, and logged response time. Expert qualitative scoring is underway. Results: We completed dataset generation and model-output acquisition for 252 outputs across 21 diagnoses. Case-generation retries occurred in 17/84 outputs (20.2%), all first-pass refusals were from Llama 3.2, all resolved with repeat prompting. Case-analysis retries were rare (1/168, 0.6%; single no-response, resolved), with no persistent failures. Balanced allocation was achieved. Median case-analysis response times for cloud-based models were 11s (ChatGPT 5.2 Thinking) and 4s (GPT-o3). Conclusions: Preliminary findings demonstrate a feasible, scalable, and reproducible workflow for comparative LLM evaluation in fever in the returning traveller assessment. Model-specific refusal behaviour is an important implementation consideration. Ongoing expert qualitative scoring will determine comparative reliability and efficacy. Current conclusions are limited to operational feasibility and dataset generation reliability.