Large language models generate diagnostic likelihood ratios with low mean bias but wide dispersion.
Accurate, context-appropriate likelihood ratios (LRs) are needed for Bayesian diagnosis, but empirical LRs are sparse because diagnostic accuracy studies are costly and context-dependent. We evaluated whether large language models (LLMs) can generate LR outputs that align with literature-reported values. We compared LR...