ABSTRACT Large language models (LLMs) represent a type of generative artificial intelligence (GenAI) that generate and interpret text, with some LLMs able to process multimodal content (e.g., images, audio, video), and can be deployed as part of agents to perform users' tasks. LLMs can perform natural language processi...
J. Gwinnutt, R. D. de Oliveira, Miriam J. Haviland et al.· Pharmacoepidemiology and Dru...· 0 citations
We propose a scalable, validity-oriented pipeline for evaluating biomedical LLM judges when high-quality human judgments are scarce. First, we augment existing human-labelled biomedical benchmarks with deterministic, metric-grounded mutations that produce auditable preference pairs. Second, we evaluate judges beyond ag...
R. D. de Oliveira, Federico Pittino, J. Gwinnutt et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.