Performance, Failures, and Oversight of a Large Language Model Agent for Clinical Data Analysis: Evaluation Study
Abstract Background Large language model (LLM) agents capable of generating and executing statistical code from natural language may broaden access to clinical data analysis, yet which pipeline stages they perform reliably and which require expert oversight remain poorly defined. Objective This study aimed to evaluate...