Skip to content
#explainable ai Open access

PREreview of "The widening evaluation gap in medical large language model research 2023 to 2026"

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/22734670. What the paper claims This is an evidence map of 11,628 PubMed-indexed records on generative language models in healthcare, January 2023 to June 2026, built to answer one question: is the clinical literature keeping pace with the systems it evaluates? The central construct is evaluation lag, the number of quarters between the release of the newest model family a study names and that study's own publication quarter. Mean lag widens from 1.33 to 6.08 quarters. Because a discontinued family ages at exactly one quarter per quarter, the authors benchmark that rise against a counterfactual holding 2023 model composition fixed, and find that migration to newer systems offsets 56.2% of the drift (95% CI 49.9-65.2, 2,000 bootstrap resamples). By design, randomised controlled trials evaluate models 4.62 quarters older than other empirical designs (95% CI 3.62-5.63), yet among the records naming a family still receiving releases no design differs from any other; what separates designs is the probability of studying a discontinued family at all, 62.1% for randomised trials against 15.6% for preprints. The framing is right and the execution is unusually careful for a bibliometric paper. Reporting the mechanical ageing rate as the null instead of zero, validating the decomposition by showing that discontinued-only records widen at 1.062 quarters per quarter, and running a pessimistic sensitivity analysis that assumes the earliest release of every named family are all things this literature normally omits. The version-specification result in Section 2 deserves to be read on its own: randomised trials specify the model version in 50.0% of cases, the lowest of any empirical design, which means half of the highest-tier evidence cannot be attributed to a determinate system. My comments below are about three places where the reported precision outruns the design. Does the conclusion hold on the data? 1. Evaluation lag is measured on a differentially selected half of the corpus, and the headline comparison is between designs that differ in exactly that selection. Lag is computable for 5,775 of 11,628 records, and coverage by design runs from 87.0% for comparative evaluations and 69.0% for randomised trials down to 36.2% for preprints and 27.0% for reviews. The paper states this plainly, which is to its credit, but the between-design contrast is its central result and the missingness is both large and design-dependent. The direction is not obvious either: if preprints name a model in the abstract mainly when that model is the novelty on sale, the preprint advantage is inflated; if trials omit the name when the system is an incidental component, the trial penalty may be understated. Two things would settle it. Worst-case bounds in the Manski sense across the plausible range would show whether the ordering survives at all, and a full-text audit of a random sample, a few hundred records stratified by design, would give an empirical estimate of how the non-naming records differ. Without one of those, the contrast rests on an assumption of ignorable missingness that the paper does not state. 2. The abstract quotes the specification that maximises the effect and omits the one that does not survive correction. The median regression gives 4.62 quarters at P = 3.4 x 10^-19; ordinary least squares on the same specification gives an adjusted difference of 1.57 quarters, P = 0.021, and after Holm correction across the seven design contrasts P = 0.084. Median regression is the defensible choice for a strongly right-skewed distribution and the paper says so, so this is not a case of specification shopping. It is a case of a reader meeting one number and not the other. The magnitude differs by a factor of about three between estimators, and the abstract should carry that range, or at minimum state that the estimate is estimator-sensitive while the ordering is not. 3. The mechanism claim is carried by 22 records, and the abstract states it as a finding rather than as the absence of evidence. Among studies naming an actively developed family, randomised trials differ from reference by 0.02 quarters (95% CI -0.50 to 0.55, n = 22). Section 2 is careful here, noting that this establishes that the gradient operates through model selection rather than demonstrating the absence of a residual timeline effect. The abstract is not: "among studies naming a model still under development no design differed from any other" reads as an equivalence result. With n = 22 the interval cannot exclude differences of half a quarter in either direction, which is not nothing when the effect being explained is a few quarters. State the equivalence bounds the sample supports, and match the abstract to the caveat already in the body. 4. The counterfactual is anchored on a very thin quarter. The 56.2% offset depends on holding model composition at its 2023 value, and 2023-Q1 contains 52 records with five distinct families, a Herfindahl-Hirschman index of 5,424 and 72.0% of mentions on one family. The paper already shows the sensitivity by reporting 64.7% when the first quarter alone is the baseline, which is a nine-point swing from the headline. Report the offset as a function of the baseline window, at least 2023-Q1, 2023-H1 and 2023 in full, each with its bootstrap interval, so the reader can see how much of the 56% is a property of the literature and how much is a property of one small quarter. Reproducibility The paper claims that harvesting, screening, classification and analysis execute end to end from deposited code, and that is the right standard for an evidence map that is meant to be regenerated as the literature grows. I could not find the archive identifier in the text I read, so the first thing to add is a repository DOI with a commit hash, alongside the verbatim PubMed query strings for all fifteen query blocks and the harvest date. The literature moves; a map without its query date cannot be reproduced even by its own authors. Two derived artefacts deserve to be released as data in their own right. The first is the model-family release-date table that defines the lag variable: every number in the paper depends on it, and it would be reusable by anyone studying model currency in any field. The second is the regular expression set used for model detection, together with the precision audit. The audit reports that 9.9% of matches were flagged for the eleven most exposed families, or 2.0% of the corpus, which is a precision estimate. Recall is not estimated anywhere, and it matters more here, because a study that names its model only in the methods section is invisible to an abstract-level regular expression, and whether that happens is plausibly design-dependent. A recall estimate from the same full-text sample proposed in point 1 would cover both issues at once. What to fix Add worst-case bounds or a full-text audit for the 50.3% of records without a computable lag, reporting coverage-adjusted design contrasts; state the missingness assumption explicitly. Put both the median-regression and the Holm-corrected OLS estimates for the randomised-trial contrast in the abstract, or state that the magnitude is estimator-sensitive. Report equivalence bounds for the active-family subgroup and align the abstract with the n = 22 caveat already given in the body. Report the counterfactual offset across several baseline windows with intervals, rather than one figure from a 52-record quarter. Give the repository DOI and commit hash, the fifteen query strings verbatim, and the harvest date. Release the model-family release-date table and the detection patterns; add a recall estimate for model detection, reported by design. State how the 12.0% of records with unattributable first affiliation are handled in the regional comparison, since the trial-grade shares compared there are 1.5% to 3.9% and that missing share could move them. The finding that computer science venues and preprint servers contain no trial-grade record should carry the PubMed-coverage caveat in the same sentence, not only in the limitations paragraph, since it is the kind of line that travels alone. Consider replacing the 45-fold growth headline with the incidence rate ratio of 1.272 per quarter (95% CI 1.257-1.289). The fold change compares two single quarters and the base quarter has 52 records; the IRR is the estimate that carries an interval. A P value of 3 x 10^-19 conveys nothing beyond the effect and its interval; consider reporting the estimate with its CI and dropping the exponent. Recommendation This should be published after revision. The central quantity is well chosen, the counterfactual is the right way to handle mechanical ageing, and the version-specification result is a concrete, actionable finding that the reporting guidelines the paper cites could enforce tomorrow. What needs work is the treatment of the half of the corpus where the key variable is undefined, and the alignment of the abstract with the uncertainty the body already acknowledges. The conclusion I would defend on this evidence is slightly narrower than the one stated, and no less interesting: the highest tier of clinical evidence is systematically about superseded systems, and the gap is driven by which systems trials select rather than by how long trials take. Competing interests: none. Evgenii Arsentev, PhD Competing interests The author declares that they have no competing interests. Use of Artificial Intelligence (AI) The author declares that they used generative AI to come up with new ideas for their review.

View source

Similar papers

#artificial intelligence Conference Open access Apr 2020

ECCOLA - a Method for Implementing Ethically Aligned AI Systems

The method, ECCOLA, is presented, which aims at making the high-level AI ethics principles more practical, making it possible for developers to more easily implement them in practice.

Ville Vakkuri, Kai-Kristian Kemell, P. Abrahamsson · 64 citations · ⚡6
#computer vision Review Apr 2024

AI-powered Code Review with LLMs: Early Results

The goal is to not only refine the accuracy of the LLM-based tool but also to underscore its potential in streamlining the software development lifecycle through proactive code improvement and education.

Z. Rasheed, Malik Abdul Sami, Muhammad Waseem et al. · 62 citations · ⚡3
#computer vision Open access Mar 2024

LLM-based agents for automating the enhancement of user story quality: An early report

The use of large language models to automatically improve the user story quality in Austrian Post Group IT agile teams is explored, with a reference model for an Autonomous LLM-based Agent System developed and implemented at the company.

Zheying Zhang, M. Rayhan, Tomas Herda et al. · 48 citations · ⚡4
#computer vision Review Mar 2024

System for systematic literature review using multiple AI agents: Concept and an empirical evaluation

This paper introduces a novel multi-AI-agent system designed to fully automate SLRs, and demonstrates how it substantially reduces the time and effort traditionally required for SLRs while maintaining comprehensiveness and precision.

Abdul Malik Sami, Z. Rasheed, Kai-Kristian Kemell et al. · 44 citations · ⚡2
#computer vision Feb 2024

Can Large Language Models Serve as Data Analysts? A Multi-Agent Assisted Approach for Qualitative Data Analysis

The proposed LLM-based multi-agent system automates qualitative data analysis process, creating opportunities for researchers and practitioners, and future improvements focus on enhancing multilingual performance and integrating continuous expert feedback.

Z. Rasheed, Muhammad Waseem, Aakash Ahmad et al. · 41 citations
#artificial intelligence Conference Open access Jun 2018

The Key Concepts of Ethics of Artificial Intelligence

It is suggested that the focus on finding keywords is the first step in guiding and providing direction for future research in the AI ethics field.

Ville Vakkuri, P. Abrahamsson · 39 citations · ⚡2

Related blog posts

Google DeepMind Blog Sep 30, 2026

Introducing SynthID Bio

Proof of concept for watermarking AI-generated proteins while preserving biological function.

MIT News · Artificial Intelligence Sep 30, 2026

This game-playing AI is the new champ at Stratego

Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.