Skip to content
Open access

Evidence Use and Identifier-Conditioned Prior Knowledge in Large Language Model Classification of Oncology Trials Assessed Through Progressive Content Removal and Counterfactual Testing: Comparative Analysis

Jul 2026 · JMIR AI · Vol 5, pp. e95565-e95565 · 1 citation · 18 references
Medicine

TL;DR

Testing whether oncology randomized trial success classification is driven by abstract evidence or by identifier-conditioned prior knowledge, and whether models follow counterfactual outcome evidence when it conflicts with original trial identifiers showed that identifiers can carry predictive signal and occasionally compete with textual evidence.

Abstract

Abstract Background Large language models (LLMs) can accurately classify biomedical documents, but strong benchmark performance does not establish that predictions are grounded in the supplied text. In biomedical literature tasks, titles, abstracts, digital object identifiers (DOIs), journal metadata, and trial identifiers may have been seen during pretraining and can trigger parametric knowledge or learned associations. Objective This study aimed to test whether oncology randomized trial success classification is driven by abstract evidence or by identifier-conditioned prior knowledge, and assess whether models follow counterfactual outcome evidence when it conflicts with original trial identifiers. Methods We evaluated 250 two-arm oncology randomized controlled trials from 7 major journals published between 2005 and 2023, each with a single primary endpoint and previously adjudicated positive or negative ground-truth label. The corpus included 58.4% (146/250) positive and 41.6% (104/250) negative trials. GPT-5.2, Gemini 3 Flash, and Claude Opus 4.5 were queried via vendor APIs under default settings using a single-token output instruction. For each trial, we created 5 deterministic input conditions: title+abstract, title only, DOI only, counterfactual title+abstract in which the primary endpoint outcome statement was minimally flipped, and the same counterfactual input paired with the original DOI to create an identifier-text conflict. Performance was assessed using valid format rate, accuracy, sensitivity, specificity, and F1-score. Results The models showed high format adherence, with valid prediction rates of 97.2% to 100%. In the title+abstract condition, all models achieved high and balanced performance (accuracy and F1-score=0.96-0.97; sensitivity=0.96-0.97; specificity=0.96-0.98). Removing evidence reduced performance stepwise: title-only accuracy and F1-score fell to 0.79 to 0.88, and DOI-only performance fell to 0.63-0.67, exceeding the 58.4% majority class baseline but indicating limited identifier-driven signal. Counterfactual edits were concentrated in outcome-bearing text, with the Results and Conclusions sections modified for all trials, whereas the titles and Methods sections required edits in only 5.2% (13/250) and 1.6% (4/250) of trials. Against inverted labels, models followed counterfactual evidence with near-ceiling performance (accuracy and F1-score=0.96-0.99). Reintroducing the original DOI caused little change for GPT-5.2 (accuracy and F1-score=0.99) but modestly reduced F1-scores for Gemini (0.97) and Claude (0.95), mainly through lower sensitivity. Conclusions The evaluated LLMs robustly followed explicit end point statements in abstracts, including when those statements contradicted original trial outcomes. However, above-chance title-only and DOI-only performance, together with small decrements under counterfactual DOI conflicts, showed that identifiers can carry predictive signal and occasionally compete with textual evidence. Progressive content removal combined with counterfactual identifier-text conflicts offers a practical, reproducible audit for grounding in biomedical LLM evaluations.

Read PDF

Similar papers

Open access Sep 2026

More signal versus more noise: comparing full text and abstract as inputs for large language model-based classification of oncology trial eligibility criteria

Abstract Objectives Large language models (LLMs) offer significant potential for automating clinical trial classification by eligibility criteria. However, the optimal input data remain unclear: while abstracts provide a condensed signal, full-text articles contain substantially more information. Whether this additiona...

J. Weyrich, F. Dennstädt, Robert Förster et al. · 0 citations
Review 2026

Enhancing Clinical Trial Analysis through Large Language Models for Multi-Evidence Natural Language Inference

It is demonstrated that modern LLMs with reasoning capabilities can effectively support real-time clinical evidence synthesis without task-specific fine-tuning, offering a pathway toward scalable automated systems for clinical trial interpretation that could substantially reduce the evidence-to-practice gap in medical...

Shobanapriyan Chandrasegaran, Amal Htait · 0 citations
#artificial intelligence Preprint Sep 2026

CLEAR: Cross-Source Evidence Adjudication for Large Language Models in Medicine

CLEAR is proposed, an agentic framework for cross-source evidence adjudication in LLMs in medicine that independently generates candidate answers from three complementary pathways---parametric knowledge, locally curated corpora, and dynamically retrieved evidence---reflecting three common sources of information availab...

Shuai Wang, Yi-Ze Zhao, Qing-Yu Chen · 0 citations
Review Open access Aug 2026

Evaluating Clinical Concept Extraction and Evidence-Bounded Terminology Linking: Multisite Model Comparison and Pilot Ablation Study

Measured extraction performance varied substantially by matching definition, whereas exact-link decisions varied with the availability of matched terminology evidence, support separate evaluation of extraction, retrieval, evidence-grounded linking, and extension candidacy.

Y. Chen, M. Popescu · 0 citations
#small language model Review Open access Aug 2026

Toward Automating the Selection of Articles Reporting EQ-5D Data for Systematic Literature Reviews Using Large Language Models: Algorithm Development and Evaluation Study

The models reproduced human screening tendencies despite the small dataset size, demonstrating the technical feasibility of LLM-assisted article selection and providing the first demonstration of LLM-assisted identification of EQ-5D data in biomedical literature.

Gábor Kertész, J. Czere, Z. Zrubka et al. · 0 citations
Review Open access Jul 2026

Future promise, current clinical ambiguity: a systematic review of machine learning algorithm outputs predicting risk of cardiovascular disease

Abstract Objective To examine whether the outputs of machine learning algorithms designed to predict risk of cardiovascular disease (CVD) address known deficiencies of the Framingham Risk Score (FRS) and improve risk estimates. Methods For this critical review, Medline, Embase and IEEE were searched from inception to 1...

S. Phillips, Han Han · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.