Skip to content
Open access

An explainable transformer model learning from entire treatment timelines for pan-cancer risk profiling across healthcare systems

Jul 2026 · medRxiv · 0 citations
Medicine

TL;DR

Chronicle is introduced, an explainable transformer that learns from entire patient trajectories, predicts diverse clinical outcomes throughout the disease course while capturing both short- and long-term temporal dependencies, providing a scalable framework to support individualized treatment decisions.

Abstract

Cancer outcomes vary widely between individual patients, each accumulating an irregular record of treatments, diagnoses, measurements, and complications. Current prognostic models reduce this complexity into a single snapshot, focus on narrow clinical settings, and rarely generalize across hospitals. Here we introduce Chronicle, an explainable transformer that learns from entire patient trajectories, predicts diverse clinical outcomes throughout the disease course while capturing both short- and long-term temporal dependencies. Trained on 53.7 million longitudinal data points from 51,711 patients spanning 67 cancer types, Chronicle operates natively on irregular data without imputation and jointly predicts eight endpoints within a flexible framework adaptable to additional outcomes. Chronicle outperformed cross-sectional models for overall survival prediction (C-index 0.84 vs 0.76-0.79), stratified patients more accurately than established prognostic systems, including TNM stage, and predicted seven adverse event and transfusion endpoints (AUC 0.80-0.92). Applied without retraining to 69,341 patients in Germany, Switzerland, and the United States, Chronicle generalized across healthcare systems and improved further with local fine-tuning. Integrated explainability traced each risk update to patient-specific clinical factors, revealing distinct temporal persistence of prognostic information, with relevance half-lives ranging from weeks for therapies to nearly one year for baseline characteristics. These findings demonstrate that learning from hospital-wide patient trajectories enables interpretable and continuously updated predictions, providing a scalable framework to support individualized treatment decisions.

Read PDF

Similar papers

Preprint Aug 2026

Explainable Transformer Models for Clinical Prediction Tasks on Structured Electronic Health Records

BERT-LER is presented, a BERT-style model for coded EHR timelines pretrained and fine-tuned from a de-identified EHR dataset of 75 million patients, that encodes laboratory test results as discrete tokens while retaining graded information through percentile-based binning, paired with Integrated Gradients for token-level attributions grounded in the input EHR sequence.

Jun-Ni Du, Lukas Adamek, Maxim A Kryukov et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Scalable Clinical Data Infrastructure and Comparative ML Evaluation for Hospitalisation Risk Prediction in Elderly Patients with Multiple Long-Term Conditions using CPRD

It is argued that LASSO, not the highest-discriminating model, is the model best suited to direct clinical deployment, and lessons for the machine learning and healthcare community regarding data infrastructure, model selection, and value of calibration and interpretability in high-stakes decision support are presented.

Asra Aslam, Volodymyr Chapman, M. O'Connell et al. · 0 citations
Open access Sep 2026

Generative model of patient health states and pan-cancer risk stratification

While large language models are powerful generators of new text, forecasting disease progression from longitudinal health histories remains a challenging problem. We introduce GenEHR, an autoregressive generative model trained on electronic health records (EHRs) from millions of patients that explicitly represents the irregular time intervals between visits when forecasting future clinical events. We combine the general-purpose patient representation learned during foundational training with parameter-efficient supervised adaptation for the task of pan-cancer risk stratification. In five large EHR cohorts supervised adaptation substantially improved prediction performance of a first cancer diagnosis within a five year horizon window. Our retrospective results support the evaluation of GenEHR-CancerRisk as a prospective clinical decision-support tool for prioritizing patients for risk-based screening for aggressive cancer types, such as pancreatic and ovarian cancer.

A. Khan, D. T. Forster, M. Harsh et al. · 0 citations
Open access Aug 2026

A novel methodology for predictive modeling of patient outcomes using multi-modal transformer networks and SHAP models

Precision medicine requires predictive models that can exploit genomic, clinical, imaging and continuously monitored physiological data at the same time, yet most existing models operate on a single modality and behave as black boxes. This paper proposes the Integrated Multi-Modal Contextual Network (IMCN), a predictive modelling framework that combines multi-modal transformer networks for cross-modal fusion, recurrent neural networks with attention for real-time sequential signals, pre-trained autoencoder networks for dimensionality reduction of high-dimensional genomic data, and a context-aware multi-task learning network for personalised risk and treatment predictions. SHapley Additive exPlanations (SHAP) are integrated to provide global and local feature attributions, so that clinicians can see which genomic markers, clinical variables and contextual factors drive each prediction. Across breast, lung, colorectal, cardiovascular, diabetic and chronic kidney disease cohorts derived from The Cancer Genome Atlas, the framework reports higher AUC, precision, sensitivity, recall and F1-score than the literature-reported benchmarks used for comparison, together with a reduction in false positives. Limitations, including the use of simulated physiological monitoring signals and the absence of independently re-implemented baselines, are stated explicitly.

V. R, Madala Guru Brahmam, Alagiri I · 0 citations
Preprint Aug 2026

A Multimodal Foundation Model for Longitudinal Patient Representation and Scalable Insight Generation in Oncology

The oFM is introduced, a foundation model developed on a real-world oncology cohort of 1.67 million cancer patients that integrates clinical trajectories with DNA, RNA, and H&E pathology and achieves a three-fold higher pooled and scale-normalized treatment-benefit AUTOC than baseline features.

E. Vorontsov, Yi-Kan Wang, A. Bozkurt et al. · 0 citations

Explainable AI for analyzing cancer outcomes using large-scale genome sequencing data

A multi-tier, explainable AI framework designed to risk-stratify patients and predict overall survival using clinical and genomic covariates is developed and demonstrates that explainable machine learning models can robustly predict survivability and highlight actionable features for oncology dashboards.

P. Nalela · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.