Skip to content

PRomop: A Decision-Ready Longitudinal Patient Health Record on the OMOP Common Data Model

Jul 2026 · arXiv.org · Vol abs/2607.13947 · 0 citations · 21 references
Computer Science

TL;DR

A flattened, decision-ready projection over a standards-based longitudinal record is a deployed pattern for turning fragmented data into actionable infrastructure, while remaining OMOP-conformant.

Abstract

Objective: Health systems and biopharma face a gap between holding patient data and acting on it: records are fragmented, manually mapped, and structured for storage rather than decisions, so every application re-derives patient state. We present PRomop, an open-source longitudinal record that closes this gap. Materials and Methods: PRomop builds on the OMOP Common Data Model (CDM 5.4) with oncology extensions and adds PatientRecord, a flattened projection collapsing each patient's longitudinal history into a single decision-ready 304-column row. State derivations - lines of therapy, disease status, normalized biomarkers - are computed once at projection time, so analytics, trial matching, and standard-of-care evaluation read one substrate. Results: PRomop is deployed by two oncology organizations - the independently governed HealthTree Foundation (~14,000 patients) and CancerBot (~3,500), a HealthKey-owned deployment - matching against 19,500 recruiting trials across five cancer types. A 20-criterion eligibility search requiring 27-39 joins over raw OMOP reduces to zero against the projection. On a synthetic 1000-patient breast-cancer cohort, eligibility screening averaged 0.30 ms via PatientRecord versus 11.0 ms from raw OMOP, a ~36.8x speedup. Discussion: The projection's significance is as a foundation for other applications: it lowers each one's marginal cost by computing error-prone clinical derivation once and removing it from every consumer. Line-of-therapy inference showed decision-readiness demands embedded clinical reasoning, and that the projection is a living artifact requiring maintenance. Conclusion: A flattened, decision-ready projection over a standards-based longitudinal record is a deployed pattern for turning fragmented data into actionable infrastructure, while remaining OMOP-conformant. Benchmarks measured a ~36.8x eligibility-screening speedup.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

ObGynLongBench: Revealing the Evidence-to-EHR Gap in Longitudinal EHR Decision-Making

The application of large language models (LLMs) to personalized medical assistants has garnered growing interest. However, existing medical benchmarks largely rely on static question answering with pre-selected evidence, leaving unclear whether LLMs can make reliable clinical decisions from real longitudinal electronic...

Jun Xiang, Zhi-Jie Bao, Rong Hu et al. · 0 citations
Jul 2026

ClinLens: Towards Long-Horizon Coding Agents for Longitudinal Multimodal Clinical Data Science

CLINLENS is introduced, a benchmark of 200 executable tasks over five linked MIMIC resources spanning structured electronic health records, notes, electrocardiograms, chest radiographs, and echocardiograms, which exposes a substantial gap between runnable submissions and correct clinical analyses.

Yuan Zhu, Ethan B. Liu, Frank Nie et al. · 0 citations
Open access Jul 2026

From fragmented records to living evidence: health system-governed, artificial intelligence-driven, continuously updated real-world clinical data from Truveta

Abstract Objectives Real-world data (RWD) have historically suffered from fragmentation, delayed availability, variable data quality, and limited analytic utility. Truveta developed an artificial intelligence (AI)-enabled data platform to address these longstanding challenges in using RWD for clinical research. This pa...

Hannah A. Burkhardt, S. Blach, Srinivasa R Burugapalli et al. · 0 citations
Open access Sep 2026

Leveraging Synthetic Clinical Data for Validation and Operational Readiness in Clinical Trials

Background/Objectives: Obtaining timely access to detailed clinical trial data is not always straightforward. Privacy requirements, governance processes, and study-specific eCRF configurations can delay access, particularly during study start-up, when teams need realistic data to develop and test validation rules, repo...

Szymon Musik, Jacek Zalewski, Julia Jurkowska et al. · 0 citations
Book Open access Aug 2026

OneEHR: Reproducible and AI Agent-Ready Longitudinal EHR Analysis Toolkit

Electronic health records support a wide spectrum of clinical prediction and decision-support studies, but reproducible EHR research now requires more than training a single predictive model. As the field expands from machine learning and deep learning to LLM-based and agentic AI, differences in cohort construction, te...

Yinghao Zhu, Zi-Xiang Wang, Lei Gu et al. · 0 citations
Open access Aug 2026

The Need for Fully-Effective Federated Analytics of Data Sources for Clinical Trials

IntroductionClinical trials are usually analysed in a single environment allowing for flexible analysis including adjustment or stratification by subgroup: `one-stage' analysis of individual-level data. Health systems datasets, often distributed across geography and providers, can streamline clinical trials. Data are i...

Stella Maris Fabiane, Sharon B. Love, D. Fisher et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.