A flattened, decision-ready projection over a standards-based longitudinal record is a deployed pattern for turning fragmented data into actionable infrastructure, while remaining OMOP-conformant.
Abstract
Objective: Health systems and biopharma face a gap between holding patient data and acting on it: records are fragmented, manually mapped, and structured for storage rather than decisions, so every application re-derives patient state. We present PRomop, an open-source longitudinal record that closes this gap. Materials and Methods: PRomop builds on the OMOP Common Data Model (CDM 5.4) with oncology extensions and adds PatientRecord, a flattened projection collapsing each patient's longitudinal history into a single decision-ready 304-column row. State derivations - lines of therapy, disease status, normalized biomarkers - are computed once at projection time, so analytics, trial matching, and standard-of-care evaluation read one substrate. Results: PRomop is deployed by two oncology organizations - the independently governed HealthTree Foundation (~14,000 patients) and CancerBot (~3,500), a HealthKey-owned deployment - matching against 19,500 recruiting trials across five cancer types. A 20-criterion eligibility search requiring 27-39 joins over raw OMOP reduces to zero against the projection. On a synthetic 1000-patient breast-cancer cohort, eligibility screening averaged 0.30 ms via PatientRecord versus 11.0 ms from raw OMOP, a ~36.8x speedup. Discussion: The projection's significance is as a foundation for other applications: it lowers each one's marginal cost by computing error-prone clinical derivation once and removing it from every consumer. Line-of-therapy inference showed decision-readiness demands embedded clinical reasoning, and that the projection is a living artifact requiring maintenance. Conclusion: A flattened, decision-ready projection over a standards-based longitudinal record is a deployed pattern for turning fragmented data into actionable infrastructure, while remaining OMOP-conformant. Benchmarks measured a ~36.8x eligibility-screening speedup.
The application of large language models (LLMs) to personalized medical assistants has garnered growing interest. However, existing medical benchmarks largely rely on static question answering with pre-selected evidence, leaving unclear whether LLMs can make reliable clinical decisions from real longitudinal electronic...
Jun Xiang, Zhi-Jie Bao, Rong Hu et al.· 0 citations
CLINLENS is introduced, a benchmark of 200 executable tasks over five linked MIMIC resources spanning structured electronic health records, notes, electrocardiograms, chest radiographs, and echocardiograms, which exposes a substantial gap between runnable submissions and correct clinical analyses.
Yuan Zhu, Ethan B. Liu, Frank Nie et al.· arXiv.org· 0 citations
Abstract Objectives Real-world data (RWD) have historically suffered from fragmentation, delayed availability, variable data quality, and limited analytic utility. Truveta developed an artificial intelligence (AI)-enabled data platform to address these longstanding challenges in using RWD for clinical research. This pa...
Hannah A. Burkhardt, S. Blach, Srinivasa R Burugapalli et al.· JAMIA Open· 0 citations
Background/Objectives: Obtaining timely access to detailed clinical trial data is not always straightforward. Privacy requirements, governance processes, and study-specific eCRF configurations can delay access, particularly during study start-up, when teams need realistic data to develop and test validation rules, repo...
Szymon Musik, Jacek Zalewski, Julia Jurkowska et al.· Healthcare· 0 citations
Electronic health records support a wide spectrum of clinical prediction and decision-support studies, but reproducible EHR research now requires more than training a single predictive model. As the field expands from machine learning and deep learning to LLM-based and agentic AI, differences in cohort construction, te...
Yinghao Zhu, Zi-Xiang Wang, Lei Gu et al.· Proceedings of the 32nd ACM...· 0 citations
IntroductionClinical trials are usually analysed in a single environment allowing for flexible analysis including adjustment or stratification by subgroup: `one-stage' analysis of individual-level data. Health systems datasets, often distributed across geography and providers, can streamline clinical trials. Data are i...
Stella Maris Fabiane, Sharon B. Love, D. Fisher et al.· International Journal of Pop...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.