OncoRAG enables accurate clinical phenotyping from multilingual oncology notes using a locally deployable mid-size model, without model weight fine-tuning or external data sharing.
Abstract
Oncology notes contain the richest clinical detail, yet they remain largely inaccessible at scale because extracting structured phenotypes requires either substantial language model infrastructure, curated training data, or cloud computing under regulatory constraints. We developed OncoRAG, combining ontology enrichment, knowledge graph construction, graph-diffusion reranking, and structured prompting with a locally deployed 14B-parameter language model without model weight fine-tuning. Applied to three cohorts—triple-negative breast cancer (TNBC; 104 patients, 42 features; primary development), recurrent high-grade glioma (RiCi; 191 patients, 19 features; cross-lingual and cross-disease evaluation with cohort-specific configuration), and MIMIC-IV (100 patients, 10 features; limited external evaluation on overlapping features)—OncoRAG achieved F1 scores of 0.80, 0.79, and 0.84, improving over direct large language model (LLM) prompting and naive retrieval-augmented generation (RAG) baselines by 0.19–0.22 and 0.17–0.19 F1, and outperforming direct prompting with a 5× larger 70B model by 0.09–0.10 F1. In an exploratory survival analysis (12 events), both feature sets showed close point estimates of the C-index (0.77 vs 0.76), but equivalence cannot be statistically confirmed given the limited event count. OncoRAG enables accurate clinical phenotyping from multilingual oncology notes using a locally deployable mid-size model, without model weight fine-tuning or external data sharing.
OncoGenRAG is a research framework that combines a parameter-efficiently fine-tuned BioBERT classifier with an entity-aware retrieval system over a curated, multi-source oncology knowledge base and provides a transparent design for evidence retrieval and abstention.
Amaan Arif, J. V. dos Santos· bioRxiv· 0 citations
The rapid expansion of the oncology literature has outpaced manual curation of clinically relevant gene-cancer-drug associations and oncogenic driver evidence. Existing automated approaches often lack transparency or are difficult to scale across heterogeneous data sources. To address this gap, we developed megaMine, a...
Muhammad Junaid, K. Prazanowska, Ha-Eun Jeong et al.· bioRxiv· 0 citations
LLMs can accurately extract clinical features from EHRs with careful document selection and prompt design and were the most accurately extracted biomarkers across these cohorts, MGMT and IDH were the most accurately extracted biomarkers.
Anna Erickson, Luke R Jackson, S. Chappidi et al.· Discover medicine· 0 citations
Phenotype-driven diagnostic benchmarks usually report the rank of the reference disease, but they rarely reveal which plausible alternatives are ranked above it or what evidence a tool-using model examines before making its decision. We introduce GraphRareBench, a provenance-preserving benchmark containing 2,365 ontolo...
LLM-extracted symptom data recover clinically coherent, reproducible SCs from unstructured discharge notes that carry independent prognostic value for mortality and readmission, supporting the clinical validity of automated, EHR-derived symptom profiling in CRC.
Y. Lee, I. Dinov, X. Hu et al.· medRxiv· 0 citations
Results indicate that VARION’s GIS-weighted centroid architecture enables individual-patient molecular subtyping that outperforms existing NBS and graph-learning clustering approaches, with high sensitivity for clinically actionable rare subtypes and robust cross- platform generalization.
Taesoo Kwon, Y. Park, J. Choi· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.