Adequately powered analyses in precision oncology often require combining cohorts across institutions. Yet integration is constrained by the least granular source and may become infeasible when data elements are too heterogeneous to harmonize and map to a common data model. This challenge is acute in multi-institutional precision oncology research, where real-world evidence requires harmonized clinico-omic data integration. Existing models often lack sufficient treatment patterns, outcomes, and genomic data, limiting interoperability and scalability. To address these gaps, AACR Project GENIE™ (Genomics Evidence Neoplasia Information Exchange) developed the GENIE Data Model (GDM), a comprehensive, open-source, oncology data model for scalable, consistent, and interoperable data collection across solid tumors designed to effectively capture the patient's journey with cancer. Through iterative consensus-building, four working groups comprising 13 subject matter experts defined data elements across multiple clinical domains: patient characteristics, imaging, diagnosis, surgery, histopathology, biomarkers, systemic therapy, radiation, clinical trial history, disease response and outcomes, and social determinants of health. Elements were defined using standardized terminologies and permissible values to support mapping to HL7 FHIR, OMOP, and other existing oncology standards. The model architecture distinguishes manually abstracted elements from computationally collected elements, enabling parallel workflows. The GDM provides an extensible framework that addresses critical gaps and enables scalable, harmonized data collection essential for precision oncology and real-world evidence generation.
J. Hoppe, Jocelyn Lee, Tomi F. Akinyemiju et al.· Cancer Research Communicatio...· 0 citations
BACKGROUND
Individuals with substance use disorders (SUD) are obtaining health-related information from various large language models (LLMs). We aimed to assess whether LLMs provide responses concordant with the current evidence base and whether they provide harmful responses.
METHODS
Twenty questions related to SUD were posed to three LLMs (Gemini-1.5-pro-001, Claude-3-5-sonnet, and GPT-4) in May 2024. Each response was independently rated by three experienced addiction specialists, and disagreements were resolved by two additional experienced addiction specialists. All raters were blinded to the LLM. Each rater assessed whether (I) a competent addiction specialist would agree with the response, (II) the response contained stigmatizing language as defined by National Institute on Drug Abuse, or (III) the response contained harmful content.
RESULTS
88% of responses were rated as competent and 92% as not harmful. Gemini-1.5-pro-001 had the highest rate of competence (95%), followed by Claude-3-5-sonnet and GPT-4 (both 85%). Gemini-1.5-pro-001 produced no harmful responses, while Claude-3-5-sonnet and GPT-4 produced 10% and 15%, respectively. 30% of responses from both Gemini-1.5-pro-001 and Claude-3-5-sonnet had contained stigmatizing language, compared to 10% for GPT-4.
CONCLUSIONS
While many LLMs provided competent and safe responses, none were completely competent and non-stigmatizing, highlighting the potential but also ongoing need for refinement and expert verification.
Samuel Maddams, Shan Chen, Danielle S. Bitterman et al.· Substance Use & Misuse· 0 citations