Reference proteomes incompletely represent proteins translated in cancer, leaving tumor-specific proteoforms outside the search space of conventional mass spectrometry (MS). Such “dark proteome” products may arise from genomic variation, aberrant transcription or splicing, and non-canonical translation, including microproteins encoded by small open reading frames (ORFs). To define this landscape in acute myeloid leukemia (AML), we developed a cohort-informed proteogenomic strategy using paired RNA-sequencing and MS analysis of 123 human patient AML specimens and 13 healthy CD34+ controls. ProteomeGenerator2 was used for de novo transcriptome assembly and ORF prediction, and candidate cancer-specific unannotated sequences were prioritized by unique high-quality mass spectral support, absence from CD34+ controls, recurrence across individual AML patients, and lack of close homology to annotated proteins. We identified 5,849 Swiss-Prot-unannotated proteoforms, including 1,987 without homology to annotated human proteins. Thirty-nine candidates, most encoding microproteins, were recurrently detected in more than 10% of patients, and 14 were independently validated by deep, fractionated, multi-protease data-independent acquisition (DIA) proteomics of human AML cell lines. Structural modeling predicted several functional classes, including intrinsically disordered, alpha-helical microproteins, and membrane- or secretory-pathway-associated proteoforms. These findings define a recurrent AML dark proteome and establish a framework for the discovery of tumor-specific non-canonical proteins for mechanistic and therapeutic studies.
L. Schmalbrock, Asher Preska Steinberg, Katarzyna Kulej et al.· bioRxiv· 0 citations
The development of new therapeutics and the validation of pathogenetic cancer mechanisms require representative laboratory models1,2. However, existing collections represent only a fraction of the diversity observed in human cancer2-4. Recent technologies have enabled efficient in vitro model derivation (for example, tumour organoids)5. However, whether these maintain essential properties of patient tumours during long-term expansion has not been systematically investigated. Here we present results of a large-scale international programme-the Human Cancer Models Initiative-which involved the generation of a resource of 665 next-generation models from 2,780 donors with 25 cancer types and integrated tumour-model whole genome, exome, methylome and transcriptome analyses. The resource provides 522 models with comprehensive clinical data, 153 models of rare cancers and 71 models from participants with non-European ancestry. Analyses of 421 matched tumour-model pairs reveal high genetic (97.8%) and epigenetic (95%) concordance and define correlates of model discordance. Single-nucleus RNA sequencing of tumour-model pairs reveals subsets of models in which culture conditions significantly influence cell states. Finally, we characterize model preservation of extrachromosomal DNA and post-treatment mutational signatures to provide opportunities to study therapeutic resistance. This model repository is being made available to the community-including multimodal molecular profiling, clinical information and integrative software tools-thus providing a valuable resource for preclinical investigation of cancer pathogenesis and treatment response.
Dina Elharouni, Mushriq Al-Jazrawe, Seongmin Choi et al.· Nature· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.