Ai-Driven Compact Gene Expression Biomarkers for Multi-Cancer Diagnosis
Abstract
Gene expression profiles can tell one cancer from another, but a profile has tens of thousands of genes, and most of them add little to a diagnosis. A large panel is also hard to read and expensive to run in a clinic. This paper asks a simple question. How few genes are enough to separate five common tumor types with high accuracy. We use the public TCGA pan-cancer RNA-Seq dataset, which has 801 samples across breast, kidney, colon, lung, and prostate tumors and more than 20000 genes. We rank genes with a univariate filter and train three standard classifiers, a logistic regression, a linear support vector machine, and a random forest, on panels of growing size. To keep the results honest, the ranking is redone inside every cross-validation fold, so no gene is chosen using the test data. Accuracy climbs fast and then flattens. A panel of only 20 genes reaches about 99% accuracy, and a panel of 10 genes is already above 90%. A linear support vector machine on the $\mathbf{2 0}$-gene panel reaches $\mathbf{9 9. 4 \%}$ accuracy with a macro F1 of 0.994. We also compare three ways of ranking genes and report the selected panel. The result is a small and reproducible biomarker panel that keeps the accuracy of the full profile while using a tiny fraction of the genes.