Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Sep 2026

Machine learning models for predicting liver cancer: a real-world cohort study in China

Liver cancer has high incidence and mortality worldwide, and timely identification is important for improving prognosis. However, prediction models based on routinely available clinical data remain insufficiently evaluated in hospitalized real-world populations. This study aimed to develop and interpret a machine-learning model for liver cancer prediction using multidimensional clinical and laboratory features. This retrospective real-world cohort included 9,284 hospitalized patients from Hainan General Hospital between January 2020 and December 2023. Patients were divided into a training set (n = 6,426) and an internal validation set (n = 2,858). Seventy-four demographic, clinical, and laboratory variables were collected. Feature selection was performed using least absolute shrinkage and selection operator regression in the training set. Six models were developed: logistic regression, random forest, support vector machine, k-nearest neighbors, extreme gradient boosting (XGBoost), and Elastic Net. Discrimination was assessed primarily using the area under the receiver operating characteristic curve (AUC). Decision curve analysis evaluated clinical net benefit, and SHapley Additive exPlanations (SHAP) were used for model interpretation. Liver cancer accounted for 48.7% of the training cohort and 48.8% of the validation cohort. LASSO retained 57 nonzero model terms at the minimum-error penalty. XGBoost achieved the highest AUC among all models, with an AUC of 0.909 (95% CI, 0.903–0.915) in the training set and 0.802 (95% CI, 0.788–0.817) in the validation set. At a threshold of 0.5, validation accuracy, sensitivity, specificity, and F1 score were 0.730, 0.758, 0.698, and 0.753, respectively. Decision curve analysis showed that XGBoost provided the greatest net clinical benefit across a broad range of threshold probabilities. SHAP analysis identified serum sialic acid, the aspartate aminotransferase-to-alanine aminotransferase ratio, alkaline phosphatase, basophil percentage, and monocyte-to-lymphocyte ratio as the leading predictors. XGBoost showed good discrimination and interpretability for liver cancer prediction in a hospitalized real-world cohort using routinely available clinical and laboratory data. The findings support the feasibility of leveraging large-scale inpatient laboratory data for risk-stratification model development. Serum sialic acid, the aspartate aminotransferase-to-alanine aminotransferase ratio, and alkaline phosphatase were important predictors. Prospective validation in outpatient, first-visit, high-risk, and external populations is required before broader clinical or screening application.

Ce-Xiong Fu, Fang Li, Shi-Bing Li et al. · 0 citations
Open access Jan 2026

From Germline Variants to Tumor Outcome: GWAS‐Based Functional Genomics Prioritizes Colorectal Cancer Susceptibility Genes and Links SMAD9 to Prognosis

Introduction Genome‐wide association studies (GWAS) have identified over 200 germline risk loci for colorectal cancer (CRC), yet the causal variants and genes behind most GWAS signals remain unknown and the link between inherited risk and tumor outcome is largely unexplored. Connecting germline single nucleotide polymorphisms (SNPs) to gene expression through expression quantitative trait loci (eQTL) and to clinical outcome is needed to interpret this inherited risk. Methods We combined CRC GWAS summary data (73,149 cases and 112,467 controls of European ancestry) with GTEx v8 eQTL from colon sigmoid, colon transverse, small intestine terminal ileum, and whole blood, integrating causal transcriptome‐wide association study (cTWAS) with SuSiE fine‐mapping, Bayesian colocalization, MAGMA SNP‐to‐gene analysis, and AlphaGenome variant‐effect prediction. Prioritized proteins were assessed by western blot. Four experimentally selected regulatory variants were genotyped in HCT116, SW480, RKO, and NCM460 cells; genotype‐protein associations were tested across cell‐line means, and cis‐regulatory effects were assessed by allele‐specific expression (ASE) and reference‐versus‐alternate dual‐luciferase assays. Prognostic relevance was evaluated in TCGA colorectal tumors. Results cTWAS identified seven genes with posterior inclusion probability (PIP) > 0.50, and colocalization across 78 gene‐tissue pairs revealed 20 associations with PP.H4 > 0.80. Four genes showed convergent evidence: SMAD9 (PIP = 0.916, PP.H4 = 0.981), MAP3K2 (PIP = 0.762, PP.H4 = 0.827), FADS1 (PIP = 0.632, PP.H4 = 0.942), and ACTR1B (PIP = 0.566, PP.H4 = 0.994). All four proteins were reduced in CRC cells. Alt‐allele dosage was inversely associated with SMAD9 (Pearson r = −0.985, BH‐adjusted p = 0.029) and FADS1 protein abundance (r = −0.995, BH‐adjusted p = 0.020), but not with MAP3K2 or ACTR1B. Reporter and ASE assays detected the clearest allele‐specific effects at SMAD9 and FADS1, whereas MAP3K2 and ACTR1B were null in the tested systems. Lower SMAD9 expression was nominally associated with poorer survival (log − rank p = 0.022 to 0.050), but these associations did not survive Benjamini–Hochberg correction across 12 tests. Conclusions Integrating statistical genetics, regulatory prediction, and locus‐directed experiments prioritized SMAD9, MAP3K2, FADS1, and ACTR1B as CRC susceptibility genes. Functional evidence was strongest and most directionally coherent for SMAD9, demonstrated allele‐specific but context‐dependent regulation at FADS1, and placed experimental bounds on the proposed MAP3K2 and ACTR1B mechanisms.

Chengguang Hu, Guang Yang, Han Xiong et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.