Synthetic Colorectal Cancer Cohort for Machine Learning Research: An Open In Silico Dataset of 10,000 Virtual Patients
Overview: The dataset provides a fully synthetic, computer-generated cohort of 10,000 virtual patients with colorectal cancer (CRC). No actual patient data was accessed, collected, or utilized at any stage of this research.Variables: The dataset includes 39 variables organized across six clinical domains. These domains cover demographics, presenting symptoms, laboratory values, tumor pathology, TNM staging, treatment, and short-term postoperative outcomes.Methodology: The dataset was built using a causal-chain design programmed in Python. It also incorporates realistic missing-at-random patterns for three specific laboratory values: CA19-9, CRP, and albumin.Validation: The synthetic data's aggregate statistics (such as microsatellite instability prevalence, postoperative complications, and 30-day mortality) closely align with published real-world benchmarks. It successfully preserves expected directional clinical relationships.Intended Use: The data is designed to be an open resource for machine learning benchmarking, clinical informatics software testing, and educational purposes. It allows researchers to bypass patient privacy barriers for pipeline development.Restrictions: The dataset is explicitly not meant for direct clinical decision-making or to validate real-world predictive performance.