Collection: Benchmark Datasets for the Evaluation of Data Imputation Algorithms
Abstract
Complete and labeled datasets were gathered from the web and processed to simulate missing completely at random (MCAR) and missing at random (MAR) missingness mechanisms, which are inspired by common missing-data patterns. From the original 12 datasets, we generated 2400 benchmark datasets with missing data using the MCAR and the MAR missing data patterns. The missing data rates in the generated datasets range from 5% to 50% of the available data. This work describes the original datasets and shares the processed benchmark datasets with the community. In addition, we also describe and share the algorithm we employed to amputate the data and produce MCAR and MAR missing data patterns. We hope that our benchmark datasets can be used as a standard benchmark for evaluating new data imputation algorithms.</p> <p><bold>IEEE SOCIETY/COUNCIL</bold> Computational Intelligence Society (CIS)</p> <p><bold>DATA TYPE/LOCATION</bold> CSV Tables; Python script</p> <p><bold>DATA DOI/PID</bold> <ext-link ext-link-type="doi" xlink:href="https://doi.org/10.34740/kaggle/dsv/16287910">10.34740/kaggle/dsv/16287910</ext-link>