Extending GouDa: Generation of Universal Datasets with (and without) Errors for Data Quality Benchmarking
Synthetic data is extremely important in areas such as data quality, data cleaning, and machine learning. It enables the analysis of use cases in which real data is insufficient, unavailable, or distorted. However, generating synthetic data also presents challenges: The data must be as realistic as possible, but at the...