Skip to content

Author

Yinyihong Liu

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Aug 2026

Synthetic Data: A Tool for Privacy Protection and Model Empowerment

Synthetic data, i.e., data simulated from some statistical model, are an important tool for both privacy protection and artificial intelligence pipelines. In the privacy context, synthetic data enable agencies to disseminate record-level information while reducing disclosure risks. In the artificial intelligence context, synthetic data allow analysts to augment training sets, increase coverage of rare cases, and support experimentation when genuine data are scarce. For each usage, we discuss key considerations and methods for generating synthetic data, including sequential modeling, differentially private synthesis, deep generative models, and large language models. Throughout, we highlight key trade-offs between data usefulness, privacy protection, and model reliability. We conclude by outlining some open research challenges and future directions for synthetic data development.

Yinyihong Liu, Jerome P. Reiter · 0 citations