Fabricated Foundations
Аннотация
Synthetic data generation (SDG) has emerged as a critical strategy for addressing challenges of bias, representativeness, and privacy in machine learning datasets. This chapter examines the theoretical foundations, technical methods, and ethical dimensions of SDG, exploring how fabricated data can both mitigate bias and introduce risks, including model collapse, amplified stereotypes, and reduced accountability. Drawing on advances in generative adversarial networks (GANs), variational autoencoders (VAEs), oversampling methods, and large language model–based synthesis, it evaluates the promise and risks of replacing or augmenting real-world data with algorithmic proxies. Ethical frameworks for fairness are considered alongside regulatory perspectives and domain-specific implications in healthcare and AI systems. Ultimately, the chapter argues that responsible SDG requires rigorous fidelity–utility–privacy evaluation, transparent governance, and a sustained commitment to equity across the data lifecycle.
Перевод пока недоступен