What is Synthetic Data Generation?
Synthetic data generation is the creation of artificial datasets that closely replicate the statistical properties and patterns of real-world data without using actual individual records. This data is crafted through specialized algorithms and mathematical models, providing a valuable resource for training machine learning models, testing software, and simulating scenarios while protecting privacy and addressing data scarcity. By preserving essential relationships within data, synthetic data generation enables robust analysis and innovation across fields like healthcare, finance, and technology without risking sensitive information.
The Basic Idea
Imagine you wanted to conduct a study testing the effectiveness of a new drug. To accurately measure this, you would need to hold a clinical trial and receive access to in-depth medical records, including patient demographic information, medical history, health conditions, and treatment outcomes. Unfortunately, this is sensitive information that could pose a significant risk to privacy. Additionally, it would be difficult and timely to collect enough valid data to make sure your conclusions reflect the entire population.
This is a hypothetical scenario where synthetic data generation could come in handy. Synthetic data is artificially generated data that mimics real-world information.1 By collecting data from reliable sources and databases, a computer generates synthetic data that mirrors the original numbers. The generated data will not be exactly like the real data, and can even be generated from aggregate data in order to reflect similar patterns while maintaining privacy.
A computer can analyze synthetic data to identify relationships and patterns, enabling the development of algorithms and statistical models that can later be applied to real-world data. In the case of a clinical trial, the computer would use existing data on patient characteristics, medical history, and treatment outcomes to create synthetic data, learn how this data relates to patient outcomes, and build algorithms that will be able to predict similar outcomes from other data sources.
There are three main methods for generating synthetic data:
- Generative Adversarial Networks (GANs): A deep learning structure where two neural networks work against one another to generate more authentic data. One network will generate synthetic data, while the other is trained to identify whether it is real or synthetic. If the latter can recognize that any data points are synthetic, the initial neural network will make adjustments until the second network cannot differentiate between real and synthetic data.3
- Variational Autoencoders (VAEs): A deep learning model that learns the patterns of real data to create synthetic data with similar characteristics. VAEs reconstruct the original input and then use variational inference to generate additional datasets that retain key features from the original.4
- Data Augmentation: A technique of synthetic data generation for incomplete datasets. Data augmentation analyzes the original data to create the missing data points and generate a more complete dataset that can be better analyzed.5
Synthetic data generation is used in various fields outside of medicine as well. Such scenarios include autonomous vehicle testing when it is too dangerous to simulate driving conditions, finance to help build algorithms to detect fraud without using private and personal financial records, and robotics to train machines in simulated environments. In short, synthetic data generation is helpful in any field where data is private or difficult to obtain.
Although technically “fake,” synthetic data replicates actual properties and characteristics. This method is an effective way to overcome challenges associated with difficulty obtaining real data, or in cases where the real data is sensitive. It is cheap, easy to produce, and since it is artificially generated, it can be easier to align the data with what you are looking for for your specific study. While it can be a pain to sift through medical histories to find a particular data point you are after, such as the outcomes of a drug intervention, with synthetically generated data, you can create domain-specific data points.2
Imagine if it were possible to produce infinite amounts of the world’s most valuable resource, cheaply and quickly. What dramatic economic transformations and opportunities would result? That is a reality today. It is called synthetic data.
— Rob Towes, Venture Capitalist at Radical Ventures, and contributor to Forbes.6
Key Terms
Machine Learning: A subset of artificial intelligence that uses statistical techniques to enable machines to learn from data and improve their models and algorithms over time.
Generative AI: A subset of machine learning that learns from existing data to identify patterns and structures, allowing it to produce original content that can sometimes be synthetically generated data.
Entity Cloning: A privacy-preserving technique related to synthetic data generation in which particular aspects of a dataset are copied, while other sensitive details are omitted. For example, in healthcare, a researcher might clone characteristics like gender, health conditions, and intervention outcomes, but not the rest of a patient’s data that would make them identifiable.7
Data masking: A privacy-preserving technique related to synthetic data generation that only changes the format of real data rather than the values. This method allows for the creation of data that cannot be linked back to sensitive information. Methods of data masking include character shuffling, word or character substitution, and data encryption.8
Synthetic Population: An artificially created model, developed using synthetic data and statistical methods, that samples or resembles a real population. Synthetic populations are generated to study population behaviors without having to access individual-level data.
Computational Social Science: An interdisciplinary field that leverages mathematical algorithms, advanced data analysis, and computational modeling to study and predict human behavior and social dynamics. Computational social science often leverages synthetic data generation to simulate social environments and test hypotheses.
History
While synthetic data generation has gained popularity in recent years with the rise of machine learning and increased focus on data privacy, the concept was first developed in the 1960s. In its early days, synthetic data generation was used to create artificial drawings. Artificial intelligence pioneers, such as Maxwell Clowes and David Huffman, created synthetic images to train computers to recognize 3D shapes and interpret visual images in a manner similar to human perception. Real images were too complex for the available technology, so synthetic data enabled tasks like edge detection, 3D modeling, and line labeling. This was an important step in giving computers the ability to perceive the world similarly to humans, which later supported different machine learning applications such as autonomous vehicles, medical imaging, and virtual reality.
In 1993, Donald Rubin, a Harvard statistics professor, realized the potential for synthetic data generation to overcome data privacy concerns. Rubin wanted to use population censuses, which include both incomplete and sensitive data, to analyze demographic and social trends. However, he found it difficult to come up with a way to anonymize the census data, which led him to realize that synthetic data generation could create an artificial census dataset that was complete and avoided privacy concerns. Through synthetic data generation, Rubin developed synthetic populations, an artificial model resembling the real population.9
The next development in synthetic data generation emerged in the 2000s as machine learning applications evolved. In order to create more complex and sophisticated AI and machine learning tools, you would need a vast amount of data to train them. Synthetic data provided a solution by allowing the generation of large, diverse datasets that could be used to improve the accuracy and functionality of these tools. This era marked the beginning of the widespread use of synthetic data in various applications, from enhancing computer vision to optimizing predictive models.9
As generative AI exploded in popularity, so too has the use of synthetic data generation. Generative AI tools are able to quickly and efficiently produce vast amounts of synthetic data based on a few simple prompts. One such tool is generative adversarial networks (GAN), developed by American computer scientist Ian Goodfellow. GANs are a popular method of synthetic data generation, using complex statistical analysis of the elements that make up a photograph (real-world data) to then create images by themselves (synthetically generated data).10
The applications and implications of synthetic data generation go beyond being able to create fun artificial images. This method is at the forefront of various industries, anywhere from powering advancements in autonomous vehicles to detecting diseases and creating advanced fraud detection financial algorithms. Synthetic data generation fuels machine learning, allowing for more robust, diverse, and scalable training environments. This has positioned synthetic data as a key driver of progress in AI-driven research and development.
People
Maxwell Clowes
An early pioneer of artificial intelligence, Clowes was one of the first computer scientists to use synthetic data to improve computers’ ability to recognize 3D shapes. Clowes was a founding member of the Cognitive Studies Program at the University of Sussex, where he brought together a variety of approaches to the study of the mind, including psychology, linguistics, philosophy, and artificial intelligence. Clowes believed that by teaching computers to recognize 3D shapes, computers could develop a type of “vision” akin to humans.11
David Huffman
Alongside Maxwell Clowes, Huffman was one of the first to use artificial images to help computers recognize 3D shapes from 2D inputs, described in his 1971 paper, “Impossible Objects as Nonsense Sentences.” Huffman completed his doctorate at MIT in 1953, where he developed Huffman coding, a data compression method that avoided the loss of data. This method is still popular today in many file formats like JPEG and MP3. Later in his career, Huffman took a deeper interest in computer vision and AI, which is where he made contributions to the field of synthetic data generation.12
Donald Rubin
Harvard Statistics Professor Rubin revolutionized development economics and randomized field experiments. While consulting for the US Educational Testing Service, Rubin was interested in developing causal models that could predict outcomes when an individual or group’s environment changes. To build this model, Rubin wanted to use census data but encouraged data privacy concerns as well as issues with incomplete datasets, which he overcame through the use of synthetic data generation.13
Ian Goodfellow
Often referred to as “the GANfather,” Goodfellow is an American computer scientist best known for his research in deep learning and the invention of generative adversarial networks (GAN). During his PhD at the Universite de Montréal, Goodfellow made a number of other significant contributions to the field of deep learning, including maxout networks (a neural network layer that improves the way a model detects complex relationships)14 and multi-prediction deep-Boltzmann machines (a computer model that consists of multiple hidden layers to understand hierarchical relationships between data).15
behavior change 101
Start your behavior change journey at the right place
Impacts
The capabilities lent by synthetic data generation pose a counter to the challenges associated with real-world data and can drive innovation in our tech-enabled world. Synthetic data allows researchers to ensure privacy, reduce bias, and train smarter computer machines.
Overcoming Privacy Concerns
Data is one of the most valuable resources in the digital era. As users, we provide new data businesses can leverage whenever we browse the web or scroll through social media. The availability of big data, when combined with advanced statistical tools like artificial intelligence, can accelerate research and innovation, allowing decision-makers to make data-informed decisions to improve our lives.
However, developments in big data have also led to increased privacy concerns, as many large-scale data sets include sensitive information. Synthetic data generation is able to overcome these privacy concerns by creating artificial data that mimics real-world data without revealing sensitive personal information. This allows researchers to make breakthroughs in fields like healthcare, finance, and technology, fields where datasets often include a vast amount of sensitive information, without compromising individual privacy.9
Reducing Bias
Real-world data often has representation or sample bias due to challenges in data collection. Limited time, incomplete datasets, and the inability to capture all scenarios can result in data that represents only a small segment of the population. For example, healthcare research usually skews heavily towards urban populations as these populations more regularly visit hospitals and doctors.16 Any conclusions drawn from that data would therefore not be reflective of individuals living in rural or remote areas. Synthetic data generation, however, can create synthetic populations that fit various demographics, which can help reduce bias.
Oftentimes, real data is missing data points. For example, financial records of new businesses or borrowers might have limited credit histories to analyze. Using synthetic data generation allows researchers to fill in the missing pieces through techniques like data augmentation, and create more fulsome datasets that are more likely to be representative of the population.17
Training AI Algorithms
Once a computer learns from synthetically generated data, it creates algorithms that can be applied to real-world data. Synthetic data is used to train computers on which patterns and relationships to look for, which then makes them more effective tools when applied to actual information. For example, if you create synthetically generated financial record data and label some of the data as fraud, a computer algorithm will learn how to identify fraudulent financial activity once it is applied to real financial records.
Controversies
While we’d like to believe that synthetically generated data is the solution to all of the challenges related to using real-world data, this method can still fall victim to many challenges.
Inability to Capture Real Life Complexities
Life can be random and unpredictable. While synthetic data generation can replicate patterns and find causal relationships, it sometimes falls short of capturing the randomness of real life. As a result, the insights gathered from synthetic data may not represent actual conditions.18
Even in instances where the generated data seems to represent reality, it’s difficult to validate and ensure that the models trained on synthetic data will be accurate when applied to real-world data.19 For example, synthetic data capturing retail customer behavior to inform stores what they should have in stock may not account for a particularly hot winter, which would result in fewer customers purchasing coats.
Dependence on Real-World Data
In order to create synthetically generated data, we must still rely on real-world data to train computers that build models and algorithms. While synthetic data generation can fill in some holes in data, it can still replicate the bias that exists in real-world data.
For example, say a company uses synthetic data to train an AI model to sort through real resumes to determine who to give an interview to. If the synthetic data is built on existing profiles of staff in a company that is male-dominated, there is a strong chance that the synthetic data will do the same.20 That’s why it is important for synthetically generated data to be built from reliable data sources, and continuously updated to reflect changes in real populations.
Lack of Privacy Policies Around Synthetic Data Use
Although synthetic data generation has been positioned as a tool to overcome privacy concerns related to accessing sensitive data, there ironically aren’t many clear standards or regulations on how it can be used because its widespread use is relatively new. Since the synthetic data is generated from real-world data, it may still contain sensitive information, especially if it is very specific and can be traced back to real people.19
Although there are persisting concerns about the use of synthetic data generation, it has deep-reaching potential to greatly impact the way we conduct research. Mitigating issues like data scarcity and providing a more efficient way to train AI models is a powerful tool that allows us to make more data-informed decisions. As technology and regulations continue to evolve, the future of synthetic data generation will likely involve finding more robust methods to address privacy concerns and ensure its accuracy.
Case Study
Autonomous Vehicles
To ensure that autonomous vehicles operate safely, they need to be able to perform in a variety of driving conditions. Think about all the different factors you may encounter while driving—there may be pedestrians, speeding cars, changes in speed limits, dangerous elements like icy roads, and so on. It’s really difficult to collect a large number of training examples in the real world that account for the wide variety of driving environments, while equally difficult and dangerous to replicate those tests in the real world.
Synthetic data generation can create 3D simulations of driving footage to train autonomous vehicles on how to perform in diverse situations and create algorithms to recognize objects like pedestrians, adapt to difficult driving conditions, and predict the movement of vehicles around them.21 The first use of synthetic data in this context was for a project called ALVINN (Autonomous Land Vehicle in a Neural Network) presented at a conference in 1989. Dean Pomerleau spearheaded the project using synthetic images that depicted a variety of conditions and trained an algorithm to detect which direction should travel. When the vehicle was put on the road, equipped with a camera that could analyze what was ahead, the algorithm helped it to maneuver based on what it could detect.
Today, synthetic data generation continues to be used for autonomous vehicles. For example, Tesla uses synthetic data to simulate scenarios that are rare in real-world environments or to train a vehicle on how to operate if sensors are broken.22 Still, the synthetic data may not reflect all the potential environments a vehicle will find itself in, as there are external factors like weather and other drivers. However, synthetic data fills in some of those gaps in atypical conditions.
Encapsulating Diversities
Think about the way that you speak—your accent, the slang you use, your conversational style. Now consider how it matches up to people in your life. There are likely differences between the way you talk, the way your grandfather talks, and the way your coworker talks. Now imagine those differences being magnified when you compare it to someone living in a different country with a completely different background than you.
Customer-care chatbots analyze the data—or, in this case, words—that you ask it to come up with a response and assist you in your query. However, it has to account for countless variations in speech. In order to learn the nuances of every customer request, it would take thousands of different real-world inputs to be able to respond to each effectively. Siri, for example, isn’t great at recognizing bilingual speech, likely because it was never trained on bilingual data.23
It would be very time consuming and difficult to collect that data to teach chatbots, but synthetic generation can overcome those challenges. YAMADA, an algorithm developed by IBM Research, generates fake sentences to train a chatbot to be able to understand inputs and be able to come up with an answer, no matter what way the user speaks!24
Related TDL Content
Bridging the Gap Between Data and Policy
You might not realize it, but a lot of your behavior is dictated by public policymakers, such as what food is available at a grocery store, the cost of healthcare, and whether or not you qualify for a loan. That’s why it’s important for policymakers to make data-informed decisions, but challenges with real-world data can make that difficult. In this article, we look at how the Nesta Innovation Mapping Team is helping policymakers by creating tools that synthesize and simplify large amounts of data to support decision-making.
Artificial General Intelligence
Generative AI systems currently depend on the real or synthetic data they are trained on. While analyzing more data improves their effectiveness, they remain limited by their training data. However, we may be approaching Artificial General Intelligence (AGI), a theoretical form of AI capable of human-like thinking, reasoning, and learning. In this article, author Kira Warje explores the potential of AGI.



















