The pitch for synthetic data is compelling. You do not need to find real contributors. You do not need to manage collection logistics across geographies. You do not need to worry about privacy, consent, or demographic gaps in your contributor pool. You generate what you need, at the volume you need, at a fraction of the cost of collecting it from humans. The economics look transformative and the industry has responded accordingly, with significant investment in synthetic data generation tools and growing enthusiasm for replacing real data collection with generated alternatives.
The problem is that the pitch describes synthetic data at its best and does not account for where it fails. From the inside of running real data collection and annotation projects, those failure points are not edge cases. They are the situations that determine whether a model performs reliably in production or collapses on the inputs that matter most.
What synthetic data is actually good for
Synthetic data has genuine value in specific applications and it is worth being precise about what those are before making the case against its overuse.
It works well for augmentation: taking a real dataset and expanding it to improve volume or balance class distributions that were skewed in the original collection. It works well for simulating rare events that would be difficult or dangerous to collect in real conditions: a self-driving vehicle encountering an unusual road configuration, a medical imaging model trained on a rare pathology presentation that appears in real data too infrequently to train on effectively. It works for privacy-sensitive domains where real data cannot be used without significant compliance overhead. And it works for padding volume in situations where the core diversity already exists in real data and you simply need more instances of patterns already well-represented.
These are real and useful applications. The problem begins when synthetic data is positioned not as a supplement to real data but as a replacement for it.
Synthetic data reflects the assumptions of the system that generated it. Which means it inherits every bias and every gap from the real data that system was built on.
The distribution gap
Synthetic data is generated by a model. That model was trained on real data. Whatever biases, gaps, and distributional skews existed in that real data are now embedded in the generative model, and they will appear in every dataset it produces. Synthetic data does not correct for the problems in the original data. It reproduces them at scale, with the additional problem that the reproduction looks clean and complete in ways that make the underlying issues harder to detect.
This is the fundamental limit of synthetic data as a solution to demographic diversity gaps in training sets. If the generative model was trained on data that underrepresents certain accents, certain skin tones, or certain linguistic patterns, the data it generates will underrepresent those same groups. Generating more data from a biased source produces a larger biased dataset. The gap does not close. It just becomes harder to see.
Real diverse data collection is operationally difficult precisely because it requires reaching contributor populations that are not well-represented in existing pools. That difficulty is also what makes it valuable. You cannot generate your way out of a diversity problem. You have to go and collect the data from the people the model needs to learn from.
The edge case problem
Real-world edge cases are real because they were not anticipated. They fall outside the distribution of what was expected, which is exactly why they are valuable for training a model that needs to handle them. Synthetic data generation is, by definition, a process of generating instances within a known distribution. It can produce variations within the space it understands. It cannot produce what it did not know to expect.
This means that the edge cases a model will encounter in production, the unexpected inputs, the novel combinations, the situations that fall outside the training distribution, are precisely the ones that synthetic data is least equipped to provide. A model trained heavily on synthetic data will perform well within the distribution it was trained on and degrade unpredictably outside it. In production, outside the training distribution is where users live.
Real data collection, by contrast, captures the genuine variation of the world as it is rather than as a generative model expects it to be. The noise, the inconsistency, the unexpected, the edge cases that nobody thought to generate: all of it is present in real data in a way that makes models trained on it more robust to the full range of inputs they will actually encounter.
The perceptual realism gap
For speech, image, and video models, synthetic data faces a realism problem that current generation technology has not resolved. Synthesised speech does not fully replicate the acoustic variation of real human speech: the microphone conditions, the background noise, the breathiness, the hesitations, the full range of vocal characteristics that vary across age, health, emotional state, and recording environment. Synthesised images do not fully replicate the lighting variation, compression artefacts, motion blur, and contextual noise of real-world photography. Synthesised video does not fully replicate the temporal consistency and physical plausibility of real recorded motion.
Models trained predominantly on synthetic data for these modalities learn to recognise patterns in the synthetic distribution. When they encounter real data in production, the gap between the synthetic distribution they trained on and the real distribution they are operating in produces performance degradation that is difficult to predict and hard to diagnose without going back to the training data.
Works: Volume augmentation on top of a diverse real dataset. Rare event simulation for conditions difficult to collect safely. Privacy-sensitive domain coverage where real data carries compliance risk. Class balancing where the core patterns already exist in real data.
Does not work: Replacing real demographic diversity. Generating genuine edge cases outside the training distribution. Replicating real-world perceptual variation for speech, image, and video models. Providing the domain judgment that specialist annotation requires.
The test: If you are using synthetic data to avoid collecting real data from a population your model will serve, the synthetic data is not solving the problem. It is obscuring it.
The domain expertise gap
In specialist domains, the failure of synthetic data is most acute. Medical AI, legal AI, financial AI: these are domains where the value of training data comes from human judgment that took years to develop. A synthetic dataset of medical imaging annotations reflects what a generative model thinks a correct annotation looks like. A dataset annotated by practising radiologists reflects what a correct annotation actually is.
The gap between those two things is not a technical problem waiting for a better generative model. It is a fundamental limit on what can be synthesised. Genuine domain expertise is not a pattern in a dataset. It is accumulated judgment about what matters, what to look for, and what a deviation from normal actually means. That judgment cannot be generated. It can only come from people who have developed it through years of practice in the field.
As AI moves deeper into specialist verticals, the proportion of training data that requires genuine human expertise to produce is increasing, not decreasing. The direction of travel in the industry is toward more reliance on real expert judgment in training data, not less.
What a realistic data strategy looks like
A data strategy that uses synthetic data well treats it as one tool among several, deployed where it genuinely helps and not used to substitute for the harder work of collecting real data from real people in the real world.
Real data collection remains the foundation. It provides the demographic diversity, the genuine edge cases, the perceptual realism, and the domain expertise that synthetic data cannot replicate. Synthetic data augments that foundation: adding volume, balancing distributions, covering rare events, and reducing cost in the specific situations where its limitations do not apply.
At ConsultBae, we work with clients who have tried to solve data collection problems with synthetic generation and encountered the failure modes described here in production. The fix, in every case, requires going back to real data collection: finding the right contributors, in the right geographies, with the right demographic profile, and collecting what the model actually needs to learn from. Synthetic data delayed that work. It did not replace it.
The case for synthetic data is real within its limits. The industry's enthusiasm for it has consistently run ahead of those limits, and the models paying the price for that enthusiasm are the ones that performed well in development and failed in the world.
Mohit Singh Katewa leads the AI Data vertical at ConsultBae, overseeing data collection, annotation, and quality operations across 100+ countries.
Relying on synthetic data and seeing gaps in production?
ConsultBae builds real data collection pipelines across 100+ countries. Let us talk about what your model actually needs.
Talk to us


