AI Training Data Quality: What Model Collapse Research Shows
7 min readA 2024 Nature paper found AI models degrade when trained on AI-made data. What the evidence shows, who it affects and what teams can do about it.
Zunkiree Labs Team
· Updated
In short: In July 2024, Nature published a paper by Ilia Shumailov and colleagues showing that models trained again and again on content generated by earlier models lose information, a process the authors call model collapse. The rare, unusual cases in the data go first. The experiments used small models, so the paper is a warning about data quality and provenance, not a forecast that AI will stop improving.
Key Takeaways
- The paper, "AI models collapse when trained on recursively generated data," appeared in Nature, volume 631, on July 24, 2024.
- The authors report that "indiscriminate use of model-generated content in training causes irreversible defects," starting with loss of the tails of the data distribution.
- In their language-model test, keeping 10% of the original data led to only minor degradation, according to the paper.
- They conclude that access to real human-produced data matters most where rare cases matter, and call for coordination on data provenance.
- Their experiments used a small model, OPT-125m, which the authors note as a limit.
The Evidence
The anchor is a single peer-reviewed paper.
- Nature, July 24, 2024. Shumailov, Shumaylov, Zhao, Papernot, Anderson and Gal, AI models collapse when trained on recursively generated data. We read the open-access copy on PubMed Central.
We did not repeat the experiments. Everything below about results comes from the paper.
What Is Model Collapse?
Most large models learn from text and images gathered from the internet. As AI-written content spreads online, later models may learn from the output of earlier ones. The paper asks what happens when that loop repeats.
The authors describe two stages. In early model collapse, models "start losing information about the tails of the distribution," meaning the rare and unusual cases. In late model collapse, the model converges to a distribution "that carries little resemblance to the original one."
They name three sources of error that build up across generations: statistical approximation error (learning from a finite sample), functional expressivity error (a model cannot represent everything) and functional approximation error (limits of the learning procedure). They show the effect in simple mathematical models and in a language-model experiment.
An everyday way to picture it: copy a photocopy of a photocopy. Each pass loses a little detail, and the faint details disappear first.
What Did the Experiments Find?
In the language-model test, the authors fine-tuned Meta's OPT-125m on the wikitext2 dataset. The original model fine-tuned on real data reached a mean perplexity of 34, from a zero-shot baseline of 115. Perplexity is a score where lower is better. When later generations trained on generated data, the authors report losing "some performance, from 20 to 28 perplexity points."
When they kept 10% of the original data in the mix, the paper says this "allows for better model fine-tuning and leads to only minor degradation of performance."
What Changed?
The change is the setting, not the model. Generated content is now common on the open web, so the old assumption that internet text is mostly human-written is weaker. The authors write that "the need to distinguish data generated by LLMs from other data raises questions about the provenance of content that is crawled from the Internet." We did not find a measured figure for how much of the web is AI-generated in the paper, so we do not give one.
Who Is Affected?
Teams that train or fine-tune models. The paper's main advice applies to them: keep real, human-produced data in the mix, and track where data came from. The authors say that "in learning tasks in which the tails of the underlying distribution matter, one needs access to real human-produced data."
Businesses using AI to generate data. Synthetic data is useful, but our reasoning, not a finding of the paper, is that a pipeline feeding generated records back into training with no real data and no checks is the pattern the study warns about.
Teams that only use AI through an API. Collapse is a training problem, so it affects them indirectly. The practical question for them is how their own data is cleaned and documented, which we cover in what the Spider 2.0 benchmark shows about AI over business data.
Limitations
- Small models. The authors used OPT-125m because of compute cost, and say so. We cannot conclude from this paper how a frontier model behaves.
- Worst-case setup. The collapse experiments train on generated data across generations. The authors' own 10% test shows that mixing in real data changes the result.
- Not a forecast. The paper does not predict that AI progress will stall. It identifies a risk that depends on how data is collected and filtered.
- Our inference is marked. Links to business practice in this post are our reasoning.
What Happens Next
This section is expectation, not fact. The authors call for "community-wide coordination" so the groups building and deploying models can share what is needed to resolve questions of provenance. For a company, the practical steps are modest: keep a record of where training and evaluation data comes from, label synthetic data as synthetic, keep a core set of verified real examples, and test models on rare cases, not just average ones. For another case where data quality and human checks decide the outcome, see how AI helps scientists find what humans could miss.
The Short Version
A 2024 Nature paper shows that models trained repeatedly on their own generated output lose rare information first and can degrade badly. Keeping some real data helped in the authors' test. The experiments used a small model, so treat the result as a reason to care about data provenance and quality, not as a prediction about any specific product.
Frequently asked questions
What is model collapse?
Model collapse is the degradation that happens when AI models are trained on content generated by earlier models. According to a 2024 Nature paper, models first lose information about rare cases in the data, and later converge to a distribution that carries little resemblance to the original.
Does this mean AI will stop improving?
No. The paper does not make that prediction. It tested a small model, OPT-125m, and showed that keeping 10% of the original data led to only minor degradation. It is a warning about data quality and provenance.
Is synthetic data always bad for training?
The paper studies training indiscriminately on model-generated content, not all synthetic data. The authors say access to real human-produced data is important when rare cases matter, and that data provenance should be tracked.
What should a company do about it?
If you train or fine-tune models, keep a record of data sources, label synthetic data, keep a core set of verified real examples and test on rare cases. If you only use models through an API, collapse affects you indirectly.
Related Insights
- How AI Helps Scientists Find What Humans Could Miss
- AI Over Business Data: What the Spider 2.0 Benchmark Shows
- How AI Is Changing Everyday Life: What Usage Data Shows
- How Satellite Data and AI Help Farmers See Crop Risks Earlier
Sources
- Shumailov et al., AI models collapse when trained on recursively generated data, Nature 631, July 24, 2024
- Open-access copy on PubMed Central (PMC11269175)