Training models on synthetic data

Small models trained partly on text written by bigger models can beat much larger ones. Synthetic data can lower privacy risk, but it isn't anonymous by default.

Written Updated 3 min read

The papers

Two papers from 2022 and 2023 showed that text written by a language model can be good training data for another model, or even for itself.

Self-Instruct, from the University of Washington, the Allen Institute for AI and others, started with 175 hand-written tasks and had GPT-3 generate about 52,000 new instructions with inputs and answers, filtered out the invalid and repetitive ones, and then fine-tuned GPT-3 on the result. The tuned model improved by 33 points (absolute) on a benchmark of unseen tasks, putting it roughly level with OpenAI's InstructGPT-001, which had been trained on private user data and human annotations. It was an early, cheap demonstration of a model bootstrapping its own instruction data.

Microsoft's phi-1, described in "Textbooks Are All You Need", applied the idea to code with a focus on quality. The 1.3-billion-parameter model was trained on about 6 billion tokens of web code filtered for "textbook quality" plus about 1 billion tokens of textbooks and exercises written by GPT-3.5. It scored 50.6% on the HumanEval coding benchmark, well ahead of StarCoder (33.6%), a 15.5-billion-parameter model trained on a trillion tokens. Phi-2 extended the approach to general language: 2.7 billion parameters, trained on 1.4 trillion tokens "from multiple passes on a mixture of Synthetic and Web datasets", which Microsoft reported as matching or beating models up to 25 times its size on complex benchmarks (Microsoft Research).

The lesson was that quality beat volume, and a strong model could supply the quality. Both phi models also relied on a lot of carefully filtered real data, which tends to get lost in the retelling.

Why I think it matters

I think the advantage shifts from who has the most data to who can generate the right data for a particular use, which favours domain expertise over data hoarding. A team that understands its problem well can describe the examples it needs and have a model produce them, then check them.

Quality matters even more with synthetic data than with real data. Training models on the output of earlier models, generation after generation, makes them lose the rarer parts of the original data, an effect the authors named model collapse (Nature, 2024). My practical reading is to keep real data in the mix and filter hard.

The privacy argument needs more care than it usually gets. Synthetic data can reduce how much sensitive data you train on, but it isn't anonymous by default: a generator trained on personal records can reproduce parts of them. The European Data Protection Board's view is that an AI model trained on personal data counts as anonymous only if it is "very unlikely" both to identify the people in its training data and to give up their data through queries (EDPB Opinion 28/2024). In July 2026 the EDPB added guidelines on anonymization with three tests: no record isolation, no linkage and no inference (EDPB). I'd expect a synthetic dataset built from personal data to be judged against the same tests.

Since then

Synthetic data became standard practice rather than a bet. Phi-4 (14 billion parameters, December 2024) "strategically incorporates synthetic data throughout the training process", and Microsoft reports that it surpassed its teacher model, GPT-4, on STEM-focused questions (report). Reasoning models are routinely trained on generated reasoning: DeepSeek R1 used about 800,000 samples mostly generated by the model itself, and s1 trained on 1,000 traces written by Gemini.

There's also a twist. DeepSeek says it deliberately used no synthetic data to pretrain its V3 base model, but found that some web pages "contain a significant number of OpenAI-model-generated answers" (DeepSeek-R1, revised version). Model-written text is now part of the web that everyone trains on.

Microsoft's own framing has become more measured too. Its March 2026 write-up on a multimodal reasoning model calls generated data "a useful augmentation to high-quality real datasets", and says it is not a replacement for them (Microsoft Research).

What I don't know is how far this goes for organizations with sensitive data: whether synthetic records built from, say, customer files will be good enough to train on and safe enough to share. For now I'd treat that as something to test case by case. Agents that generate their own training data from experience are a related idea, covered in early experience.