Post

Generating Synthetic Data with LLMs in Python

Generate synthetic training data with LLMs in Python — structured prompts, diversity controls, quality filtering, and how to avoid the traps that poison a dataset.

Generating Synthetic Data with LLMs in Python
🚀 Find Better AI Tools Faster Submit Your AI Tool & Reach Thousands of Builders Get Started →

When You Don’t Have Enough Data

Real labeled data is expensive and slow to collect. LLMs can generate it — synthetic examples for training a smaller model, testing a pipeline, or bootstrapping a fine-tuning dataset. Done carelessly, synthetic data is repetitive and skewed. This guide shows how to generate it well in Python.

1
pip install openai

Generating Structured Examples

Ask for JSON so the output drops straight into a dataset. Here we generate labeled customer-support intents.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
import json
from openai import OpenAI

client = OpenAI()

def generate_batch(intent, n=5):
    prompt = f"""Generate {n} realistic customer support messages with intent "{intent}".
Vary length, tone, and phrasing. Return a JSON list of objects with keys
"text" and "intent". Return only JSON."""
    raw = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": prompt}],
        temperature=1.0,   # high temperature for diversity
    ).choices[0].message.content
    return json.loads(raw)

data = generate_batch("refund_request", n=5)

Note temperature=1.0. Low temperature gives near-duplicate outputs — the opposite of what a dataset needs.

Forcing Diversity

A single prompt produces samey data. Vary the seed conditions: personas, contexts, edge cases. Loop over dimensions instead of asking for one big batch.

1
2
3
4
5
6
7
8
9
10
11
12
13
personas = ["a frustrated first-time buyer", "a calm long-time customer",
            "someone in a hurry", "a non-native English speaker"]

dataset = []
for persona in personas:
    prompt = f"""Write 3 refund_request messages as {persona}.
JSON list of text. Only JSON."""
    out = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": prompt}],
        temperature=1.0,
    ).choices[0].message.content
    dataset.extend(json.loads(out))

Explicit diversity axes beat hoping the model varies on its own.

Filtering for Quality

Generated data has duplicates and off-target examples. Filter before using it. A cheap first pass: drop near-duplicates by normalized text.

1
2
3
4
5
6
7
8
9
10
11
def dedup(rows):
    seen, out = set(), []
    for r in rows:
        key = r["text"].strip().lower()
        if key not in seen:
            seen.add(key)
            out.append(r)
    return out

clean = dedup(dataset)
print(f"{len(dataset)} -> {len(clean)} after dedup")

For semantic duplicates (same meaning, different words), embed each text and drop pairs above a similarity threshold — see embeddings.

The Traps

Synthetic data inherits the generator’s biases and blind spots. If you train a model only on data from one LLM, you cap it at that LLM’s distribution — sometimes worse, as errors compound across generations. Mix in real data, validate label accuracy on a sample by hand, and never ship a dataset you haven’t eyeballed. Synthetic data is a supplement, not a free lunch.

Frequently Asked Questions

What temperature should I use for synthetic data?

High — around 0.9 to 1.2. Low temperature produces near-identical outputs, which defeats the purpose. Diversity is the whole reason to generate rather than copy.

Can I train a model entirely on synthetic data?

Sometimes, for narrow tasks, but it’s risky. The model inherits the generator’s biases and can degrade if errors compound. Blending synthetic with real data is safer and usually stronger.

How do I check synthetic data quality?

Deduplicate, spot-check labels by hand on a random sample, and measure diversity with embeddings. If a downstream model trained on it generalizes to real inputs, the data is doing its job.

Takeaways

  • Generate structured JSON at high temperature for usable, varied data.
  • Force diversity with explicit personas and contexts, not one big prompt.
  • Deduplicate and hand-check — synthetic data supplements real data, never replaces it.

Feed your cleaned dataset into an SFT pipeline to train a specialized model.

Khushal Jethava
Khushal Jethava

Machine Learning Engineer at Codiste, specializing in Generative AI, NLP, and Computer Vision. Building production AI systems with Python.

🚀 Find Better AI Tools Faster Submit Your AI Tool & Reach Thousands of Builders Get Started →
This post is licensed under CC BY 4.0 by the author.
🚀 Find Better AI Tools Faster Submit Your AI Tool & Reach Thousands of Builders Get Started →
🚀 Find Better AI Tools Faster Submit Your AI Tool & Reach Thousands of Builders Get Started →