Post

Generating Synthetic Data with LLMs in Python

Generate synthetic training data with LLMs in Python โ€” structured prompts, diversity controls, quality filtering, and how to avoid the traps that poison a dataset.

Generating Synthetic Data with LLMs in Python
๐Ÿš€ Find Better AI Tools Faster Submit Your AI Tool & Reach Thousands of Builders Get Started โ†’

When You Donโ€™t Have Enough Data

Real labeled data is expensive and slow to collect. LLMs can generate it โ€” synthetic examples for training a smaller model, testing a pipeline, or bootstrapping a fine-tuning dataset. Done carelessly, synthetic data is repetitive and skewed. This guide shows how to generate it well in Python.

1
pip install openai

Generating Structured Examples

Ask for JSON so the output drops straight into a dataset. Here we generate labeled customer-support intents.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
import json
from openai import OpenAI

client = OpenAI()

def generate_batch(intent, n=5):
    prompt = f"""Generate {n} realistic customer support messages with intent "{intent}".
Vary length, tone, and phrasing. Return a JSON list of objects with keys
"text" and "intent". Return only JSON."""
    raw = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": prompt}],
        temperature=1.0,   # high temperature for diversity
    ).choices[0].message.content
    return json.loads(raw)

data = generate_batch("refund_request", n=5)

Note temperature=1.0. Low temperature gives near-duplicate outputs โ€” the opposite of what a dataset needs.

Forcing Diversity

A single prompt produces samey data. Vary the seed conditions: personas, contexts, edge cases. Loop over dimensions instead of asking for one big batch.

1
2
3
4
5
6
7
8
9
10
11
12
13
personas = ["a frustrated first-time buyer", "a calm long-time customer",
            "someone in a hurry", "a non-native English speaker"]

dataset = []
for persona in personas:
    prompt = f"""Write 3 refund_request messages as {persona}.
JSON list of text. Only JSON."""
    out = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": prompt}],
        temperature=1.0,
    ).choices[0].message.content
    dataset.extend(json.loads(out))

Explicit diversity axes beat hoping the model varies on its own.

Filtering for Quality

Generated data has duplicates and off-target examples. Filter before using it. A cheap first pass: drop near-duplicates by normalized text.

1
2
3
4
5
6
7
8
9
10
11
def dedup(rows):
    seen, out = set(), []
    for r in rows:
        key = r["text"].strip().lower()
        if key not in seen:
            seen.add(key)
            out.append(r)
    return out

clean = dedup(dataset)
print(f"{len(dataset)} -> {len(clean)} after dedup")

For semantic duplicates (same meaning, different words), embed each text and drop pairs above a similarity threshold โ€” see embeddings.

The Traps

Synthetic data inherits the generatorโ€™s biases and blind spots. If you train a model only on data from one LLM, you cap it at that LLMโ€™s distribution โ€” sometimes worse, as errors compound across generations. Mix in real data, validate label accuracy on a sample by hand, and never ship a dataset you havenโ€™t eyeballed. Synthetic data is a supplement, not a free lunch.

Frequently Asked Questions

What temperature should I use for synthetic data?

High โ€” around 0.9 to 1.2. Low temperature produces near-identical outputs, which defeats the purpose. Diversity is the whole reason to generate rather than copy.

Can I train a model entirely on synthetic data?

Sometimes, for narrow tasks, but itโ€™s risky. The model inherits the generatorโ€™s biases and can degrade if errors compound. Blending synthetic with real data is safer and usually stronger.

How do I check synthetic data quality?

Deduplicate, spot-check labels by hand on a random sample, and measure diversity with embeddings. If a downstream model trained on it generalizes to real inputs, the data is doing its job.

Takeaways

  • Generate structured JSON at high temperature for usable, varied data.
  • Force diversity with explicit personas and contexts, not one big prompt.
  • Deduplicate and hand-check โ€” synthetic data supplements real data, never replaces it.

Feed your cleaned dataset into an SFT pipeline to train a specialized model.

Khushal Jethava
Khushal Jethava

Machine Learning Engineer at Codiste, specializing in Generative AI, NLP, and Computer Vision. Building production AI systems with Python.

๐Ÿš€ Find Better AI Tools Faster Submit Your AI Tool & Reach Thousands of Builders Get Started โ†’
This post is licensed under CC BY 4.0 by the author.
๐Ÿš€ Find Better AI Tools Faster Submit Your AI Tool & Reach Thousands of Builders Get Started โ†’
๐Ÿš€ Find Better AI Tools Faster Submit Your AI Tool & Reach Thousands of Builders Get Started โ†’