← all articles

Model distillation explained: shrinking a big model down for one job

What distillation actually is

Model distillation is the process of training a smaller “student” model to imitate a larger “teacher” model on a specific task. You’re not compressing the big model’s weights or pruning its layers. You’re using the big model’s behavior as training data to teach a different, smaller model how to act.

The classic version, from Hinton, Vinyals, and Dean’s 2015 paper, trains the student on the teacher’s soft output distribution rather than just the final answer. If a classifier is 70% confident it’s a cat and 25% confident it’s a dog, that ratio carries information a hard label (“cat”) throws away. The student learns from that shape, not just the label.

For LLMs today, most teams doing this in practice aren’t matching logit distributions. They’re generating a dataset of teacher outputs (prompts in, completions out) and fine-tuning a smaller model on those pairs. That’s closer to supervised fine-tuning on synthetic data than the original Hinton-style distillation, but people still call it distillation because the goal is the same: bake a big model’s competence on a narrow task into a small model that’s cheaper and faster to run.

The math that makes people bother

Nobody does this for fun. They do it because a frontier model is overkill for one repeated job and the API bill or latency shows it.

Say you’re running a classifier that tags support tickets by category, or a function that extracts structured fields from an invoice, or a router that decides which tool to call. These are narrow, repetitive tasks. You’re paying frontier-model prices and eating frontier-model latency (often 1-3 seconds for a short completion) for a job that doesn’t need broad reasoning or world knowledge. A model with a fraction of the parameter count, fine-tuned specifically on that task, can hit similar accuracy on that narrow slice while running in tens of milliseconds on cheaper hardware, sometimes even on a CPU or a small GPU you already own.

The tradeoff is upfront work. You have to generate a training set, fine-tune, evaluate, and maintain the pipeline. That only pays off if the task runs often enough and consistently enough that the fixed cost of building the student amortizes. A one-off script you run twice a month isn’t a candidate. A classifier sitting in the hot path of every incoming request, running thousands of times a day, is.

How the training data actually gets made

In practice this looks like: take a representative set of inputs for your task (real support tickets, real invoices, real user queries), send each one to the teacher model, and log the input-output pair. You now have a supervised dataset without hand-labeling it yourself.

A few things matter here that get skipped in the hand-wavy version of this story:

Coverage matters more than volume. A thousand examples that span the actual variety of inputs your system sees in production beats ten thousand examples that are all slight variations of the easy case. If your ticket classifier never sees an ambiguous, multi-category ticket during training, it will fall apart on the ones that actually matter.

You need to check the teacher’s outputs, not just trust them. The teacher model is not ground truth. It makes mistakes, and if you train a student on those mistakes, you’ve distilled the errors along with the competence. Some teams sample and manually review a subset of the generated labels before they train on them. Skipping this step is how you end up with a fast, cheap model that confidently repeats the teacher’s blind spots.

Terms of service matter, and they vary by provider. Several major API providers restrict using their model’s outputs to train a competing general-purpose model. Using outputs internally to build a narrow classifier for your own product is a different situation than trying to reconstruct a general assistant, but the exact line depends on the provider’s current terms, not on what feels reasonable. Read the terms for whichever API you’re using before you build a data pipeline around it, because this is a contract question, not a technical one.

What you actually give up

The student model is narrower on purpose, and that has real costs beyond “it’s slightly less accurate.”

It generalizes worse. A model distilled to classify support tickets into eight categories will do that well and will do almost nothing else well. If your inputs drift, say a new product line introduces a ninth category or the phrasing in tickets shifts because your product changed, the student won’t adapt the way a frontier model would. You’re trading flexibility for speed and cost, and that trade is invisible until the distribution shifts under you.

There’s a real gap between teacher and student, and it doesn’t close with more training alone. A student with far fewer parameters has less capacity to represent whatever function the teacher learned. Past a point, throwing more distilled examples at it gives diminishing returns because the bottleneck is capacity, not data. If your task is genuinely hard, a small model may just not be able to hold the pattern, no matter how good the training set is.

You now own a second model. That means monitoring it for drift, re-running the distillation pipeline when the underlying task changes, and having an eval set you trust to catch regressions. A frontier model gets better over time because the vendor updates it. A distilled student only gets better when you retrain it. That’s a maintenance commitment, not a one-time project.

Evaluation has to be honest. It’s tempting to eyeball a handful of outputs and call it good. Build a held-out test set before you start, one the student never sees during training, and score the student against the teacher on it using whatever metric actually matters for the task (exact match, F1, pass rate, whatever). Don’t publish or act on comparisons you didn’t actually run. If you haven’t measured it, you don’t know it, no matter how confident the student sounds.

When it’s worth doing, and when it isn’t

Distillation earns its keep when three things are true at once: the task is narrow and stable, the volume is high enough that per-call savings add up, and you have (or can generate) enough representative examples to teach the pattern.

It’s usually not worth it for tasks that need broad reasoning, that involve open-ended generation where quality is subjective, or that run infrequently enough that the frontier model’s bill was never the bottleneck in the first place. It’s also a bad fit for tasks that change shape often. A student trained on last quarter’s ticket categories doesn’t know this quarter’s category was added yesterday.

A reasonable way to think about it: distillation converts an ongoing API cost into an upfront engineering cost plus an ongoing maintenance cost. Whether that trade is worth it depends entirely on your volume and how stable the task is, not on whether distillation sounds impressive.

A rough shape for doing this yourself

If you’re weighing it, the steps look roughly like this: pick a task narrow enough that a small model has a real shot, collect a representative set of real inputs, generate teacher outputs for them and spot-check a sample for correctness, hold out a test set before touching the student, fine-tune a small open-weight model on the rest, and then score the student against the teacher on the held-out set using the metric that actually matters for your job. If the gap is small enough and the volume justifies the maintenance, ship it. If not, keep calling the big model and revisit later when volume grows.

None of this is exotic. It’s supervised fine-tuning with a synthetic dataset generated by a bigger model instead of hand labeling. The part that actually takes engineering judgment is deciding whether your task is narrow and stable enough to be worth the ongoing cost of owning a second model.

If you’re building with LLMs day to day and want more of this kind of unglamorous, ship-it-yourself detail on what actually works, check out more from AI Tool Gazette.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →