Model Distillation
A 7-billion-parameter model running on your laptop can match a 671-billion-parameter model on specific tasks. The technique that makes this possible is distillation: training a small student model to replicate what a large teacher model knows.
The teacher generates training data. The student learns from it. The result: a model 100x smaller that handles the same tasks at zero per-query cost.
Train a small model to mimic a large one โ not by copying weights, but by learning from its behavior.
Distillation is a training technique where you take a large, capable model (the teacher) and use it to teach a smaller model (the student). Instead of training the student on raw internet text from scratch, you train it on the teacher's outputs โ its answers, its reasoning steps, even its probability distributions over possible next tokens.
The student never sees the teacher's internal weights. It only sees what the teacher produces. The key insight: the teacher's step-by-step reasoning traces teach the student not just what the answer is, but how to think through the problem. This is why distilled reasoning models (like DeepSeek-R1-Distill) can punch far above their weight class.
Three reasons to care about making models smaller.
Cost
DeepSeek's distilled 7B model runs on a laptop and costs nothing per query. The 671B teacher costs real money per token through an API. For high-volume workloads โ customer support, content moderation, data extraction โ the cost difference is the difference between viable and not.
Latency
A 7B model generates 100+ tokens per second locally on a consumer GPU. A 1.6T MoE model generates 20-30 tokens per second through a cloud API. For real-time applications โ chatbots, code completion, voice assistants โ that 3-5x speed difference is the difference between feeling instant and feeling sluggish.
Privacy
Local models keep data on-device. No API calls. No data leaving your machine. No third-party logging. For legal, medical, financial, and defense applications where data sovereignty is non-negotiable, a distilled local model is often the only compliant option.
Three steps from teacher to student.
Generate Training Data
Run thousands of prompts through the teacher model. Save both the final answer and the teacher's chain-of-thought reasoning. For DeepSeek-R1-Distill, the R1 teacher generated millions of reasoning traces covering math, code, and logic problems.
Fine-Tune the Student
Train the smaller model on this teacher-generated data. The student learns to match the teacher's outputs token by token โ it absorbs not just what answers look like but the reasoning patterns that produced them.
Reinforcement Learning (Optional)
Let the student attempt tasks on its own. Compare its outputs to what the teacher would have produced. Reward closer matches. This step can squeeze out another 5-15% quality improvement on targeted benchmarks.
Frontier labs now ship distilled versions alongside their flagship models.
1.5B to 70B, MIT licensed
The most complete distilled family. R1 (671B teacher) generated reasoning traces; students from 1.5B to 70B learned to replicate them. The 7B and 32B variants are particularly strong โ competitive with much larger models on math and coding benchmarks while running on consumer hardware via Ollama or LM Studio.
8B and 70B, open weights
Meta's Llama 3.x distilled variants use the largest Llama as the teacher. Fine-tuned on curated instruction datasets. Strong general-purpose performance โ chat, summarization, extraction โ at a fraction of the full model's cost and hardware requirements.
14B, trained on synthetic data
Microsoft trained Phi-4 partly on outputs from GPT-4. The result: a 14B model that competes with models 3-5x its size on reasoning benchmarks. Demonstrates that synthetic teacher data can produce capable small models.
2B and 7B, open weights
Google's open small models, trained with distillation techniques from larger Gemini teachers. Designed for on-device inference โ phones, laptops, edge hardware. Good for text generation and classification tasks where latency and privacy matter more than raw capability.
Distillation is replication, not innovation.
A distilled model is only as good as its teacher. If the teacher hallucinates on legal questions, the student learns to hallucinate on legal questions. If the teacher has a systematic bias, the student inherits it. The student cannot discover capabilities the teacher never demonstrated.
Distilled models are also narrower. They excel at the specific tasks the teacher was prompted on during data generation, but may fail on unrelated tasks the teacher handles easily. A DeepSeek-R1-Distill-7B trained on math and code reasoning traces will not suddenly become a strong creative writer or medical diagnostician.
"Distillation doesn't make models smarter. It makes smart models smaller โ and sometimes that's all you need."
Default to the smallest model that meets your quality bar. A distilled 7B model that handles 90% of your workload at zero cost is better than a 671B model you're afraid to call.
Distillation inherits the teacher's blind spots. Test the student on YOUR data and YOUR tasks โ not just the benchmarks the teacher was optimized for.