DeepSeek released R1 alongside a family of distilled smaller models — ranging from 1.5B to 70B parameters — specifically so practitioners could run powerful reasoning locally without a data centre. The choice between 14B, 32B, and 70B is not simply “bigger is better”: it involves your hardware, your tolerance for latency, and how much of R1’s full reasoning capability your tasks actually require. This guide breaks down the tradeoffs so you can make a practical choice rather than defaulting to the largest model that fits.
What “Distilled” Means Here
The R1 distilled models are not simply smaller versions of the 671B MoE base model. They are separate models — built on the Qwen-2.5 architecture (for 1.5B, 7B, 14B, 32B) and the LLaMA-3 architecture (for 70B) — that were trained to replicate R1’s reasoning behaviour by learning from R1’s chain-of-thought outputs. The training process transfers the reasoning style and capability from the large model into the smaller architecture. This means R1-Distill-14B thinks in the same extended, step-by-step style as the full R1, but within the capacity of a 14B parameter model. The distillation is highly effective: the smaller models genuinely reason rather than just pattern-match, which is what makes them competitive with proprietary models of similar size.
Hardware Requirements
The practical first question is what hardware you have. Running these models comfortably requires keeping most or all of the model in GPU VRAM to avoid slow CPU offloading. At 4-bit quantisation (Q4_K_M with llama.cpp or similar), approximate VRAM requirements are: 14B needs around 9–10GB, which fits on an RTX 3080 10GB, RTX 4070, or any 12GB consumer GPU; 32B needs around 20–22GB, requiring an RTX 3090 or 4090 (24GB), or two 12GB GPUs; 70B needs around 40–45GB, requiring an A100 40GB or two 24GB consumer GPUs. At full 16-bit precision, requirements roughly double. If you want fast interactive inference rather than slow batch processing, staying within VRAM is essential — CPU offloading on a 32B model reduces generation speed to 2–5 tokens per second, which makes it unsuitable for real-time use.
Performance by Size
On the benchmarks that matter for reasoning models, the performance gap between sizes is meaningful but not as large as parameter count alone might suggest. On AIME 2024 (competition mathematics), R1-Distill-70B scores around 70%, 32B around 72%, and 14B around 69% — the 32B is actually slightly ahead of the 70B here, likely due to the Qwen vs LLaMA architecture difference rather than size alone. On MATH-500, all three score above 93%, with 70B at ~95%, 32B at ~94%, and 14B at ~93%. On coding benchmarks (HumanEval, LiveCodeBench), the 70B leads more clearly. The practical takeaway is that the 32B is the sweet spot for most reasoning tasks — it nearly matches 70B performance while fitting on a single high-end consumer GPU.
Figure 1 — DeepSeek R1 distilled models: hardware needs vs benchmark performance
Running with Ollama
The fastest way to try any of these models locally is Ollama, which handles downloading, quantisation, and serving through a clean CLI:
ollama run deepseek-r1:14b # pulls and runs 14B
ollama run deepseek-r1:32b # pulls and runs 32B
ollama run deepseek-r1:70b # pulls and runs 70B — needs 40GB VRAM
# Check GPU utilisation
nvidia-smi
# Run via API (OpenAI-compatible endpoint)
curl http://localhost:11434/v1/chat/completions -H "Content-Type: application/json" -d '{"model":"deepseek-r1:14b","messages":[{"role":"user","content":"Prove that sqrt(2) is irrational"}]}'
The 14B Sweet Spot for Many Practitioners
If you have a single consumer GPU with 10–12GB of VRAM — an RTX 3080, 3080 Ti, or 4070 — then R1-Distill-14B is your practical ceiling, and it is a genuinely capable model. On the tasks where reasoning models are most useful (mathematics, step-by-step problem solving, logical analysis), it performs within a few percentage points of models that were considered frontier-level just one year earlier. The 7B is noticeably weaker on hard multi-step reasoning but still better than any standard 7B model at structured problem-solving. If the 14B does not produce satisfactory results on your specific task, that is a signal to consider the 32B or the full R1 API rather than accepting poor output quality. The 14B is particularly strong on structured reasoning tasks — step-by-step mathematics, logic puzzles, and short code problems — but shows its limits on problems requiring very long reasoning chains (10+ steps) or broad factual recall combined with reasoning. For those tasks, the jump to 32B is meaningful and worth the hardware upgrade if the 14B is consistently falling short.
The 32B as the Practical Top End for Consumer Hardware
The R1-Distill-32B requires a GPU with at least 24GB VRAM — currently only the RTX 3090 and RTX 4090 among consumer cards, though older Quadro and A-series cards in workstations also qualify. At 4-bit quantisation it fits comfortably in 24GB and runs at around 15–25 tokens per second depending on generation length and context size. The performance difference between 32B and 70B is small enough that the 32B is the recommended local reasoning model for anyone with a 3090 or 4090 — the 70B requires datacenter hardware and is only worth the investment if you have benchmarked the 32B and found it consistently insufficient on your tasks.
Multi-GPU for the 70B
Running the 70B across multiple consumer GPUs is possible using llama.cpp’s tensor parallelism or Ollama’s multi-GPU support. Two RTX 3090s (48GB total) can run the 70B at 4-bit comfortably. Generation speed will be lower than a single A100 due to the PCIe communication overhead between GPUs, but for batch processing or tasks where quality matters more than latency, this is a viable setup. The more common path is to use the DeepSeek R1 API (either DeepSeek’s own API or via providers like Fireworks.ai and Together.ai) for 70B-level quality when you need it, and self-host the 14B or 32B for the bulk of queries.
Context Length Considerations
All distilled variants support a 128k token context window, matching the full R1 model. In practice, reasoning models consume significant context during the thinking phase — a hard mathematics problem might generate 2,000–8,000 thinking tokens before the final answer. Keep this in mind when sizing your context budget: at 128k tokens, even very extended reasoning chains leave ample room for long input documents, but if you are running on hardware where memory is tight, shorter context lengths (32k) will improve throughput significantly without affecting most practical use cases.
Choosing in Practice
The decision is straightforward once you know your hardware. If you have 8–12GB VRAM: use 14B. If you have 24GB VRAM: use 32B — it is the best quality-to-hardware ratio in the family. If you have 40GB+ VRAM or two 24GB cards: 70B is available, but verify the quality improvement justifies the inference cost on your actual tasks. If you have no suitable GPU, use DeepSeek’s API, which prices R1 at a fraction of OpenAI o1’s cost and gives you full 671B MoE quality. The distilled models are primarily for practitioners who need local inference for privacy, latency, or cost reasons — if those constraints do not apply, the full R1 API is the better choice.
The R1 distilled family makes frontier-level reasoning available on hardware that most ML practitioners already own. The 14B runs on a gaming GPU bought in 2020; the 32B runs on a GPU bought for gaming in 2022. That accessibility — not just performance — is why the distilled models matter. Open-weights reasoning capability at consumer hardware prices represents a genuine democratisation of what was state-of-the-art AI just a year before these models were released.
Figure 1 — R1 distilled model sizes: VRAM requirements at 4-bit quantisation
The R1 distilled family makes frontier-level reasoning available on hardware most ML practitioners already own. The 14B runs on a 2020 gaming GPU; the 32B on a 2022 flagship. Open-weights reasoning capability at consumer hardware prices represents a genuine shift in access — what was state-of-the-art API-only AI just a year before R1’s release is now something you can run offline, privately, and at zero marginal cost per query.
Speed matters too: the 14B at 4-bit runs at 40–60 tokens per second on an RTX 4070, making it fast enough for interactive use. The 32B runs at 15–25 tokens per second on a 4090 — still interactive. The 70B on two 3090s drops to 8–12 tokens per second due to inter-GPU communication overhead. If generation speed is a hard requirement for your use case, factor these numbers into your hardware choice alongside benchmark quality.