DeepSeek-R1 Distill Qwen 14B VRAM Requirements
14.77B parameters, 48 layers, 8 key-value heads. At Q4_K_M with a 8K context it needs 10.0 GB — 29 of 30 accelerators here can hold it.
Running DeepSeek-R1 Distill Qwen 14B
Runs comfortably on a GeForce RTX 4090
- Weights
- 7.7 GB
- KV cache
- 1.5 GB
- Headroom
- 12.1 GB
- Max context
- 72K
Where the memory goes
| Component | Memory |
|---|---|
| Model weights (Q4_K_M) | 7.7 GB |
| KV cache at 8K tokens | 1.5 GB |
| Runtime allowance | 0.8 GB |
| Total | 10.0 GB |
Weights are only half the question
A 14.8B model at Q4_K_M is 7.7 GB of weights. That part is fixed — it is the file on disk, and it does not change while the model runs. What changes is the key-value cache, which grows with every token of context you give it.
DeepSeek-R1 Distill Qwen 14B spends 192.0 KB per token. At 8K that is 1.5 GB, comfortably smaller than the weights. At its full 128K it is 24.0 GB — larger than the weights themselves. That is the trap: the model loads fine and then fills the card once a conversation gets long.
The reason the cache is as small as it is comes down to grouped-query attention. DeepSeek-R1 Distill Qwen 14B has 40 query heads but only 8 key-value heads, and the cache is sized by the latter. Calculators that use the query head count overstate it 5-fold, which is where the wildly pessimistic numbers people quote usually come from. Despite the name this is a Qwen2 architecture — R1's reasoning distilled into Qwen2.5 14B, not a DeepSeek model shape. That matters for planning: reasoning models spend thousands of tokens thinking before they answer, so the KV cache, not the weights, is what usually decides whether a long session fits. Its config carries a sliding_window value but sets use_sliding_window to false, so none is applied here.
By quantisation
| Format | Bits/weight | Weights | Total at 8K |
|---|---|---|---|
| FP16 / BF16 | 16 | 27.5 GB | 29.8 GB |
| Q8_0 | 8.5 | 14.6 GB | 16.9 GB |
| Q6_K | 6.6 | 11.3 GB | 13.6 GB |
| Q5_K_M | 5.5 | 9.5 GB | 11.8 GB |
| Q4_K_M | 4.5 | 7.7 GB | 10.0 GB |
| Q3_K_M | 3.9 | 6.7 GB | 9.0 GB |
By context length
| Context | KV cache | Total |
|---|---|---|
| 2K tokens | 0.4 GB | 8.9 GB |
| 4K tokens | 0.8 GB | 9.3 GB |
| 8K tokens | 1.5 GB | 10.0 GB |
| 16K tokens | 3.0 GB | 11.5 GB |
| 32K tokens | 6.0 GB | 14.5 GB |
| 64K tokens | 12.0 GB | 20.5 GB |
| 128K tokens | 24.0 GB | 32.5 GB |
Will it run on your hardware?
| Accelerator | Memory | Runs it | Headroom | Max context |
|---|---|---|---|---|
| GeForce RTX 5090 | 32 GB | Yes | 19.4 GB | 111K |
| GeForce RTX 4090 | 24 GB | Yes | 12.1 GB | 72K |
| GeForce RTX 3090 | 24 GB | Yes | 12.1 GB | 72K |
| GeForce RTX 5080 | 16 GB | Yes | 4.7 GB | 33K |
| GeForce RTX 4080 SUPER | 16 GB | Yes | 4.7 GB | 33K |
| GeForce RTX 5070 Ti | 16 GB | Yes | 4.7 GB | 33K |
| GeForce RTX 4070 Ti SUPER | 16 GB | Yes | 4.7 GB | 33K |
| GeForce RTX 5070 | 12 GB | Yes | 1.0 GB | 13K |
| GeForce RTX 4070 | 12 GB | Yes | 1.0 GB | 13K |
| GeForce RTX 3060 12GB | 12 GB | Yes | 1.0 GB | 13K |
| GeForce RTX 4060 Ti 16GB | 16 GB | Yes | 4.7 GB | 33K |
| GeForce RTX 3080 | 10 GB | No | — | 4K |
| Radeon RX 7900 XTX | 24 GB | Yes | 12.1 GB | 72K |
| Radeon RX 7900 XT | 20 GB | Yes | 8.4 GB | 53K |
| Radeon RX 9070 XT | 16 GB | Yes | 4.7 GB | 33K |
| 2× RTX 3090 (48 GB) | 48 GB | Yes | 34.2 GB | 128K |
| 2× RTX 4090 (48 GB) | 48 GB | Yes | 34.2 GB | 128K |
| 4× RTX 3090 (96 GB) | 96 GB | Yes | 78.3 GB | 128K |
| Mac (M4 Max, 64 GB) | 64 GB | Yes | 38.0 GB | 128K |
| Mac (M4 Max, 128 GB) | 128 GB | Yes | 86.0 GB | 128K |
| Mac (M4 Pro, 48 GB) | 48 GB | Yes | 26.0 GB | 128K |
| Mac (M4, 24 GB) | 24 GB | Yes | 6.1 GB | 40K |
| Mac (M3 Max, 128 GB) | 128 GB | Yes | 86.0 GB | 128K |
| Mac Studio (M2 Ultra, 192 GB) | 192 GB | Yes | 134.0 GB | 128K |
| Mac (M1 Max, 32 GB) | 32 GB | Yes | 11.4 GB | 69K |
| Mac (M2, 16 GB) | 16 GB | Yes | 0.7 GB | 12K |
| NVIDIA A100 40GB | 40 GB | Yes | 28.0 GB | 128K |
| NVIDIA A100 80GB | 80 GB | Yes | 66.0 GB | 128K |
| NVIDIA H100 80GB | 80 GB | Yes | 66.0 GB | 128K |
| NVIDIA L40S 48GB | 48 GB | Yes | 35.6 GB | 128K |
At Q4_K_M with a 8K context. "Only just" means the model uses more than 95% of usable memory — it will load and then fall over as soon as anything else touches the device.
Frequently asked questions
How much VRAM does DeepSeek-R1 Distill Qwen 14B need?
What GPU do I need to run DeepSeek-R1 Distill Qwen 14B?
Does context length change how much memory DeepSeek-R1 Distill Qwen 14B needs?
Why is the cache smaller than other calculators say?
Which quantisation should I use?
Architecture: Hugging Face config.json. GGUF K-quant effective bit widths as produced by llama.cpp.
An estimate of memory, not a benchmark. It counts model weights, the key-value cache at the context you choose, and a fixed runtime allowance. Actual usage moves with the runtime, the batch size, whether flash attention is on, and how much the operating system has already taken from a shared memory pool. Treat a result within a gigabyte of your card's capacity as 'probably not' rather than 'just fits'.
Data on this page last verified .