Phi-4 VRAM Requirements
14.66B parameters, 40 layers, 10 key-value heads. At Q4_K_M with a 8K context it needs 10.0 GB — 29 of 30 accelerators here can hold it.
Running Phi-4
Runs comfortably on a GeForce RTX 4090
- Weights
- 7.7 GB
- KV cache
- 1.6 GB
- Headroom
- 12.1 GB
- Max context
- 16K
Where the memory goes
| Component | Memory |
|---|---|
| Model weights (Q4_K_M) | 7.7 GB |
| KV cache at 8K tokens | 1.6 GB |
| Runtime allowance | 0.8 GB |
| Total | 10.0 GB |
Weights are only half the question
A 14.7B model at Q4_K_M is 7.7 GB of weights. That part is fixed — it is the file on disk, and it does not change while the model runs. What changes is the key-value cache, which grows with every token of context you give it.
Phi-4 spends 200.0 KB per token. At 8K that is 1.6 GB, comfortably smaller than the weights. At its full 16K it is 3.1 GB — still under the weights, which is unusual for a long-context model. That is the trap: the model loads fine and then fills the card once a conversation gets long.
The reason the cache is as small as it is comes down to grouped-query attention. Phi-4 has 40 query heads but only 10 key-value heads, and the cache is sized by the latter. Calculators that use the query head count overstate it 4-fold, which is where the wildly pessimistic numbers people quote usually come from.
By quantisation
| Format | Bits/weight | Weights | Total at 8K |
|---|---|---|---|
| FP16 / BF16 | 16 | 27.3 GB | 29.7 GB |
| Q8_0 | 8.5 | 14.5 GB | 16.9 GB |
| Q6_K | 6.6 | 11.3 GB | 13.6 GB |
| Q5_K_M | 5.5 | 9.4 GB | 11.7 GB |
| Q4_K_M | 4.5 | 7.7 GB | 10.0 GB |
| Q3_K_M | 3.9 | 6.7 GB | 9.0 GB |
By context length
| Context | KV cache | Total |
|---|---|---|
| 2K tokens | 0.4 GB | 8.9 GB |
| 4K tokens | 0.8 GB | 9.3 GB |
| 8K tokens | 1.6 GB | 10.0 GB |
| 16K tokens | 3.1 GB | 11.6 GB |
Will it run on your hardware?
| Accelerator | Memory | Runs it | Headroom | Max context |
|---|---|---|---|---|
| GeForce RTX 5090 | 32 GB | Yes | 19.4 GB | 16K |
| GeForce RTX 4090 | 24 GB | Yes | 12.1 GB | 16K |
| GeForce RTX 3090 | 24 GB | Yes | 12.1 GB | 16K |
| GeForce RTX 5080 | 16 GB | Yes | 4.7 GB | 16K |
| GeForce RTX 4080 SUPER | 16 GB | Yes | 4.7 GB | 16K |
| GeForce RTX 5070 Ti | 16 GB | Yes | 4.7 GB | 16K |
| GeForce RTX 4070 Ti SUPER | 16 GB | Yes | 4.7 GB | 16K |
| GeForce RTX 5070 | 12 GB | Yes | 1.0 GB | 13K |
| GeForce RTX 4070 | 12 GB | Yes | 1.0 GB | 13K |
| GeForce RTX 3060 12GB | 12 GB | Yes | 1.0 GB | 13K |
| GeForce RTX 4060 Ti 16GB | 16 GB | Yes | 4.7 GB | 16K |
| GeForce RTX 3080 | 10 GB | No | — | 4K |
| Radeon RX 7900 XTX | 24 GB | Yes | 12.1 GB | 16K |
| Radeon RX 7900 XT | 20 GB | Yes | 8.4 GB | 16K |
| Radeon RX 9070 XT | 16 GB | Yes | 4.7 GB | 16K |
| 2× RTX 3090 (48 GB) | 48 GB | Yes | 34.2 GB | 16K |
| 2× RTX 4090 (48 GB) | 48 GB | Yes | 34.2 GB | 16K |
| 4× RTX 3090 (96 GB) | 96 GB | Yes | 78.3 GB | 16K |
| Mac (M4 Max, 64 GB) | 64 GB | Yes | 38.0 GB | 16K |
| Mac (M4 Max, 128 GB) | 128 GB | Yes | 86.0 GB | 16K |
| Mac (M4 Pro, 48 GB) | 48 GB | Yes | 26.0 GB | 16K |
| Mac (M4, 24 GB) | 24 GB | Yes | 6.1 GB | 16K |
| Mac (M3 Max, 128 GB) | 128 GB | Yes | 86.0 GB | 16K |
| Mac Studio (M2 Ultra, 192 GB) | 192 GB | Yes | 134.0 GB | 16K |
| Mac (M1 Max, 32 GB) | 32 GB | Yes | 11.4 GB | 16K |
| Mac (M2, 16 GB) | 16 GB | Yes | 0.7 GB | 11K |
| NVIDIA A100 40GB | 40 GB | Yes | 28.0 GB | 16K |
| NVIDIA A100 80GB | 80 GB | Yes | 66.0 GB | 16K |
| NVIDIA H100 80GB | 80 GB | Yes | 66.0 GB | 16K |
| NVIDIA L40S 48GB | 48 GB | Yes | 35.6 GB | 16K |
At Q4_K_M with a 8K context. "Only just" means the model uses more than 95% of usable memory — it will load and then fall over as soon as anything else touches the device.
Frequently asked questions
How much VRAM does Phi-4 need?
What GPU do I need to run Phi-4?
Does context length change how much memory Phi-4 needs?
Why is the cache smaller than other calculators say?
Which quantisation should I use?
Architecture: Hugging Face config.json. GGUF K-quant effective bit widths as produced by llama.cpp.
An estimate of memory, not a benchmark. It counts model weights, the key-value cache at the context you choose, and a fixed runtime allowance. Actual usage moves with the runtime, the batch size, whether flash attention is on, and how much the operating system has already taken from a shared memory pool. Treat a result within a gigabyte of your card's capacity as 'probably not' rather than 'just fits'.
Data on this page last verified .