Llama 3.2 1B VRAM Requirements
1.24B parameters, 16 layers, 8 key-value heads. At Q4_K_M with a 8K context it needs 1.7 GB — 30 of 30 accelerators here can hold it.
Running Llama 3.2 1B
Runs comfortably on a GeForce RTX 4090
- Weights
- 0.6 GB
- KV cache
- 0.3 GB
- Headroom
- 20.4 GB
- Max context
- 128K
Where the memory goes
| Component | Memory |
|---|---|
| Model weights (Q4_K_M) | 0.6 GB |
| KV cache at 8K tokens | 0.3 GB |
| Runtime allowance | 0.8 GB |
| Total | 1.7 GB |
Weights are only half the question
A 1.2B model at Q4_K_M is 0.6 GB of weights. That part is fixed — it is the file on disk, and it does not change while the model runs. What changes is the key-value cache, which grows with every token of context you give it.
Llama 3.2 1B spends 32.0 KB per token. At 8K that is 0.3 GB, comfortably smaller than the weights. At its full 128K it is 4.0 GB — larger than the weights themselves. That is the trap: the model loads fine and then fills the card once a conversation gets long.
The reason the cache is as small as it is comes down to grouped-query attention. Llama 3.2 1B has 32 query heads but only 8 key-value heads, and the cache is sized by the latter. Calculators that use the query head count overstate it 4-fold, which is where the wildly pessimistic numbers people quote usually come from. Ties input and output embeddings, which at a 128k vocabulary saves about a fifth of the total parameters.
By quantisation
| Format | Bits/weight | Weights | Total at 8K |
|---|---|---|---|
| FP16 / BF16 | 16 | 2.3 GB | 3.4 GB |
| Q8_0 | 8.5 | 1.2 GB | 2.3 GB |
| Q6_K | 6.6 | 0.9 GB | 2.0 GB |
| Q5_K_M | 5.5 | 0.8 GB | 1.8 GB |
| Q4_K_M | 4.5 | 0.6 GB | 1.7 GB |
| Q3_K_M | 3.9 | 0.6 GB | 1.6 GB |
By context length
| Context | KV cache | Total |
|---|---|---|
| 2K tokens | 0.1 GB | 1.5 GB |
| 4K tokens | 0.1 GB | 1.6 GB |
| 8K tokens | 0.3 GB | 1.7 GB |
| 16K tokens | 0.5 GB | 1.9 GB |
| 32K tokens | 1.0 GB | 2.4 GB |
| 64K tokens | 2.0 GB | 3.4 GB |
| 128K tokens | 4.0 GB | 5.4 GB |
Will it run on your hardware?
| Accelerator | Memory | Runs it | Headroom | Max context |
|---|---|---|---|---|
| GeForce RTX 5090 | 32 GB | Yes | 27.7 GB | 128K |
| GeForce RTX 4090 | 24 GB | Yes | 20.4 GB | 128K |
| GeForce RTX 3090 | 24 GB | Yes | 20.4 GB | 128K |
| GeForce RTX 5080 | 16 GB | Yes | 13.0 GB | 128K |
| GeForce RTX 4080 SUPER | 16 GB | Yes | 13.0 GB | 128K |
| GeForce RTX 5070 Ti | 16 GB | Yes | 13.0 GB | 128K |
| GeForce RTX 4070 Ti SUPER | 16 GB | Yes | 13.0 GB | 128K |
| GeForce RTX 5070 | 12 GB | Yes | 9.3 GB | 128K |
| GeForce RTX 4070 | 12 GB | Yes | 9.3 GB | 128K |
| GeForce RTX 3060 12GB | 12 GB | Yes | 9.3 GB | 128K |
| GeForce RTX 4060 Ti 16GB | 16 GB | Yes | 13.0 GB | 128K |
| GeForce RTX 3080 | 10 GB | Yes | 7.5 GB | 128K |
| Radeon RX 7900 XTX | 24 GB | Yes | 20.4 GB | 128K |
| Radeon RX 7900 XT | 20 GB | Yes | 16.7 GB | 128K |
| Radeon RX 9070 XT | 16 GB | Yes | 13.0 GB | 128K |
| 2× RTX 3090 (48 GB) | 48 GB | Yes | 42.5 GB | 128K |
| 2× RTX 4090 (48 GB) | 48 GB | Yes | 42.5 GB | 128K |
| 4× RTX 3090 (96 GB) | 96 GB | Yes | 86.6 GB | 128K |
| Mac (M4 Max, 64 GB) | 64 GB | Yes | 46.3 GB | 128K |
| Mac (M4 Max, 128 GB) | 128 GB | Yes | 94.3 GB | 128K |
| Mac (M4 Pro, 48 GB) | 48 GB | Yes | 34.3 GB | 128K |
| Mac (M4, 24 GB) | 24 GB | Yes | 14.4 GB | 128K |
| Mac (M3 Max, 128 GB) | 128 GB | Yes | 94.3 GB | 128K |
| Mac Studio (M2 Ultra, 192 GB) | 192 GB | Yes | 142.3 GB | 128K |
| Mac (M1 Max, 32 GB) | 32 GB | Yes | 19.7 GB | 128K |
| Mac (M2, 16 GB) | 16 GB | Yes | 9.0 GB | 128K |
| NVIDIA A100 40GB | 40 GB | Yes | 36.3 GB | 128K |
| NVIDIA A100 80GB | 80 GB | Yes | 74.3 GB | 128K |
| NVIDIA H100 80GB | 80 GB | Yes | 74.3 GB | 128K |
| NVIDIA L40S 48GB | 48 GB | Yes | 43.9 GB | 128K |
At Q4_K_M with a 8K context. "Only just" means the model uses more than 95% of usable memory — it will load and then fall over as soon as anything else touches the device.
Frequently asked questions
How much VRAM does Llama 3.2 1B need?
What GPU do I need to run Llama 3.2 1B?
Does context length change how much memory Llama 3.2 1B needs?
Why is the cache smaller than other calculators say?
Which quantisation should I use?
Architecture: Hugging Face config.json (public mirror of the gated original). GGUF K-quant effective bit widths as produced by llama.cpp.
An estimate of memory, not a benchmark. It counts model weights, the key-value cache at the context you choose, and a fixed runtime allowance. Actual usage moves with the runtime, the batch size, whether flash attention is on, and how much the operating system has already taken from a shared memory pool. Treat a result within a gigabyte of your card's capacity as 'probably not' rather than 'just fits'.
Data on this page last verified .