Gemma 2 27B VRAM Requirements
27.23B parameters, 46 layers, 16 key-value heads. At Q4_K_M with a 8K context it needs 17.9 GB — 18 of 30 accelerators here can hold it.
Running Gemma 2 27B
Runs comfortably on a GeForce RTX 4090
- Weights
- 14.3 GB
- KV cache
- 2.9 GB
- Headroom
- 4.2 GB
- Max context
- 8K
This model uses sliding-window attention on some layers, which makes its real cache smaller than shown at long context. The figure here is an upper bound.
Where the memory goes
| Component | Memory |
|---|---|
| Model weights (Q4_K_M) | 14.3 GB |
| KV cache at 8K tokens | 2.9 GB |
| Runtime allowance | 0.8 GB |
| Total | 17.9 GB |
Weights are only half the question
A 27.2B model at Q4_K_M is 14.3 GB of weights. That part is fixed — it is the file on disk, and it does not change while the model runs. What changes is the key-value cache, which grows with every token of context you give it.
Gemma 2 27B spends 368.0 KB per token. At 8K that is 2.9 GB, comfortably smaller than the weights. At its full 8K it is 2.9 GB — still under the weights, which is unusual for a long-context model. That is the trap: the model loads fine and then fills the card once a conversation gets long.
The reason the cache is as small as it is comes down to grouped-query attention. Gemma 2 27B has 32 query heads but only 16 key-value heads, and the cache is sized by the latter. Calculators that use the query head count overstate it 2-fold, which is where the wildly pessimistic numbers people quote usually come from.
By quantisation
| Format | Bits/weight | Weights | Total at 8K |
|---|---|---|---|
| FP16 / BF16 | 16 | 50.7 GB | 54.4 GB |
| Q8_0 | 8.5 | 26.9 GB | 30.6 GB |
| Q6_K | 6.6 | 20.9 GB | 24.6 GB |
| Q5_K_M | 5.5 | 17.4 GB | 21.1 GB |
| Q4_K_M | 4.5 | 14.3 GB | 17.9 GB |
| Q3_K_M | 3.9 | 12.4 GB | 16.0 GB |
By context length
| Context | KV cache | Total |
|---|---|---|
| 2K tokens | 0.7 GB | 15.8 GB |
| 4K tokens | 1.4 GB | 16.5 GB |
| 8K tokens | 2.9 GB | 17.9 GB |
Will it run on your hardware?
| Accelerator | Memory | Runs it | Headroom | Max context |
|---|---|---|---|---|
| GeForce RTX 5090 | 32 GB | Yes | 11.5 GB | 8K |
| GeForce RTX 4090 | 24 GB | Yes | 4.2 GB | 8K |
| GeForce RTX 3090 | 24 GB | Yes | 4.2 GB | 8K |
| GeForce RTX 5080 | 16 GB | No | — | — |
| GeForce RTX 4080 SUPER | 16 GB | No | — | — |
| GeForce RTX 5070 Ti | 16 GB | No | — | — |
| GeForce RTX 4070 Ti SUPER | 16 GB | No | — | — |
| GeForce RTX 5070 | 12 GB | No | — | — |
| GeForce RTX 4070 | 12 GB | No | — | — |
| GeForce RTX 3060 12GB | 12 GB | No | — | — |
| GeForce RTX 4060 Ti 16GB | 16 GB | No | — | — |
| GeForce RTX 3080 | 10 GB | No | — | — |
| Radeon RX 7900 XTX | 24 GB | Yes | 4.2 GB | 8K |
| Radeon RX 7900 XT | 20 GB | Only just | 0.5 GB | 8K |
| Radeon RX 9070 XT | 16 GB | No | — | — |
| 2× RTX 3090 (48 GB) | 48 GB | Yes | 26.3 GB | 8K |
| 2× RTX 4090 (48 GB) | 48 GB | Yes | 26.3 GB | 8K |
| 4× RTX 3090 (96 GB) | 96 GB | Yes | 70.4 GB | 8K |
| Mac (M4 Max, 64 GB) | 64 GB | Yes | 30.1 GB | 8K |
| Mac (M4 Max, 128 GB) | 128 GB | Yes | 78.1 GB | 8K |
| Mac (M4 Pro, 48 GB) | 48 GB | Yes | 18.1 GB | 8K |
| Mac (M4, 24 GB) | 24 GB | No | — | 3K |
| Mac (M3 Max, 128 GB) | 128 GB | Yes | 78.1 GB | 8K |
| Mac Studio (M2 Ultra, 192 GB) | 192 GB | Yes | 126.1 GB | 8K |
| Mac (M1 Max, 32 GB) | 32 GB | Yes | 3.5 GB | 8K |
| Mac (M2, 16 GB) | 16 GB | No | — | — |
| NVIDIA A100 40GB | 40 GB | Yes | 20.1 GB | 8K |
| NVIDIA A100 80GB | 80 GB | Yes | 58.1 GB | 8K |
| NVIDIA H100 80GB | 80 GB | Yes | 58.1 GB | 8K |
| NVIDIA L40S 48GB | 48 GB | Yes | 27.7 GB | 8K |
At Q4_K_M with a 8K context. "Only just" means the model uses more than 95% of usable memory — it will load and then fall over as soon as anything else touches the device.
Frequently asked questions
How much VRAM does Gemma 2 27B need?
What GPU do I need to run Gemma 2 27B?
Does context length change how much memory Gemma 2 27B needs?
Why is the cache smaller than other calculators say?
Which quantisation should I use?
Architecture: Hugging Face config.json (public mirror of the gated original). GGUF K-quant effective bit widths as produced by llama.cpp.
An estimate of memory, not a benchmark. It counts model weights, the key-value cache at the context you choose, and a fixed runtime allowance. Actual usage moves with the runtime, the batch size, whether flash attention is on, and how much the operating system has already taken from a shared memory pool. Treat a result within a gigabyte of your card's capacity as 'probably not' rather than 'just fits'.
Data on this page last verified .