Mistral 7B v0.3 VRAM Requirements
7.25B parameters, 32 layers, 8 key-value heads. At Q4_K_M with a 8K context it needs 5.6 GB — 30 of 30 accelerators here can hold it.
Running Mistral 7B v0.3
Runs comfortably on a GeForce RTX 4090
- Weights
- 3.8 GB
- KV cache
- 1.0 GB
- Headroom
- 16.5 GB
- Max context
- 32K
Where the memory goes
| Component | Memory |
|---|---|
| Model weights (Q4_K_M) | 3.8 GB |
| KV cache at 8K tokens | 1.0 GB |
| Runtime allowance | 0.8 GB |
| Total | 5.6 GB |
Weights are only half the question
A 7.3B model at Q4_K_M is 3.8 GB of weights. That part is fixed — it is the file on disk, and it does not change while the model runs. What changes is the key-value cache, which grows with every token of context you give it.
Mistral 7B v0.3 spends 128.0 KB per token. At 8K that is 1.0 GB, comfortably smaller than the weights. At its full 32K it is 4.0 GB — larger than the weights themselves. That is the trap: the model loads fine and then fills the card once a conversation gets long.
The reason the cache is as small as it is comes down to grouped-query attention. Mistral 7B v0.3 has 32 query heads but only 8 key-value heads, and the cache is sized by the latter. Calculators that use the query head count overstate it 4-fold, which is where the wildly pessimistic numbers people quote usually come from. A 32k vocabulary against Llama's 128k. Smaller embedding tables are most of why it is lighter than Llama 3.1 8B despite an otherwise identical shape.
By quantisation
| Format | Bits/weight | Weights | Total at 8K |
|---|---|---|---|
| FP16 / BF16 | 16 | 13.5 GB | 15.3 GB |
| Q8_0 | 8.5 | 7.2 GB | 9.0 GB |
| Q6_K | 6.6 | 5.6 GB | 7.4 GB |
| Q5_K_M | 5.5 | 4.6 GB | 6.4 GB |
| Q4_K_M | 4.5 | 3.8 GB | 5.6 GB |
| Q3_K_M | 3.9 | 3.3 GB | 5.1 GB |
By context length
| Context | KV cache | Total |
|---|---|---|
| 2K tokens | 0.3 GB | 4.8 GB |
| 4K tokens | 0.5 GB | 5.1 GB |
| 8K tokens | 1.0 GB | 5.6 GB |
| 16K tokens | 2.0 GB | 6.6 GB |
| 32K tokens | 4.0 GB | 8.6 GB |
Will it run on your hardware?
| Accelerator | Memory | Runs it | Headroom | Max context |
|---|---|---|---|---|
| GeForce RTX 5090 | 32 GB | Yes | 23.8 GB | 32K |
| GeForce RTX 4090 | 24 GB | Yes | 16.5 GB | 32K |
| GeForce RTX 3090 | 24 GB | Yes | 16.5 GB | 32K |
| GeForce RTX 5080 | 16 GB | Yes | 9.1 GB | 32K |
| GeForce RTX 4080 SUPER | 16 GB | Yes | 9.1 GB | 32K |
| GeForce RTX 5070 Ti | 16 GB | Yes | 9.1 GB | 32K |
| GeForce RTX 4070 Ti SUPER | 16 GB | Yes | 9.1 GB | 32K |
| GeForce RTX 5070 | 12 GB | Yes | 5.4 GB | 32K |
| GeForce RTX 4070 | 12 GB | Yes | 5.4 GB | 32K |
| GeForce RTX 3060 12GB | 12 GB | Yes | 5.4 GB | 32K |
| GeForce RTX 4060 Ti 16GB | 16 GB | Yes | 9.1 GB | 32K |
| GeForce RTX 3080 | 10 GB | Yes | 3.6 GB | 32K |
| Radeon RX 7900 XTX | 24 GB | Yes | 16.5 GB | 32K |
| Radeon RX 7900 XT | 20 GB | Yes | 12.8 GB | 32K |
| Radeon RX 9070 XT | 16 GB | Yes | 9.1 GB | 32K |
| 2× RTX 3090 (48 GB) | 48 GB | Yes | 38.6 GB | 32K |
| 2× RTX 4090 (48 GB) | 48 GB | Yes | 38.6 GB | 32K |
| 4× RTX 3090 (96 GB) | 96 GB | Yes | 82.7 GB | 32K |
| Mac (M4 Max, 64 GB) | 64 GB | Yes | 42.4 GB | 32K |
| Mac (M4 Max, 128 GB) | 128 GB | Yes | 90.4 GB | 32K |
| Mac (M4 Pro, 48 GB) | 48 GB | Yes | 30.4 GB | 32K |
| Mac (M4, 24 GB) | 24 GB | Yes | 10.5 GB | 32K |
| Mac (M3 Max, 128 GB) | 128 GB | Yes | 90.4 GB | 32K |
| Mac Studio (M2 Ultra, 192 GB) | 192 GB | Yes | 138.4 GB | 32K |
| Mac (M1 Max, 32 GB) | 32 GB | Yes | 15.8 GB | 32K |
| Mac (M2, 16 GB) | 16 GB | Yes | 5.1 GB | 32K |
| NVIDIA A100 40GB | 40 GB | Yes | 32.4 GB | 32K |
| NVIDIA A100 80GB | 80 GB | Yes | 70.4 GB | 32K |
| NVIDIA H100 80GB | 80 GB | Yes | 70.4 GB | 32K |
| NVIDIA L40S 48GB | 48 GB | Yes | 40.0 GB | 32K |
At Q4_K_M with a 8K context. "Only just" means the model uses more than 95% of usable memory — it will load and then fall over as soon as anything else touches the device.
Frequently asked questions
How much VRAM does Mistral 7B v0.3 need?
What GPU do I need to run Mistral 7B v0.3?
Does context length change how much memory Mistral 7B v0.3 needs?
Why is the cache smaller than other calculators say?
Which quantisation should I use?
Architecture: Hugging Face config.json. GGUF K-quant effective bit widths as produced by llama.cpp.
An estimate of memory, not a benchmark. It counts model weights, the key-value cache at the context you choose, and a fixed runtime allowance. Actual usage moves with the runtime, the batch size, whether flash attention is on, and how much the operating system has already taken from a shared memory pool. Treat a result within a gigabyte of your card's capacity as 'probably not' rather than 'just fits'.
Data on this page last verified .