Which GPU for local LLMs?
Model-by-model VRAM requirements for local LLM inference (Q4 and Q8 quantization), plus gaming VRAM by resolution — from the same registry the planner uses.
VRAM per model
How much GPU memory each model needs to serve, at 8k context. Larger contexts add roughly 0.05 GB per 1k tokens.
| Model | Q4 (4-bit) | Q8 (8-bit) |
|---|---|---|
| Mistral 7B | 8 GB VRAM | 12 GB VRAM |
| Llama 3.1 8B | 8 GB VRAM | 12 GB VRAM |
| Qwen 2.5 7B | 8 GB VRAM | 12 GB VRAM |
| Phi-4 14B | 12 GB VRAM | 20 GB VRAM |
| Qwen 2.5 14B | 12 GB VRAM | 20 GB VRAM |
| Codestral 22B | 20 GB VRAM | 32 GB VRAM |
| Qwen 2.5 32B | 24 GB VRAM | 40 GB VRAM |
| Llama 3.1 70B | 80 GB VRAM | 80 GB VRAM |
| Qwen 2.5 72B | 80 GB VRAM | 96 GB VRAM |
| DeepSeek R1 8B (distill) | 8 GB VRAM | 12 GB VRAM |
Estimate: parameters × bytes-per-quant plus runtime and context overhead, rounded up to the nearest real GPU size.
Gaming VRAM by resolution
Comfortable VRAM floors for modern titles at high settings — approximate, not per-title tuned.
1080p
8 GB VRAM
1080p high settings — 8GB is the comfortable floor for modern titles.
1440p
12 GB VRAM
1440p high settings — 12GB avoids texture-streaming hitches.
4k
16 GB VRAM
4K high settings — 16GB for texture headroom.