Local AI inference: hardware requirements guide
Running LLMs or other models locally (Ollama, vLLM, llama.cpp).
Hardware requirements by tier
Three tiers for every workload: the minimum that works, the recommended sweet spot, and the comfortable headroom level. These are the same tiers the WisePC decision engine uses when it plans a build around your goal.
| Tier | CPU cores | RAM | Storage | GPU | VRAM | Network |
|---|---|---|---|---|---|---|
| Minimum | 8 cores | 16 GB | 1 TB (ssd) | ai | 12 GB VRAM | 1 GbE |
| Recommended | 12 cores | 32 GB | 1 TB (ssd) | ai | 16 GB VRAM | 2.5 GbE |
| Comfortable | 16 cores | 64 GB | 2 TB (ssd) | ai | 24 GB VRAM | 10 GbE |
| Overkill | 26 cores | 103 GB | 4 TB (ssd) | ai | 24 GB VRAM | 10 GbE |
Beyond Comfortable — 26 cores, 103 GB RAM, 4 TB (SSD): headroom you will never use for this workload alone. Put the difference into storage, backup or silence.
Recommended system shape
Derived by the decision engine from the recommended tier — the same logic that plans full builds.
- Form factor
- tower
- Build style
- diy
- Drive plan
- 2× 2TB NVMe
Requirements (12 threads, 32GB RAM, 4TB storage, dedicated GPU) need a custom tower build.
One box or a separate machine?
Whether this workload deserves its own system or should live on the main server.
Keep on the main server
GPU-bound compute is expensive to duplicate — a second GPU box roughly doubles cost for little resilience gain. Keep it on the main machine.
Storage arrays, priced
Five physical strategies for the recommended capacity — the same planner, drive model and reference prices as the storage planner tool.
| Strategy | Drives | Usable | Cost | Idle |
|---|---|---|---|---|
| Cheapest | 1× 4TB hdd | 4 TB · none | ≈ €100 | 5 W |
| Balanced | 1× 16TB hdd | 16 TB · none | ≈ €240 | 6 W |
| Redundant | 3× 8TB hdd | 8 TB · parity2 | ≈ €450 | 15 W |
| Expansion-friendly | 1× 20TB hdd | 20 TB · none | ≈ €300 | 7 W |
| PerformanceRecommended | 2× 4TB nvme | 4 TB · parity1 | ≈ €480 | 2 W |
For a fast working tier, go all-flash — the performance strategy protects speed without the HDD rebuild pain.
Backup is not storage
The extra copy that survives drive failure, deletion and ransomware — the same engine as the backup planner.
Protect 1TB with at least two copies on separate hardware, ideally one offsite. The recommended start is a dedicated NAS.
Recommended: Separate NAS — A dedicated box that only stores backups · ≈ €278
- No offsite copy (high): Add a second copy outside the building — cloud or a drive at a friend's house.
- No versioning (medium): Use a versioning-aware tool (restic, borg, Backblaze) that keeps history.
- Unscheduled backups (medium): Schedule backups automatically and monitor them.
- Backup on the same machine (medium): Keep the backup on separate hardware or offsite.
What will it cost?
Market-neutral EUR bands from the reference drive model (≈ = estimated street price). Live component prices resolve on the build page for your region.
- Storage array
- ≈ €100 – €480
- Array energy / year
- ≈ €81/yr (∼37 W idle)
Compute (CPU, board, RAM, GPU) is not included — see the recommended build for live prices.
Growth outlook
What happens when the data keeps growing — the same engine as the growth simulator.
This architecture becomes inadequate in year 2 — 2 bays needed. Plan the next step (bigger array, NAS or DAS, or a stronger platform) before then.
Exceeds max 2TB by 1TB — need larger array/DAS/NAS. Fresh Performance array for 3TB: 2 bays, ~480€.
Also covered in this guide
Code assistant (local LLM)
Local code completion and chat models (Ollama, Continue).
Which tier do you need?
- Pick Minimum (8 cores, 16 GB RAM) only for a single-purpose machine on a tight budget — expect little headroom.
- Recommended (12 cores, 32 GB RAM) is the sweet spot: enough for the workload plus the usual side-services, without overspending.
- Pick Comfortable (16 cores, 64 GB RAM) when this workload shares the machine with others or will grow — you pay for headroom, not for anxiety.
- GPU rule for this workload: ai.GPU rule for this workload: ai (16 GB VRAM).
Frequently asked questions
How much RAM does local ai inference need?
16 GB is the sensible minimum, 32 GB covers most real setups, and 64 GB gives comfortable headroom for growth and extra services.
How many CPU cores does local ai inference need?
A 8-core CPU is the minimum, 12 cores is the recommended sweet spot, and 16 cores is comfortable when it shares the machine with other workloads.
Does local ai inference need a dedicated GPU?
A dedicated GPU is strongly recommended — this workload does AI compute. Needs 12–24 GB VRAM (12 minimum, 16 recommended, 24 comfortable).
What storage and network does local ai inference expect?
Storage: 1 TB of SSD is the recommended baseline (1 TB minimum, 2 TB comfortable). Network: 2.5 GbE is the recommended baseline.
What runs well alongside local ai inference?
It pairs naturally with: Image generation, AI experimentation, Databases.
How much power does local ai inference use?
The storage array idles around 37 W (≈ 81 €/year at 0.25 €/kWh). The full system adds CPU, board and fans on top — see the cost section for the honest bands.