llama3.2:3b: memory needed and real speed (benchmark)
llama3.2:3b measured on a home PC: about 2.3GB of memory, about 190 tokens/sec on GPU and 27.3 tokens/sec on CPU only. Which PCs can run it, by VRAM and RAM.
Summary
- Memory used: about 2.3GB (measured 2026-10-04)
- GPU (16GB VRAM): about 189.7 tokens/sec Instant
- CPU only: about 27.3 tokens/sec Comfortable
Speed by setup
| Setup | Tokens/sec | Feel |
|---|---|---|
| GPU, 16GB VRAM (RTX 4080) | 189.7 | Instant |
| 12GB VRAM (simulated) | 189.7 | Instant |
| 8GB VRAM (simulated) | 189.7 | Instant |
| CPU only (no GPU) | 27.3 | Comfortable |
Which PCs can run it?
- 8GB VRAM GPUFits in VRAM · about 189.7 tokens/secInstant
- 12GB VRAM GPUFits in VRAM · about 189.7 tokens/secInstant
- 16GB VRAM GPUFits in VRAM · about 189.7 tokens/secInstant
- 8GB RAM, no GPURuns on CPU · about 27.3 tokens/secComfortable
- 16GB RAM, no GPURuns on CPU · about 27.3 tokens/secComfortable
- 32GB RAM, no GPURuns on CPU · about 27.3 tokens/secComfortable
We assume about 4GB of RAM for the OS and 1GB of VRAM headroom.
Compared with other models
FAQ
How much memory does llama3.2:3b need?
In our test it used about 2.3GB when loaded. For fast GPU inference, plan on 8GB of VRAM or more; without a GPU, 8GB of system RAM or more.
Can llama3.2:3b run without a GPU?
Yes. On CPU only it generated about 27.3 tokens/sec (comfortable).
Measured on the site owner's PC (Ryzen 7 9800X3D / 31GB RAM / RTX 4080 (16GB VRAM)). "8GB / 12GB VRAM" values were measured on the RTX 4080 with the usable VRAM limited to that amount; real 8GB or 12GB cards are slower GPUs, so actual speeds may be lower. How we measure