gemma4:12b-it-q8_0: memory needed and real speed (benchmark)
gemma4:12b-it-q8_0 measured on a home PC: about 14.2GB of memory, about 43 tokens/sec on GPU and 4.6 tokens/sec on CPU only. Which PCs can run it, by VRAM and RAM.
Summary
- Memory used: about 14.2GB (measured 2026-10-04)
- GPU (16GB VRAM): about 43.5 tokens/sec Instant
- CPU only: about 4.6 tokens/sec Slow
Speed by setup
| Setup | Tokens/sec | Feel |
|---|---|---|
| GPU, 16GB VRAM (RTX 4080) | 43.5 | Instant |
| 12GB VRAM (simulated) | 17.2 | Comfortable |
| 8GB VRAM (simulated) | 8.2 | Short wait |
| CPU only (no GPU) | 4.6 | Slow |
Which PCs can run it?
- 8GB VRAM GPUPartly on CPU (slower) · about 8.2 tokens/secShort wait
- 12GB VRAM GPUPartly on CPU (slower) · about 17.2 tokens/secComfortable
- 16GB VRAM GPUFits in VRAM · about 43.5 tokens/secInstant
- 8GB RAM, no GPUNot enough memory
- 16GB RAM, no GPUNot enough memory
- 32GB RAM, no GPURuns on CPU · about 4.6 tokens/secSlow
We assume about 4GB of RAM for the OS and 1GB of VRAM headroom.
Compared with other models
FAQ
How much memory does gemma4:12b-it-q8_0 need?
In our test it used about 14.2GB when loaded. For fast GPU inference, plan on 16GB of VRAM or more; without a GPU, 32GB of system RAM or more.
Can gemma4:12b-it-q8_0 run without a GPU?
Yes. On CPU only it generated about 4.6 tokens/sec (slow).
Does gemma4:12b-it-q8_0 run on an 8GB VRAM GPU?
Yes, but it does not fit entirely, so part of it runs on the CPU: about 8.2 tokens/sec with 8GB VRAM versus about 43.5 tokens/sec with 16GB.
Measured on the site owner's PC (Ryzen 7 9800X3D / 31GB RAM / RTX 4080 (16GB VRAM)). "8GB / 12GB VRAM" values were measured on the RTX 4080 with the usable VRAM limited to that amount; real 8GB or 12GB cards are slower GPUs, so actual speeds may be lower. How we measure