gemma4:12b-it-q8_0: memory needed and real speed (benchmark)

gemma4:12b-it-q8_0 measured on a home PC: about 14.2GB of memory, about 43 tokens/sec on GPU and 4.6 tokens/sec on CPU only. Which PCs can run it, by VRAM and RAM.

Summary
  • Memory used: about 14.2GB (measured 2026-10-04)
  • GPU (16GB VRAM): about 43.5 tokens/sec Instant
  • CPU only: about 4.6 tokens/sec Slow

Speed by setup

SetupTokens/secFeel
GPU, 16GB VRAM (RTX 4080)43.5Instant
12GB VRAM (simulated)17.2Comfortable
8GB VRAM (simulated)8.2Short wait
CPU only (no GPU)4.6Slow

Which PCs can run it?

We assume about 4GB of RAM for the OS and 1GB of VRAM headroom.

Compared with other models

Tokens/secGPU (RTX 4080)CPU onlygemma3:4b2.8GB142.021.5qwen3.5:9b6.0GB96.411.2★ gemma4:12b-it-q8_014.2GB43.54.6
Measured on Ryzen 7 9800X3D / 31GB RAM / RTX 4080 (16GB VRAM).

FAQ

How much memory does gemma4:12b-it-q8_0 need?

In our test it used about 14.2GB when loaded. For fast GPU inference, plan on 16GB of VRAM or more; without a GPU, 32GB of system RAM or more.

Can gemma4:12b-it-q8_0 run without a GPU?

Yes. On CPU only it generated about 4.6 tokens/sec (slow).

Does gemma4:12b-it-q8_0 run on an 8GB VRAM GPU?

Yes, but it does not fit entirely, so part of it runs on the CPU: about 8.2 tokens/sec with 8GB VRAM versus about 43.5 tokens/sec with 16GB.

Measured on the site owner's PC (Ryzen 7 9800X3D / 31GB RAM / RTX 4080 (16GB VRAM)). "8GB / 12GB VRAM" values were measured on the RTX 4080 with the usable VRAM limited to that amount; real 8GB or 12GB cards are slower GPUs, so actual speeds may be lower. How we measure