llama3.2:3b: memory needed and real speed (benchmark)

llama3.2:3b measured on a home PC: about 2.3GB of memory, about 190 tokens/sec on GPU and 27.3 tokens/sec on CPU only. Which PCs can run it, by VRAM and RAM.

Summary
  • Memory used: about 2.3GB (measured 2026-10-04)
  • GPU (16GB VRAM): about 189.7 tokens/sec Instant
  • CPU only: about 27.3 tokens/sec Comfortable

Speed by setup

SetupTokens/secFeel
GPU, 16GB VRAM (RTX 4080)189.7Instant
12GB VRAM (simulated)189.7Instant
8GB VRAM (simulated)189.7Instant
CPU only (no GPU)27.3Comfortable

Which PCs can run it?

We assume about 4GB of RAM for the OS and 1GB of VRAM headroom.

Compared with other models

Tokens/secGPU (RTX 4080)CPU only★ llama3.2:3b2.3GB189.727.3gemma3:4b2.8GB142.021.5qwen3.5:9b6.0GB96.411.2
Measured on Ryzen 7 9800X3D / 31GB RAM / RTX 4080 (16GB VRAM).

FAQ

How much memory does llama3.2:3b need?

In our test it used about 2.3GB when loaded. For fast GPU inference, plan on 8GB of VRAM or more; without a GPU, 8GB of system RAM or more.

Can llama3.2:3b run without a GPU?

Yes. On CPU only it generated about 27.3 tokens/sec (comfortable).

Measured on the site owner's PC (Ryzen 7 9800X3D / 31GB RAM / RTX 4080 (16GB VRAM)). "8GB / 12GB VRAM" values were measured on the RTX 4080 with the usable VRAM limited to that amount; real 8GB or 12GB cards are slower GPUs, so actual speeds may be lower. How we measure