Local LLMs that run on 16GB VRAM (benchmarks)
Which local LLMs run on a GPU with 16GB of VRAM, and how fast, based on real measurements.
Models that fit in VRAM (fast)
- gemma4:12b-it-q8_014.2GB · about 43.5 tokens/secInstant
- qwen3.5:9b6.0GB · about 96.4 tokens/secInstant
- qwen3.5:4b3.1GB · about 132.1 tokens/secInstant
- gemma3:4b2.8GB · about 142.0 tokens/secInstant
- qwen3.5:2b2.3GB · about 170.4 tokens/secInstant
- llama3.2:3b2.3GB · about 189.7 tokens/secInstant
- gemma3:1b0.9GB · about 216.9 tokens/secInstant
Models that run partly on the CPU (slower)
None of the models we measured fit here yet.
FAQ
Which local LLMs run on a 16GB VRAM GPU?
In our tests, 7 models fit entirely in 16GB of VRAM. The largest is gemma4:12b-it-q8_0 (about 14.2GB, about 43 tokens/sec).
Other sizes: 8GB VRAM / 12GB VRAM / No GPU (by RAM)
Measured on the site owner's PC (Ryzen 7 9800X3D / 31GB RAM / RTX 4080 (16GB VRAM)). "8GB / 12GB VRAM" values were measured on the RTX 4080 with the usable VRAM limited to that amount; real 8GB or 12GB cards are slower GPUs, so actual speeds may be lower. How we measure