Local LLMs that run on 8GB VRAM (benchmarks)
Which local LLMs run on a GPU with 8GB of VRAM, and how fast, based on real measurements.
Models that fit in VRAM (fast)
- qwen3.5:9b6.0GB · about 96.4 tokens/secInstant
- qwen3.5:4b3.1GB · about 132.1 tokens/secInstant
- gemma3:4b2.8GB · about 142.0 tokens/secInstant
- qwen3.5:2b2.3GB · about 170.4 tokens/secInstant
- llama3.2:3b2.3GB · about 189.7 tokens/secInstant
- gemma3:1b0.9GB · about 216.9 tokens/secInstant
Models that run partly on the CPU (slower)
- gemma4:12b-it-q8_014.2GB · about 8.2 tokens/secShort wait
FAQ
Which local LLMs run on a 8GB VRAM GPU?
In our tests, 6 models fit entirely in 8GB of VRAM. The largest is qwen3.5:9b (about 6.0GB, about 96 tokens/sec).
What happens if a model does not fit in 8GB of VRAM?
It still runs, but the overflow is processed on the CPU, so it slows down. For example, gemma4:12b-it-q8_0 (about 14.2GB) ran at about 43.5 tokens/sec with 16GB VRAM, about 8.2 with 8GB, and about 4.6 on CPU only.
Other sizes: 12GB VRAM / 16GB VRAM / No GPU (by RAM)
Measured on the site owner's PC (Ryzen 7 9800X3D / 31GB RAM / RTX 4080 (16GB VRAM)). "8GB / 12GB VRAM" values were measured on the RTX 4080 with the usable VRAM limited to that amount; real 8GB or 12GB cards are slower GPUs, so actual speeds may be lower. How we measure