Local LLMs that run on 8GB VRAM (benchmarks)

Which local LLMs run on a GPU with 8GB of VRAM, and how fast, based on real measurements.

Models that fit in VRAM (fast)

Models that run partly on the CPU (slower)

FAQ

Which local LLMs run on a 8GB VRAM GPU?

In our tests, 6 models fit entirely in 8GB of VRAM. The largest is qwen3.5:9b (about 6.0GB, about 96 tokens/sec).

What happens if a model does not fit in 8GB of VRAM?

It still runs, but the overflow is processed on the CPU, so it slows down. For example, gemma4:12b-it-q8_0 (about 14.2GB) ran at about 43.5 tokens/sec with 16GB VRAM, about 8.2 with 8GB, and about 4.6 on CPU only.

Other sizes: 12GB VRAM / 16GB VRAM / No GPU (by RAM)

Measured on the site owner's PC (Ryzen 7 9800X3D / 31GB RAM / RTX 4080 (16GB VRAM)). "8GB / 12GB VRAM" values were measured on the RTX 4080 with the usable VRAM limited to that amount; real 8GB or 12GB cards are slower GPUs, so actual speeds may be lower. How we measure