Local LLMs that run on 12GB VRAM (benchmarks)

Which local LLMs run on a GPU with 12GB of VRAM, and how fast, based on real measurements.

Models that fit in VRAM (fast)

Models that run partly on the CPU (slower)

FAQ

Which local LLMs run on a 12GB VRAM GPU?

In our tests, 6 models fit entirely in 12GB of VRAM. The largest is qwen3.5:9b (about 6.0GB, about 96 tokens/sec).

What happens if a model does not fit in 12GB of VRAM?

It still runs, but the overflow is processed on the CPU, so it slows down. For example, gemma4:12b-it-q8_0 (about 14.2GB) ran at about 43.5 tokens/sec with 16GB VRAM, about 17.2 with 12GB, and about 4.6 on CPU only.

Other sizes: 8GB VRAM / 16GB VRAM / No GPU (by RAM)

Measured on the site owner's PC (Ryzen 7 9800X3D / 31GB RAM / RTX 4080 (16GB VRAM)). "8GB / 12GB VRAM" values were measured on the RTX 4080 with the usable VRAM limited to that amount; real 8GB or 12GB cards are slower GPUs, so actual speeds may be lower. How we measure