Local LLMs that run on a 32GB RAM PC without a GPU (benchmarks)

Local LLMs that run on a PC with 32GB of RAM and no GPU, with measured speeds. We assume the OS and browser use about 4GB.

Models that run

FAQ

Which local LLMs run on a 32GB RAM PC without a GPU?

In our tests, 7 models ran, and 5 of them were comfortable for chat (15 tokens/sec or more). The largest is gemma4:12b-it-q8_0 (about 4.6 tokens/sec).

Can I use local AI without a GPU?

Yes. Smaller models are fast enough for chat on a CPU alone. Larger models get slower, so check whether a model fits in your GPU's VRAM if you want speed.

Other sizes: 8GB RAM / 16GB RAM / With a GPU (by VRAM)

Measured on the site owner's PC (Ryzen 7 9800X3D / 31GB RAM / RTX 4080 (16GB VRAM)). "8GB / 12GB VRAM" values were measured on the RTX 4080 with the usable VRAM limited to that amount; real 8GB or 12GB cards are slower GPUs, so actual speeds may be lower. How we measure