Local LLMs that run on a 8GB RAM PC without a GPU (benchmarks)

Local LLMs that run on a PC with 8GB of RAM and no GPU, with measured speeds. We assume the OS and browser use about 4GB.

Models that run

  • qwen3.5:4b3.1GB · about 19.3 tokens/secComfortable
  • gemma3:4b2.8GB · about 21.5 tokens/secComfortable
  • qwen3.5:2b2.3GB · about 27.7 tokens/secComfortable
  • llama3.2:3b2.3GB · about 27.3 tokens/secComfortable
  • gemma3:1b0.9GB · about 61.7 tokens/secInstant

Not enough memory

FAQ

Which local LLMs run on a 8GB RAM PC without a GPU?

In our tests, 5 models ran, and 5 of them were comfortable for chat (15 tokens/sec or more). The largest is qwen3.5:4b (about 19.3 tokens/sec).

Can I use local AI without a GPU?

Yes. Smaller models are fast enough for chat on a CPU alone. Larger models get slower, so check whether a model fits in your GPU's VRAM if you want speed.

Other sizes: 16GB RAM / 32GB RAM / With a GPU (by VRAM)

Measured on the site owner's PC (Ryzen 7 9800X3D / 31GB RAM / RTX 4080 (16GB VRAM)). "8GB / 12GB VRAM" values were measured on the RTX 4080 with the usable VRAM limited to that amount; real 8GB or 12GB cards are slower GPUs, so actual speeds may be lower. How we measure