Local LLMs that run on a 16GB RAM PC without a GPU (benchmarks)
Local LLMs that run on a PC with 16GB of RAM and no GPU, with measured speeds. We assume the OS and browser use about 4GB.
Models that run
- qwen3.5:9b6.0GB · about 11.2 tokens/secShort wait
- qwen3.5:4b3.1GB · about 19.3 tokens/secComfortable
- gemma3:4b2.8GB · about 21.5 tokens/secComfortable
- qwen3.5:2b2.3GB · about 27.7 tokens/secComfortable
- llama3.2:3b2.3GB · about 27.3 tokens/secComfortable
- gemma3:1b0.9GB · about 61.7 tokens/secInstant
Not enough memory
- gemma4:12b-it-q8_0 (about 14.2GB)
FAQ
Which local LLMs run on a 16GB RAM PC without a GPU?
In our tests, 6 models ran, and 5 of them were comfortable for chat (15 tokens/sec or more). The largest is qwen3.5:9b (about 11.2 tokens/sec).
Can I use local AI without a GPU?
Yes. Smaller models are fast enough for chat on a CPU alone. Larger models get slower, so check whether a model fits in your GPU's VRAM if you want speed.
Other sizes: 8GB RAM / 32GB RAM / With a GPU (by VRAM)
Measured on the site owner's PC (Ryzen 7 9800X3D / 31GB RAM / RTX 4080 (16GB VRAM)). "8GB / 12GB VRAM" values were measured on the RTX 4080 with the usable VRAM limited to that amount; real 8GB or 12GB cards are slower GPUs, so actual speeds may be lower. How we measure