Local LLM benchmarks: speed and memory on a home PC

Every model below was downloaded and measured on the same home PC (Ryzen 7 9800X3D / 31GB RAM / RTX 4080 (16GB VRAM)), with the GPU and on CPU only.

Will it run on your PC?

Pick your RAM and GPU to see which of the models we measured you can run, and how fast.

    Speeds were measured on Ryzen 7 9800X3D / 31GB RAM / RTX 4080 (16GB VRAM); a different GPU or CPU will be faster or slower. We assume about 4GB for the OS and 1GB of VRAM headroom.

    All results

    ModelMemory usedWith GPUFeelCPU onlyFeelMeasured
    gemma3:1b0.9 GB216.9Instant61.7Instant2026-10-04
    llama3.2:3b2.3 GB189.7Instant27.3Comfortable2026-10-04
    qwen3.5:2b2.3 GB170.4Instant27.7Comfortable2026-10-04
    gemma3:4b2.8 GB142.0Instant21.5Comfortable2026-10-04
    qwen3.5:4b3.1 GB132.1Instant19.3Comfortable2026-10-04
    qwen3.5:9b6.0 GB96.4Instant11.2Short wait2026-10-04
    gemma4:12b-it-q8_014.2 GB43.5Instant4.6Slow2026-10-04

    Numbers are generation speed in tokens per second.

    How we measure

    • Each model answers the same question through Ollama; we record generation speed (tokens/sec) and memory used.
    • "With GPU" uses the RTX 4080 (16GB VRAM). "8GB / 12GB VRAM" values limit how many layers go on the GPU so that only that much VRAM is used.
    • Feel: 30+ tokens/sec is instant, 15–30 comfortable, 5–15 a short wait, under 5 slow.

    By VRAM: 8GB / 12GB / 16GB · By RAM (no GPU): 8GB / 16GB / 32GB