qwen3.5:4b: memory needed and real speed (benchmark)
qwen3.5:4b measured on a home PC: about 3.1GB of memory, about 132 tokens/sec on GPU and 19.3 tokens/sec on CPU only. Which PCs can run it, by VRAM and RAM.
Summary
- Memory used: about 3.1GB (measured 2026-10-04)
- GPU (16GB VRAM): about 132.1 tokens/sec Instant
- CPU only: about 19.3 tokens/sec Comfortable
Speed by setup
| Setup | Tokens/sec | Feel |
|---|---|---|
| GPU, 16GB VRAM (RTX 4080) | 132.1 | Instant |
| 12GB VRAM (simulated) | 132.1 | Instant |
| 8GB VRAM (simulated) | 132.1 | Instant |
| CPU only (no GPU) | 19.3 | Comfortable |
Which PCs can run it?
- 8GB VRAM GPUFits in VRAM · about 132.1 tokens/secInstant
- 12GB VRAM GPUFits in VRAM · about 132.1 tokens/secInstant
- 16GB VRAM GPUFits in VRAM · about 132.1 tokens/secInstant
- 8GB RAM, no GPURuns on CPU · about 19.3 tokens/secComfortable
- 16GB RAM, no GPURuns on CPU · about 19.3 tokens/secComfortable
- 32GB RAM, no GPURuns on CPU · about 19.3 tokens/secComfortable
We assume about 4GB of RAM for the OS and 1GB of VRAM headroom.
Compared with other models
FAQ
How much memory does qwen3.5:4b need?
In our test it used about 3.1GB when loaded. For fast GPU inference, plan on 8GB of VRAM or more; without a GPU, 8GB of system RAM or more.
Can qwen3.5:4b run without a GPU?
Yes. On CPU only it generated about 19.3 tokens/sec (comfortable).
Measured on the site owner's PC (Ryzen 7 9800X3D / 31GB RAM / RTX 4080 (16GB VRAM)). "8GB / 12GB VRAM" values were measured on the RTX 4080 with the usable VRAM limited to that amount; real 8GB or 12GB cards are slower GPUs, so actual speeds may be lower. How we measure