2026-10-05Translated from the Japanese original

Does Limiting CPU Threads Affect Speed? The Identity of Local LLM Bottlenecks

CPU performance is important when running AI on your own PC. However, I discovered that the common belief that "the more cores and threads you have, the faster the generation speed" does not always apply. In this post, I will explain in detail the "relationship between CPU thread limits and generation speed" that I actually verified on this account.

Key points
  • Even when limiting the number of threads by half, almost no change was seen in generation speed
  • The cause is highly likely to be the "speed of reading data from memory" rather than computing power
  • It is important to understand that there are situations where the CPU's power cannot be fully utilized

Experimental Results: Speed Remains Unchanged Even When Cutting Threads

On this account, I measured the impact of limiting the "CPU number of threads (something like a path for calculations performed simultaneously)" when running AI models in a local environment.

To get straight to the conclusion, the results showed that even when significantly reducing the number of threads, the generation speed remained almost the same. Let's look at the actual measurement data below to see exactly how much difference there was.

  • For gemma4:12b model: - When using 8 CPU cores: 4.6 tokens/sec (a speed where wait time is noticeable) - When limiting to 4 CPU threads: 4.6 tokens/sec (a speed where wait time is noticeable)
  • For qwen3.5:9b model: - When using 8 CPU cores: 11.2 tokens/sec (a usable speed with a short wait) - When limiting to 4 CPU threads: 11.0 tokens/sec (a usable speed with a short wait)

Why Is Computing Power Not Being Fully Used?

It is natural to think, "If I have a high-performance CPU, it should be able to process faster using more threads." However, it is believed that other factors strongly influence the operation of current AI (Large Language Models).

In official information and technical backgrounds, it is stated that the process of generating text tends to be influenced more by the "speed of reading data from memory" than by the "speed of calculation" itself.

When running an AI model, it is necessary to take out and calculate a vast amount of data (parameters) from memory one after another. Because this "data movement" becomes the bottleneck (processing delay) that determines the overall movement, even if you only boost the CPU's computing power, it is difficult to achieve higher speeds.

  • Memory bandwidth (the width of the path for exchanging data) becomes more important than computing performance
  • Therefore, even if you intentionally limit the number of threads, the impact on performance may remain minimal
  • It is important not to have excessive expectations for the number of CPU cores and to be conscious of the memory environment as well

Summary: Perspectives When Running Local AI

What we learned from this experiment is that to run AI comfortably on your own PC, it is important to be conscious of the overall flow of data rather than just pursuing the number of CPU threads.

If you feel "the generation speed is slow on my PC," it would be good to consider the possibility that the cause is the data transfer speed from memory rather than a lack of CPU cores.

In the world of local AI, understanding these hardware characteristics one by one through actual measurements is the shortcut to building an optimal environment.

  • Be aware that there are behaviors that do not rely too much on the number of threads
  • Check memory-related performance as well
  • Measure in your own environment "which model produces what level of speed"
Useful for
  • People who want to run AI on their own PC but are unsure which specs are important
  • People who want to know specifically about the differences in CPU core counts and thread counts
  • People seeking accurate information based on actual measurement data in a local environment
Glossary
Thread
The unit of calculation that a CPU can process simultaneously.
Token
The minimum unit when AI handles text. In Japanese, it corresponds to roughly 1–2 characters.
Bottleneck
The slowest part that limits the speed of the entire process.

FAQ

Why is it okay to reduce the number of threads?

Because in AI processing, the speed of reading data from memory often becomes more important than the calculation itself. Therefore, even if you limit the threads for calculation, there may be no significant difference in speed.

In the end, is choosing a good CPU meaningless?

No, it is not meaningless. However, the point of this post is that it is important to consider memory speed and bandwidth along with core counts and thread counts, rather than focusing solely on those.

Summary

From the actual measurement results showing no change in generation speed even when limiting CPU threads, I explained how local AI operation may depend more on "reading data from memory" than on "calculation." This provides a new perspective for hardware selection.

This article was translated from Japanese by AI; numbers and model names were automatically checked against the original. The original was written with AI from the sources cited there.