2026-10-05Translated from the Japanese original

[Measurement Report] Speed and VRAM Consumption Benchmarks when Running Image Generation AI (Stable Diffusion) on Your Own PC

The worry "I want to run image generation AI on my own PC, but I don't know what kind of specs are required" is something often heard when starting with local AI. This time, using the experimental environment of this account, I will explain in detail the results of measured generation speeds and VRAM (Video Memory) usage for two major model systems.

Key points
  • Standard generation sizes differ between SDXL-based and SD1.5-based systems
  • GPU VRAM capacity is a crucial factor that determines the size of the model you want to run
  • Depending on the tool used (Forge), it may reserve more memory than the actual required amount

Experimental Environment for This Test

To grasp accurate figures, verification was conducted in the following environment.

CPU: AMD Ryzen 7 9800X3D 8-Core Processor

Memory: 31GB

GPU: NVIDIA GeForce RTX 4080 (VRAM 16376 MiB)

Tool: Stable Diffusion WebUI Forge (*One of the interfaces for running image generation AI)

Settings: Sampler Euler a, CFG 5, 25 steps (*The number of calculations to complete an image)

  • novaAnimeXL (SDXL-based): Approx. 5.8 seconds (4.3 steps/sec) at 1024x1024 size
  • novaAnimeXL (SDXL-based): Approx. 5.8 seconds (4.3 steps/sec) at 832x1216 size
  • Realistic Vision 5.1 (SD1.5-based): Approx. 1.7 seconds (15.1 steps/sec) at 512x512 size
  • Realistic Vision 5.1 (SD1.5-based): Approx. 2.1 seconds (11.7 steps/sec) at 512x768 size

Differences in "Standard Size" by Model

Image generation AI broadly exists in models of different generations. In this experiment, we compared two representative types: the "SDXL-based" and "SD1.5-based" systems.

First, an important point is that the recommended standard output size differs for each model. According to official information, SD1.5-based systems are considered standard at around 512px, while SDXL-based systems are standard at around 1024px.

This difference in size significantly impacts generation speed and VRAM usage.

  • SD1.5-based: Lightweight and fast, making it easy to handle even in relatively low-spec environments
  • SDXL-based: High quality, but requires larger computational resources

VRAM (Video Memory) Consumption and Points of Caution

The amount of "VRAM" used by the GPU during generation is the most important metric for running it comfortably. The usage recorded in this experiment is as follows.

For SDXL-based models, about 10.3GB to 10.4GB of VRAM was consumed. On the other hand, consumption for SD1.5-based models remained at about 7.9GB to 8.3GB.

There is one point of caution here. The tool "Forge" used in this measurement has a property of pre-allocating more memory if there is available capacity in the system.

Therefore, there is a good possibility that models can run even on GPUs equipped with less VRAM than the measured figures.

  • GPU with 8GB VRAM: There is a high possibility of running SD1.5-based systems realistically
  • GPU with 16GB VRAM: In this experimental environment (16GB), consumption was just over about 10GB even for SDXL-based systems. Looking at this figure, it seems that if you have more than this capacity, it falls within the range where it can be run with relative ease.

Summary

Through this experiment, it became clear that the required specs are distinctly divided depending on the model you want to run. If you mainly enjoy SD1.5-based systems, there is room to try even with relatively modest VRAM capacity. If you want to operate high-definition SDXL-based systems smoothly, the key point is to secure a generous amount of VRAM capacity.

If you are unsure which model will run on your PC environment, start by checking your GPU's VRAM capacity first.

  • Choose a model according to your goals (speed-oriented or high-quality oriented)
  • One way to start is by trying low-load SD1.5-based systems first
  • Keep in mind that memory allocation amounts fluctuate depending on the characteristics of the tool
Useful for
  • People who don't know if they can run it on their own PC
  • People who want to run AI locally without relying on paid services like ChatGPT
  • People who want to start with AI illustrations
Glossary
VRAM (Video Memory)
Memory dedicated to image processing installed on a GPU (graphics board).
Stable Diffusion WebUI Forge
One of the software packages for operating the image generation AI "Stable Diffusion" via a browser.
SDXL-based / SD1.5-based
Refers to the generations (types) of image generation AI models. SDXL is a newer, high-definition model, while SD1.5 is an older, fast, and highly versatile model.

FAQ

What should I do if I am told that VRAM is insufficient?

According to official information, memory consumption fluctuates depending on the tool and settings used. Since there was a tendency for more to be allocated due to the characteristics of the Forge tool in this experiment as well, it is possible to run it if you have slightly more margin than the measured values.

Summary

A report verifying the speed and VRAM consumption by model system when running image generation AI in a local environment. The measured data showed that SD1.5-based systems are fast and memory-efficient, while SDXL-based systems offer high quality but require more resources.

This article was translated from Japanese by AI; numbers and model names were automatically checked against the original. The original was written with AI from the sources cited there.