A Japanese office worker puts AI to work, and measures everything.
I'm a shachiku (an overworked Japanese office worker) who runs AI on the side. This site measures how fast local AI really runs on a home PC (Ryzen 7 9800X3D / 31GB RAM / RTX 4080 (16GB VRAM)), and documents experiments like building 3D characters with AI.
Articles
Creating a 3D Character from an Image with AI | Failures and Solutions in Attempting a VRChat Avatar with Hunyuan3D-2.1
A record of the challenge to turn a design sheet drawn with image generation AI into a 3D model using Hunyuan3D-2.1 to aim for a VRChat avatar. I summarize why side profiles break, 6 common 'AI tropes' failures in 3D conversion, the results of automatic rigging (SkinTokens), and my next strategy.
An AI Coach Teaches You in Real-Time! The Mechanism of the Browser-Based "AI Mahjong Dojo"
A blog post that explains, in a way easy for beginners to understand, the introduction of the "AI Mahjong Dojo" posted on Threads. We will take a detailed look at the mechanism where an AI provides real-time corrections for your discards by utilizing local LLMs.
ComfyUI v0.39.0 Update Explanation: Support for Latest Models and Key Functional Improvements
A major update has arrived for the standard image generation AI tool 'ComfyUI,' including support for major new models. We will explain in detail the partner nodes added in this update, as well as improvements regarding video production and high-quality exports.
Facial Expression Instructions in Image Generation AI: How a Single Word Changes a Character's Impression
I verified how words related to "expressions" added to prompts (instructions for the AI) affect the details of the generated images.
[Measurement Report] Speed and VRAM Consumption Benchmarks when Running Image Generation AI (Stable Diffusion) on Your Own PC
For those who want to run image generation AI in a local environment, this article provides a detailed explanation of specific GPU specs and performance verification results for each model.
On the Phenomenon of "Fake URLs" Generated by Local AI and Its Background
An analysis of model-specific behaviors (hallucinations) and security considerations that can occur when running AI in a local environment.
More Accurate AI Performance on Mac: Explanation of Tokenizer Improvements in the Latest Ollama Development Build
For those using Ollama on Macs with Apple silicon, we provide a detailed explanation of the important fixes included in the latest development build (v0.40.0-rc1).
Does Limiting CPU Threads Affect Speed? The Identity of Local LLM Bottlenecks
When running generative AI (local LLMs) on your own PC, it is common to assume that the more CPU cores and threads you have, the faster it will be. However, experimental results showed that even when the number of threads was halved, there was almost no change in speed. We explore why this phenomenon occurs.
The Role of the "Director" in Moving a Story: The Mechanism for Realizing Natural Conversations with Local AI
I will explain in detail the mechanism behind the scene progression by AI, which is used in 'Narratable,' a tool actually used by the operator of this account.
[Full Local] I Built 'Narratable' to Weave Stories with AI Characters
Introducing a case study on constructing a multi-person roleplay chat that completes entirely within your own PC without depending on an internet connection.
The First Step to Starting Local AI Most Easily: How to Utilize Ollama
An explanation of Ollama, the simplest entry point for entering the world of "Local AI" where you run AI directly on your own PC.
The Biggest Advantages of Running AI on Your Own PC are "Peace of Mind" and "Freedom"
Explaining the major benefits of implementing local AI, such as the protection of confidential information, and specific use cases.
On the Trade-off Between 'Model Size' and 'Execution Environment' in Local AI
Reflections on the balance between performance, speed, and power consumption when running large-scale AI models at home.
Points to Note When Loading Long Documents into AI: The Relationship Between Context Length and Memory
Explaining the mechanism of 'context length,' which is important when running AI in a local environment, and its relationship with PC memory consumption.
The Key to Running Local AI Lies in "VRAM" Capacity! Explaining the Relationship Between GPU Memory and Speed
For those who want to run AI comfortably in a local environment, one of the most important specs is "VRAM (Video Memory)." We will explain in detail why VRAM is important, including the dramatic differences in speed when actually running models.
Benchmarks
See all →| Model | Memory used | With GPU | Feel | CPU only | Feel | Measured |
|---|---|---|---|---|---|---|
| gemma3:1b | 0.9 GB | 216.9 | Instant | 61.7 | Instant | 2026-10-04 |
| llama3.2:3b | 2.3 GB | 189.7 | Instant | 27.3 | Comfortable | 2026-10-04 |
| qwen3.5:2b | 2.3 GB | 170.4 | Instant | 27.7 | Comfortable | 2026-10-04 |
| gemma3:4b | 2.8 GB | 142.0 | Instant | 21.5 | Comfortable | 2026-10-04 |
| qwen3.5:4b | 3.1 GB | 132.1 | Instant | 19.3 | Comfortable | 2026-10-04 |
Numbers are generation speed in tokens per second.
By VRAM: 8GB / 12GB / 16GB · By RAM (no GPU): 8GB / 16GB / 32GB
About this site
This is the English edition of a Japanese blog, 「AIに働かせたい社畜の実験室」. Posts, articles and benchmarks are produced by an automated pipeline: news collection, drafting with AI, fact checks against sources, and publishing. English articles are translated from the Japanese originals by a local AI, and every number is automatically checked against the original.
The site shows ads (Google AdSense). Questions and corrections are welcome via the contact form (in Japanese, but English messages are fine).