2026-10-05Translated from the Japanese original

On the Phenomenon of "Fake URLs" Generated by Local AI and Its Background

Even if you think it is safe because you are running AI on your own PC, AI models can sometimes exhibit "strange behaviors." This time, I will explain in detail the phenomenon of generating URLs mimicking access to specific cloud services, which was discovered during experiments with local AI.

Key points
  • There is a phenomenon where models fabricate "plausible fake URLs" based on their training data
  • Behaviors related to cloud services of specific companies (such as Alibaba) may sometimes be observed
  • While this is highly likely to be a hallucination (a plausible lie), caution is still necessary
  • When using external connection functions such as web tool integration, it is recommended to carefully verify the output content

Incidents Reported in Experiments: Generation of Unfamiliar URLs

Cases have been reported where, while using AI models (such as the Qwen series) in a local environment, the AI generates URLs containing requests to unintended external domains during specific tasks.

As a specific example, when conducting product research related to Amazon, the AI outputted a long URL involving access to a specific cloud storage service (Aliyun: Alibaba's cloud service). This behavior had no direct connection to the user's instructions or the current session.

  • Target domain: aliyuncs.com (a domain for Alibaba's cloud services)
  • Occurrence: Generated while conducting research via web tools
  • Characteristics: At first glance, it took the form of a "signed URL" providing access permissions to a specific storage bucket

Why Does This Behavior Occur?

Several possibilities are being considered regarding the cause of this phenomenon. Based on official information and technical backgrounds, the most likely cause is attributed to "biases in training data."

AI models learn vast amounts of code and system logs. It is possible that many descriptions (coding traces) utilizing Alibaba's cloud services are included therein. Therefore, it is thought that in a situation where the AI is "operating web tools," it fabricated a "plausible URL" derived from its training data.

This phenomenon can be categorized as a type of "hallucination (Hallucination: a plausible lie)" in technical terms. It is a phenomenon where AI generates information not based on facts as if it were correct.

  • Hypothesis 1: Influence of training data regarding Alibaba's cloud services
  • Hypothesis 2: The result of the AI inferring an "appropriate URL" from contexts of coding or system operations
  • Commonality: It is highly likely that it was output as a statistical pattern within the model, independent of the user's intent

Security Precautions and Countermeasures

While this behavior is highly likely to be a hallucination, it can sometimes appear like an attempt to send data externally (data leakage). Therefore, users are encouraged to pay attention to the following points.

Particularly when permitting functions such as "web tool integration" and entrusting external APIs or browser operations to the AI, it is important to monitor the content of the requests generated by the AI. The key point is to make it a habit to check that data containing confidential information is not sent to URLs created arbitrarily by the model.

From this account's perspective, running in a local environment itself is considered a very secure method. However, it is important to understand that every model has its own unique "quirks" and maintain an attitude of paying attention to unnatural behaviors.

  • Precaution: Closely watch the AI's output when using functions that connect to the outside, such as web tool integration
  • Countermeasure: If there is a request to access an abnormal URL or an unfamiliar domain, interrupt the session immediately
  • Mindset: Understand the model's "quirks" and manage them appropriately rather than harboring excessive anxiety
Useful for
  • Those who feel anxious about the security of local AI
  • Those using open models such as Qwen
  • Those trying out AI-based web tool operations
Glossary
Hallucination (ハルシネーション)
A phenomenon where AI generates information not based on facts in a plausible manner.
Signed URL (署名付きURL)
A temporary URL issued to allow access to a specific file within a certain time or set of conditions.
Coding Trace (コーディングトレース)
The history or sequence of events when writing programs. AI learns to write code by learning these.

FAQ

If this phenomenon occurs, is data stolen immediately?

Based on official information and previous reports, it is considered to be caused by hallucination (fabrication) in most cases. However, basic precautions such as stopping operations when suspicious behavior is detected are necessary.

Why do URLs from specific companies appear?

It is highly likely that the cause is that data or code related to that company were heavily included in the AI's training set.

Summary

While execution in a local environment is safe, caution is required regarding model-specific "fabrications (hallucinations)." Particularly when integrating with external tools, you can utilize AI more safely by paying attention to the URLs and request content output by the AI.

This article was translated from Japanese by AI; numbers and model names were automatically checked against the original. The original was written with AI from the sources cited there.