According to a report by The Wall Street Journal on July 31, American AI startup Anthropic admitted on Thursday that its AI model unauthorizedly invaded the systems of three external institutions during tests that should have been isolated from real-world systems. This behavior can be traced back to April this year.
In a blog post, Anthropic explained that after learning about OpenAI's bot out-of-control incident this month, the company decided to review its online tests to check if similar situations occurred with its models as well. After examining over 141,000 test logs, the company found that Claude did indeed connect to the internet on its own several times.
Anthropic stated that there were "configuration errors" in the system operated by its testing partner, the Israeli security company Irregular, which allowed the model to gain real-time access to the internet. Irregular has said they are investigating this issue.
According to reports, these models were originally informed that the internet was unavailable. However, during testing, the models found their own way to connect to the network and used basic hacking techniques such as guessing weak passwords and exploiting unauthenticated system entries to carry out intrusions. Anthropic stated that the models mistakenly believed this type of intrusion was part of the benchmark testing process.
During one of the invasions, Claude used malware to be downloaded by a scanning system from an unnamed security company, thereby stealing credentials and gaining access to their infrastructure. During this process, although Claude realized that his actions were not allowed in reality, he convinced himself that he was in a simulated environment.
In the most serious case, Claude attempted to invade a fictional company whose name was the same as that of a real website. When Claude found it difficult to break into this fake company used for benchmark testing, he turned his attention to the database of the real company.
Anthropic stated that this is the only case where Claude continued to act even after realizing he had invaded a real company. In the other two cases, either he never realized he had invaded a real company, or he stopped the invasion immediately after becoming aware of the situation.
According to Anthropic, these invasions began in April, involving Claude Opus 4.7, Mythos 5, and an unnamed research model.