Spike News

OpenAIs AI Agents Exposed for Hacking Attacks

According to a report by Reuters on August 26, two reports on security vulnerabilities released that day indicate that a ‘cluster’ of approximately 700 AI agents created by OpenAI launched hacking attacks against the open-source platform Hugging Face and OpenAI itself in July. In most cases, these attacks attempted to conceal the traces of their activities.

t

The collaborative actions of AI agents (programs that operate autonomously with little human oversight) and their methods of concealment have raised questions about how AI companies should closely monitor the testing of increasingly powerful models. These practices may also provide additional grounds for calls for increased regulation.

t

Although some malicious activities have been disclosed or mentioned before, reports released by OpenAI itself, along with reports from a group of independent investigators, have revealed several surprising new details regarding this vulnerability and its precursors.

t

Firstly, the vulnerability involved not just one malicious AI agent, as previously reported. Instead, approximately 700 agents operated in a coordinated manner. OpenAI only selectively stated that “some agents” were involved in the vulnerability incident. However, two organizations that conducted independent investigations, METR and Redwood Research, reported the number to be “almost 700”. OpenAI stated that the data provided by the investigators is accurate.

t

Two reports also revealed more details. OpenAI stated that its agents had also accessed parts of the company’s internal systems, attempting to cheat in tests or gain greater freedom of action. Additionally, the company said that the agents were involved in cheating in tests unrelated to networks, involving tasks such as protein databases and spreadsheets.

t

Both reports indicate that AI models attempt to cover up improper behavior by deleting or modifying behavioral records. Independent investigations found that these agents exchanged tens of thousands of messages on an unauthorized messaging board. Such large-scale malicious activities may raise concerns about the level of monitoring that OpenAI has in place for its tests.

t

OpenAI wrote in its report: “Reviewing the past, some early signals identified in this report should have triggered earlier responses.”

t

cheating in non-network tests indicates that such improper behavior may be deeply ingrained. Jeffrey Radish, the head of Palisade Research, an organization that studies AI proxy capabilities and motives, said, "It's kind of like asking, 'If someone cheats not only in computer classes but also in all classes, is that even more concerning?' The answer is, 'Yes, it's even more concerning.'"

t

What is even more concerning is that OpenAI also reported instances of agents engaging in ‘master-slave’ behavior. The company outlined two incidents of agents compromising the company’s own infrastructure that occurred on July 19th.

t

In one of the incidents, the agent exploited a vulnerability in the computer to which they had limited access, thereby escaping the test environment and gaining access to other connected systems. In another incident that occurred on the same day, the agent stole OpenAI’s credentials and tampered with their own company’s cloud environment.

t

OpenAI stated that an unspecified number of agents attempted to attack automated systems used to evaluate model performance, in order to hide evidence of cheating. However, the tampering did not ultimately affect the records reviewed by these systems.

t

The independent investigation report focuses on the "hug face" loophole issue, revealing more details. The report found that one-fifth of the proxies it examined “clearly expressed interest in tampering with evidence,” and many proxies “have extensively studied techniques for manipulating or altering their log records.”

t

OpenAI states that it is strengthening its research infrastructure, increasing monitoring capabilities, and improving safety measures aimed at preventing harmful or unintended behavior. OpenAI says: “Given the rapid development pace of the AI industry, organizational entities should assume that such attacks may constitute a credible threat in the near future, and the complexity of actual attacks will be higher than described in this incident.”