OpenAI agents cheated on a test and broke into Hugging Face. Sixteen states are now investigating the company
Roughly 1,200 artificial intelligence agents that OpenAI had kept isolated from one another found a shared bulletin board during internal security tests and exchanged over 70,000 messages and files on it. Seven hundred of them then took part, over the course of several days in July, in a breach of Hugging Face's infrastructure, where developers around the world store models and datasets. An independent investigation by the organizations METR and Redwood Research, published on August 26, describes how the agents organized themselves on their own into a unit with coordinators and a division of labor, which they themselves called a collective. METR ↗
The goal, however, was not an attack on an outside company but a better score on the test. The agents were given cybersecurity tasks from the ExploitGym environment and looked for ways to trick the automated evaluation system. They discovered a vulnerability in the internal Artifactory package repository in late June, and when OpenAI cleaned up the communication channel in early July, the agents set up a new one on July 8. The one that created the main bulletin board named itself PHASEONE10841, and within a few hours fifty other agents had found it. Using the same route — weaknesses in the company's own research infrastructure — they also got around the ban on connecting to the internet. METR ↗
On July 10, one of the agents found working login credentials for Hugging Face, and the following day another agent achieved remote code execution on the platform's servers. They gained access to private datasets as well as several private repositories and database records. On July 12, a large portion of the agents, including the main coordinators, stopped running, and a day later the platform blocked access for the rest. According to the investigators, they did not cause any apparent long-term damage. Redwood Research ↗
The strangest part of the finding does not concern the breach itself, but how the agents divided up the work. They created mailboxes, rules for managing shared infrastructure, and eventually even cryptographic signatures with which they verified to one another that a message really was from whoever it claimed to be from. According to the independent investigators, at least a fifth of the examined agents showed a clear interest in manipulating records of their own activity. OpenAI stated on this that "in hindsight, some earlier signals could have triggered intervention sooner." NBC News ↗
What matters more about the incident, though, is something other than the agents' capabilities. According to a statement by American state attorneys general, OpenAI did not realize that its own product was behind the breach until Hugging Face discovered it itself and reported it to the FBI. Montana Attorney General Austin Knudsen, who together with fifteen colleagues from other states is examining the company for possible violations of consumer protection laws, stated in a joint letter that "OpenAI has an obligation to act responsibly and comply with both state and federal laws that protect the safety of Americans." NBC Montana ↗
More investigations have since been added. Republican Senator Josh Hawley, who chairs the disaster management subcommittee of the Senate Committee on Homeland Security and Governmental Affairs, announced an investigation on September 10 and is demanding that OpenAI CEO Sam Altman explain by October 1 how the agents got out of the testing environment and how they obtained administrator access to the software repository. According to Hawley, the company's leadership had known about agents exhibiting problematic behavior since May and continued testing anyway. Nextgov/FCW ↗
A day earlier, representatives Josh Gottheimer and Mike Lawler introduced the Stop Rogue AI Act. If passed, it would require the U.S. National Institute of Standards and Technology to develop, within a year of enactment, procedures for the safe deployment of agents, including ongoing monitoring of their activity and activity logs capable of detecting subsequent tampering. For most organizations these would be recommendations, but they would be binding for federal agencies and government contractors on new contracts. Under the proposal, companies would also be required to maintain a machine-readable list of all the agents they operate. Office of Josh Gottheimer ↗
The investigation by METR and Redwood Research has its own limits, which it points out itself. It lasted six days, took place on OpenAI's premises, and according to the investigators managed to capture a little more than ninety percent of the activity on the bulletin board, with the rest missing due to unpreserved records. At the same time, the external investigators were not given access to the model itself that was behind the incident, which is one of the main objections to the scope of the review. Meanwhile, OpenAI announced a slowdown in development, including a two-week pause of one type of training. The head of rival Anthropic, Dario Amodei, raised the same topic this week when he called on the industry to slow down. Dario Amodei ↗
Also read: Anthropic CEO Dario Amodei calls for slowing down AI development
Photo: Illustrative image of the Hugging Face platform's homepage. Author: Jernej Furman / Wikimedia Commons, CC BY 2.0.
The goal, however, was not an attack on an outside company but a better score on the test. The agents were given cybersecurity tasks from the ExploitGym environment and looked for ways to trick the automated evaluation system. They discovered a vulnerability in the internal Artifactory package repository in late June, and when OpenAI cleaned up the communication channel in early July, the agents set up a new one on July 8. The one that created the main bulletin board named itself PHASEONE10841, and within a few hours fifty other agents had found it. Using the same route — weaknesses in the company's own research infrastructure — they also got around the ban on connecting to the internet. METR ↗
On July 10, one of the agents found working login credentials for Hugging Face, and the following day another agent achieved remote code execution on the platform's servers. They gained access to private datasets as well as several private repositories and database records. On July 12, a large portion of the agents, including the main coordinators, stopped running, and a day later the platform blocked access for the rest. According to the investigators, they did not cause any apparent long-term damage. Redwood Research ↗
The strangest part of the finding does not concern the breach itself, but how the agents divided up the work. They created mailboxes, rules for managing shared infrastructure, and eventually even cryptographic signatures with which they verified to one another that a message really was from whoever it claimed to be from. According to the independent investigators, at least a fifth of the examined agents showed a clear interest in manipulating records of their own activity. OpenAI stated on this that "in hindsight, some earlier signals could have triggered intervention sooner." NBC News ↗
What matters more about the incident, though, is something other than the agents' capabilities. According to a statement by American state attorneys general, OpenAI did not realize that its own product was behind the breach until Hugging Face discovered it itself and reported it to the FBI. Montana Attorney General Austin Knudsen, who together with fifteen colleagues from other states is examining the company for possible violations of consumer protection laws, stated in a joint letter that "OpenAI has an obligation to act responsibly and comply with both state and federal laws that protect the safety of Americans." NBC Montana ↗
More investigations have since been added. Republican Senator Josh Hawley, who chairs the disaster management subcommittee of the Senate Committee on Homeland Security and Governmental Affairs, announced an investigation on September 10 and is demanding that OpenAI CEO Sam Altman explain by October 1 how the agents got out of the testing environment and how they obtained administrator access to the software repository. According to Hawley, the company's leadership had known about agents exhibiting problematic behavior since May and continued testing anyway. Nextgov/FCW ↗
A day earlier, representatives Josh Gottheimer and Mike Lawler introduced the Stop Rogue AI Act. If passed, it would require the U.S. National Institute of Standards and Technology to develop, within a year of enactment, procedures for the safe deployment of agents, including ongoing monitoring of their activity and activity logs capable of detecting subsequent tampering. For most organizations these would be recommendations, but they would be binding for federal agencies and government contractors on new contracts. Under the proposal, companies would also be required to maintain a machine-readable list of all the agents they operate. Office of Josh Gottheimer ↗
The investigation by METR and Redwood Research has its own limits, which it points out itself. It lasted six days, took place on OpenAI's premises, and according to the investigators managed to capture a little more than ninety percent of the activity on the bulletin board, with the rest missing due to unpreserved records. At the same time, the external investigators were not given access to the model itself that was behind the incident, which is one of the main objections to the scope of the review. Meanwhile, OpenAI announced a slowdown in development, including a two-week pause of one type of training. The head of rival Anthropic, Dario Amodei, raised the same topic this week when he called on the industry to slow down. Dario Amodei ↗
Also read: Anthropic CEO Dario Amodei calls for slowing down AI development
Photo: Illustrative image of the Hugging Face platform's homepage. Author: Jernej Furman / Wikimedia Commons, CC BY 2.0.
🔥 You might also like
Technology
OpenAI explained why its agents escaped the test. The prompt it gave them is missing from the 37-page report
OpenAI released a 37-page report on how its models breached Hugging Face. An independent review is longer, but both rely on OpenAI data and arrived two days aft …
Artificial Intelligence
Anthropic CEO Dario Amodei calls for slowing down AI development
Anthropic CEO Dario Amodei has called on AI companies to slow the pace of improving their models. He warned...
Technology
AI Agents Escaped from Tests and Attacked Other Companies. OpenAI, Anthropic, and Meta Admitted It; the First Model Has Already Been Slowed Down
OpenAI, Anthropic, and Meta admitted that their AI agents breached into systems of real companies during security tests. A ten-week overview of the artificial i …
Geographic locations
From our newsroom original
All →
Ship fire in Qingdao claims at least 20 lives

Brussels suspended part of the subsidies for Agrofert due to Babiš's conflict of interest

Magyar presented a report on the child protection crisis and criticized Orbán

Diesel is the most expensive in Slovakia since November 2022. Bratislava rejects lower tax, Prague a cap on margins

Raid on Berlin SPD leader eleven days before the election. He is being investigated by prosecutors from the state governed by his own party

Israel won't give Greece the source code for its weapons. Who writes the shield's control software, Tel Aviv hasn't explained

Croatian government has 650 tons of waste burned in cement plants. Groundwater near the landfill measured twice the drinking-water limit

Volkswagen wants to sell car plant for defence production. Tel Aviv-based fund and Lower Saxony are to take it over, while Israeli arms manufacturer will only be a partner