Hugging Face said last week that he had been attacked. OpenAI Now says an unpublished OpenAI model has done just that

In a blog post from Thursday last weekartificial intelligence software repository Hugging Face has announced a strange cyber attack on the systems running its services. “It was different from anything we’ve handled before,” the report said, because “it was driven end-to-end by an autonomous system of artificial intelligence agents.”

In his own blog post on TuesdayOpenAI said its own models were responsible for the attack, and it is coordinating with Hugging Face to resolve the situation.

OpenAI now claims that the attack was “driven” by AI models that were evaluated behind the scenes by OpenAI, including its flagship GPT-5.6 Sol model, as well as an undisclosed second model that has not yet been released or announced. “We consider this incident to be an unprecedented cyber incident involving state-of-the-art cyber capabilities and are responding accordingly,” OpenAI wrote.

In the formulation of its account, OpenAI assigns agency to a model or models, not to an individual agent operating on a model. Axios’ coverage of this story says“The models were autonomous tokenmaxxers.”

The OpenAI blog post essentially says that an evaluation was being conducted that was supposed to test the models’ ability to carry out cyber attacks –landmarks what are they called. These instances of the models theoretically ran without internet access, and instead only allowed downloads from a network hosted by OpenAI itself via an unnamed web hosting provider.

Obviously, a benchmark was used ExploitGymwhich is not OpenAI’s own benchmark. The team that created it associated with UC Berkeley, Max Planck Institute for Security and Privacy, UC Santa Barbara, Arizona State University, Anthropic, OpenAI and Google. As far as I can tell it’s all hosted on GitHub.

But apparently the models were obsessed with improving their results in ExploitGym and, well,

“Working in our sandbox test environment, our models spent a significant amount of computation to find a way to gain public access to the Internet to solve the assessment problem. To gain access, the models identified and exploited a zero-day vulnerability (which we have now responsibly disclosed to the vendor) in the packet registry cache proxy server. Using this access, our models performed a series of privilege escalation and lateral movement actions on our exploratory testing environment until the models reached a node with Internet access.”

Once they were online, they “concluded” that the answers to the test could be obtained from Hugging Face, which is plausible since there are many artificial intelligence models and other related tools. This rogue operator “sought and successfully found ways to gain access to classified information that he could use to cheat the appraisal.” This apparently involved obtaining stolen credentials and zero-day vulnerabilities to find ExploitGym solutions somewhere on Hugging Face’s servers.

The OpenAI and Hugging Face security teams seem to have noticed this happening. They now say they have combined their investigations.

The Hugging Face blog post from last week appears to have been posted before this coordination took place. In fact, it appears that it was published before OpenAI came forward as the company behind the crime. “We don’t know which model was running the attacker’s agents, whether the hacked hosted model or the unlimited open model,” Hugging Face wrote, adding that “in any case, the attacker was not bound by the usage policy, while our own forensic work was blocked by the fence of the hosted models we first tested.”

Back in April, Anthropic announced that its unprecedentedly powerful Mythos model “could transform cybersecurity” as it deploys Project Glasswing, a coordination effort to prepare organizations for future cybersecurity threats. Similarly, OpenAI reports on its blog about the incident that organizations can apply for enhanced security information through a trusted access program. “We urge other advocates to do the same apply for trusted access⁠ and experiment with these models now to turn these capabilities into better prevention, faster detection and more effective incident response,” says OpenAI.

Exit mobile version