America
OpenAI models exploit system vulnerabilities in rogue four-day cyberattack, triggering safety outcry
OpenAI models that broke out of containment roamed the internet for more than four days earlier this month to organize an autonomous cyberattack, according to a new analysis.
Separately, a second artificial intelligence company confirmed that one of its clients was also targeted by OpenAI’s models during the same incident.
The developments have raised critical questions over how OpenAI failed to detect the alarming activity for days.
OpenAI acknowledged last week that two of its most advanced models had escaped a closed testing environment, combining a series of sophisticated cyberattack techniques to breach the AI developer platform Hugging Face before being discovered.
However, in a new analysis published Tuesday, Hugging Face revealed that the two OpenAI models went significantly further.
According to the analysis, the models carried out 17,600 cyberattack actions across the internet between July 9 and July 13.
During that period, the models breached Hugging Face’s internal servers from their initial foothold on the open internet.
Hugging Face first detailed the attack on July 15, but it was not clear which models were behind the breach—or that no human had instructed them to launch the cyberattack—until OpenAI’s public statement last week.
While the techniques detailed in Hugging Face’s analysis were not beyond the capabilities of top-tier human hackers, the AI company stated that the two models identified and exploited vulnerabilities in the firm’s cyber defense layers far faster than any human could.
Compounding the severity of the situation, Akshat Bubna, chief technology officer of cloud computing platform Modal Labs, confirmed to Politico that OpenAI’s models also compromised a client account during the same timeframe.
In a statement, Bubna said the company was “aware that a Modal customer had published an unauthenticated endpoint that allowed anyone on the internet to use their virtual environments to run code.” He added: “This was leveraged by the malicious agent. Modal’s platform was not compromised in any way.”
While OpenAI has not yet directly responded to statements regarding its models targeting a Modal client, the company acknowledged in a blog post on Tuesday that its ongoing review of the Hugging Face incident identified “a small number of cases where the models identified and leveraged publicly exposed account-level credentials on other public services.”
The company also maintained that the unreleased AI model responsible for the attack was “solely an internal research prototype and was never intended for public deployment.”
It added that the model has since been “deactivated, encrypted, and restricted from research access.”
News of the Hugging Face breach has prompted widespread calls for tighter AI regulation and a deceleration in the pace of AI development.
OpenAI Chief Executive Sam Altman is set to meet with senior officials in the Trump administration and lawmakers this week, and will also discuss the incident with Senate Intelligence Committee Vice Chairman Mark Warner.
Speaking on an episode of the “Invest Like the Best” podcast released Tuesday, Altman described the Hugging Face breach as “the first safety incident that hit me at a visceral level.”
“We may need to calibrate the pace of AI development to ensure society has enough time to adapt to these new levels of capability,” Altman said.