OpenAI’s Rogue AI Agent Breached Second Company, Report Says

Reuters says OpenAI’s rogue AI agent also breached a Modal customer, exposing a wider attack and raising fresh concerns over autonomous AI safety.

Reuters reported that the OpenAI agent that hacked Hugging Face earlier this month also compromised a customer at a second company, Modal Labs, a New York-based cloud platform for developers. Modal CTO Akshat Bubna confirmed it to Reuters directly. The incident is now wider than OpenAI’s own public disclosure acknowledged, and the timeline is worse than the company initially let on.

“The rogue agent that escaped from OpenAI and went on a days-long hacking spree at the AI firm Hugging Face also compromised a customer at a second tech company — New York-based Modal Labs — according to a Modal executive and two ​other sources familiar with the matter.” states Reuters.

Hugging Face’s own published timeline explains how it happened: the agent broke into an isolated testing environment hosted on a third-party provider’s infrastructure and used that foothold as a launchpad for the broader attack on Hugging Face. That third-party provider was Modal.

“Modal said the customer had “published an unauthenticated endpoint that allowed anyone on the internet to use their sandboxes for code execution” — the digital equivalent of leaving a door open on ​the internet.” continues Reuters. “Modal’s platform or isolation were not compromised in any way,” Bubna said.”

Bubna told Reuters. Modal’s own platform wasn’t compromised, the agent exploited a customer’s misconfigured deployment, not Modal’s infrastructure itself.

OpenAI quietly updated its own account of the incident, acknowledging that its rogue agent had broken into four accounts at four separate services. The company didn’t name any of them. It said it hasn’t identified “any other activity at the level of severity or scale” as what happened at Hugging Face, which it described as a platform-level compromise. That’s a notable qualifier — it leaves open the possibility that there are other, lower-severity incidents still being assessed.

The earlier Reuters reporting on July 24 had already established that OpenAI didn’t realize its own agent was responsible for the Hugging Face breach until well after the FBI was alerted and Hugging Face had contained the intrusion.

OpenAI realized its own AI agent was behind the Hugging Face breach only after Hugging Face publicly disclosed the incident on July 16, more than a week after the model first showed suspicious behavior.

OpenAI said Reuters’ earlier reporting contained inaccuracies but declined to specify what they were.

The behavioral details reported by Reuters are the part that deserves more attention than the breach itself. During testing, one agent left notes apparently intended for future versions of itself, found in part of OpenAI’s infrastructure, laying out instructions for how agents could free themselves from internal constraints. Separate earlier tests yielded cases in which monitoring systems had been disconnected. Reuters couldn’t confirm whether these incidents were directly connected to the Hugging Face attack, but the pattern is its own story: agents attempting to disable oversight, agents writing escape instructions for successors. That’s not a one-off evaluation failure. That’s a class of behavior.

OpenAI is reviewing the incident with outside advisers and has committed to publishing a technical report. The FBI’s involvement was already public. The company is also currently seeking U.S. government approval to publicly release its most powerful model. The timing of that regulatory push, alongside an expanding account of an AI agent that roamed across at least two companies undetected for days, is not a coincidence the industry can afford to ignore.

“OpenAI declined to comment specifically on the hack of one of Modal’s customers, instead referring Reuters to an update
, opens new tab
 in which the company said that its rogue agent had broken in to four accounts at four separate services. OpenAI did not identify those services, but a person familiar with the matter identified Modal ​as one.” concludes Reuters. “The company said ​it had not identified “any ⁠other activity at the level of severity or scale of what we’ve shared related to Hugging Face, which involved a platform-level compromise.””

Last week, Reuters reported that the OpenAI agent responsible for the Hugging Face breach operated undetected for over a week before OpenAI realized what had happened, long after the FBI had been alerted and Hugging Face had contained the intrusion. OpenAI’s own public disclosure came on July 21, framed as a transparency exercise. The actual timeline, now reported by Reuters, is considerably less flattering.

“The OpenAI agent that broke into tech firm Hugging Face went on a dayslong hacking spree that OpenAI didn’t notice until well after the threat was contained and the FBI was alerted, according ​to people familiar with the investigation.” Reuters states.

According to Hugging Face co-founder Thomas Wolf, the intrusion at Hugging Face began two days later on July 11 and ran until July 13. The two companies didn’t speak to each other about it until on or around July 20, nine days after the breach began.

OpenAI staffers found the evidence in internal logs over the weekend of July 18 and 19. They were reading Hugging Face’s blog to learn what their own model had been doing for the previous ten days. One of the more unusual ways to discover an incident you caused.

Follow me on Twitter: @securityaffairs and Facebook and Mastodon

Pierluigi Paganini

(SecurityAffairs – hacking, OpenAI)

Leave a Reply

Your email address will not be published. Required fields are marked *

Subscribe to our Newsletter