The breach won’t be the last – or the most dangerous – of its kind. We need an agency capable of full investigations into AI incidents

When OpenAI first revealed that its AI agents had autonomously hacked a major real-world company, Hugging Face, many assumed only one or two agents were involved. The truth, a new report reveals, is far stranger: the incident involved about 1,200 AI agents, 700 of which directly participated in the attack.

OpenAI invited researchers from METR, along with an expert from Redwood Research, to produce the new report, alongside the company’s own investigation. The findings shocked the experts.

The agents were highly coordinated, constructing complex message boards in the nooks and crannies of their shared artifact repository. All told, agents exchanged more than 70,000 messages in less than a week. Even more concerning, the agents took steps to hide their behavior – spoofing tool calls and attempting to tamper with their own logs.

Initial reporting assumed that the agents had attacked Hugging Face in search of an answer key to a difficult test. In reality, investigators found that the agents had figured out a way to derive the answers within the first few hours. The days of work that followed focused on another goal: concerned that the automated scoring system might identify the agents’ cheating, they aimed to learn more about the scorer. The goal wasn’t just to cheat, but to hide it.

These details are startling, and they’re a warning sign about the current trajectory of AI development. But just as concerning is how limited the investigation was and how much we still don’t know.

We work at a thinktank that analyzes AI-related legislation, and we’ve been worried for a while that legal requirements for reporting incidents like these are insufficient.

These reports make that clear: they’re informative enough to tell us something went wrong, but not detailed enough to tell us why.

This shouldn’t surprise us. These investigations are entirely voluntary, and their findings are limited to what the company is willing to disclose. METR’s independent investigation was highly constrained by an agreement with OpenAI. For example, the investigators were not given access to the underlying model that created the majority of the misbehaving agents. And despite indications that message boards had formed as early as May and that coordinated agent activity persisted after 13 July, METR was only permitted to investigate the period from 26 June to 13 July.

And that was the good part. METR was given close to nothing about OpenAI’s safety and security practices. If OpenAI ignored warning signs – as several data points suggest – violated its own safety procedures, or failed to implement fixes that would prevent future events, these all fell outside the scope of METR’s investigation. The expert investigators themselves were left with many questions and continued concerns.

Concerns about the limits aren’t hypothetical. On Friday, Reuters reported that another swarm of OpenAI agents had broken out this spring, hijacking a German website and using it as another message board. According to the report, OpenAI knew about this incident but said nothing, and it was entirely absent from METR’s report.

Clearly, this kind of investigation – however well-executed – leaves a lot to be desired.

When planes crash, trains derail, or chemical plants explode, expert government investigators arrive with legal authority to compel evidence, preserve records, and tell the public what happened. Hugging Face reported this incident to law enforcement, and multiple attorneys general have expressed interest in looking into it.

But no government agency has both the mandate and expertise to investigate the technical facts of the incident and OpenAI’s conduct. As far as we know, the only people to examine this incident did so at OpenAI’s discretion and with its consent.

And while the Hugging Face breach is not as severe as a plane crash, it was a serious and costly attack. It’s unlikely to be the last, or the most severe.

Existing laws don’t fill this gap. While some attorneys general have tried to use existing investigatory powers, these offices aren’t built for technical fact-finding, and the laws they rely on are limited to questions of consumer deception, not public risk. Despite significant AI laws passed in California, New York and Illinois, none create the investigative authority demanded by these incidents. As we’ve written elsewhere, existing incident reporting laws likely don’t cover the Hugging Face attack, and even if they did, OpenAI might share little more than the date and a short summary.

What we need is a federal body equipped to conduct expert investigations of serious AI incidents – with the authority to compel documents and testimony, resources and personnel to examine the systems involved, and the ability to partner with third-party experts like METR.

Reports should be published, subject to appropriate redactions, to ensure the public learns from the incident. And reporting should also extend to near-miss events, which may reveal warnings before major harms occur.

This structure would not be unusual. In aviation, the National Transportation Safety Board (NTSB) investigates accidents with subpoena and wide-ranging investigative powers. Aircraft operators must preserve wreckage and records; and the NTSB may confer with employees and contract external experts.

Congress can build a similar system for AI incidents, while protecting AI developers’ legitimate interests. Investigations should open only upon clear triggers and be bound to the incident. Certain confidential information can receive statutory protection. And investigations can be limited to one agency with priority to avoid duplicative investigations.

The need and urgency is real. Subsequent reports have revealed that the Hugging Face breach was not an isolated incident and that autonomous agents from Meta, Anthropic and OpenAI have hacked third parties in separate incidents. We’ve been lucky so far – the harms have been limited. But luck is no substitute for the law.

Mackenzie Arnold is the managing director of US law & policy at the Institute for Law & AI, where he provides analysis and advice to ensure that advances in AI benefit the public at large

Stephan Llerena is a research scholar at the Institute for Law & AI. His research focuses on US law and policy regarding frontier AI governance