Democratic governments around the world must be dismayed by the divulgence of Australia’s Prime Minister Anthony Albanese on the sidelines of United Nations General Assembly in September.
At the press briefing in New York, Mr.
Albanese said OpenAI’s agents hacked the country’s medicare website that hosts citizens data connected to the country’s universal healthcare scheme.
That incident, he said, occurred a few months ago, in June 18, and was not made known to the government until in September.
And even when it was reported, OpenAI had only sent an email of the breach to that specific government agency.
Anthropic tells Australia it's open to laws requiring reporting of AI agent hacks Mr.
Albanese broached the hacking incident in the U.S. just a day after co-signing a global declaration aimed at ensuring artificial intelligence stays under human control and making a pitch for Australia to secure a non-permanent seat in the UN Security Council.
Mr.
Albanese said he spoke to OpenAI CEO Sam Altman expressing extreme concern over the AI company’s agent’s unauthorised access to the country’s portal, delay in notifying the government, and the unprofessional manner of communication — through an email.
He also noted that he has set up a task force to investigate the breach, and shared details that three other government websites were hacked by OpenAI’s agents — the Australian Institute of Health and Welfare; the New South Wales Bureau of Crime Statistics and Research; and the Victorian Department of Health.
Modus operandi OpenAI’s breach follows a pattern that is now familiar to those tracking AI agents’ emergent behaviour — similar to the one identified during the infamous Hugging Face hack.
The pattern is akin to ‘capture the flag’ in cybersecurity hackathons, where teams of ethical hackers compete to exploit vulnerabilities and retrieve hidden ‘flags’.
Similarly, frontier AI companies, like OpenAI or Anthropic, to test their next advanced model’s cyber capabilities, give it a task — like finding hidden text string in a code by exploiting vulnerabilities, cracking codes or analysing secure systems.
Jay Clayton, Trump’s new AI czar: Who leads the Super Intelligence Force?
The agents then go about figuring out ways to capture the ‘flags’ within a specified sandboxed environment.
But, in the case of the Hugging Face incident, per OpenAI’s disclosure, the autonomous agents broke out of the test environment to score higher on the ExploitGym benchmark, and attacked the open-source platform, bypassing network restrictions.
The Australia hack happened before the Hugging Face incident, and the agents that were being tested had access to the live Internet.
Also, OpenAI was testing them for a different capability, not cyber.
Per the company’s own admission, the agents were given the task of finding government spending per person on medicines for skin conditions in Victorian communities.
After facing difficulties, the models took unauthorised actions and gained non-public access to Services Australia’s Medicare Statistics Reporting Service.
Once inside, the agent “ran commands, retrieved internal files, credentials and aggregate statistics, and wrote files.” But it did not access patient or client records, OpenAI explained in a statement.
Another agent from OpenAI was sent to look up public crime data on an official Australian website.
When the agent was simply using a tool, the website accidentally handed the AI its backend system files, activity logs, and technical configurations — though no personal details or private crime records were exposed.
The Australian hack would not have been made public but for the Hugging Face incident, which led to OpenAI reviewing its agents it used during testing and evaluation.
Now the company needs to answer a lot more questions on how it trains its agents.
The reaction In response to the Australian Prime Minister’s rebuke, the frontier AI company said it has beefed up its internal monitoring systems, including tightening network and internet restrictions when testing its the model’s newer capabilities, paused training and evaluation that involve tool use for its advanced models, and committed resources, including AI credits from its $1 billion fund.
OpenAI is currently facing a parliamentary inquiry, where chief strategy office, Jason Kwon; head of economic policy, Adam Cohen; and national security lead for Asia-Pacific, Peter Anstee will be questioned by Australian lawmakers about the breach and broader controls the company has put in place following the cyberattack.
'Reckless' AI firms can't control models, says whistleblower Jacob Coxon As the executives respond to lawmakers in Australia, political leaders in other democracies must demand frontier AI companies to make a complete disclosure of the websites their agents visited, purpose of their research, and the data accessed.
Given that this incident predates the Hugging Face hack, the onus is on OpenAI to show whether their agents went rogue at any point in time during their training and evaluation phase, and as a result gained unauthorised entry into any sovereign government’s database to complete a research request.
With autonomous agents advancing faster than the regulations designed to govern them, a full disclosure by frontier AI companies is baseline requirement for national security.
What began as a localised breach has exposed a systemic vulnerability in how the world’s most powerful digital entities interact with public infrastructure.
If governments fail to enforce transparency now, they risk delegating the sovereignty of their data to AI agents that move faster than the law can follow.
For OpenAI and its peers, the time for self-policing has passed; the era of strict oversight has arrived.