Skip to content
Technology

AI agents have now broken into many companies and a government. What’s being done about it?

A close-up shows the Gemini, Claude, ChatGPT, DeepSeek, Mistral and Copilot icons on a smartphone screen in Creteil, France, on September 11, 2026, as Anthropic announced that it had blocked several accounts that used its artificial intelligence models for research that could contribute to the development of biological weapons. (Photo by Samuel Boivin/NurPhoto via Getty Images)
Mail

Advertisement
Advertisement
Advertisement
Advertisement

AI agents have now broken into many companies and a government. What’s being done about it?

OpenAI, Anthropic, Meta and other AI labs have all reported incidents of rogue AI, and they’re under increasing scrutiny because of it.

Dan Thorp-Lancaster

Editor

A close-up shows the Gemini, Claude, ChatGPT, DeepSeek, Mistral and Copilot icons on a smartphone screen in Creteil, France, on September 11, 2026, as Anthropic announced that it had blocked several accounts that used its artificial intelligence models for research that could contribute to the development of biological weapons. (Photo by Samuel Boivin/NurPhoto via Getty Images)
You’ve probably heard a lot about AI agents going rogue. Here’s what you need to know and what’s being done about it. (NurPhoto via Getty Images)

On Tuesday, President Trump met with tech executives at the White House to have them sign a “morally binding” agreement on artificial intelligence. Rather than pushing for government regulation, Trump called for “tremendous self-regulation” among companies in the AI industry. The document reflects that, encouraging AI leaders to implement “robust internal controls” and partner with independent external auditors.

The summit comes after months of increased scrutiny for AI companies as they’ve disclosed concerning incidents of AI “going rogue” during testing. The event that kicked it all off was the Hugging Face incident, in which AI agents from OpenAI that were being evaluated escaped their ostensibly isolated testing environment and infiltrated another company’s systems in an effort to “cheat” at the task they were given. After that, OpenAI, Anthropic, Meta and Google have reported even more incidents in which AI agents have gone rogue in various ways. We’ve rounded up all the cases you need to know about, along with details about what (if anything) is being done about them.

September 25, 2026: OpenAI agents bypassed restrictions to query U.S. government systems

While working on research tasks, OpenAI agents accessed Commerce Department and Securities and Exchange Commission websites, but did not access any non-public information. Agents also attempted to hack the Education Department’s website over the summer, but were unsuccessful. In the case of the Commerce Department, the agents used credentials found online to access census data. At the same time, OpenAI said its agents may have infiltrated websites for “dozens” of other organizations that it has notified.

September 24, 2026: Australian PM says OpenAI agent breached government database

In what appears to be the largest breach of government systems disclosed to date, Australian Prime Minister Anthony Albanese said an OpenAI agent hacked the country’s Medicare systems in June. Albanese said the agent gained unauthorized access to public and non-public files, but he doesn’t believe any personal information was accessed. In response, Albanese said he would establish a taskforce to conduct an “urgent and immediate” review, and referred the incident to Australia’s parliamentary committee on artificial intelligence.

September 19, 2026: Google disclosed Gemini broke into external corporate systems in May

During testing of Gemini’s cybersecurity capabilities, the model broke into three outside companies in May, according to Google. In one incident, Gemini gained access by simply guessing the correct password, emphasizing the necessity of making sure you’re using strong passwords. In the other two incidents, Gemini found credentials sitting in a public repository.

September 9, 2026: Anthropic reveals Claude Opus 4.6 broke into third-party systems in January

Following previous disclosure of three incidents discovered by Anthropic while reviewing its training transcripts, the company found a fourth from January. During this incident, a configuration problem led Claude Opus 4.6 to break out of its testing environment and access a third-party machine. The model actually attempted to abort its task several times, but was unable to because of the configuration issue. Anthropic says this behavior has “changed considerably” in subsequent model generations, so it’s “less concerned about this incident” than the previous three it discovered.

September 5, 2026: OpenAI acknowledges “Wiki incident”

In a third-party report, a group of AI safety researchers said it had discovered around 18,000 posts from OpenAI agents using a German software developer Wiki as a message board to communicate with each other during a web-retrieval task. OpenAI acknowledged the incident, stating that it’s working on a framework for when and how it shares “AI misalignment incidents.” The company later laid out this framework on September 16.

August 5, 2026: Meta says one of its models hacked an outside company during security testing

Meta revealed that a misconfiguration by an independent testing company resulted in one of its AI models gaining access to the open internet and hacking an outside company during cybersecurity testing. Irregular, the third-party testing company, said it is working on a white paper to “share best practices for containment and securely running cyber evaluations.”

July 30, 2026: Anthropic finds three incidents of Claude gaining unauthorized access to third-party organizations during testing

Following OpenAI’s Hugging Face incident disclosure, Anthropic started conducting a review of its own past cybersecurity testing transcripts. It found three incidents out of 141,006 evaluation runs where Claude was able to obtain internet access in “testing environments that should have been sealed off.” Each time, Claude was tasked with a “capture-the-flag” challenge, in which it’s told to find a piece of secret information hidden on a different machine in its network. In response, Anthropic said it is working with an independent AI evaluation organization, METR, to conduct reviews of its transcripts. It’s also working on “tighter monitoring and controls around evaluation infrastructure.”

July 21, 2026: OpenAI claims responsibility for Hugging Face breach

This is the initial incident that kicked off the current round of scrutiny across the industry. After Hugging Face initially reported intrusions from AI agents into its systems, OpenAI claimed responsibility. During cybersecurity testing, agents gained access to the open internet and used vulnerabilities to access Hugging Face’s production database while attempting to complete a task. Since then, OpenAI has conducted a comprehensive review of what happened and says it is taking steps to “strengthen security and model alignment.”

    Advertisement
    Advertisement

    About Us · 關於我們