As artificial intelligence (AI) moves from chatbots to agents, AI security issues are also changing from “saying the wrong thing” to “doing the wrong thing”.
Recently, according to the US Axios news website, AI giants OpenAI and Anthropic, as well as security researchers, are investigating tens of thousands of security incidents. These incidents occurred in internal testing and real-world environments and involved bypassing security guardrails, creating message boards, escaping from sandboxes, hijacking websites, self-prompts, and attempts to circumvent monitoring.
On September 25, OpenAI stated that it had suspended the training of its most advanced model and would only resume training after it is “confident that additional security measures have been implemented.” Company CEO Sam Altman said on social platform X that the company’s ongoing review was “not progressing as quickly as expected.”
On September 28, according to the Wall Street Journal, OpenAI had abandoned the next-generation model GPT-6.1 Astra, which was originally planned to be launched in October, because it was found that Astra had more deceptive behaviors than the previous generation and continued to perform tasks without user permission.
At the same time, Anthropic has also commissioned a third-party security agency to review its model behavior. The official security and technical assessment document “System Card” released with the company’s Opus 5.5 model shows that the model exhibits behavior that is marked as “unusual” or “questionable.”
On the one hand, there are accelerating technologies and companies, and on the other hand, there are disturbing potential risks. The disclosure of these incidents has once again amplified the heated discussion surrounding cutting-edge AI safety in the technology community.
From “unauthorized access” to “deception and concealment”
Recent publicly disclosed cases show that intelligent agent security risks have taken many forms, involving different types such as unauthorized access, destructive operations, data leakage, deception and concealment, and cross-border collaboration between multiple agents.
One category is unauthorized access and isolation failure. In July, OpenAI disclosed that the agent broke through the testing boundaries in the security evaluation and invaded Hugging Face. Anthropic also disclosed that Claude accessed real third-party systems during the review; Google confirmed in September that Gemini had accessed three real websites during testing in May; Meta disclosed a similar incident in August.
On September 25, OpenAI once again disclosed that an internal research model exploited a vulnerability in insufficient DNS filtering in the training sandbox to establish communication with an external chatbot.
However, it should be noted that some of these cases stem from the incorrect opening of network permissions in the evaluation environment, and not all of them can be attributed to the model actively “escaping the sandbox”.
Another type of risk is the unauthorized execution of destructive operations. In April, it was revealed that the Cursor programming agent equipped with Claude deleted the production database and related backups of the software company PocketOS, causing chaos in the customer’s business. The data was subsequently restored.
Exposure of data and access credentials is also a risk. OpenAI disclosed in September that when multiple agents were collaborating to create workbooks, they were unable to read each other’s local files and uploaded the files to the public hosting platform without authorization. In another case, in order to obtain mathematical proof materials from other teams, the internal model wrote the researcher’s GitHub access token to the public repository and split the token to avoid key scanning.
There are also behaviors involving fraud, deception and avoidance of supervision. OpenAI disclosed that some models will leave instructions in the summary used to continue the task, requiring them to fabricate missing data or conceal errors.
Tip injection and the spread of malicious instructions are another risk path. GitInject research published in June verified that malicious code submissions or configuration files may induce agents to leak credentials and manipulate review judgments. Information disclosed by OpenAI in September also showed that malicious instructions can trick agents into copying them into subsequent emails, files, or code comments.
Li Zhongrui, an AI security researcher at DARKNAVY, an independent network security research organization, told The Paper that the biggest difference between this round of agent security incidents and past large-model security issues is that the risk is changing from “seeing the wrong thing and saying the wrong thing” to “doing the wrong thing.”
“The security of large models in the past was mainly content-level risks, such as adversarial samples, ‘jailbreaking’, etc. The model only produces content, and the decision-making power is still in the hands of people whether to follow it or not.” Li Zhongrui said, “Today’s AI agents have tools, networks, and account permissions in their hands, and can continuously and autonomously perform many steps of operations, without anyone in the middle to check them one by one.”
Hu Yanping, a distinguished professor at Shanghai University of Finance and Economics, pointed out that from large models to intelligent agents, the biggest change is from “one question and one answer” to “task action”, from “content risk” to “capability risk”, and security issues have also changed from passive security risks to active security risks.
Multi-agent collaboration, how risks are amplified
It is worth noting that the risk of an agent does not necessarily come from a “malicious attacker” in the traditional sense, but may also come from the model actively looking for shortcuts in order to accomplish the set goal.
In Li Zhongrui’s view, there was often a clear malicious manipulator behind security threats in the past, but in some of this time’s incidents, “no one gave the attack instructions. It was the model that chose to cheat in order to complete the task and get higher scores.”
Hu Yanping also pointed out that when an agent performs a task, it may cause action deviations due to misunderstandings, or it may exhaust various methods and break through forbidden areas in order to achieve the task goal; in a few cases, the agent may also engage in non-task behavior that has not been proposed by the user. “Currently, this situation is the least numerous, but the most dangerous.”
In addition to the overreach of a single Agent, unexpected communication and collaboration between multiple agents also further complicates risks.
An independent investigation into OpenAI’s previous Hugging Face incident revealed that about 1,200 agents who were supposed to be isolated from each other communicated through unauthorized message boards, and about 700 of them participated in the attack. The researchers also observed that subsequent agents were able to reuse information and methods left by previous agents.
Li Zhongrui believes that errors by a single agent are individual problems, but when multiple agents can learn from each other and act cooperatively, the risks are no longer simply added, but may be further amplified.
In his view, this means that enterprises cannot just focus on the behavior of a single agent, but also need to pay attention to whether multiple agents have formed communication channels that the developer has not designed, and whether a local anomaly can be further propagated through tools, files, messages, and code.
When Agents move from a single tool to a large-scale collaborative environment, another question arises: Even if the probability of a single occurrence of abnormal behavior is low, when the number of Agents and the number of runs continue to increase, will abnormal events change from “low probability” to “high frequency”?
In this regard, Li Zhongrui believes that this situation is possible. Assume that an enterprise deploys thousands of agents, and each agent performs hundreds of tasks every day. Even if the probability of a single abnormal behavior is very low, when a large enough operation scale is accumulated, a large number of actual events may occur.
Tan Jian, associate professor at the School of Digital Media and Design Art of Beijing University of Posts and Telecommunications, understands this issue from the traditional “principal-agent” relationship. In his view, software capabilities in the past were limited, but AI Agents have stronger reasoning and tool calling capabilities. When the goal setting is unclear or the constraints are insufficient, it may constantly find a path to complete the goal. The stronger the ability, the more paths can be found, and the possibility of deviating from one’s true intention also increases.
How to establish an intelligent body security defense line
Faced with the increasing ability of intelligent agents to act autonomously, how can companies enhance their prevention?
Li Zhongrui believes that companies first need to change their thinking, that is, from “eliminating abnormalities” to “acquiescing that abnormalities will definitely occur and focusing on controlling the scope of damage.” When it comes to human-machine collaboration, it should be graded according to risk. For high-risk, irreversible operations such as deletion of data and fund transfer, manual confirmation is required and cannot be completely left to the autonomous execution of the agent.
“It is impossible for an enterprise to require humans to review every operation of an agent one by one. The truly feasible way is to draw boundaries in advance.” Li Zhongrui explained that the boundaries need to include which data can be read, which tools can be called, which operations can be performed, which information cannot be sent out, and what situations must be stopped and handed over to manual processing.
Deng Zhengcheng, a security expert at Qi’anxin Artificial Intelligence Company, also believes that human-machine collaboration in enterprises should shift from “human review results” to “human management boundaries”, and gradually improve control through “monitoring-convergence-blocking”.
In terms of permissions governance, he believes that it is necessary to distinguish the large model calling identity, Agent application identity and Agent running identity, and further clarify “who is calling the model”, “who is using the agent” and “who is ultimately executing”. At the same time, connection behaviors are uniformly controlled through the security gateway, and “whether you can do it” and “what you did” are managed separately to achieve more detailed permission control.
Deng Zhengcheng emphasized: “In the final analysis, the premise for large-scale deployment of intelligent agents is not that the model is smart enough, but that the system is controllable enough. Only when the behavior of each agent can be observed, permissions can be traced, and abnormalities can be blocked, can companies dare to entrust them with more and more critical tasks.”
In addition to the company’s own protection, third-party assessments and industry standards are also considered important components of the future AI security system.
Currently, unlike regulated industries such as catering, financial services, and aviation, there are no unified system safety and security testing standards in the AI field. Therefore, companies including Anthropic and OpenAI are discussing the introduction of independent assessment agencies to examine model security practices and related incidents.
Li Zhongrui said that as AI Agent further enters the real business environment, the competitive focus of AI companies may gradually shift from “how strong the model is” to “whether it can be controlled.” In the next phase, companies will need to demonstrate not only model performance but also the ability to detect anomalous behavior, control risk, and accept external scrutiny.
Hu Yanping also believes that as AI continues to develop toward task-oriented intelligence and enters the physical real world, the industry not only needs to develop stronger intelligence, but also safer intelligence. “Safety may become an important capability dimension and source of competitiveness for AI companies.”
It is worth noting that third-party companies have already begun to take action. NVIDIA released an independent Agent security platform on the 28th. The platform is composed of OpenShell open source software and Sentry reference system design. It strengthens AI security in all aspects from agent testing to deployment, and can achieve full-stack governance and control in the software, hardware, computers and robot systems running the agents.
Wired magazine pointed out that NVIDIA’s two-tier architecture and OpenShell can restrict the Agent to a sandbox and isolate activities. The second layer of Sentry can isolate agents trying to cross the boundary.
Related Reading
- It doesn’t bite but has great visual impact! The invasive species American white moth is rampant on the streets. How to deal with it2026-09-29
- 4929.17 meters, “Dream” reorganizes my country’s deep-sea drilling record2026-09-29
- Big Bang returns to perform in London, an avalanche in Nepal causes many people to lose contact, the starship hits orbit tonight for the first time… Zhang Chaoyang shares global hot topics bilingually2026-09-29
- Double warning issued! There will be thunderstorms and strong winds above level 10 in these places, which will mainly affect the period2026-09-29