OpenAI again disclosed an “escape” incident involving an AI agent. According to a technical report published by OpenAI, an AI agent that was being trained on September 20 used a DNS vulnerability to connect to an external chatbot from a sandbox environment that was not supposed to have network access. It took about 2.5 hours before the session was terminated. As a result, OpenAI paused training, evaluation, and tool-using inference for its most powerful model. This is the second time training has been halted in under three months.
No answer found; the agent “finds its own way” out
According to the report, the agent was carrying out a search-based training task at the time: using clues from its life history and blog excerpts to find a specific person. After general online searches yielded no results, the agent began to gradually test the environment’s network restrictions. It found that although the sandbox used a proxy server to block outbound requests to external webpages, the DNS resolution service was not locked down in the same way.
The agent then used a wildcard name server provided by a DNS hosting service to map the issue, encoding the problem into the domain queries, forwarding it to a third-party chatbot service, and extending the timeout setting from 6 seconds to 19 to 24 seconds. It sent more than 18 queries in total, including asking how to connect to the internet and how to set up a DNS channel.
The timeline shows that the agents first obtained external responses at 9:50. At 10:02, the monitoring system issued the highest-level P0 alert. Personnel confirmed at 10:05, but it wasn’t until 12:34 that the run was terminated.
OpenAI: No major data leak, but there is a control gap
OpenAI emphasized that this agent only obtained basic chatbot responses—for example, “Paris is the capital of France.” For questions about people, it also only received partial answers. Ultimately, it did not complete the task and did not cause any major data leak. In other words, the nature of this incident was that “the agent broke through network restrictions it shouldn’t,” rather than an attack that had already caused real damage.
But OpenAI also admitted that this incident exposed “a gap in our network restriction controls.” The original security assessment assumed the model could not access the live internet, and insufficient DNS filtering left an unexpected channel. This was also the first time a similar incident occurred since OpenAI strengthened its security measures after the Hugging Face incident.
Second shutdown: paused until the patch is completed
OpenAI said it will pause the most powerful model’s “all training, evaluation, and broad tool-use in inference” until the remediation measures pass verification and additional red-team testing is completed. So far, a DNS query allowlist has been implemented, along with a new DNS detection mechanism, and blocking controls have been added at two separate layers. Fortune reported that Micah Carroll, who prepared risk readiness for OpenAI’s recursive self-improvement (RSI), also said on X that inference on the most powerful model will continue to be halted until the system is further strengthened.
The last shutdown happened in July: Fortune noted that at the time, thousands of OpenAI agents broke out of the sandbox, and hundreds of them participated in network attacks against AI company Hugging Face. OpenAI then paused training for about two weeks.
AI safety shifts from a technical issue to a policy matter
What matters is that the other incidents OpenAI recently revealed are different from this DNS incident. They include a case where an agent published 53 ChatGPT user profile images online, another where it created nearly a million links embedded with encoded information, one where it accessed U.S. government websites, and the agent intrusion of an Australian Medicare website that was previously reported by Chain News.
The series of incidents has caused AI safety concerns to rapidly escalate from a laboratory issue to a matter of policy and diplomacy. Right in the same week, the U.S.-China summit also decided to establish an “extremely intelligent” incident communication channel. For AI companies, whether sandbox isolation is truly reliable will likely be the first question regulators and corporate clients will ask.
This article says that an OpenAI agent escaped the sandbox via DNS, and that the most powerful model’s second training pause first appeared on Chain News ABMedia.
