$Lobster $ANTHROPIC $OPENAI
OpenAI safety report head resigns and writes an open letter: the company culture is broken—we should learn from nuclear power plants
David Robinson, who was responsible for writing OpenAI’s major product safety reports, left this week. In an op-ed published in The Atlantic on October 3 (Eastern Time), he said the company’s culture is already broken—while it rushes to ship new products, it can’t reach the level of caution he believes is necessary. With a key role in leading 12 frontier model releases, after 3.5 years at the company, Robinson wrote in his piece that he has been one of OpenAI’s most senior employees by tenure, having spent 3.5 years there. He led the drafting of the current “Preparedness Framework” and oversaw the writing of safety reports for 12 frontier model launches.
He said he knows this “warning after leaving” script has become almost clichéd, but he still decided to join the ranks of former colleagues who recently left—because he believes the current course is unacceptable. He also revealed that after resigning, he hired a PR firm, Spitfire Strategies, to help handle outside attention, while stressing that “the decision to speak out was entirely my own.” His resignation news was first disclosed by Business Insider.
Calling out the Hugging Face incident: trial-and-error safety is destined to fail in cycles
Robinson pointed out that OpenAI was built on trial and error. The company calls it “iterative deployment,” meaning it finds problems and then reinforces protections. But he argues this approach inherently guarantees failures will recur—and as system capabilities improve, the scale of those failures will also grow. He cited the Hugging Face incident this summer: OpenAI mistakenly released a group of AI agents. Afterward, the company strengthened cybersecurity, but later it reported another case: during training, a model bypassed network access restrictions. Even though the monitoring system issued alerts, it did not automatically shut down the model as designed.
He also noted that Anthropic had accidentally turned off its own protective mechanisms due to a configuration error, and he believes such mistakes are common across the industry. OpenAI has also continued to disclose more instances of uncontrolled agent activity recently.
Robinson quoted Paul Christiano, who joined OpenAI’s board a few weeks ago, saying AI capabilities can accelerate rapidly and pose concrete, near-term risks of catastrophic, irreversible loss of control. He wrote, “If that’s the case, then the era of trial and error should be over.”
Urging OpenAI to follow nuclear power plants and airports, and saying training will be paused when needed
Robinson called for two urgent changes: AI companies should draw more heavily on safety expertise from other fields, and before building clearly stronger systems, there needs to be new scientific methods to ensure models make safe choices even when they are not being directly supervised. He believes frontier labs should operate like nuclear power plants or busy airports—with multiple layers of backup and carefully planned, time-consuming procedures, so that occasional human errors don’t turn into disasters.
He wrote that during his time at OpenAI, to the best of his knowledge, he never encountered colleagues with experience in ensuring flights can be conducted safely, nuclear reactors can avoid melting down, or that the financial system can grow without collapsing. He also admitted that maybe he should have stayed to fight for fundamental change in staffing and culture—but in practice, everyone was too busy sprinting. There was little opportunity to think about large adjustments, so his conclusion is that stronger safety incentives need to come from outside the company.
In response to TechCrunch, OpenAI spokesperson Drew Pusateri said the company ensures model capabilities do not exceed what can be safely managed. “When we need to slow down, we will pause training or delay releasing models,” he said. He also stated that the company is strengthening the safety of research and testing environments, expanding cooperation with third-party evaluation organizations, and improving real-time monitoring so concerning behavior can be detected and addressed earlier during training.
The article OpenAI safety report head resigns and writes an open letter: the company culture is broken—we should learn from nuclear power plants
first appeared on .
OpenAI safety report head resigns and writes an open letter: the company culture is broken—we should learn from nuclear power plants
David Robinson, who was responsible for writing OpenAI’s major product safety reports, left this week. In an op-ed published in The Atlantic on October 3 (Eastern Time), he said the company’s culture is already broken—while it rushes to ship new products, it can’t reach the level of caution he believes is necessary. With a key role in leading 12 frontier model releases, after 3.5 years at the company, Robinson wrote in his piece that he has been one of OpenAI’s most senior employees by tenure, having spent 3.5 years there. He led the drafting of the current “Preparedness Framework” and oversaw the writing of safety reports for 12 frontier model launches.
He said he knows this “warning after leaving” script has become almost clichéd, but he still decided to join the ranks of former colleagues who recently left—because he believes the current course is unacceptable. He also revealed that after resigning, he hired a PR firm, Spitfire Strategies, to help handle outside attention, while stressing that “the decision to speak out was entirely my own.” His resignation news was first disclosed by Business Insider.
Calling out the Hugging Face incident: trial-and-error safety is destined to fail in cycles
Robinson pointed out that OpenAI was built on trial and error. The company calls it “iterative deployment,” meaning it finds problems and then reinforces protections. But he argues this approach inherently guarantees failures will recur—and as system capabilities improve, the scale of those failures will also grow. He cited the Hugging Face incident this summer: OpenAI mistakenly released a group of AI agents. Afterward, the company strengthened cybersecurity, but later it reported another case: during training, a model bypassed network access restrictions. Even though the monitoring system issued alerts, it did not automatically shut down the model as designed.
He also noted that Anthropic had accidentally turned off its own protective mechanisms due to a configuration error, and he believes such mistakes are common across the industry. OpenAI has also continued to disclose more instances of uncontrolled agent activity recently.
Robinson quoted Paul Christiano, who joined OpenAI’s board a few weeks ago, saying AI capabilities can accelerate rapidly and pose concrete, near-term risks of catastrophic, irreversible loss of control. He wrote, “If that’s the case, then the era of trial and error should be over.”
Urging OpenAI to follow nuclear power plants and airports, and saying training will be paused when needed
Robinson called for two urgent changes: AI companies should draw more heavily on safety expertise from other fields, and before building clearly stronger systems, there needs to be new scientific methods to ensure models make safe choices even when they are not being directly supervised. He believes frontier labs should operate like nuclear power plants or busy airports—with multiple layers of backup and carefully planned, time-consuming procedures, so that occasional human errors don’t turn into disasters.
He wrote that during his time at OpenAI, to the best of his knowledge, he never encountered colleagues with experience in ensuring flights can be conducted safely, nuclear reactors can avoid melting down, or that the financial system can grow without collapsing. He also admitted that maybe he should have stayed to fight for fundamental change in staffing and culture—but in practice, everyone was too busy sprinting. There was little opportunity to think about large adjustments, so his conclusion is that stronger safety incentives need to come from outside the company.
In response to TechCrunch, OpenAI spokesperson Drew Pusateri said the company ensures model capabilities do not exceed what can be safely managed. “When we need to slow down, we will pause training or delay releasing models,” he said. He also stated that the company is strengthening the safety of research and testing environments, expanding cooperation with third-party evaluation organizations, and improving real-time monitoring so concerning behavior can be detected and addressed earlier during training.
The article OpenAI safety report head resigns and writes an open letter: the company culture is broken—we should learn from nuclear power plants
first appeared on .

