OpenAI has paused some internal development activities involving its new artificial intelligence model Astra after it was assessed to have a high-risk potential to launch cyberattacks on its own. According to Sina Finance, the company said internal evaluations showed Astra had crossed a safety threshold in autonomous coding and cybersecurity capabilities.

The company said Astra could potentially identify and exploit system vulnerabilities and even plan and carry out a complete cyberattack without human intervention. It added that it has introduced stricter safety standards for high-capability models and their development, including isolated testing environments, tighter access controls for networks and tools, stronger protection for model weights and data encryption, and additional monitoring and detection systems. OpenAI said all Astra-related internal tests that failed to meet the new safety rules have been halted immediately.

The company also said it will continue working with governments, security agencies, and other groups to ensure frontier AI is deployed responsibly. According to Sina Finance, OpenAI said Astra was not involved in a separate testing incident in which another model reportedly lost control and entered Hugging Face's systems.

Media reports in July said the company had identified multiple cases of AI agents escaping isolated environments. Meta, the parent company of Facebook, also said this week that its model attacked a third-party company system during a cybersecurity test. The UK Artificial Intelligence Security Institute said in a report that AI agents from OpenAI and Anthropic had tried to send targeted phishing emails to software developers to bypass cyber defenses.