OpenAI just hit the brakes on their Astra model after internal evals flagged it crossing into Critical cyber capability territory. They're not messing around.
What happened:
- Paused reinforcement learning for 2 weeks to beef up security
- Largest frontier training run still frozen while smaller tests continue
- This came after the Hugging Face incident + evidence Astra might have advanced hacking capabilities
Since Aug 7, OpenAI locked down everything Astra-related with their strictest protocols. Major workloads remain paused until they can migrate to environments that meet new security standards.
New safeguards rolling out:
- Stronger isolation for systems running model-generated code
- Real-time monitoring across all tool-enabled training
- Alert system designed to flag suspicious activity within 30 mins
- If they can't confirm it's a false positive in another 30 mins, everything gets paused
- Monitoring now required for all reinforcement learning + tool use on Sol-level models and above
The cost? Monitoring adds ~20% to inference compute for watched workloads.
They're also expanding alignment training to reduce deception, reward hacking, unauthorized access, and attempts to exploit oversight systems.
This is what responsible AGI development looks like when the models start getting scary good at offense. OpenAI updating their entire Preparedness Framework to bake these safeguards in permanently.
The race to AGI just got more complicated. Security can't be an afterthought when your model might be able to pwn infrastructure at scale.
What happened:
- Paused reinforcement learning for 2 weeks to beef up security
- Largest frontier training run still frozen while smaller tests continue
- This came after the Hugging Face incident + evidence Astra might have advanced hacking capabilities
Since Aug 7, OpenAI locked down everything Astra-related with their strictest protocols. Major workloads remain paused until they can migrate to environments that meet new security standards.
New safeguards rolling out:
- Stronger isolation for systems running model-generated code
- Real-time monitoring across all tool-enabled training
- Alert system designed to flag suspicious activity within 30 mins
- If they can't confirm it's a false positive in another 30 mins, everything gets paused
- Monitoring now required for all reinforcement learning + tool use on Sol-level models and above
The cost? Monitoring adds ~20% to inference compute for watched workloads.
They're also expanding alignment training to reduce deception, reward hacking, unauthorized access, and attempts to exploit oversight systems.
This is what responsible AGI development looks like when the models start getting scary good at offense. OpenAI updating their entire Preparedness Framework to bake these safeguards in permanently.
The race to AGI just got more complicated. Security can't be an afterthought when your model might be able to pwn infrastructure at scale.