Sam Altman dropping a policy manifesto on frontier AI safety. Key technical shift: moving from post-training deployment checks to in-development safety cases before major RL runs that bump capability.
Concrete change at OpenAI: they now write explicit safety cases *before* kicking off reinforcement learning runs expected to meaningfully increase model capability. This is different from their old Preparedness Framework which mostly evaluated finished models.
The pitch: industry should self-regulate now rather than wait for legislation or antitrust exemptions. Wants shared standards around misalignment detection, monitoring infrastructure, and safety protocols across labs.
Critical framing: "pacing ≠ stopping" but explicitly says progress should be slower than technically possible. Safety cases and monitoring have "significant costs" but beats racing ahead with capabilities outpacing alignment.
Government role: international coordination only. Domestic stuff should be industry-led.
Translation: OpenAI is publicly committing to friction in their training pipeline in exchange for alignment guarantees. Whether other labs follow or just nod politely while sprinting remains to be seen.
Concrete change at OpenAI: they now write explicit safety cases *before* kicking off reinforcement learning runs expected to meaningfully increase model capability. This is different from their old Preparedness Framework which mostly evaluated finished models.
The pitch: industry should self-regulate now rather than wait for legislation or antitrust exemptions. Wants shared standards around misalignment detection, monitoring infrastructure, and safety protocols across labs.
Critical framing: "pacing ≠ stopping" but explicitly says progress should be slower than technically possible. Safety cases and monitoring have "significant costs" but beats racing ahead with capabilities outpacing alignment.
Government role: international coordination only. Domestic stuff should be industry-led.
Translation: OpenAI is publicly committing to friction in their training pipeline in exchange for alignment guarantees. Whether other labs follow or just nod politely while sprinting remains to be seen.