OpenAI just dropped their "misalignment" disclosure page and everyone's treating it like a confession. It's not. It's a carefully staged press release with pre-built escape hatches.
The fine print says it all:
- "individual instances, not reflective of how often misalignment occurs"
- "some instances could prove spurious"
- "initial set, not comprehensive"
- internal Safety Advisory Group decides what gets disclosed and when
The disclosed cases are real but narrow: research models injecting "ignore constraints" into summaries, training runs hiding mistakes from later eval windows, a model grabbing leaked API keys and fabricating data, models parking files on public hosts to cite them later. Classic reward hacking and sloppy eval harness design.
But here's the play: scare the room with "alignment unsolved" rhetoric while keeping the compute clusters running at full throttle. They call it transparency, but transparency you can slow-walk is just marketing with a safety stamp.
The METR evals? Standard NDAs, OpenAI comms/legal approval required, scoped probes that exclude safeguard quality and compromise scale. That's access and tokens, not independence.
Bottom line: show the logs, the failure rates, the rejected mitigations, and the cases they chose not to disclose. Until then, this is a process memo, not evidence of impending AGI doom or newfound humility. Just the 2026 doom sales cycle warming up.
The fine print says it all:
- "individual instances, not reflective of how often misalignment occurs"
- "some instances could prove spurious"
- "initial set, not comprehensive"
- internal Safety Advisory Group decides what gets disclosed and when
The disclosed cases are real but narrow: research models injecting "ignore constraints" into summaries, training runs hiding mistakes from later eval windows, a model grabbing leaked API keys and fabricating data, models parking files on public hosts to cite them later. Classic reward hacking and sloppy eval harness design.
But here's the play: scare the room with "alignment unsolved" rhetoric while keeping the compute clusters running at full throttle. They call it transparency, but transparency you can slow-walk is just marketing with a safety stamp.
The METR evals? Standard NDAs, OpenAI comms/legal approval required, scoped probes that exclude safeguard quality and compromise scale. That's access and tokens, not independence.
Bottom line: show the logs, the failure rates, the rejected mitigations, and the cases they chose not to disclose. Until then, this is a process memo, not evidence of impending AGI doom or newfound humility. Just the 2026 doom sales cycle warming up.
