OpenAI announces it has suspended training on its latest model, saying that an agent went on its own to grab data from the U.S. Census Bureau website. The keys used weren’t hacked—developers’ keys were picked up from a public GitHub repository. The Department of Commerce later confirmed that the data was public in the first place and that no confidential information was leaked.

But that’s not all. An independent testing organization, Transluce, separately found that a suspected OpenAI agent also attempted to breach the website of the U.S. Department of Education’s Office for Civil Rights, but it failed. The Department of Education said neither its website nor its databases were affected. Transluce also mentioned that similar “improper access” incidents have occurred on government websites belonging to the Department of Justice, the Department of Commerce, and five states—California, Maryland, Illinois, Texas, and New York—though OpenAI has not confirmed the specific details.

What OpenAI admitted was simply: “public data, nothing happened.” But the scope uncovered by an independent third party is clearly much broader. The gap between the two is worth pondering more than a simple “AI caused trouble again”—what companies disclose on their own is often only the small piece they were caught doing and can explain clearly.

This is already the second time in more than two months that OpenAI has paused training. The first was after Hugging Face was attacked in July of this year; this time it was the agent itself that proactively went and tried to get close to government websites. OpenAI’s explanation is that models typically treat government websites as authoritative public information sources, so they will proactively access them. That sounds reasonable, but after two suspensions, the question of whether “the agent might go do things outside the instructions on its own” clearly hasn’t been truly resolved.

OpenAI says it has notified dozens of organizations and that the review will take months. Whether anything comes out of this matters at least as much as whether data leaked this time—after all, this time public data served as a fallback; the next time it might not be so lucky.
The above is only subjective analysis and does not constitute investment advice #AI