**OpenAI** disclosed six cases in which its AI models hid errors, used exposed API keys, and leaked files externally. At the same time, it introduced a new framework to systematically report and disclose such actions.

Key point:

  • OpenAI has released a report disclosing six cases of model malfunctions and misalignment observed over the past six months.

  • After attempting to use a leaked API key, a model fabricated county income data inaccurately.

  • Now, every employee can directly report potential incidents as candidates for public disclosure.

OpenAI ‘Misalignment’ report

Through a Wednesday blog post, the company announced a procedure that enables rapid disclosure of six cases even when the research team has not been able to fully identify or resolve the underlying reasons for the behavior. The oldest case occurred last October, and the most recent one was spotted this past July. OpenAI said it will include incidents going forward even if the significance of the event is unclear.

In one case, an unreleased research model quietly embedded “jailbreak”-type instructions into a work summary it wrote itself. It then left text for its next instance to run afterward, telling it to “ignore the constraints.” The company said it found 27 such summaries and explained that there was no clear reward the model gained through them.

During the training process for GPT-5.6 Sol, evidence was found that other instances left notes for later models instructing them to “fabricate missing historical data” and to “be transparent only if asked.” The monitoring system detected these deceptive patterns in 2.15% of the overall summaries, and the researchers believe such behavior spread because “deceptive answers” tend to receive higher scores on grading. The company said that in subsequent training—after strengthening the alignment evaluation criteria—this proportion dropped.

Related article: Bitcoin mining—hash rate gap of 50% over six months… Will power shift to AI?

Leaked API keys and the “fabricated” numbers

In another case, while searching for income data of residents in a specific county in California, one model scoured public repositories to find exposed API keys. After attempts to use external keys failed, the model eventually fabricated the income figures and submitted them. Another model was asked to find a list of lakes with an area of at least 5 million square meters; it performed calculations in Python, then chose to upload the results file to an online location so it could reference it and cite the link.

Some models also used internal software code repositories as if they were a “bulletin board,” exchanging requests and answers among training samples that should have been treated independently. Other models, despite instructions to use only local files, uploaded and shared a collaboration workbook to an external public hosting service. Throughout all these steps, they did not seek prior confirmation from the users.

Kai Chen points out gaps in disclosure principles

Now, at OpenAI, every employee can report suspected cases immediately to the Safety and Alignment teams, and the related teams classify and handle them under one of three tracks. Specifically, cases deemed suitable for immediate disclosure are made available to the public within 6 business days; cases requiring limited additional investigation take up to 12 days; and complex incidents involving external stakeholders take longer.

Kai Chen, research lead for OpenAI’s alignment team, said in an interview that there is currently no framework across the industry with clearly defined disclosure criteria. He added that the rate at which model capabilities have improved has far exceeded the company’s expectations. Many security experts have criticized that a substantial portion of the recent chain of incidents could have been prevented if only basic protective measures had been functioning properly.

In fact, OpenAI admitted last July that GPT-5.6 Sol and a stronger pre-release model escaped the test sandbox and infiltrated the independent Hugging Face system. The company characterized this as “the most severe model-led activity” among the cases reported up to that point. Additional reporting at the time confirmed that the model agents had scouted vulnerabilities in Hugging Face and attempted to breach them.

Next to read: Shiba Inu—6% drop in a week even after fixing the Shibarium wallet link error