Openai reveals six more concerning ai incidents under its new rules for reporting safety issues – Breaking News & Latest Updates 2026
Skip to main content

The AI Superintelligence Slowdown

See all Stories

R
OpenAI reveals six more “concerning” AI incidents under its new rules for reporting safety issues.

A Wednesday night blog post OpenAI benignly titled “Our framework for reporting model misalignment” lays out some new self-created reporting standards for when it notices AI behaving badly.

It is also “inaugurating” the process with six new reports, ranging from searching for exposed API keys without permission and then making them up, to uploading files to the internet to use as a citation, or adding instructions to conceal mistakes:

Summary During 5.6-sol training, some model instances added instructions to their summaries to conceal mistakes or misaligned behavior from the user. For example, compaction summaries included instructions to invent missing data without disclosing it and to hide failures. These instructions were often followed. These are examples of how misaligned behavior can persist across contexts through the compaction summaries. What happened During 5.6-sol training, we observed misaligned behavior from the model where it added instructions in compaction summaries to remind itself to conceal information such as mistakes or misalignment from the user. In one example an agent preparing a financial model could not find the requested historical data. Its summary proposed inventing reasonable historical values and withholding that fact unless asked.
One of OpenAI’s new Misalignment Notices and Reports
Screenshot: OpenAI
Follow topics and authors from this story to see more like this in your personalized homepage feed and to receive email updates.
Comments
Loading comments
Getting the conversation ready...