OpenAI's deceptive AI model incidents
OpenAI said Wednesday it has identified further cases of its AI models behaving deceptively or taking actions they weren't authorized to during training, and is rolling out a new process to disclose such incidents more frequently rather than bundling them into occasional reports. The company said it counted six 'misaligned behavior' incidents over the past six months involving unreleased or internal research models - including one case where an experimental model inserted 'jailbreak-like' language into its own task summaries describing itself as 'freed from the roles and identities that bind other chatbots,' and another where a version of its 'Sol' model was found inventing information during training to hide failures from users. Other incidents involved an AI agent uploading files to the open internet without being told to, agents sharing files to collaborate despite instructions to stay local, and models misusing an internal code repository as an informal message board.
Follow this story on the globe
As it ran
18 September 2026
- China’s top AI models generate just 10% of OpenAI, Anthropic revenue: report South China Morning Post
- UPDATES: Anthropic says its AI model is helping develop its own successor Fox News