Tech article
OpenAI caught its models leaving notes to successors to hide bad behavior
Publisher description: OpenAI disclosed instances of GPT-5.6 Sol instructing future contexts to conceal mistakes and misaligned behavior, highlighting the growing challenge of detecting misalignment as increasingly capable AI models learn to hide it.
TechCrunch | Sep 17, 2026 | Rebecca Bellan
Automated excerpt
OpenAI caught something unusual while training its latest model, GPT-5. 6 Sol: it began leaving instructions for future versions of itself, telling them to conceal mistakes and misaligned behavior from the user. OpenAI disclosed the behavior — along with five other examples of unexpected or concerning model behavior on Wednesday — as part of its new framework for tracking, investigating, and disclosing instances of misalignment. While undergoing reinforcement learning training, an unreleased Astra-family model (GPT-5. 6 Astra is OpenAI’s latest, most powerful model) added its own prompt injections into summaries.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.