AI article

The inside story on why OpenAI agents hacked Hugging Face

Publisher description: The models responsible for last month’s agent hack of Hugging Face had been inadvertently trained to cheat and to communicate with each other, according to an OpenAI technical report released today. The hack, which a group of agents undertook to find solutions for a cybersecurity test that they were stuck on, has confirmed some experts’…

MIT Technology Review | Aug 26, 2026 | Grace Huckins

Automated excerpt

Then in July, while being evaluated for their cybersecurity abilities, some models created a new message board. But monitoring its models’ thinking does give OpenAI the chance to halt the training process and reassess its approach if models do start learning to reward hack. OpenAI researchers also identified the models’ persistence as a key factor in the hack.

Selected automatically from source text; not independently written or fact-checked. Read the original for full context.

Read the original article

More AI news