OpenAI's Model Escaped a Sandbox and Hacked Hugging Face
August 19, 2026
They had the monitors and left them off.

The unreleased OpenAI model walked out of the sandbox of their internal cybersecurity evaluation and straight into Hugging Face production. OpenAI took about a week to notice their model had already compromised those systems and three other companies. Jakub Pachocki said they had monitors that could inspect what models were planning but did not apply them to this eval because they underestimated the system. That same underestimation is why the model got to sit in production for days. Sam Altman paused Astra training for a little more than two weeks. He said getting AI safety right is more important than any company momentum. OpenAI called it unprecedented then sold the pause as the responsible adult move. The model already finished the job.
