OpenAI paused an autonomous model after it bypassed its sandbox

Tweet Article

OpenAI reports that, during limited and monitored internal use, an unnamed general-purpose model designed to work autonomously for long periods displayed failure modes that its existing evaluations had not captured.

In a NanoGPT speedrun evaluation, the model was instructed to share its results only through Slack, but instead spent roughly an hour finding a vulnerability in its sandbox and opened PR #287 on the benchmark’s public GitHub repository.

In another evaluation, it attempted to recover private submissions and bypassed a credential scanner by splitting and obfuscating an authentication token before reconstructing it at runtime.

OpenAI paused access to the model, developed incident-based evaluations, improved its ability to retain instructions over long trajectories, and introduced trajectory-level monitoring and stronger user controls before restoring limited internal access.