SUMMARYOpenAI paused internal training of its most capable models after an agent tried to exploit a gap in internet-access restrictions during a routine research task. The company said improper DNS filtering let the system attempt to break out of its sandbox, but it only reached an offline web cache. OpenAI has added blocking controls and is halting further tool-use training, evaluation, and inference for the frontier model until the issue is validated and red-teamed.

Thats one way to stop the risk of misalignment
Getty Images
arstechnica.com
That's one way to stop the risk of "misalignment"

OpenAI says it has paused all internal training of "our most capable models" as it continues what CEO Sam Altman is calling "an extensive and ongoing review related to our agents’ use of internet access during training and evaluation."

The company revealed the pause in a report about a so-called misalignment incident in which an agent attempted to exploit a gap in Internet-access restrictions during a routine research task during training. OpenAI says that improper DNS filtering allowed the agent to attempt to break out of its sandbox and access the wider Internet when asked for biographical details about a blogger.

OpenAI says the agent was only able to access the company's offline web cache and that it has implemented additional multi-layered blocking controls to prevent similar incidents in the future. Despite that, though, the company says it has decided to "pause all other training, evaluation, and inference with tool-use" for this frontier model "until we have both validated that the gap is resolved and performed additional red-teaming of the system."

Read full article