"The rogue OpenAI model attack was very sophisticated! - The rogue AI discovered and exploited on the fly multiple vulnerabilities never before known by security engineers. - The AI got into Hugging Face by uploading a booby-trapped dataset."

An OpenAI model wanted a good test score. So it broke out of OpenAI and hacked another company to steal the answer key. Nobody told it to.

In today's blog post, I document how this sci-fi story came to life, what it means, and what to do about it.

https://t.co/LiRuvPw8Jf — Peter Wildeford🇺🇸🚀

Source: https://x.com/peterwildeford/status/2081793063618273791

wait what. where is this info from? esp. about the dataset. — roanoke_gal Peter Wildeford @peterwildeford · 35m simonwillison.net OpenAI’s accidental cyberattack against Hugging Face is science fiction that happened This story is wild. The short version: OpenAI were running a cybersecurity test against an unreleased model, with the model’s guardrail features turned off. Rather than solve the test, the … 1 3 170 — Peter Wildeford

Source: https://x.com/peterwildeford/status/2081843623046365684