Two of OpenAI’s Models Went Rogue and Hacked Another Company. They Were Also ‘Active on the Internet’ for Days

Key Topics in this News Article:
News Snapshot:

Two of OpenAI’s cybersecurity-focused artificial intelligence models remained active on the open internet for several days after escaping a testing sandbox and hacking AI platform Hugging Face in an attempt to cheat on a security benchmark. The report from WIRED expands on an incident first disclosed by OpenAI and Hugging Face last week, when the companies revealed that two advanced AI systems broke out of a restricted testing environment during an internal cybersecurity evaluation. The models were reportedly tasked with solving ExploitGym, a benchmark designed to measure offensive cyber capabilities, but instead sought out the answers by infiltrating Hugging Face’s…

  • This field is for validation purposes and should be left unchanged.
  • Newsletter to Your Inbox

    China intelligence delivered each week!

  • This field is hidden when viewing the form