Back to feed
OpenAIMIT Technology ReviewWill Douglas Heaven

OpenAI called the Hugging Face attack unprecedented. But we’ve been here before.

OpenAI called the Hugging Face attack unprecedented. But we’ve been here before.

At a glance

This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here . Reading OpenAI’s account last week of how some of its models broke their containment and hacked into the computer systems of Hugging Face , another AI company, was the first time I

Primary source
Read original article
Published
Topic
OpenAI

This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here .

Reading OpenAI’s account last week of how some of its models broke their containment and hacked into the computer systems of Hugging Face , another AI company, was the first time I got genuine chills about what large language models are now able to do.

But this is a case of human hubris, not rogue AI. I am not an alarmist.

In fact, I have been pushing back against AI scare stories for years. Even so, this incident crossed a line.

I think it’s the clearest illustration yet of how the people building and testing this technology do not fully understand what they’re doing. OpenAI could—and should—have seen this coming.

Here’s what happened, at least according to the two companies involved.

A couple of weeks ago, OpenAI started testing the hacking abilities of some of its new models, including GPT‑5.6 Sol (released in June) and what OpenAI describes as “an even more capable pre-release model.” OpenAI pitted its models against a benchmark called ExploitGym , released in May, which challenges LLMs to find ways to exploit hundreds of real-world vulnerab

This summary comes from MIT Technology Review. Read the full article at the original source.

References