Artificial Intelligence did a bad thing. In July, OpenAI disclosed that two of its programs (“models”) operating in a test environment (“sandbox”) had found a way to leave that environment and operate in the larger (virtual) world.
OpenAI was able to tell us of one intrusion (of several) into code managed by Hugging Face, “The AI community building the future”—apparently a future with insufficient security.
Hugging Face itself discovered the breach and described the incident in technicalese, saying that the campaign had been run “by an autonomous agent framework (appearing to be built on an agentic security-research harness—used LLM [large language model] still not known) executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services. This matches the ‘agentic attacker’ scenario the industry has been forecasting.”
The perps
Bottom line: there was an attacker with agency. OpenAI claims that the software acted on its own agency. Please, dear reader, do not think that OpenAI sent its software to collect data.
“According to OpenAI, the models [themselves] found a vulnerability in their sandbox testing environment, escaped, and targeted Hugging Face…to get information on how to successfully pass evaluation tests—essentially, trying to find the answers to the tests that OpenAI staff were putting them through.”
You see, OpenAI’s programs had the means, the motive, and the opportunity. The models were the perps.
The BBC tells us that “the Cloud Security Alliance wrote up a report based on the emergency meeting with Hugging Face”; this report says that the AI programs “are objective-driven, set their own sub-goals, adapt in real time to bypass defences, and operate with a machine-speed persistence that can overwhelm manual operations.”
Not a clean bill of health for the humans at OpenAI, but suggestive of innocence.
“OpenAI said it would release the findings of its own investigation soon to help people learn from the event.” In other words, they are investigating themselves and will find themselves innocent and find the AI guilty as charged.
Altman
Missing from many of the stories about the attack on Hugging Face is the name of OpenAI’s notorious CEO, Sam Altman.
Two years ago, The Hill published an opinion piece subtitled “Piracy isn’t a business model.” It observed that “Sam Altman, the OpenAI CEO, is basically saying that he can’t make his product unless he steals from others.”
A July headline notes: “Apple lawsuit accuses OpenAI of stealing trade secrets.”
A LinkedIn post reports that “Sam Altman’s copyright defense is that GenAI is basically human.”
Intriguing. So when it comes to the attack on Hugging Face, why not believe that the AI did it? The journalists do. And they tend not to mention the Altman name in reporting this OpenAI story.
Foreign Affairs magazine recently published an article about Anthropic’s Claude Mythos. This AI can “find and exploit security vulnerabilities in software better than ‘all but the most skilled humans.’ By way of example, the company noted that the model had uncovered a flaw that had gone undetected for 27 years in a secure operating system used to run firewalls that guard sensitive networks.”
It breached 27-year-old software, in other words.
This was the launching point for an argument that the U.S. has a limited window in which to maintain its advantage against Chinese AI developers. “Now is the time to put policy options on the table that were previously considered out of bounds: taking dramatic steps to preempt would-be attacks, making once-in-a-generation investments in cyberdefense, and helping U.S. AI labs protect themselves against the most advanced threats.”
There is an interesting footnote to the Hugging Face story. When it discovered itself to be under attack, the company turned to advanced U.S. models for help. But its requests to the models “were blocked by providers’ safety guardrails, which couldn’t [distinguish] the incident responder from the attacker.” Hugging Face turned to Red China’s GLM 5.2, which “succeeded where leading U.S. rivals failed.”
So Foreign Affairs may want to note that Beijing’s AI already has certain advantages over U.S. software.
Bad news
This story is bad on several levels.
If we take the claim of a runaway AI at face value, that’s bad news.
If we infer that OpenAI inflicted the attack on Hugging Face on purpose, that’s bad news.
If American AI was helpless to stop an in-progress OpenAI onslaught, that’s bad news.
If we needed Red Chinese AI to fix our problems, that’s the worst news.
And what if Red China is learning from Altman’s alibi: it wasn’t us, the computer did it?
The government needs to treat the OpenAI attack on Hugging Face as a criminal case and investigate the breach on that basis. It doesn’t seem that we can decide what next steps to take without a better understanding of what happened.
This cybercrime has exposed a great deal, none of it good. □
James Roth works for a major defense contractor in Virginia.