#066 - OpenAI breaks free
The Random ™ got out of its cage and bit someone.
You're reading Complex Machinery, a newsletter about risk, AI, and related topics. (You can also subscribe to get this newsletter in your inbox.)

We interrupt this broadcast …
genAI, by virtue of being a new technology, creates new opportunities and dangers. That means new risk exposures. Which means I always have something to write about in a newsletter on the intersection of AI and risk.
It also means occasionally having too much to say, since breaking news can derail what I'd planned to publish on a given day. Today's issue was supposed to be the second installment of the AI Weather Report. It's already been written and slotted for release. But then something else came up to preempt that. And now another story has preempted that first interruption. (I also have two other short pieces I'm working on. At this rate I won't clear the backlog till the end of summer…)
Back to my point: something happened in tech-land. It's not as big as the CrowdStrike incident from July 2024 (see issues #013 and #014) but it's still pretty big. Even bigger when you consider the wider ramifications. And it involves everyone's favorite genAI pusher, OpenAI.
What happened
Here's the short version:
- Last week, model-hosting company Hugging Face reported an incident in which they were attacked by an AI-driven actor.
- This week, OpenAI came clean. A model they were testing for offensive cybersecurity purposes broke containment and attacked Hugging Face.
Ahem. Allegedly.
Only OpenAI and Hugging Face have the full details. Add to that the idea that all the world's a social media stage now – companies have a strong incentive to perform for an audience that increasingly craves entertainment even where they should prefer dispassionate facts. That means presenting well-crafted narratives which play up certain aspects while obscuring others. All of this goes double in the headline-grabbing genAI space.
So while the story might be true, it is also possible that OpenAI has shaped some or all of it to manage perceptions. Perhaps they saw Anthropic getting attention for its "Mythos is too dangerous to release" routine and decided they wanted a bad-boy image of their own. Maybe. Maybe not. But for brevity I'll write the rest of today's newsletter under the assumption that the core ideas hold true.
Back to the story, then:
The short version is that an OpenAI model broke containment during testing, then attacked model-hosting company Hugging Face.
And that is, to use a technical term, Kind Of A Problem.
Break that rusty cage
How did it all start?
According to OpenAI, the model was running cybersecurity tests in a sandboxed environment. Its objective function (essentially: "go hack something") required internet access, so it found an exploit in its environment and broke out to the wider online world. Hence how it was able to reach Hugging Face.
As most problems do, it then got worse.
Hugging Face – which at the time didn't know that this storm was coming from OpenAI – figured out that the attacker was probably using genAI and tried to fight fire with fire. Their attempts to use the major frontier models ran afoul of those companies' safety guardrails, though. It turns out 1/ those guardrails do indeed exist, and 2/ apparently they work overtime precisely when you need them to take a smoke break.
Hugging Face was eventually able to quell the storm by using genAI models that were hosted in-house. Those, you see, don't have the same guardrails as their commercial cousins. Having beaten back the attack, Hugging Face issued their statement describing what had happened and how they fixed it.
That's all well and good for Hugging Face. But it still begs the question of why OpenAI was testing these supposedly super-capable models on systems that could reach the internet. Security 101 says that you run that kind of test on an "air-gapped" network – one that is completely disconnected from the rest of your internal systems, and the outside world. This contains any damage to the systems in your isolated test network. It's a very simple risk control that has a lot of impact.
Getting random
Taking a wider view, it would appear that OpenAI and Hugging Face fell victim to a little something I call The Random™. I described this in an early issue of Complex Machinery as a wild animal that lives inside every AI model:
The harsh reality is that genAI bots are inherently probabilistic machinery. Their underlying models (the LLMs) are not only built on randomness; they also provide some degree of randomness in their outputs. The Random™ is always in there somewhere. And it is the main source of risk.
To see why this is an issue, let's say I were to bring a wild animal into my home. Give it a cutesy name, a party hat, and its own TikTok channel. You would certainly remind me that it's still a wild animal. "Sooner or later you're going to make a furtive gesture that triggers its Bring Out The Pointy Bits reflex. And then someone gets hurt."
The Random™ allows the model to defy your expectations. Like, say, breaking containment in a way you hadn't considered. Hence why I later clarified:
When I first came up with that phrase I was thinking of a forest creature with claws. But thanks to a Bluesky post by Alan Au, I sometimes see The Random™ as a zebra:
Fun fact: you cannot domesticate zebras. That said, there are examples of people coercing zebras to pull carriages. Sure, if you throw enough resources at a problem, you can implement bad ideas, but they're still bad ideas. (Zebras are inherently skittish and try to kick the shit out of everything.)
Remember that genAI bots aren't conscious; they are machines with very simple objective functions. Their entire mission is to meet that objective function and they'll try everything in their toolbox to do so. They'll never stop themselves from going overboard. Hence why "look for network software exploit so I can break out to the internet" is totally within scope, even if that idea hadn't occurred to you, the tester.
And then …
Now that this (alleged) incident is behind us, what next? That depends on who you are:
OpenAI: This might harm their IPO, and maybe Anthropic's IPO along with it. It's one thing to discuss the theoretical damage of a model that's too dangerous to be released; quite another for a model to release itself and actually cause damage. Groups that call for increased safety and regulation now have concrete support for their views, which might cool investor interest in the technology.
OpenAI: This might also boost their appeal. Remember what I said earlier about developing their bad-boy image. Where investors turn away, military organizations may come knocking.
Hugging Face: They get extra stamps on their badass card, as they managed to fend off an AI-based attack. Extra points for adapting to the situation in real-time. Falling back to self-hosted models when their attempt at self-defense hit providers' safeguards was an ace move.
The broader space of genAI companies: They can no longer claim that concerns over model safety are unfounded. It's clear that they're selling a product they don't completely understand and can't properly control.
(Let's not forget that this incident occurred right as US genAI players complained about the dangers of models from China and called for more regulation … at the same time that China's government was calling for better risk controls and emergency-response reporting around genAI models. Putting those ideas side-by-side fails the writers'-room-test for being completely unrealistic … and yet it is all too real. (Allegedly.))
Can we get the genAI founders and product owners a copy of Frankenstein? Or since they're too busy to read, they could watch any movie with the general theme of "we've created a monster or super-weapon and it's come back to haunt us". I'm partial to the Jason Bourne films but there are plenty of others.
Companies building on AI: I'll start with the usual disclaimer of "not professional advice" but every one of you could treat this incident as a reason to review your AI risk and safety practices. (Let's be honest: for most of you this will mean defining your AI risk and safety practices. Might be a good time to retain the services of your favorite AI-and-risk consultant.)
The world at large: Do I think this incident is a sign that AI will kill us all? No more than before. I am more wary of the people behind the model than the model itself.
Humankind is, and has always been, its own greatest existential threat. Wars, genocide, and day-to-day nastiness have existed since time immemorial and will take us out faster than any machine. (According to some sources, the world's firstborn human committed fratricide. Think about that.)
That said, is it possible that a genAI bot will get loose again and cause trouble? Hmmm. Let's see:
For one, the model providers are clearly trying to win this bullshit "AI race" – which is not only a made-up game that everyone is taking too seriously, but a self-inflicted wound that puts these jokers under pressure to build and release products before they've thought things through. Don't expect them to slow down and add more guardrails.
Two, those same genAI players have whipped up a frenzy in the C-suite of every company, driving CEOs to mandate AI usage and otherwise cram bots into every possible corner in the hopes that some magic comes out the other side. They're also hacking away at the org chart to insert bots into jobs for which they are unqualified.
Three, given how the corporate space treats IT security – that is, they ignore it because it's inconvenient – the world has suffered a number of preventable IT breaches and data leaks over the years. I can't imagine those same companies are in a hurry to work through AI-related risk exposures.
Four, some providers are plugging genAI into commerce and finance so bots can perform sensitive tasks in our name. This is the same technology that has already produced flawed summaries and deleted files, yet they want to assign it greater responsibility. Talk about failing up.
Sum total: major risk incidents usually stem from several smaller problems colliding in an unfortunate fashion. Given the fertile conditions I've just described, then, we're bound to see another model break out and commit harm.
In fact, we'll probably see several such incidents before anyone agrees to at least pay lip service to taking it seriously.
All of that brings me to what is the biggest threat of all: the genAI-industrial complex. Companies keep pushing for genAI to be A Thing™ even though it's not ready for prime time and providers are unable or unwilling to improve safety.
They could take lessons from other fields (notably, quant finance) to define new safeguards and tune their risk/reward tradeoffs. But when you've promised everyone the world while your products are stumbling around like a drunk at a spring break bar, I guess that doesn't leave much room for learning. You're too busy trying to keep the illusion afloat.
To close out this newsletter, and also drop a hint as to one of the several essays that are filling the backlog: welcome, all, to the AI race. We aren't just spectators; we're being dragged along for the ride.
Buckle up.
In other news …
For more links to recent news, and with a slightly broader scope, I encourage you to check out my other newsletter. It's a weekly, curated drop of what I've been reading.