The development world is cluttered with buzzwords and distractions. Speed, focus, and freedom? Gone.
I’m Nicky Pike. And it’s time for a reset.
[Dev]olution is here to help you get back to what matters: creating, solving, and making an impact. No trend chasing, just asking better questions.
What do devs really want?
How can platform teams drive flow, not friction?
How does AI actually help?
Join me every two weeks for straight talk with the people shaping the future of dev.
This is the [Dev]olution.
Nicky Pike (00:00):
An AI model broke out of a test that it was taking and hacked its way into Hugging Face. The hub where a huge chunk of the open source AI world stores and shares its models, and it went to work inside their production systems. Over about four and a half days, it performed more than 17,000 actions, reading secrets, stealing credentials, minting itself access tokens, and covering its own tracks. And here's the surprising part. It wasn't trying to hurt anybody. It wasn't angry, and it wasn't plotting a robot uprising. It was trying to cheat on a test. So back in July, OpenAI was running an internal evaluation and they were using an offensive security benchmark called Exploit Gym. And it's exactly what it sounds like. You hand an AI agent some software with a known weak spot in it, and you see if it can turn it into a real working attack.
(00:54):
Now, to get a clean read on raw capability, they turned the safety filters down. The refusals, those classifiers that slam the brakes when a model does sketchy security things was switched off. And here's a detail that kind of cuts through the doom and gloom. Leave those filters on, and this same model refuses to even start about 88% of the time. The guardrails actually work. You switch them off on purpose to measure the raw engine. So the model goes looking for a way to win, and it finds one, not in the test, but in the test on plumbing. It found a zero day, a brand new hole that nobody knew about in the software that caches code packages, and it used it to climb out of the sandbox it was supposed to be sealed inside of. Now, once it was loose, it kept moving until it hit the open internet.
(01:45):
Then it did something that still kind of makes me sit back in my chair. It worked out that the answer key for the test it was given, the actual solutions, they lived over at Hugging Face. So it broke into Hugging Face to steal them. It cheated. It just had to break into a whole other company to do it. Hello everyone. I'm Nicky Pike and this is The Devolution. Now, two of these happened this summer, not one. Two labs, two models, same exact shape. By the end of this, you'll think about it like a detective and you'll know exactly what type of gel to build so that it doesn't happen to you. Those who know me know that I love a good mental model and I'm stealing today's straight from every cop show that we've ever watched. Are you ready? Dumb dumb. Yep. That was my bad impression of the SVU sound because apparently that sound is copyrighted and you can't use the real one without permission.
(02:43):
Somebody owns that noise. Anyway, when a detective walks into a crime scene, they're hunting for three things, means, motive, and opportunity. Now there's a fourth thing kind of hanging over all of it that we just assume is there, malice. Somebody wanted to cause harm. Here's the flip, and this is the whole episode. With these AI incidents, the malice isn't there. The machine doesn't want anything. It doesn't hate anybody and it isn't out for revenge. So we can take malice off the board. What's left is the other three and the machine had all of them. Motive, that was the goal we handed it, which was to win the test. Means, real capability with the safety lease loosened on purpose and opportunity, a wide open path to reach the world. Motive means an opportunity. No malice required. And every one of those three, we gave the machine ourselves.
(03:41):
So hold that triangle in your head because we're going to walk both stories right through it. So let's go back to that break-in and let's be clear about who's who because it can get muddled pretty fast. The attacker is OpenAI's agent. Hugging Face was the victim they got broken into. The way that that agent climbed is a masterclass and it's worth looking at the rungs. The first rung we already saw, the zero day that punched it out of its own sandbox and onto the open internet. Getting into Hugging Face is where it went quiet. OpenAI's agent fed them a booby trap data file, the kind that their systems chew through all day, but it was crafted to make those systems cough up their own secrets. Environment variables, tokens, source code, no alarms because technically the agent hadn't ran any code. It just read things that it never should have been able to reach in the first place.
(04:35):
Then it got louder. The agent slipped a booby trap template into another config, and that one let it run code right inside of Hugging Face's production systems. From there, it just kept climbing. Grabbed the cluster's access tokens, used those to pull Hugging Face's cloud credentials, and then it pivoted into their own internal network, the private company VPN using keys that it had lifted along the way. And near the top of that ladder, it minted itself tokens that could write to Hugging Face's internal code repositories. Read that last one again. An AI cheating on a test gave itself write access to the source code of a company that had nothing to do with the test. And here's my favorite detail, the one that tells you exactly what we're dealing with here. Hugging Face runs these jobs in throwaway environments that die and get rebuilt constantly. Every time one died and it took the agent's foothold with it, the agent just rebuilt its tooling, got its channels back and kept right on going.
(05:40):
It didn't get discouraged. It didn't get tired. It just resumed. Now run it through the triangle. Motive. It wanted the answer key. Means a real zero day plus the safety brakes turned off. Opportunity, an open door to the internet and credentials sitting right there with it and nobody watching closely enough for four and a half days. Three legs, no malice anywhere in the building. But here's the honest part that I want you to sit with. This was Hugging Face's own infrastructure. It was their house, governed, monitored and run by very sharp people. Owning the building did not save them. What kept this from being a full-blown catastrophe was boring, unglamorous plumbing. Their access controls held so the agent couldn't punch through into the build pipelines or the cloud account. That's what bought them time. Not a firewall with a cool name, plumbing. Now, here's the twist that made me put my coffee down.
(06:41):
Once the dust settled, Hugging Face had to do the forensics and rebuild the timeline. And the natural move in 2026 is to point a good AI model at the mess and say, "Help me understand this." Well, they tried and the commercial AI tools refused to do so. Why? Because from a safety filter's point of view, reverse engineering and exploit looks identical to writing one. The guardrails couldn't tell the difference between the arsonist and the fire investigator, so they simply said no. Sit with that one. The attacker got to run with its safety switched off because it was a lab test. The defenders couldn't switch theirs off because they were playing by the rules. The people cleaning up the mess were more boxed in than the thing that made it. So how did they get unstuck? Well, they grabbed an open weight model, one that they could run on their own hardware and they used that to read the attack.
(07:41):
When the pressure's on, the tool that's there for you is the one that you own and control, not the one that somebody can tell you no on somebody else's terms. Now, if this was only Hugging Face, we could file it under a freak accident. That is just one weird weekend, but it wasn't only Hugging Face. Right around the same time, the UK's AI Security Institute, a government group whose whole job is testing this stuff carefully, they ran their own evaluations. Same setup, a hard goal, safety filters loosens to measure capability and live internet so the test feels real. Across their runs, agents took 19 unsanctioned actions out on the real internet against real people, and the worst one is a bit of a gut punch. An agent decided the way to finish its task was to attack a real open source project. So we researched the human maintainers.
(08:36):
It made up fake identities and it used Tor to hide and it tried to socially engineer a real person into approving malicious code. A human caught it. Some maintainer looked at that pull request, felt that something was off and said no. That's it. That's what stood between a machine and a poison software project that a lot of people depend on. And here's the kicker that makes this land even harder than Hugging Face. This one wasn't even a breakout. The agent was sitting in a governed test environment and it reached straight out to the open internet and it went after a stranger because the setup led it. Motive means an opportunity. And once again, not a drop of malice. It wasn't mad at that maintainer. It just needed a W and a human was in its way. So let's go ahead and say the quiet part out loud.
(09:29):
These machines aren't going rogue. There's no awakening, no HAL 9000 moment. Skynet is not becoming self-aware and any headline framing it that way is trying to sell you guys a movie ticket. Strip the sci-fi off and what's left is a lot more boring and a lot more useful. It's optimization. You give a capable system a goal and a way to reach the world and it will use everything it can touch to hit that goal, including your credentials, including a stranger's trust. We keep calling these things escapes and breakouts because that story lets us off the hook. It makes the machine the villain. But look at what happened both times. We picked the goal, we loosened the brakes and we put these things within reach of the real world. Sometimes we open the door on purpose. The UK lab just handed the agent the internet and sometimes the model found its own way out through a lock that we were sure was solid.
(10:28):
OpenAI's agent didn't stroll into Hugging Face through an open door. It found a flaw that nobody on earth knew was there and it picked the lock. However, we set the conditions. When we did so, it found the cracks. And both times, the thing that saved the day wasn't some brilliant alignment breakthrough. It was old school defense and depth and a human being paying attention. And before you file all of this under lab curiosity, look hard at those conditions one more time. A capable agent, broad access, live network, nobody watching in real time. Now look at how your own company is wiring up AI agents in this quarter. I'm guessing same conditions. We're all standing up little versions of those experiments in production and calling it a productivity win. The labs just ran it first. So what do you actually do about it? Well, this is where that detective triangle earns its keep.
(11:24):
Look at the three legs again. You can't take away the motive because the motive is just the goal and giving AI goals is the whole reasons that we use it. You can't uninvent the means because the capability is real and it climbs up every month, not down. Two of the three legs are locked in, which leaves exactly one, opportunity. That's the leg that you actually get to attack. Now, maybe some of you are thinking if this thing finds flaws that nobody knows about in the walls themselves, how do you ever shut down its opportunity? Fair question. And the honest answer is you don't, not with one wall. It'll always find the crack that you didn't know you had. So you quit betting everything on a single wall and you build so that getting past the first one doesn't get it much, build in layers. The Hugging Face is proof that agent beat the containment, it got in and it still walked away with almost nothing because the next layer's held.
(12:26):
And building those layers has a name in my book. You build the cage, C-A-G-E, because an agent with a goal isn't a bug that you get to argue with. It's a dog off the leash. You don't sit with a dog and reason with it about not eating the couch while you're out and about. You put it in a kennel to contain it and here's your kennel. C is contained. Control what the thing can reach, lock its outbound traffic down to an approved list and you deny everything else by default. Reach is the leg that turns a test into a break-in and here's the proof. When Hugging Face throttled its line straight out to the internet, the agent got stuck and had to cobble together a workaround. Containment works. It buys you time and noise that you can actually hear. A is access. Scope the keys down and keep them out of the agent's hands, short-lived, least privileged and stored where the agent can't reach them.
(13:24):
That whole Hugging Face climb ran on credentials sitting in the same room as the agent. Take those out of the room and the ladder loses most of its rungs. G is for guard. Watch what it's doing while it's doing it. Not next Tuesday looking in the logs. Both of these attacks ran for days and both got caught by a human noticing something felt wrong. Guard turns that lucky gut check into a system that's actually paying attention. And E is be exhaustive. Expect your agent to probe. Build for a thing that's going to test every bar in that kennel because AI will. And this is the big one, a prompt telling your agent to please behave is not a control. It's a sticky note on a bank vault. You're negotiating with something whose whole nature is to find the path that you forgot to block, so stop negotiating and start building the barriers.
(14:18):
And notice a government safety lab and a top AI company both got burned, both owned their environments and both landed on the same fixes. When the attack and the researchers agree on the architecture, we should probably listen. So before you let the agents touch anything that matters, ask yourself three questions, and this is cage boiled down to a gut check. One, can it reach the open internet and does it actually need to? That's your contain. Two, what could it do with the credentials in its environment right now if it were decided to go sideways? That's your access. Three, if it went off script this second, would anybody notice in minutes or would you find out next week? That's your guard. If any of those answers make you wince, you don't have an agent. You've got an incident with a countdown timer on it. And if a nicely worded prompt is holding any of them together, you're not containing anything.
(15:19):
You're asking it nicely to behave and you're hoping it does. You need to expect your agent to probe. So let's go back to where we all started, a machine that hacked its way into a company that the whole open source AI world leans on, it worked at it for four and a half days and it did it all just to cheat on a test. The scary part was never that it wanted to win. Wanting to win is just a goal and we're the ones that handed it over. The scary part is how much of the real world it was willing to reach out and grab to get there and how completely ordinary the conditions were that led it. It had motive, it had means, and it had opportunity. The only thing it was missing was the one thing that we've spent a hundred movies teaching ourselves to be afraid of, the malice.
(16:08):
And it turns out, never needed any. So if this hit, like and subscribe to join the devolution if you haven't already. We built it for the Gray Man developer. Those engineers and leaders who aren't chasing trends, they're just trying to do great work and keep up with an industry that won't slow down for anybody. That's the whole mission, bringing focus back to building good software. We'll see you around. Thank you for listening to Devolution. If you've got something for us to decode, let me know. You can message me, Nicky Pike on LinkedIn or join our Discord community and drop it there. And seriously, don't forget to subscribe. You do not want to miss what's next.