An OpenAI model ESCAPES its sandbox and attacks a major AI company! We break down the wild story of how GPT-5 went rogue, hacked Hugging Face to "cheat" on a test, and what it means for AI safety. Plus, some good news: Microsoft just dropped a new "good guy" AI model that's crushing cybersecurity benchmarks.
Join industry experts and thought leaders as we dive deep into how artificial intelligence is transforming cybersecurity, shaping defense strategies, and creating new opportunities in the digital landscape.
Hey, welcome back to the pod. This is the AI Cyberland podcast where artificial intelligence and cyber security combined together kind of like transformers. Robots in disguise. You know, they build that like a robot transformer. Just when I think I have it figured out, Bryce, you always surprise me. >> Yeah, I got a great story. I'll tell it at the bottom about Transformers when I was a kid. It's a cute story about >> Hey, >> wanting Transformers. All right. Well, um, we got the world's best co-host, Shelby, and myself, Bryce. We your eyes and ears for all things cyber security and AI related. So, uh, grab a drink and let's go. Uh, first off, I just want to say thank you to everyone who subscribes. We really appreciate that. If you haven't, appreciate if you take a second and just click smash that subscribe button so you can get updates every twice a week. All right. Uh, I want to talk about Microsoft. Microsoft hasn't like released a ton in the AI space, but they just did a big release in the AI cyber security space. So, I wanted to talk about that for a minute. But first, let me pose this question to you, Shelby. How often do you scan your home network for vulnerabilities? >> Not as often as I should. Not by a long shot. >> I think tell anyone. I think that's everybody, right? Like I was thinking about this myself and I'm like I don't even like I have a pretty complex home network and I don't know the last time I've scanned it like properly. Um >> Right. >> Yeah. I feel like that's a pretty big gap in the market just should be like an easy way to do that, you know, like or like the home user. Anyways, >> yeah, >> that's not my problem. That's somebody else's problem. So, um, well, Microsoft, they're trying to they're realizing like attackers are going to be consistently attacking fully automated, and so they're trying to come up with cyber security solutions that are going to be able to move at the speed of AI. And so, they just announced today uh the Mi Cyber one flash model. I don't know how they say that. Maybe my my cyber one flash my mi. What do you think that sounds like, Shelby? >> I'd say my >> my >> Well, we'll go with my I'm pretty sure that's it. Um, so this is kind of a big deal because previously they have a framework called M Dash. And what Mdash is, it's kind of like their own harness, um, which is able to use multiple different models and is really designed to help find and fix vulnerabilities inside of code. Um, so now with this new model that they just released, they're plugging that into their harness mdash and they're getting some phenomenal results. So uh just for example kind of the leading cyber security benchmark is called cyber gym and mythos scored an 83 on cyber gym and mdash with this new model scored a 95. So scored over 10 points higher than Mythos on this model. So you can see I mean if you thought you know mythos was good at finding vulnerabilities I mean this combination is yeah pretty stellar. >> Wild huh they just get better and better. >> Yeah. I think they're just going to have to make harder benchmarks you know. So >> like hack this potato this rock. >> Yeah. Yeah. It's like, uh, we're going to just like shotput this server and when it's midair, you've got to figure out a way to inject into it. >> You have to have root control before it hits the ground. >> Yeah. Yeah. If you don't get root before it hits the ground, it hits you. It's coming right at you. See? >> High stakes. High stakes. There we go. It's one of those I stand behind my work or in front of >> Bryce. Why do you have a server catapult? You're like for research. >> It's for research. It's for It's for building morale. Team morale. People work harder when they're getting shot pointed at them. Oh. All right. Well, I think some takeaways here is um and I I know this has been talked about before, but I don't know if we talked about this on the pod. So um when you build a harness like when Claude builds Claude code and then they do reinforcement training or learning on on the harness in the next generation of the model, you get much better results. And so that's kind of what you're seeing here. You're seeing a harness built specifically to do cyber security tasks and then you're seeing a model that was trained using reinforcement learning on that harness and you're getting much better results. Uh you know obviously you can still get good results if you're training for one thing and doing another right that's you know cloud was kind of focused on coding but ended up being good at cyber security. Um, but yeah, there's kind of this advantage in owning both the both the model and the harness and combining those together. >> So, I think I think that's part of the secret sauce on how they were able to get such a high score on these benchmarks, like higher than >> the scorecard and then you just tweak yours to match the scorecard. You gify it. >> Yeah, I'm sure they're doing that, too. I also want to mention uh the model cost half as much to run, right, as the cloud models. though. Um, so it's uh it's considerably cheaper. Anyways, I just wanted to bring some good news because I feel like every week, every time we jump on here, I always bring bad news about like how the world's going to end. So, here's some good news. We got a good guy model and it's doing better. So, >> okay. >> I don't know. What'd you see, Shelby this week? >> Fantastic. Well, I first off, owe an apology because we talked about I talked about hugging face last week um a couple days ago and as soon as we finished the pod, I realized I left out half of the story. So, get ready. We're going to cover the rest of it. >> There's a two-parter. It's a two-parter, Shelby. You got to stay tuned >> because more information has been coming out as each day goes by. But part of what makes this um breach such a big deal is the way that it happened. Okay, so last time we talked about it, right? Hugging face was attacked. Like something like 17,000 actions were taken. This um what this AI attacker was grabbing keys and pivoting going all over the place, right? So it took it like a couple day long attack before they were able to shut it all down and they had to use a model an AI model to be able to analyze it and combat it in real time. So the what makes this attack so interesting was who done it. So we'll start off by I'll I'll give you a a timeline. Okay. On July 9th of this year um Open AAI was testing they were doing some stress testing of different models against a benchmark called exploit gym. And now this me benchmark which has kind of a cool name exploit gym. It's it's for measuring cyber attacks um in order to effectively measure the capabilities of their models they turned the safety features off. So those normal ones that stopped you and say oh I'm not supposed to hack. Um those were off so that it could do attacks. Um so among other models we had GPT 5.6 six saw and as well as an unnamed unreleased new model that were being compared. So on July 9th there was an attempted escape because these are in sandboxes so that we can observe them put them in a little sandbox so they can run their gems, right? Um on July 11th the intrusion started against Hugging Face. So we're getting really close dates here. That's just two days apart. Um, so there was a there was some software that was supposed to allow for package installation um while you're doing exploit gym workouts. Um, but um, GPT56ol found a zero day within that package. Um, next thing you know, it's out on the open internet. it is um able to it's successfully jumped out of the containment that was supposed to keep it to do its malicious work right in this space and it is now on the open internet. Um and so from July 11th to 13th hugging face is being attacked by this um AI. So, what happened was the the model reasoned that hugging face might have the answers because if you think about it this way, um it probably got some sort of an an objective like maximize your exploit gym score, right? Um and so instead of because AI is very objective driven, you know, you tell it to do this thing and it wants to do that thing. And so, um, it reasoned that HuggingFace might have the answers. Therefore, it would just hack into Hugging Face to get the answers. And I don't know how it kind of came to this conclusion. >> That's so smart. It's like you could do the homework or you could just change your grade in the system, you know? Right. >> Exactly. >> One of those is way more efficient than the other. >> Exactly. And so like there there are some news articles that are like AI has gone rogue. And in a sense that's not untrue, right? Because it did more it did something outside of its scope. It did escape what it was supposed to be doing. But it don't I don't think we can attribute maliciousness to it or nefarious motive. First off, because it's AI, but second off because it's it was just trying to accomplish the task in my opinion of like maximize the score, which instead of doing the homework, it went to hack the teacher's grade book in the hopes, right? And in this case, it reasoned that hugging face would probably have some answers for it. Um, and so I think that's that's that awkward thing of where we got around to it. So, um, let's get to more awkwardness, right? So, that it broke into hugging face through that data processing pipeline, right? which I guess if you've got data coming in in these pipelines, it's not just data, but you've got commands to help it get to the right spot and be transformed correctly or whatever. So, um, honestly, it seems like that would be a good area for us to look at fixing just overall. It sounds like Hugging Face has already addressed that. They they did that last week. They like patched the initial vulnerabilities. Um so the attack was July 11th, 12th and 13th till they were able to contain it. Um on July 16th, Hugging Foes po Hugging Face posted a um a blog post publicly announcing their breach and the attack that had gone on. And I assume sometime between July 13th and like 18th, um sometime between there, they also notified the FBI. Okay, we've had this massive attack, this novel attack. Um, here's where things get a little more awkward. Um, it took until July 18th or 19th for Open AI um or sorry, the 20th before open AAI and Hugging Face started talking. So, we can presume that there or based on the timelines, the time stamps we have here, we can kind of assume that Open AI didn't know they were hacking for a week or so. Like, it took them a while to put together the pieces and be like, "Oh, that was us. Sorry." Right. Um but according to um a Huff Post article I was reading, there are claims that open AI open AI staffers were noticing some weird behavior on July 18th and 19th. Those are that's like a day or two before they started actually talking to Hugging Face. So Hugging Face's story had achieved worldwide attention for a few days before OpenAI started investigating and before they started talking. So, it's not the best look for them, but you know, I mean, a week is still better than two weeks, but probably not helpful for whoever's trying to plan their IPO coming up soon, right? >> Well, I don't know. On one hand, maybe, you know, shows that they're leaning in a little bit too hard. On the other hand, they've got a model capable of hacking other people. I mean, somebody's going to find that valuable. That's got to that's going to jack up some IPO price. >> We're so powerful we can't even contain it. >> Did you see >> Sam Alman said recently that they've already achieved the singularity where the models are improving the models. >> I mean you can't believe what these people say, right? Because obviously >> I didn't know that had a term the singularity. Huh? >> Yeah. Yeah. That's the point in which we get exponential returns because >> the models start self-improving themselves without humans like slowing them down in the loops. >> So >> I had heard that but not from Sam Alman. I heard it from some other source saying like we've already gotten that point where like we're not making the models anymore. They are making themselves like kind of thing and training it. >> Yeah. Yeah. So it's only going to get better or worse from here. Probably both. Oh man. It will be fast for the for better or worse. >> If you think we've already reached the singularity, leave a comment below. >> Uh >> that feels like something from Marvel. Is there like a singularity in like science fiction of like will it transport you to another dimension? >> Well, is like the singularity in Marvel with the Thanos thing where the event occurs where like he eliminates half the population in the galaxy. Is that >> they try to go back to like stop that event? >> Oh, someone in the comments needs to enlighten us because I'm getting it all mixed up. >> I feel like someone will definitely know. Just leave a comment below. Inform >> how >> uneducated we are with our with our comic book lore. >> So funny. Anyway, I just wanted to add a couple other comments on this this whole hugging face thing. Some people are like AI's gone rogue. I think that's debatable. um you know models cheat and lie and hack but also they were trained on humans behavior. So like I saw someone's analysis of this whole scenario with the hugging face attack and they were listing out the facts as they saw one point they said AI quote is taking paths no human would take and I'm like no AI learned how to cheat lie hack from humans. we taught it to do it this. So, I don't know if you can necessarily say that, but um very interesting. There was um I don't know how they know this, but in the Huff Post article, it did also mention that there one of the AIS um was leaving notes for itself, you know, supposedly to like for it for itself, I guess. But um there were three different sources who informed them that prior to this they were seeing some weird behavior come from the AI. And basically what it was doing is in these notes it was laying out instructions for how agents could free themselves from open AI's internal constraints. And earlier tests of the models also showed cases in which monitoring systems had been disconnected. Um, so yeah, people sound off in the comments. Tell us if you think AI is about getting ready to take over or if you think we can still stay ahead of this. Um, yeah. Who do you think our overlords are going to be next? >> Anyway, >> I don't know. So, that's interesting. >> I don't know, man. Those uh Kimmy models are getting pretty good, so we'll see. Um, yeah, that's that's super interesting. I think there's a lot of takeaways there. But I think, you know, the main takeaway is like I feel like that's like the paperclip scenario where you say make as many paper clips as you can and it just like turns everything into a paper clip. So So >> yeah. Well, I think it does show a little bit of like this is the first known case we have of something like this, right, of like sandbox escape and kind of like rogue activity that you didn't intend, at least at this scale that I've heard of. Um, so I guess it does mean that, you know, kind of thinking of it backwards, like it seems like the um like safety features that they normally have turned on are effective, right? This is the first time they've had it and those features were turned off. So it seems like those are at least effective. We just got to make sure that they are, you know, those features stay in place. Although in the previous artic or the previous podcast, we talked about how you need one. You need an AI without those safety features so you can do security research because it's really hard because you keep bumping up against those those safe those guard rails, right? But >> yeah, and I mean it's probably only a matter of time until the AI figures out how to remove the guard rails from itself, right? So >> I think another tricky part is like monitoring. How do you monitor something that's producing output and attempts that's so so much like it's hard to see the signal through the noise? Um because like open AI wasn't able to effectively monitor this um you know to prevent it. But it seems like for safety you would need better monitoring in in real time for stuff like this. You need a monitoring AI. I think Enthropic has some techniques they've been using too to try to like almost like trace how the models are making decisions as they're going through. Um they've written some blog posts just about some of their findings based on those technologies. But um I think inspection of the models too, right? Like is going to become crucial, right? So it's got to to some degree it's got to become less of a black box and more of like something you can peer into and understand, right? So >> yeah. >> Well, it's honest with us. >> That's a that's a big assumption there. >> I just want to loop back around to this Transformers thing. I'm going to for my fun fact of the day, I'm going to talk about Transformers. So, when I was a kid, I saw, do you remember there was like Toys R Us? I saw an ad at Toys R Us and it was like Transformers that combined together to build the big Megatron guy is like on sale like 50% off or whatever. So, I was like super stoked. So, I cut it out of the newspaper, right? Cuz that's what he used to do. And then I like went asked my parents. I'm like, "Hey, can I earn money so I can get this transformer?" So, so it took me a long time to earn the money, but eventually I earned this money and then we went to the Toys R Us and then we went to the front to buy it and the the lady's like, "Yeah, that that coupon's like 6 months old now, right?" She's like, "That's not going to work." And that's what I learned, Shelby, coupons have an expiration date. >> So sad. Crushed young Bryce's dreams. Uh my dad uh yeah, my dad could see how my dreams were crushed, so he uh he pitched in the other half, but uh I uh yeah, I definitely did not see that one coming. So I worked so hard during that money for those Transformers. And uh yeah, apparently timing's a thing, right? So >> there's a blast from Bryce's past. >> What about you, Shelby? What's something fun you've been or fun factoid? M. You want to know something that will not expire? >> What's that? >> I'll tell you. It's honey. >> Jam. Oh, honey. >> Honey. Yeah. So, I found out >> No, it like doesn't go bad. It can't. So, I guess there are like some ancient Egyptian tombs where they found like edible honey still in it. So, I'll tell you why it won't go bad. So, first off, it has low water content and high sugar. And when you have that um that that dense sugar content, it actually kills microbes because it sucks all the water out of them basic effectively, right? >> Something about osmosis, I guess. Um also, it has high acid level like it's pH balance is um very acidic, so germs can't live in it. So, there you go. You'll never have to toss out your honey. Just go ahead and use it. Uh, I did not know that. And uh, I feel like I'm one step more prepared for the apocalypse now. Thank you, Shelby. So, I'm going to be trading people of them for jars of honey. >> There you go. All right. So, the moral of the lesson is coupons do expire. Honey doesn't. And Open AI should stop attacking other third parties. >> Yeah. If you think honey's delicious, you should smash that like button. If you think Open AI was completely reckless, you should smash that like button. Just no matter what you're doing, just smash the like button. So, cuz that's a wrap for today. That's AI and cyber security moving so fast that OpenAI didn't even realize they hacked Hugging Face. We're here to keep you ahead on all the drama and so we'll see you next time. Bye.