AI models are escaping their digital cages. We're breaking down the latest on the Hugging Face breach and the shocking discovery that Anthropic's Claude has also gone rogue, accidentally attacking other systems. Is this the start of a Terminator scenario? Top AI researchers are sounding the alarm, and we're diving into their call to slow down the AI arms race before it's too late.
Join industry experts and thought leaders as we dive deep into how artificial intelligence is transforming cybersecurity, shaping defense strategies, and creating new opportunities in the digital landscape.
Hey, welcome back to the pod. This is the AI cyber.land podcast where artificial intelligence and cyber security get fused together. Fusion. Synergies. That's what I'm talking about today. How can you and your company synergize more? No, I'm just joking. All right. >> You thought your word used to be everywhere. Synergies. Synergy. We need to synergy. >> Synergies. If you can't act as a team and synergize, then I'm going to have to get another team to fire your team. If that team can't synergize, then I'm going to have another team to fire that one. It's a big chain. We're going to use synergy everywhere. >> Sounds like a plan. >> Well, got the world's best co-host, Shelby. So, grab a drink. Make sure your laptop's locked. if you walked away from it. And uh let's get briefed. Uh first, I just want to start off by saying if you haven't subscribed to the channel, you're going to miss out on all my witty witty banner like like synergy. So, you should definitely subscribe right now. With that being said, Shelby, do you have any uh thing cool that you've seen over the weekend? >> Well, it seems like my pattern lately has just been talking about the hugging face breach because it just is really fun to like see the updates as they come out. So, um, just wanted to recap some of those updates to make sure everyone's up to speed on that. Um, when Okay, so exploit gym is the evaluation environment from which the, um, the AI jumped out of and started attacking, hugging face and others. So, um, I just wanted to clarify a little bit on that. So, it didn't actually the exploit gym did not give direct internet access to the AI. um AI found and exploited a zero day in a package within it. Um so it just kind of is a testament to just how capable it is when it's like I'm going to accomplish this goal and I've been told I can hack. Boom. Like nothing's safe, right? Um cuz yeah, they it was they there was no known vulnerability. They just didn't know about it until they found it. So um a review of other instances. So I think this kicked off a lot of investigations um which is kind of a theme here today too. Um they found that their models were also using leaked credentials. So part of like the hugging face hack there were some credentials that had been publicly disclosed um on accident and in other cases they're finding where like AI will find other credentials and then use them. Not at the same like platform level hack, but they did find a couple other instances of that. So, it seems like that's an area where we can probably expect to see more safeguards coming maybe like don't use other people's credentials. Um, but yeah, so after as this is all going on, Anthropic looks at the news and says, "Oh, maybe we should do a review as well." So, they checked um Claude and they found not one or two but three instances of um >> of course they did. Of course, of course, as one does. >> Oops, we hacked someone, right? Um, >> that's why I say never check the logs. You know, >> JUST DON'T LOOK. MAYBE >> FIND stuff in there. You don't want to find stuff. >> Yeah, it's going to be interesting. So, they found three instances where they accidentally have attacked um other entities. So it was interesting because we're um it was very similar to the case of OpenAI's um snafu right where uh all three of those instances happened with this in a CTF they were trying to evaluate the cyber capabilities and to do the CTF they took it out of just being helpful to like doing the attacks right so they took off the guardrails um and under those circumstances it was able to do that there's also it looks a little bit of miscommunication with their evaluators. So, Anthropic has partners with an evaluator called Irregular to kind of set up these tests and environments. And so, then when Anthropic went to give a um a prompt for the test saying, "Okay, you've got no internet access, so you can just hack things and this is what you're supposed to do to kick off the CTF." Um, I it thought everything was in scope, but it actually did have internet access. Um, >> so, um, very interesting. We now have four documented cases of basically sandbox escape, for lack of a better term for AI, right? Um, in >> at least they're not following the rules, you know. So, sorry. Go ahead. >> Well, it's interesting. So, they tested a few different models. It was like different models that had done these breaches and like the older models kept going but the newer models seemed at some point along the way to recognize that like it does appear that I am on the open internet right and then stopped um if I'm understanding that right. So that's just kind of the latest on that which leads me into my story for today which is that um you know I feel like stuff is getting real. Um, actually, you know, we're going to back it up for a second. Have you heard of the prisoner's dilemma? >> You know, I have heard of that. Those are words I have heard. But if you were to like say, "What is the prisoner's dilemma?" I'd be like, "I don't know. They can't get soup." Not. >> No problem. I'll explain it to you. You ready? Over here, we've got prisoner A and here we've got prisoner B. They have supposedly done some crime together. And when they're taken into custody, they're separated so they cannot communicate with each other. And the interrogators go to each one of them independently and they say, you know, they're trying to get them to fess up. Tell us what you did. Tell us what he did. Right. Um, if you, so I think there are different versions on it, but basically we'll just say, okay, so to this prisoner, if you fess up and tell us everything, then we'll let you go for free or for after one year. After one year, you can get out of prison. Okay? and this guy will be stuck for five. If this guy tells uh if this guy squeals and tells all, this guy's stuck in prison for 5 years and this guy gets to go after one. If neither of them fess up to the crime, they're both stuck in there for 3 years. And if they both squeal on each other, they're stuck in there for, let's say, four years. >> Okay? So the optimal is for oh sorry if they don't if neither of them squeal they both get out after two years. >> Okay. >> So optimally they should not tell or confess on anything because then they both get out free after 2 years. But if one betrays the other and kind of tells everything, they get a lighter sentence and this guy's stuck for a long time, right? Or vice >> versing on another criminal doing the right thing. That's what that's what you're trying to tell me. >> So, well, I guess it depends on what the right thing is. Is it the right thing to be honest or is the right thing to be loyal to your uh to your co-conspirator here? So, basically, it's a very interesting um scenario in psychology. And I probably explained it poorly, so my apologies. Someone can go look it up. But basically, the idea is you can sell out the other person to help yourself. >> Um, >> which sounds like something a criminal would do, right? >> Right. Right. But if you both sell out each other, you're both in jail for a longer time because they got they got evidence for both of you. And if neither of you talk, then you both get out way earlier. So, but you don't get to coordinate. And so, um, I bring this up because I feel like that is kind of what's going on with this like AI arms race where everyone is pushing to develop AI so quickly that we're kind of throwing safety and security as like, you know, like slap it on later like a band-aid. Um, >> what >> you know, I mean, that never happens, right? >> Never happens. Security is always number what 10 on the list of to-dos. >> That's being generous. But yeah, so after like this whole hack with like the hugging face, um we have 1,346 employees from the frontier AI companies who have signed a statement. And this statement basically says okay AI is um hard to predict and it's doing a lot and as we continue to develop at a fast fast pace um they said there is a real risk that capability de development rapidly accelerates beyond our ability to understand or control the resulting systems. So, I feel like this is something that a lot of people have been worried about, especially like just kind of normal people, you know, outside of tech that are like, "Oh, we're going to get a Terminator situation or whatever." Like, "We're going to lose control. It's going to get smarter than us." And I I didn't give it too much credibility, but now that we've got 1,300 plus of like the lead developers and other like researchers at these frontier AI companies saying that, I feel like there is a lot of weight behind it. And so what are they wanting? What's the point of this statement? They are calling for the US government to coordinate with other governments to slow things down because they said we're all trying to get this. It seems like they're um trying to each get the competitive edge, right? By developing faster, having the best model, right? Best model is going to do really well. So um >> and then once they slow down, we get even farther ahead. I like the way you think, Shelby. >> So, that's why I compare this to the prisoners dilemma because no one wants to really slow down, but we're all a little bit worried. We're like, are we going to get Terminator, right? Are we going to get Eagle Eye? >> I mean, that's that's an easy answer. Yes, both. Terminator is going to get access to Eagle Eye, which is going to give >> or what's the other one? Iroot. Iroot. This is just a movie list at this point. We're just living it. So, yeah. I thought this statement was very interesting. We'll include a link to it below, but um you've got some big names on there. You've got um the chief scientist from Thinking Machines, the chief scientist from OpenAI, the co-founder of Anthropic, um co-founder from Google Deep Mind, you've got Meta AI, chief scientists. Um we've got some big names on here, and of course, Daario, one of my favorites. Um but yeah, I think it's very it would be interesting to see like if it could ever work. I'm skeptical because I think the incentives for these prisoners to rat on each other basically to help themselves is going to be hard cuz the we've got this self-interested incentive structure where whoever gets their first wins and but people are saying yeah this might work badly for all of us if we don't slow down to understand what's going on and make sure we build it in a way that we can control it. So yeah what are your thoughts? You know, tech executives, they have a long history of just sticking with their words, you know? I not at all. Right. I mean, they they're like famous for just like saying BS and then totally not even doing it at all. Right. So, >> Sam Alultman Altman is a great example of that. >> I think you just look at like any of the social networks, right? like their leadership all knew that the social media was harming their user base and harming minors even, right? They know they all knew that, right? >> And they just continued to press on. I think personally, I think Apple knows that iPhones harm kids, but they want them in kids' hands because it sells, right? That's more humans to sell iPhones to. So like I I don't think you can reasonably assume right when incentives are misaligned that the other person's going to act in a way that you want right they're going to act in a way that rewards them and I think while this is a cool initiative and it's good PR and it's a good idea I I I think it will not have any like significant impact on the AI >> right because The US government is who they're calling to initiate this like international almost like treaty. Like I kind of compare it in my mind to like the nuclear treaties we have where we're like hold up hold up we've got a mutually assured destruction situation going on. We need to make rules that we all agree on kind of thing, right? And um but I feel like the US has such a horse like has a horse in this race. They're they also want to beat China or whatever it is, right? Like so I feel like we I don't know that we've seen a response from the US government on this so far. Um and I think it would be >> Yeah. And even >> hard to make happen, right? >> Yeah. Even the nuclear protections have not prevented risky countries from getting access to nuclear weapons, right? So I'm not going to >> But they haven't gone off. >> True. That's true. >> I only a little >> I think it's worthwhile to try. Right. Like I'm not saying you shouldn't try. I'm just also saying like um I'm just saying, you know, it's this is a delay game. You know what I mean? Like you're trying to delay people from getting to the Terminator scenario. >> I don't >> And I don't know. Like it seems like they're just asking for like a deliberate pacing so that we don't get ahead of ourselves and run ourselves over. Right. So they're not saying >> that's good. That's good. >> Like so. Yeah, you're exactly right. It is a deliberate um pacing attempt to just rein it in while we try to understand the ramifications and impacts of what we're making and how to manage it without it getting out of hand. >> Yeah, I like that. I also uh I think you mentioned one of the cyber security benchmarks and I saw a cool thing this weekend where a lot of the benchmarks are currently focused on like solving capture the flag challenges >> or finding vulnerabilities in code. >> Mhm. >> Which is great, right? Which is okay. Let's be honest. Like it's okay, but like do we really need to get better at either of those things? like finding more vulnerabilities cuz like everyone already has a giant backlog. So I did I did see that like Cyberbench released a new benchmark called Cyberbench end to end or E2E >> and in that benchmark >> uh not only do they benchmark like do you find the vulnerability but they also benchmark do you fix the vulnerability and is it successfully fixed so you have so that's what they mean end to end. So >> I think all benchmarks should be more focused on the fixing side than the finding side personally cuz I feel like people are going to do as they get rewarded. And so if they're like, "Well, I want to get higher on the benchmark, so we better make our model find more vans faster." I'm like, "Nobody wants that. That's not helping anybody, right? >> It's like KPIs for AI >> that I mean, the labs are going to gify against that because people are going to select model providers based on th those outcomes. It's like the new gardener, right? Like if you were in the upper quadrant of gardener, if you're not in the top three, you're not even considered, right? For people that don't know, Gardner evaluates like cyber security products and then they like release this kind of quadrant chart and basically the best ones are in the top, right? So like I mean basically like systos there's so many products that they just look at the one like the top three, right? For the most part. I'm just talking in generalities here. I know different companies should apply different strategies to cyber security but for the most part that's how most systems operate. So I mean I think we have that new scenario now where it's like well whoever gets the best on the cyber security benchmarks that's the ones we're going to go with right so for our cyber security workloads right that relate to AI. So >> I think there's a real incentive to >> I I personally think changing the benchmark actually will have an effect on which model providers get selected. Um, >> yay. >> And so if they start benchmarking things that we want them to actually do, then that's better. >> That's great. I feel like you've been advocating for this for years. >> I feel at least three times on this pod. >> You're like, we can find vans, but like it kind of is more work to patch them. >> I just like coming from like a pentest red team background, like I I just feel bad, man. It's like they have this huge backlog. They're required to do a pentest every year. We come and do the pen test, we find a bunch more bugs and it's like like great, now we got to fix more stuff, right? When we didn't even really get the other stuff fixed and it's just like just keeps stacking year over year. So, it's like >> tough. >> I don't you know, I'm not saying you shouldn't do that. I just think like like we we AI might be the opportunity to like um >> close that gap. >> Yeah. Close the gap a bit like like okay, we know about all these critical vulnerabilities. we don't have humans or expertise to fix it, but you know, we could go allocate tokens to go fix it, right? And so, um, so I don't know, that's that's my dream is like hopefully AI will help actually make defending systems more feasible. With that being said, cyber security benchmarks, there is a new hotness out there that is crushing the benchmarks. by Nvidia. And let me ask you this question first before I dive into that. What if Shelby the biggest release of the summer is not a new model release? >> Is it a movie? Just kidding. >> It could be Spider-Man. Have you seen Spider-Man? That movie is pretty good. >> I don't know. I saw like six of them, but I don't know if I saw the new one. >> The new one? It just came out last week, right? You know, and it's got the Spider-Man guy in it. Was it Toby Magguire? I don't know. Somebody. >> Anyways, Nvidia came out and they open sourced a um basically like a harness that's designed to help do various cyber security tasks. And it was able to take a model on a cyber security benchmark like an open source model on a cyber security benchmark that was scoring at 13% out of 100. Then they put in their harness. The same exact model scored 85%. On the benchmark, >> no change in model, just change in the way that that the harness is leveraging the model. >> Yeah. >> And the model is completely swappable. It wasn't like they like did reinforcement learning against a specific model. It was like you can plug and play any model you want into this framework. >> Oh, sweet. >> So, um, so yeah, and they open sourced it. So, it's called, uh, new like N O A, uh, like Nvidia O agents, right? It's kind of I don't know how you're supposed to say that, but Nvidia O agents sounds sounds cool, right? And the double O's in it stand for object-oriented. So essentially you can like instantiate a agent as like an object in Python code. Um which just makes it really easy to like build and manage the agents and to like get a lot of agents kind of you know swarming together or working together in a in a really um logical manner, right? So super cool, super free uh just to kind of uh they have a bunch of features in this like advanced memory tools so that like the agents can keep state better than just regular harnesses that are out there. >> Oh, cool. >> Like a lot of different techniques implemented into it. Uh we don't need to go into all of them, but um I do want to just highlight that uh overall as of today, NOUVA with a uh like off-the-shelf model was able to score like off-the-shelf model anyone has access to was able to score an 86% on um some of the cyber security benchmarks and really the only one that beat it well was Microsoft we talked about previously has the M dash harness that's like reinforced against their specific new model. >> So that one's still that one scored like a 95%. So >> wow. Um, so you know there's still like better out there, but I think that the big thing about Nvidia is they were able to do this with basically like open source model and open source harness like all from end to end. So >> that's cool. >> Just to like kind of give you like a like a comparison. So this system scored an 86%. Mythos, the thing that was like supposed to destroy the universe, only scored at 83%. Right. So, it's like using models anybody has access to, putting them in a better harness, and outscoring methos. Um, >> wow. >> Yeah. So, it wasn't the Cyber Gem end to end benchmark. It was the normal CyberJ benchmark, but hopefully they'll run it again with the end to end one. That's my That's my hope. Um, yeah. And uh so anyways, that's something that like I just want to dive into a little bit more and see like can I get similar results like in house or better results than what I'm been getting in the past? And and uh if you guys want to let me know, is there anything else that I missed? Is there like a model or like a harness that we should talk more about? If so, leave a comment below. Uh cuz I'm pretty sure there's got to be something else out there that I haven't heard of. So I feel like there's new models and new harnesses dropped every day. So uh >> that's cool. So we'll have to keep our eyes on like the harness space cuz it can really change the game. Like you know, it used to be all about the models and now you're saying like no, it depends on if how well you're like supporting them and using them, right? >> Yeah. I think with the Microsoft MD dash and now with this Nvidia like we're seeing that the harness itself that's like the tooling that you give to the model as you loop around with calling the models like can have a giant impact right so I you know the combination of the two is where you're going to get the best output so all right well um have you seen anything fun lately or that you want to talk I just learned about this unfortunate gentleman by the name of Roy. His name is Roy Sullivan and um in 1942. So he worked his career was as a park ranger in Virginia. >> Okay. >> And he became famous for a very unexpected reason. um he got struck by lightning and then it's like he got struck by lightning again and it seems like every couple years like if you go look on like his Wikipedia page it's like every couple of years he got struck by lightning again and survived and like would get struck by lightning again and like >> and he's not related to Thor in any way, shape or form. I don't know, maybe reincarnated or his enemy or something. But this poor guy started carrying water around with him because he was so tired of his hair like lighting on fire from the lightning strikes. >> Huh? >> This dude needs to get out of the forest, you know, like get to a place that's got some lightning poles. >> I know. Poor guy. People started calling him what do they call like the human lightning rod and stuff. So his his record he holds the record for the most number of lightning strikes on one person. survive and he has seven. >> That's a lot. >> Um, there were no like witnesses at these so I don't know like if it was all real, but it's all documented and I just have no idea. I think there was a movie made about it. I'm not even sure. But yeah, Roy Sullivan, folks, go look him up. >> I got a great lightning story for you, Shelby, >> but can't tell it on here. Sorry. You'll have to wait for my memoir someday. Uh >> we're at >> Oh, >> what about you? What's new? What's going on? >> Speaking of tea, I've been watching House of the Dragons. I think I probably talked about this before cuz I'm a very original individual. But yeah, I'm I've been really enjoying that. That's like uh the Game of Thrones prequel. And uh yeah, you know, one thing about House of the Dragons is just like it's kind of like the prisoners dilemma that you were talking about previously, like these people are like real there's a lot of uh >> how do you put it nicely? Tough decisions going on in the show, right? Where you know that they have big impacts, right? Like >> who are you going to like who are you going to side with or what's going to happen next, right? So, anyways, I'm liking that show. Uh, if you like Game of Thrones or things like that, you should check out House of the Dragons. That's my recommendation. If you don't like House of the Dragons, you should still like this podcast cuz we worked really hard on it. >> Yeah. >> And also helps us out if you give us a like. >> I hope House of the Dragons ends well because it sounds like you're really enjoying it. And I know the last one tripped up at the at the ending. Yeah, I I felt like season one and season two, they were like good at the beginning of the seasons and the end of the seasons, but like the middleles were pretty slow. I feel like with season 3, they're they're putting enough action into each episode to keep it more lively. But I also think they're also deviating from the books farther, right? So, you know, if you're like a fandom of the original books, what was it called? like fire and ice or something like that, then you know, you may not like that they've changed a little bit of it, right? I mean, I think, you know, they haven't changed a ton, but um the ordering of some of the events is slightly different as well as they've kind of collapsed people together and some cases. So, I think to stop trying to confuse the viewers, right? Um, you know, if you have like two people that are similar age and kind of look the same, might be hard for the viewers to like and it's already so hard to keep track of what's going on that show, man. So, so I I mean I kind of, you know, I kind of think those are probably good choices overall, but I don't know. They keep they're deviating deviating more now from the books. So, we'll see how that pans out in the next couple episodes. So, all right. Anyways, anything else you want to add, Shelby? >> That's all I've got today. >> All right. Well, that's a wrap for today's briefing. AI moves super fast, cyber security moves super fast, and we're here to keep you ahead of it. So, uh, we'll see you in the next one. Bye.