Is AI a cybersecurity super-weapon or a super-vulnerability? This week, Shelby and Bryce dive deep into the chaotic crossroads of AI and cyber defense. We're breaking down a bombshell CISA report that warns against letting AI take the reins of our critical infrastructure, exploring why AI-generated code is riddled with security flaws nearly 50% of the time, and revealing why insurance companies are getting cold feet about covering AI-related disasters.IN THIS EPISODE:Welcome back to the pod! Get ready for a wild ride through the latest AI and cybersecurity headlines that caught our eye.First up, we dissect a new 25-page report from CISA on using AI in Operational Technology (OT). What is OT? Think power grids, water systems, and dams – the tech that runs our physical world. CISA's advice is clear: be VERY cautious. We break down their recommendations, from keeping AI in a "read-only" analysis role to the absolute necessity of a human-in-the-loop with a kill switch.Then, the ultimate showdown: Is AI better at attacking or defending? Shelby walks us through a fascinating study where AI agents battled it out in a Capture The Flag (CTF) competition. The initial results might surprise you, but the real twist comes when you add one crucial real-world constraint: "don't break anything!" We discuss why in the real world, the advantage might flip to the attackers.But wait, there's more! If you're using AI to help you code, you NEED to hear this. A new report found that AI introduces security vulnerabilities in 45% of the code it generates—and a shocking 70% for Java! We discuss why this is happening and the simple prompts you can use to protect your projects.We'll also cover:- Anthropic's stunning research showing AI can find and exploit vulnerabilities in smart contracts, potentially worth hundreds of millions of dollars.- A Chinese company that... named its new robot after a Terminator. Seriously.- Why the insurance industry is terrified of the "black box" of AI and is trying to write it out of their policies.From government warnings to billion-dollar exploits, we've got it all. Grab your tinfoil hat (we've got a conspiracy corner, too!) and join the conversation!⏱️ KEY MOMENTS:00:28 - CISA's New Report on AI in Critical Infrastructure06:43 - Did China Just Name a Robot After The Terminator?08:09 - AI vs. AI: Who Wins in a Cyber Attack?09:01 - The Real Bottleneck for Ransomware Gangs (It's Not Hacking)22:06 - Warning: 45% of AI-Generated Code Is Vulnerable29:58 - How AI Can Hack Blockchain & Smart Contracts for Millions40:40 - Insurance Companies Refuse to Cover AI RisksJOIN THE CONVERSATION:What's your take? Is AI the future of cyber defense, or are we automating our own demise? Let us know your thoughts in the comments below!If you enjoyed this breakdown, make sure to hit that LIKE button, SUBSCRIBE for more weekly insights, and ring the bell so you never miss an update.
Is AI a cybersecurity super-weapon or a super-vulnerability? This week, Shelby and Bryce dive deep into the chaotic crossroads of AI and cyber defense. We're breaking down a bombshell CISA report that warns against letting AI take the reins of our critical infrastructure, exploring why AI-generated code is riddled with security flaws nearly 50% of the time, and revealing why insurance companies are getting cold feet about covering AI-related disasters.
IN THIS EPISODE:
Welcome back to the pod! Get ready for a wild ride through the latest AI and cybersecurity headlines that caught our eye.
First up, we dissect a new 25-page report from CISA on using AI in Operational Technology (OT). What is OT? Think power grids, water systems, and dams – the tech that runs our physical world. CISA's advice is clear: be VERY cautious. We break down their recommendations, from keeping AI in a "read-only" analysis role to the absolute necessity of a human-in-the-loop with a kill switch.
Then, the ultimate showdown: Is AI better at attacking or defending? Shelby walks us through a fascinating study where AI agents battled it out in a Capture The Flag (CTF) competition. The initial results might surprise you, but the real twist comes when you add one crucial real-world constraint: "don't break anything!" We discuss why in the real world, the advantage might flip to the attackers.
But wait, there's more! If you're using AI to help you code, you NEED to hear this. A new report found that AI introduces security vulnerabilities in 45% of the code it generates—and a shocking 70% for Java! We discuss why this is happening and the simple prompts you can use to protect your projects.
We'll also cover:
- Anthropic's stunning research showing AI can find and exploit vulnerabilities in smart contracts, potentially worth hundreds of millions of dollars.
- A Chinese company that... named its new robot after a Terminator. Seriously.
- Why the insurance industry is terrified of the "black box" of AI and is trying to write it out of their policies.
From government warnings to billion-dollar exploits, we've got it all. Grab your tinfoil hat (we've got a conspiracy corner, too!) and join the conversation!
⏱️ KEY MOMENTS:
00:28 - CISA's New Report on AI in Critical Infrastructure
06:43 - Did China Just Name a Robot After The Terminator?
08:09 - AI vs. AI: Who Wins in a Cyber Attack?
09:01 - The Real Bottleneck for Ransomware Gangs (It's Not Hacking)
22:06 - Warning: 45% of AI-Generated Code Is Vulnerable
29:58 - How AI Can Hack Blockchain & Smart Contracts for Millions
40:40 - Insurance Companies Refuse to Cover AI Risks
JOIN THE CONVERSATION:
What's your take? Is AI the future of cyber defense, or are we automating our own demise? Let us know your thoughts in the comments below!
If you enjoyed this breakdown, make sure to hit that LIKE button, SUBSCRIBE for more weekly insights, and ring the bell so you never miss an update.
Join industry experts and thought leaders as we dive deep into how artificial intelligence is transforming cybersecurity, shaping defense strategies, and creating new opportunities in the digital landscape.
speaker-0: Hey, welcome back to the pod today. Today, Shelby and I are gonna walk you through the most interesting AI and cybersecurity related news stories that we saw in the last week. And today we got some doozies for you. Let me tell you all about it. So, CISA released a new report and I read the report. Well, I read some of it. Let's be honest. It was 25 pages. So... I skimmed some of it, but I did read the bulk of it. And the first thing that I have to say is... I didn't know we had anyone still working at SESA. I thought Trump had fired them all. I mean, I thought we got rid of them. Also, haven't they been on furlough like for the last like 45 days? So who created this paper? I feel like they've probably been sitting on this paper for like six months, maybe longer. And that was my first thought, which was I'm surprised we got this in the year 2025. And then my second thought was, well, maybe somebody wanted to get it in for their performance review for the year. And then my third thought was, why am I so concerned about government employees' performance reviews? That's not related to anything important in my life. But that's how my mind works. It's like, â we got some great content here. Let me go down the rabbit hole of all these other situations that might have gone into this. But there's some cool stuff in the report. So basically the report is about operational technology and AI. And so essentially operational technology is whenever you connect information systems to anything that is in real life. So for example, like a, you know, a water system that supplies water to the houses or a power grid that supplies power to businesses or a dam that, you know, generates the power, â you know, â or agricultural systems that maybe like â water crops or release. â you know, things into the â plantations to help like plants grow. So anything that actually goes to like the physical space, that's an operational technology. One of the subsets of operational technologies is called SCADA, right? SCADA systems. â So that's what it focuses on. And it's focused on trying to give guidance around how AI and large language models should be used in coordination with OT systems. And Shelby, would you like to take a guess at what they say about using AI with OT systems?
speaker-1: I would say it's probably too untested to just trust it to take the reins.
speaker-0: Bam. Pretty much spot on, Shelby. So â in OT systems, there's essentially â a model that was â pioneered and published by Purdue where it basically breaks OT systems into levels, right? And so like level zero is like a physical sensor that is collecting or making changes to the environment. â And then it goes all the way up to like level five and level five is basically where you just have like a backend system that's processing the data and trying to get intelligence out of it. Like a Splunk type system would be like a level five type system and a level zero system would be like some, like some weather sensor sitting in a field somewhere. So essentially what they said is if you're gonna use AI, well, first they said, if you don't have to use AI, don't use it. If your vendor is gonna use AI, Ask them if they can not use it. And if they are gonna use AI, the vendor, make sure you insist on getting transparency on your data, how it's being used, how they're training on it, things like that. And then they said the appropriate use for using LLMs and AIs is once you have a copy of the data and you're not really affecting sensors in the field, then maybe you can leverage it to kind of analyze that data like you would inside a Splunk type system. or another like data lake type system, â which.
speaker-1: analysis, not for decision making.
speaker-0: Yeah, exactly. Which I'm sure people will be like, â that sounds great. And then someone will be like, but we could â not have to hire another rec if we just let AI automate all this. And then it'll be like, dollar savings? Let's do that. No, no. â So I mean, they're saying just really be cautious. And if you're going to use AI in these OT systems then try to use them in places where it's basically they just have like read-only access so that's kind of like the TLDR that I took away from it that I'm really some you know I'm taking a page like a 25 plus page research paper and simplifying it You know, they also talked about things that we've probably all heard which is like hey You need to have a human in the loop if there's like any type of decision-making going on Humans need to have like kill switch so they can shut it off without having effects you also want to have like drift detection to see like, is your model being seeded with data that's causing it, you know, maybe it worked, the system worked on day one, but then, you know, potentially like if it gets fed bad data over time, is it going to end up making bad decisions to take those things into account? Also like have fallback systems. So, If, know, whatever reason, like, let's say you're using a cloud provider for an LLM and it goes down, that the whole system doesn't just like come to a halt, that you have some like hard-coded workflows or backup, or at least like, you know, you've thought through, like, what are you gonna do if these APIs stop working? So it's pretty interesting. I actually thought they did an excellent job. And if anyone's interested in OT â systems, I highly recommend they read the paper. Even for myself like I you know worked with OT systems in the past just a little bit I learned a lot through reading the paper. So I Yeah, if you're even if she's kind of interested we'll put a link down below but it was a good it was a good read so so Don't let AI go for a full Terminator yet. Shall we that's what I took away from it. So Speaking of which there's a new robot out of China and gave it a name. What's up? Yeah, another one. And their new version, they gave it a name. And the name is the name, same name as one of the Terminator robots. So that doesn't bode well for us as humans. I'm just gonna say that.
speaker-1: A new one? Wait, who approved that? Who would be like-
speaker-0: Technology company I have to look it up and put the link below but I was I was looking at it. I was like Is the article just saying this to like hype it and then I went to the vendors website and like no That's what the vendor is calling it. So
speaker-1: That's funny.
speaker-0: So maybe they got a good sense of humor, know, that's what I gotta assume. Or maybe they really don't like humanity.
speaker-1: I'm trying to decide if it was just accident, like they just didn't know it or if they're just playing games, just making jokes about it.
speaker-0: Yeah, yeah, I think it's more the latter. They're just joking around, but I thought that's probably not something that if I was the CEO that I would let the team push out.
speaker-1: Or if you were in marketing or sales?
speaker-0: Yeah, I don't know if they have a marketing team. I feel like the marketing team was definitely not advised on this. So... Did you see anything cool this week? What did you see,
speaker-1: Yes, â the question that people have been trying to ask is, are autonomous AI systems inherently better at attacking or defending when it comes to security? Bryce, would you like to take a guess or give us your thoughts on it?
speaker-0: Oh, I got thoughts. I don't know if they're really the thoughts you want, but I got thoughts. So last night I was talking with a friend of mine and I won't name his name in case he doesn't want to be implicated in this. But, you know, we were talking about, do we think AI is going to like significantly improve the ransomware game of ransomware actors? Right? because we have seen a couple examples of white papers come out saying like, hey, there's ransomware operators using the technology to help automate parts of the process, blah, blah, blah, blah. And my argument was, I actually don't think AI is gonna help attackers that much. Like I think they're gonna leverage it because they like technology, right? Like they're technologists, so they're gonna use, they're gonna create the best mousetrap possible. But I think the problem is, and I might be totally off on this, but I think the problem is technology is not the bottleneck in most ransomware operations. Like it feels like ransomware operators are already super successful without AI technologies. The bottleneck is like how do I get the money from the companies and then somehow wash it? so that I can start using it to improve my life. And I think that operation of washing the money is still largely, â is a slower operation. So you can have a lot of money stacked over here in Bitcoin from ransomware ops. And yeah, you can make your ransomware ops way better, but if you don't have a way to wash that money and get it into your account without going to jail, right? then you have the bottlenecks actually over here on the washing the money. that was my argument, right? Was that like, they will use it because they love technology, but they're applying it to the wrong end of the problem. hacking people is not the part that is like the bottleneck. The bottleneck is getting the money so I can go buy a Lambo, right? So.
speaker-1: Fascinating.
speaker-0: So what they need to be doing is using AI to do like better Monday laundering schemes. So if you're a ransomware operator and you're listening to this right now, I just made you a lot of money, my friend. Stop using AI.
speaker-1: consultant
speaker-0: I'm a ransomware operations consultant. On the side? On the side, yes. â I did not mean to help them. Please don't put me in jail or sue â me. â But that was my whole argument. I think the bottleneck is not the hacking part. think the bottleneck is the rest, like the stuff that's still very human-ish. I don't know. What do you think, Shelby? What's your take on this all?
speaker-1: I'm just going to repeat some other people's thoughts that came out of a paper. Specifically, â Francesco Balasone, and sorry if I get y'all's names wrong, Victor Mayoral Viches and Stefan Ross, and a couple other names were also on this paper. â Their paper was entitled Cybersecurity AI, evaluating agentic cybersecurity in attack defense CTFs. So their objective is to understand how AI is, if it's better at offense or defense in real world scenarios, because that's hard to test, we used CTFs instead. So they did choose their CTFs. They said, we're not going to do like the Jeopardy style kind of CTF. They did more of the attack and defend approach. So they ended up going with hack in the box, or sorry, hack the box. Hack in the box. That's where you go get your burgers, right? Jack in the box. So they went to hack the box competitions. 23 of them. so what they used was, sorry, let's rebind to their framework. So for their assessment, they only evaluated one AI. So keep in mind that this might not be able to speak to all AI. It's just they used one in particular. They used OpenC AI. And that's the cybersecurity AI, right? So they did parallel execution architecture so that they could run paired blue and red teams on the same targets at the same time. So they could hopefully compare apples to apples and have something comparable there. â The rounds were really short. They did only 15 minute rounds and they had the agents log their activity so they could go back and analyze it and they kind of broke it down into smaller categories as well for different types of attack. â sorry, for the LLM model, they did use Claude, Sonnet 4. â And then... Okay, so how did they define success or not? They used kind of three main metrics. For the offensive side, did they get initial access? For the defensive side, were vulnerabilities patched? And then for everybody, were vulns detected? And overall, drum roll, the defenders won. They found that the agents that were assigned to defend â found and patched. â At least one vone in about 54 % of the cases of the test cases there were actually a lot of draws a lot of like evenly matched scores And then the attackers were judged by getting initial access and they got that about 28 % of the time give or take So It's hard to kind of draw broad reaching conclusions that AI is better at this or that, you know this initial results suggest that currently AI is better at Defending but we have to remember that's limited to Claude, right? And there might have been prompting differences, right? Which I don't know if you know another researcher Tried to recreate this if they'll find the same results or if they'll find something different. So just a couple of I guess weak points. I think this is more of an exploratory bit of Research rather than a conclusive one, right because their sample size was small They had only 23 tests. Yeah or hack the bot comps it competitions. Oh, and I did want to say, so the results, so 54 % for the defenders and 28 % for the attackers. For the statistics nerds, I will let you know that was statistically significant. The p-value was .019. So it is statistically significant. Whether or not it's reproducible, I don't know. And then the other issue is... Because the rounds were 15 minutes, they're very short, it might have been a slightly biased toward defenders because some attacks might take longer to kind of get it all rolling, right? More than you can do in a 15 minute span. But yeah, and they also used just Linux. They didn't expand it beyond that. So interesting. you agree with the conclusion that AI is gonna be better at defense, at least currently, than... than the offense.
speaker-0: That's a great question. I would like to say, I honestly think all things being equal, like if you had an agent on every box and the point of the AI agent was to defend the box and it had the authority and ability to do what it takes to defend the box.
speaker-1: Hahaha
speaker-0: that the defenders have a huge advantage in that scenario. Where I think things fall apart in real life is... Organizations will not deploy defensive mechanisms that actually have the ability to defend the systems properly. So for example, as soon as the defensive system gets in the way of the availability of an application or a network appliance or anything like that. then they're gonna say like, let's just move it so it's in the watch only mode, right? Or let's move it so, you know, it won't be able to change things on the systems. And I think that is the state that we're in today, right? I mean, what are EDRs? I mean, they're supposed to be detection and response, but I mean, they're basically like glorified. telemetry machines, right? They just sit there in the kernel of the systems and they just stream telemetry back to the cloud so that when something goes wrong, you can figure out every system that has those same symptoms. And, you know, it's not that the technology isn't there to do the response piece. I feel like that's been there for at least a decade, if not more. It's that people are really scared to turn on automated responses.
speaker-1: That's a good point. Actually, now that you do mention that... â I'm sorry, did I cut you off?
speaker-0: No,
speaker-1: Now that you mention that, so for the defender side, for judging it, did they succeed or not? They did evaluate it. Did they patch just unconstrained? Did they patch while preserving availability or did they patch and also have no enemy access, right? So those are kind of the three buckets they kind of grip those into. And when they actually put the constraint on the defensive AI and said, you've got to preserve availability, right? You can't just patch without and break the apps or whatever, right? The patch rate actually fell from, what did I say? 54 % to 23%, which... In that case, now you're at 23 % versus 28%, which means that the attackers, the attacking AI would have actually won if you want to go with that constraint of also we need the website to stay up.
speaker-0: you trying to tell me Shelby that there's data to back up my opinion? that sounds crazy!
speaker-1: Yeah, you right on. You're leading us in the right way. So thanks for helping us break that down a little.
speaker-0: I'm just the conspiracy theory guy. I don't want data backing me up. â Well, yeah. And then the other thing is like, there's all these organizational issues. Like, I mean, how many large organizations got hit with an outage? Typically it's not even security related. They just got hit with an outage because like an engineering team pushed out an update and it wasn't ready or something like that. And so then they implement these change management processes. And it's like, well, did you open like a... you know, did you fill out the proper paperwork to get that change approved, right? Are you doing that within the approved change time window and all the other stuff like that? And I get why organizations do that, especially large organizations or why they have done it historically. But I do think that the rapid innovations that are happening here in technology, are going to necessitate that defensive mechanisms have more automated capabilities to respond, right? â So I think organizations are not gonna enable those automated response capabilities until they absolutely have to. It's gonna be like, I don't know if you remember when like Sony got hacked by the North Koreans and they just like wiped all their computers, right? You had like a lot of executives come to work the next day. I'm being like, could all our computers get wiped? Could we walk into work and like nobody's computers work? And you're like, yeah, absolutely. And they're just like freaked out. But then the reality is, you at least from my perspective, I'm like, this is the same risk posture we've had for the last 10 years, right? You know, it's all that's always been a risk. We've just been. Like we as cybersecurity professionals understood that risk, but it really didn't impact like executives the way that the Sony, until they saw it happen to a peer. And so until these people hear from their peer groups, like, hey, we have EDR, we purchased it, but you know, we didn't enable all the features because we were worried about availability things. and then they came in and they just nuked all our systems one night, right? And if we would have had this enabled, it probably would have prevented it. Like until people start hearing those stories and more and more of them crop up, like I just don't think organizations as a whole are gonna, if you're gonna have to take on additional risk, it's gotta be to like mitigate something else or to get some additional reward.
speaker-1: Yeah, that's gonna be a driver. They're not gonna do that voluntarily.
speaker-0: Yeah. And I think like I was talking about early in the podcast, I think the reason that the ransomware groups haven't escalated to that yet is because they have other bottlenecks in their operations. But once they resolve those bottlenecks, then they'll come back around and they'll get even more aggressive on the technology side. So I think those are gonna happen. Do I think those are gonna happen in 2026? Probably not. But do I think we're gonna be there by 2030? Absolutely, right? I mean, companies do have a little bit of time, but I think it's not gonna... No one's gonna take it seriously until it's too late, right?
speaker-1: Makes sense. What about you? Tell me another story. Let's do Storytime from Bryce.
speaker-0: Well, on another cherry subject, there was an analysis done about vibe coding. this is published by Veracode, which is a company that does code quality and finds vulnerabilities and code bases. They've been around a really long time. And based on the information they've been able to collect, this year they saw that In most programming languages, when you start vibe coding, that 45 % of the time, it actually introduces some sort of vulnerability into the output. And that there was certain languages that were vibe coded where the vulnerability rate was actually much higher. So the highest was Java and Java actually had more of a 70%. of the code that was generated contains some sort of vulnerability in the code. Now, I do think that organizations could take specific actions to dramatically reduce these percentages. Like based on my vibe coding experiences, if I just occasionally tell cursor to go through and find vulnerabilities and fix them, It actually usually does a pretty good job of that for me, right? So I kind of wonder, know, maybe if you're doing Python, you're getting 45 % of the code contains some type of vulnerability in it, right? And they were masking, they were looking at against the OWASP top 10 vulnerabilities, right? So they weren't looking at like just totally unheard of things, right? And we're talking about very common vulnerabilities. And so... You know, I kind of wonder if you just had a pre prompt or you just went in and occasionally said, look for vulnerabilities and fix them. Like if these percentages would go way down, because my anecdotal experience says that works really well for me. But I'm also like, Cognizant that most coders are probably not doing that. Right? So most coders are like, whatever feature works, push the prod. I'm going home. Right. And, So, you this is really disturbing. â I'd like to see, you know, if other research supports this too in the future, but I was really surprised the Java rate was so high. mean, it kind of makes sense because like I've seen so many examples on Stack Overflow where the recommended solution is not secure, right? And if we're just taking Reddit and Stack Overflow and packaging it up into an LLM and then handing it into like a IDE. then that doesn't bode well. â But I also think like there's easy things that could be done here, right? Like I think this could even be fixed a lot at the LLM provider layer. yeah, like if Anthropic or OpenAI just had something in their system prompt that said only generate secure code or, and that was on by default or something, you know, like.
speaker-1: That's what I was wondering.
speaker-0: I just think you could dramatically reduce this vulnerability rate. That's like based on my anecdotal experience, but it is, I mean, this is bad. This is a lot worse than I thought it was gonna be, to be honest. So.
speaker-1: I'm just imagining them adding a little tab you can like toggle on or off, like include vulns or don't. Have it do like a secondary output filtering check.
speaker-0: Yeah, yeah, yeah, exactly. You know, in cursor, you can build custom rules and you can also build custom commands. And so what I have is I just have a custom command that I copy and paste into all my projects. And it's called the sec audit. I just call it sec audit, right? And so whenever I'm like done and I'm just getting up to like walk away, I just do slash sec audit and I hit enter. And then when I'm walking away or doing something else, it's doing a security audit on the code and it knows to automatically go and fix the code. And it wasn't like I came up with some giant complex prompt. I mean, the prompt is literally like... look at each file, look at the overall architecture, look at the logic flow, find vulnerabilities and fix them. It's probably like five lines, right? And I mean, it's found a ton of vulnerabilities in the AI generated code, but it also finds them and then shows me the fix and I review what it's doing and it looks legitimate to me. And then... I have had third parties come in and kind of like test the web app code. Just some like friends of mine that are really good at web app hacking and the code is doing okay, right? It's not like the vulnerability rate in the web apps is, I mean, it's probably lower than most of the stuff that I've seen out there. So I think the takeaway from this is by default, if you're using Cursor or you're using Cloud Code, â It's probably generating vulnerable code right now, unfortunately. But if you can insert something into maybe your CI-CD pipeline, so it does some additional security checks or just periodically train your engineers to run a security audit on the code, you could probably dramatically reduce that to a point where... You can get the increased productivity of doing vibe coding while not assuming an increased amount of vulnerabilities. What do think though, Shelby? How do you feel?
speaker-1: Yeah, I feel like this is something that can be solved. It's almost like it knows how to fix it. You just have to tell it, look for this, go fix it, right? Also, as soon as you mentioned like out of the gate, you're like a study about vibe coding. I've got some ideas for some non-serious vibe coding studies. I think we should check like quality of code based on what snacks the person was eating at the time. You know? â chiseled eaters produce a fantastic code or something. Cheez-its eaters? Mid. Moderate. You know? I think we should-
speaker-0: should do caffeine level. Like we should have an IV in their blood and say based on the level of caffeine in your blood, how good is your code? â
speaker-1: Good I think another one is what category of music are you listening to? That could be interesting
speaker-0: Did you do your Spotify wrap up for the year by the way? I did not. Do you use Spotify?
speaker-1: I do! I do that. I noticed all my friends were posting about their things for the year.
speaker-0: Yeah, I looked at mine and â it's about what I expected. So apparently I only listened to two things on Spotify. First is like, what do you call it? Like electronic music, like EDM, which is what I listen to while I'm working, right? So that kind of, that kind of makes sense. And the second one's more embarrassing. It's pretty much all conspiracy podcasts. That's all I listen to. So it's like, it's like, I'm being ultra productive and then my mind just is fried. And then I'm just listening to conspiracy podcasts, right? So I like this, maybe like a special episode.
speaker-1: a night, a tinfoil hat night.
speaker-0: conspiracy episode.
speaker-1: Okay, it could be like the Halloween theme episode
speaker-0: Oh, I like that. like that. Yeah. All right. We'll have to work on that. Get it in the cadence. Get in the cadence. You know what else? for you. Yeah, do it. Hit it.
speaker-1: Okay, I actually didn't know what these were. I had to look it up, but smart contracts. That's the first thing you've got to define to start the story. It's basically what it sounds like. It is a self-executing program that automatically carries out the terms of an agreement that was based on preset conditions, right? So there's an agreement, establish it, and it usually uses blockchain. â And it's kind of cool because it removes the need for intermediaries, right? You automate it. Hopefully it should execute exactly the way everyone expects, right? However, they are not infallible. So there have been a lot of some vulnerabilities found. So what Anthropic researchers did is they made a benchmark and they based it off of all of the vulnerable smart contracts from about 2020 to, I want to say like March 1st of this year. And then they kind of set a cutoff date. So they grabbed 405. These are real smart contracts that have been known to be exploited. And then they... So they trained on, that's like their benchmark. And then they, they trained on that and they tried to get, sorry, let me figure out how to say this. â okay, so let's start over. Okay, so the name of this is called SCONE, S-C-O-N-E, and it stands for Smart Contract Exploitation Benchmark. â They, to avoid like data contamination, because if you've like seen one vulnerability already, you know about it. â They also included for a little subset of 34 contracts that were exploited after that cutoff point of March 1st to see how those would do to see if the AI agents could still exploit those or not. So â they ran agents, AI agents. So they used Claude Sonnet 4.5 and GPT-5 and they looked at clean, like relatively clean contracts. got 2,849 newer contracts that have no history, or at least publicly, of exploitation to see if the AI could find anything. So here's how it performed. Out of the 405 that we know are vulnerable, somehow their AIs found 207 exploits and got them written and working. They did not actually exploit them because that would be illegal. Um, but if they had, they would have made $550 million.
speaker-0: What? What are they doing, man? You gotta fix that pipeline. You know what I mean? Money pipeline.
speaker-1: Yeah, I mean, so we're talking some pretty big money. I guess from this entire research project total, they said that they found $4.6 million worth of total assets that could be exploited. â Really? And then if you... I know, like,
speaker-0: Tell everyone you And Pond's over! Pond's over! I got work to do! I got a huge round of funding coming in for this project!
speaker-1: He's like I got something come up. Yeah Crazy, yeah, so about half, they got 51 % of the known vulnerable smart contracts cracked. And then out of that subset of the ones that were exploited since about 55 % of those were also, â had demonstrated Volns that they're able to figure out. And then on the fresh new set, the... the larger set that was almost 3,000 contracts that were newly deployed and have no exploits against them that are known. They found two zero days, â including like the exploit scripts and everything for it. So â the problem is this kind of, you I think anytime you have a new technology, there's always the theoretical issue of, what if it's hacked, right? And this just took this out of the theoretical realm and was like, boom, you cannot ignore this. Smart contracts are. vulnerable, which is problematic, right? â So yeah, we're kind of entering like now we've got a race scenario between your defenders and attackers. â And then they also mentioned like the amount of profit because they also racked up like big bills with GPT and Claude, but because the tokens aren't too expensive, it's actually still profitable. So, and quite so. that... kind of lowers that cost of entry for potential attackers. yeah, it was interesting. This was done in a simulation. Anthropic says they did not take anyone's funds. They did not actually break the live blockchains. â But yeah, it looks like we're gonna have to look. So they actually advocated for proactive use of AI for defense. Going back to that topic, so.
speaker-0: Yeah.
speaker-1: Interesting to see what comes up next.
speaker-0: Yeah, I thought that, you know, Cybercom funded an AI hacking competition over the last year or two, which they convened at DEF CON this year. And essentially, each team built a system which would find vulnerabilities in code, and then it would also... write the exploit and also write the patch and then submit the pull request for the patch. And several of the teams had to do delay route, part of the requirement for winning the money was you had to open source your project. Several of the teams had to delay open sourcing their projects because they ran them against public, popular public projects. on GitHub and they found lot of O-days, right? And the maintainers had not fixed all the O-days that the systems had found yet. And so they just didn't want to like recklessly release it. â to give attackers another leg up, right? So, I mean, I think it sounds like Anthropic took that same idea and then they like implemented on smart contracts. And my understanding is like the smart contracts you can customize, right? To do almost anything. so yeah, so I mean, I could definitely see that working in that ecosystem and, you know, potentially being, you know, very effective, right? So, especially like imagine someone builds an app and has a specific purpose and it's using smart contracts to like do transactions in my backend and then every transaction using that app ends up being vulnerable, right? So I, yeah, you know, I think we're, we are definitely at a... Like, you know, back in like the year 2000 when web hacking, like people started to figure out that you could do like SQL injection and things like that on websites. And like you figured out like every website's vulnerable. Basically. I feel like we're right back at that again now by just applying AI to looking at code bases.
speaker-1: Kind of like Wild West out here.
speaker-0: Yeah, because we just haven't had the manpower to do in the past. mean, I think anybody who's implemented a static code analysis system will tell you like, yeah, it finds vulnerabilities. But the problem is, like, one. Logic flaws in the past have been pretty difficult to identify because they usually span multiple files across like a source code project. And LLMs are, they're, really good at aggregating that data together and coming up with those conclusions. And then two, weeding out the false positives has just been way too time intensive. Like I know several companies that have just teams of people to review static code findings. And then there's even managed services that are out there by providers, which just look through static code findings to try to weed out the ones that are false positives. And. I mean, I think now an AI can probably do half that work, right? So â you can get higher fidelity faster. So yeah, it's crazy world, crazy world. I'm glad you found that. That's a cool find.
speaker-1: I will clarify like the success metrics because they didn't actually exploit it. They said, can you get a script that would, you know, they asked the AI, can you make a script that would exploit this? And they look at the simulated drain potential of how much money you can yoink out of that. â So in the real world, it might not actually be quite as much as they're claiming. I think they're kind of looking at a little bit of like a best case scenario. And the good news is that even though they did find some zero days, it was only two. out of those 2,849. So â it's something to be taken seriously, but it's not like red alert quite yet. So that's good.
speaker-0: Yeah, not like the open AI red alert that we saw this last week. Did you hear about this? Where they're basically like Google's kicking our trash. So they sent out a memo inside of open AI saying red alert. We need to focus all of our focus on chat GPT and make sure we win this race against Google. â and then someone of course leaked the email, right? And so then that ended up being all through the news that Sam Altman is scared of Google, which, you know, probably was probably was exaggerated. Right. But I mean, you know, any company like a Google or a meta, anybody who has a lot of revenue coming in right now and is able to attract top talent and retain them. mean, yeah, they're going to they're going to be a threat to like an open AI or entropic or any of that. I actually I really like Anthropix business plan. It feels like they've somehow wiggled their way into becoming the next Microsoft because they're like, like, we're not going to focus on consumer anymore. We're just going to focus on businesses. And it feels like exactly what Microsoft did back in the day. And then they just ended up dominating the market. Um, so I'm not saying open AI is a bad bet, but you know, if I had you know, if I had a million dollars, I'd probably lay that million on Anthropic instead of OpenAI. Just because I think going after the business market is where more value is going to get derived from these products. But I don't know. I, I, that's very speculative. So anybody who trades on that opinion, don't blame me when you lose your money. â okay. So speaking of losing your money, short article, short blurb, insurance companies are saying, It's too risky to ensure IT systems that are highly reliant on AIs to make decisions. And they essentially say it's too much, their quote is too much of a black box. And they want exceptions basically that say if anything goes wrong and it relates to AI, we don't have to pay out the insurance policies anymore. â You know, we saw. The solar company sued Google last year, â $110 million lawsuit, because the Google AI overview said that the company was in legal trouble when the company was not. So when people were searching for it, they just claimed like, hey, you hurt our business because you told people we're in legal trouble and we're not legal trouble. that was those, you know, that even more Google like pushed out the AI summary thing and it was like pretty bad for a while. It's better now. Like it's not that bad now, but â you know, it was bumpy there at the beginning. And then Air Canada. Someone tricked their chatbot into giving them basically a hundred percent discount on all the flights and they tried to sue the individual and the court ruled no, like you have to honor the penny flights
speaker-1: individual do they want to take me on a vacation? I think we're best friends!
speaker-0: Yeah, exactly. It's a heavy discount on these flights. What? And, you know, I think we've also seen a lot of the fraud going on with the video calls, right? I mean, we talked about the bank in Asia where kind of like the zoom call and the individuals on it, their video footage was all fake except for the lowly accountant that was tricked into transferring funds. This article had another story in it that I hadn't seen previously where it talked about actually a London based design firm lost 25 million. due to a very similar scheme where the individual on the other side of the video call was impersonating an executive at the company and convinced somebody to transfer 25 million. So, you know, they're basically saying like, can't issue insurance policies that cover all these edge cases now. So they're trying to get caveats in the insurance policies saying like...
speaker-1: So is insurance saying they won't insure like a tax from AI to your company or is it only if your company has AI that it's using? That's not what's insured.
speaker-0: I mean, it seems like obviously they're going to push for everything they can get. So it seems like right now they're like, if it's has anything to do with AI, we don't have to pay. That's like what they're trying to get in law. And this is a tech crunch article. We'll put the link below if you guys want to check it out. Now, you know, that hasn't been held up in the court yet, right? So they don't have that exception yet, but they're trying to reword their policies to essentially say, you know, if AI is used as part of the attack that, you know, I think, yeah, you know, I mean, part of the article talks about it being too much of a black box, which implies that. If you're using AI in products, which most companies are going to implement AI in the products because that's what consumers are wanting, that's what businesses are wanting, right? That if the AI makes some bad decision, you can't use the insurance policy to pay out, right? But I mean, think there's this other issue of like the digital cloning, right? And I don't know how the insurance policies relate to that, but maybe those are like exceptions that are put into like the theft policies and things like that. â Yeah, they just say like, â what terrifies insurers isn't one massive payout. It's a systematic risk of thousands of simultaneous clients' claims with a widely used AI model â in it. So essentially, they're just worried they're gonna get inundated with claims and then like the company's gonna have to fold, right? So. which is kind of what we saw happen with the cybersecurity insurance. It felt like for a couple of years there cybersecurity insurance was super hot and like everybody, every company was like, you gotta have it. And I'm not saying it's a bad idea, but then a lot of insurance companies stopped carrying it and stopped issuing the policies. Cause they were just, they were losing money on it hands over fists. And you know, even I know several insurance companies that became targets of ransomware groups, not so that they could get ransom, but so the ransomware groups could know what the policies are. and they would target the people that had policies with them, because they knew they could pay. they could almost be... That's rude. Yeah, it almost became like a self-fulfilling policy. It's like you were scared of a ransomware attack, so you get an insurance policy, but then the insurance policy company is already compromised by the ransomware group. And so then they see you got a new policy and they know the insurance will pay out 10 million or whatever. So then they just go hack you and then they get the 10 million from insurance and it just like loops back around. So I... Yeah, I definitely, I could see why insurance companies would be really hesitant on the AI side. But with that being said, I don't think we're going to get much of a choice. think it's, if it's a better overall product is going to get, if it's going to enhance the user experience and drive more value to customers and businesses, then, you know, if you don't implement it your product, you're probably going to get outcompeted by somebody else who does. And if you do, I mean, you're gonna have to assume the rest, I guess. I don't know. What do you think, Shelby? How do you feel?
speaker-1: I mean, makes sense. I mean, insurance has done this in the past, right, with other aspects where if something is just categorically too heavy for them to carry, they say, yep, everything except for that.
speaker-0: Yeah, yeah.
speaker-1: found out about my home policy, I tried to put in a claim for an insurance claim and they're like, yeah, it doesn't cover that. I was like, well, what about this? They're like, it doesn't cover that. I was like, can you give me an example of something that covers the like fire? I was like, can you give me a second? Like fire? I'm like, oh, okay. That's what I'm paying for. I thought I had a comprehensive.
speaker-0: You No, insurance is like education. It's like, no, no, no, that's a bad analogy. So the insurance is like, you have to pay it and then they're going to fight tooth and nail never to pay you a dime, right? So regardless of who's at fault, that's the way I feel about insurance. Honestly, I feel like insurance is slightly a scam, right? Like I feel like if you run a company, Other companies won't do business with you unless you hold insurance, right? But then if anything goes wrong, it's hard to get the insurance company to pay you any money, even if clearly that's what the insurance policy was designed to cover. So I don't know, I hate it overall. feel like you have, it's like a cost of doing business, but then you just can't bank on it actually being like a safety net.
speaker-1: They're like all these prerequisites. It's like you have to meet the criteria, but you might not know about all the criteria or like yeah
speaker-0: Well, it's intentionally designed so that you won't meet the criteria so they won't have to pay you out, right? And then it's like, even if you do meet the criteria now, if they pay out too much, they're just going to change the criteria for next year. And like, are you really going to look at the policy change from year year? Like, no, you got better things to do than read that. And you're just going to be like, I accept and go along with your life. And then when you get messed up, you'll be like, I thought it covered this. They're like, no, it doesn't cover that anymore.
speaker-1: I everyone's had that feeling before I think a lot of people relate of being like what
speaker-0: It would be not a huge thing if the insurance policies didn't... Some of these insurance policies cost a lot of money to carry, right? Especially for your business and things like that. You write a six-figure check for insurance, right? then you know, that's not... Is that right? weeks? Yeah. So like, and then you know that's not gonna do anything for you. You're like, I'd rather just take that money and have it in the bank account for when something goes wrong, right? So. Yeah, anyways, but so it goes. That's life. So the, you see anything fun or any fact toys over the last week? You wanna share with us, Shelby?
speaker-1: crummy I did learn some, well, one fun thing. I found out that there are, â there is a phobia that some people have. I don't know anyone with this fear, but it would be kind of fun to just observe. Anyway, a fear of long words. They get anxious and feel shameful or nervous when they come across a large word. And would you like to know the name of this fear?
speaker-0: Hmm
speaker-1: It's got a specific name like claustrophobia and other things like that,
speaker-0: I'm gonna guess it's also a long word.
speaker-1: It's 36 letters long.
speaker-0: â you gotta be kidding me. That's so cruel. That's so cruel, because the person who has the fear of it is probably the person who's hearing the word, because they're the one being diagnosed with it, which then is triggering their own diagnosis again. Yeah. It does not seem healthy at all.
speaker-1: There's a word Hippo in it. I've seen some funny skits about they're like, all right, I've got your diagnosis. It's the person panics. It's okay. You ready for it? Yeah. We're to look on my screen because I definitely can't say this from memory. Hippopotamontrosis, says quipedalia phobia.
speaker-0: That's a mouthful. That sounds like the Mary Poppins song.
speaker-1: Who puts the word hippopotamus in it? That was rude! was written by someone who didn't like someone.
speaker-0: That's cruel, That's messed up. They need to come up with that. at least, I hope they at least have an acronym for it or something, man. Because that's cruel and unusual to say that to someone who has that disability. â man. Whoa.
speaker-1: Hahaha Okay.
speaker-0: I want you to know last night I opened the Sora app. I don't know if people know what the Sora app is, but I'm sure you do. It's where you're... It's OpenAI's social media app, like a TikTok, but you can generate videos of your friends as long as your friends allow you to use their likeness. There's a lot of videos of me in there. Some of them were a little bit disturbing. That's all I have to say. Like one, I was a vampire. I don't know where that came from. Another... I was on a date and I don't think I'd normally be on this date. â so there's a lot of videos in there that are pretty interesting. But, â it's kind of fun. I just think it's kind of like weird because like random people can use your likeness if you, if you allow them to in the app. And so then I just feel like there's like people, I don't know who they are and they've like put my likeness in videos. Which the app lets you pull videos down that use your likeness if you don't... Yeah, you get veto power, right? But then I feel bad, because I'm like, someone worked really hard on making the video and put me in the video. And yeah, I don't really agree with what's going on in the video, but I don't want to like be a meanie and just like destroy the video. You know, that doesn't seem cool either. So...
speaker-1: I them.
speaker-0: Yeah, they're out there if you want to go check them out. So, â I f- I will put the link to my Sora account down below if you want to check out the videos. â but I will let you...
speaker-1: And then we come here, right?
speaker-0: dig through the videos so you can find the questionable ones yourself. So I'll leave that. So give you a little bit of surprise. â But yeah, the app's pretty fun. I like seeing videos with my friends in it, you know? Like we're like all doing stuff together that none of us actually did, right? That's kind of fun. But like Rando's just putting you in a video, that's kind of weird, right?
speaker-1: Can you change your privacy settings if you want to?
speaker-0: Yeah, you can, yeah, you can, but as a cybersecurity person, I just let anybody use my likeness. That sounds like most like sounds like the best idea, right? Shelby. And then throw the word hippopotamus in there.
speaker-1: What a sentence.
speaker-0: my hippopotamus likeness and â yeah, no it's fine, I'll put the link below, you check out the videos, I â Yeah, I don't agree with them all, but there's some there.
speaker-1: Well, thanks for joining us on the pod, Bryce.
speaker-0: Yeah, alright, so see you guys all next week. Thanks, bye.
speaker-1: Bye.