AiCyber.Land

Should using a powerful, uncensored AI model be like eating a deadly pufferfish? This week on AI cyber.land, we dive into the explosive debate around "obliterated" AI models that have their guardrails removed. We explore the massive cyber security risks they pose and the high-stakes arms race between AI-powered attackers and defenders. Are we heading for a future of unstoppable, machine-speed cyber attacks?

Show Notes

Should using a powerful, uncensored AI model be like eating a deadly pufferfish? This week on AI cyber.land, we dive into the explosive debate around "obliterated" AI models that have their guardrails removed. We explore the massive cyber security risks they pose and the high-stakes arms race between AI-powered attackers and defenders. Are we heading for a future of unstoppable, machine-speed cyber attacks? --- IN THIS EPISODE: Welcome back to the pod where AI meets cyber security! This episode kicks off with a spicy question: If a world-class chef can serve you deadly fugu, should developers have access to equally dangerous, uncensored AI models? We break down the rise of "obliterated" models—AI with its safety features surgically removed. These tools can be weaponized for cybercrime at unprecedented levels, and we discuss the services already offering them on demand. This isn't theoretical anymore; a recent Unit 42 report shows ransomware gangs are already using AI to automate their attacks, moving at machine speed. But the good guys are fighting back! We look at how companies like Crowdstrike are baking their own specialized AI models to detect and remediate threats instantly. We also unpack the frustrating "haves and have-nots" dilemma, using Anthropic's new Fable 5.1 and Mythos 5.1 models as a case study to show how safety guardrails can actually nerf an AI's performance. Then, Shelby takes the lead with a deep dive into Cloudflare's latest security capabilities for MCP (Model-Client Protocol) traffic. Learn about the dangers of "Shadow MCP" and "Portal Bypass," and discover the step-by-step process for locking down your AI agent environment. From client hooks to gateway policies, we cover the essential layers of defense you need to know about. Stick around to the end for Bryce's latest Lego creation (built during a Zoom call) and Shelby's hilarious (and slightly painful) hammock mishap! --- KEY MOMENTS: ⏱️ KEY MOMENTS: 01:16 - Is Uncensored AI Like Eating a Deadly Puffer Fish? 02:24 - The Dark Side of AI: Uncensored & "Obliterated" Models 06:05 - How Ransomware Gangs Are Now Using AI to Attack Faster 08:51 - Are AI Companies Keeping the Best Tech for Themselves? 11:25 - Securing AI Agents: A Deep Dive Into Cloudflare's MCP Tools 19:44 - Key AI Security Threats: "Shadow MCP" & Portal Bypass 25:24 - The Secret to Surviving Long Zoom Calls --- JOIN THE CONVERSATION: What's your take on uncensored AI models? Should they be available to researchers and developers, or is the risk too great? Let us know your thoughts in the comments below! If you enjoy our breakdown of the latest in AI and cyber security, be sure to hit that LIKE button, SUBSCRIBE to the channel, and ring the notification bell so you never miss an update. AICyberSecurity #ArtificialIntelligence #Cloudflare #UncensoredAI #CyberDefense #Ransomware #AIThreats #MCP

What is AiCyber.Land?

Join industry experts and thought leaders as we dive deep into how artificial intelligence is transforming cybersecurity, shaping defense strategies, and creating new opportunities in the digital landscape.

Hey, welcome back to the pod. This is the AI cyber.land podcast where artificial intelligence meets cyber security. Sometimes the reverse of that, too. And uh we got the world's greatest co-host, Shelby. So, uh, you know, sit back, relax, and, uh, let's talk about all the breaches, breakthroughs, and everything else in between that's happened in the last week. Before we get rolling, I just want to say I've noticed that some of you could still click that subscribe button. So, if you could do so now, I'd really appreciate it because, you know, I just want to beat my friends. That's pretty much it. I want to get more subscribers than my friends and I'm kind of in like a competition right now and uh, I'm losing. So, if you could click the subscribe button, I'd appreciate it. Okay, the more more on that at the end. Okay, so Shelby, let me ask you a question. >> Should it be legal for an individual to eat a fish? >> Yes. >> What if I were to ask you this question? Should it be legal for an individual that goes to a world class chef at a restaurant to eat the deadly puffer fish. >> Sure. >> Do you know what the deadly puff? >> As long as it's not endangered. Yeah, the fugu. >> Yeah, the fugu. Right. Where it's like the skin, the whatever are like all toxic, so they have to like cut those out. If they don't do that correctly, then you could die within minutes. Maybe sign a waiver first. >> You're like, sign a waiver. Okay, just everyone remember Shelby thinks it's totally okay as long as you sign a waiver. All right. Well, I've been I feel like there's kind of like this storm brewing in the AI space, right? where uh people are taking the models that are, you know, have guard rails on them, right, for general public consumption and then they're removing the they're removing the guardrails. Uh one of the techniques for this is called obliteration, right? Where they basically kind of tweak some of the weights inside the model so the model will never say no, right? But then, you know, you could use the model to do some stuff that the model provider may not want you to do, like ask question, >> hugging face, >> like hack hugging face, things like that. Just little things that would theoretically never occur. Um, >> yeah. And you know, I've actually noticed that there's a couple services online that will actually provide you with access to the obliterated models on demand, like via their APIs. I'm not going to name the names of the providers here just because I don't want to like help a cyber criminal or an evil person, but I'm sure if they Googled it, they could probably find it. So um so you know some of the other model providers like open router they won't route you to these obliterated models but then you know you know for every yin there's a yang or something right so so there's other model providers that are like yeah we're more than happy to route you there like a models model you know >> so what do you think should you be a should like model providers allow you to use these uncensored models Uh, I see. Okay. So, it does seem problematic, right? Because you've just lowered the barrier for someone to be able to commit cyber crime at unprecedented levels and possibly without even knowing it. We've got some of the world's best minds who accidentally attacked other people, right? So, I do see it being highly problematic. It does seem like it could be incredibly disruptive. It It's like it's completely weaponized, which seems like it would be a problem because even a well-intended researcher with good intentions and um a legitimate, you know, research project could accidentally have it go a little A-wall, right? >> Yeah. So yeah, I can see why that it wouldn't be good for everyone to just be having these unrestricted versions of the models. >> Yeah. I mean, I think we could look at this in relationship to a couple other industries, right? Like obviously, you know, a chef could cook the puffer fish and serve it to you as long as they're doing it like carefully, right? Um like in another industry like the locksmithing industry, right? they go and they open up locks for people which could potentially be a bad thing but you know if it's your house and you own it like that's fine right so but generally they have like certain verifications or accreditations to become like a licensed locks smmith right so you're just not like a criminal um you know and then there's like you know other industries like you know to be a firefighter to like fight a fire you don't really need to know how create a fire, right? So, um, which is an an argument I've heard before, but I don't think it necessarily holds up in spy cyerspace to be honest. Um, so look, I think the the fact of the matter is, right, like one way or another, the cyber criminals are going to get access to these like unfederated models, right? So, we as cyber defenders, we're going to need to still prepare against that type of attack vector. Um, one thing that was of interest that I saw this last week was a unit 42 report that talked about how uh ransomware operations are integrating AI more into their operations to automate like the lateral movement and uh collection processes. uh so they can move faster like more at machine speed uh which is something we've talked about a lot on the channel. So I think that in combination with like these uncensored services and models becoming more and more accessible to people is you know going to greatly increase like the speed and veracity of the cyber attacks that are going to occur. Now I you know I do think there are certain cyber defense companies that are like heavily investing in capabilities to defend against these things at machine speed as well. But a lot of those capabilities are really to be determined whether they're going to be really effective or not. I know this last week Crowdstrike had their big conference. It's called Falcon in Vegas and they actually unleashed a new AI model that they had baked in house that they're going to integrate into their EDR agent that is designed specifically to kind of detect and remediate um at machine speed. So, uh they actually went and like baked their own model specifically for that. I also know that there's new startups out there which are baking models uh and that can run locally on devices that are specifically designed to figure out like is the activity that's on the system uh within the baseline of what the user typically does kind of like a next generation user behavior analytics type play. So, I mean, there is definitely like a cat and mouse game here in in the works and there's like a lot of money to be had obviously and whoever can figure out like how to defend against these attacks that are, you know, occurring today like now and going to become more prolific in the future. But, you know, I for one am like not not a big fan of like censorship. You know, I feel like, you know, if you identify who you are, like, hey, I'm Bryce and there's reasonable controls and so that way people can say like, okay, if something bad happens with that account, we can go hold price accountable, then >> in my opinion, that's usually like sufficient uh for like a type of guard rail. Like I I don't I don't love the like fact that you know the model providers collect all this data across the entire internet and then they give themselves an unfettered version and then they provide us with like a version that's one slight degraded like if we look at there's a release of uh they just released Anthropic just released Fable 5.1 as well as uh mythos 5.1 but mythos 5.1 is only available to cyber security companies. uh that are doing cyber security work. But when we actually look at the raw benchmarks between those two models, the only difference between those two models according to Anthropic is the guard rails that they implemented into them. >> Yeah. And there is a sign not I mean it's significant enough that it's obvious in the benchmarks performance decrease in fable 5.1 over mythos 5.1 because it has to waste its resources on thinking and about the guardrails. The guardrails have essentially like lowered the capability of the model. Now, it's not like so much that it's like impairing, but I mean, it is obvious on the benchmark like the only difference between these two are the guardrails, right? That's what they're saying. And if the only difference is the guardrails and one performs a lot better than the other, then obviously it's the guardrails that are impairing it. So, I I uh I don't know, that's kind of frustrating, right? And it's kind of frustrating because it goes back to that argument of like the halves and the have nots, right? So, like I feel like a lot of these companies are kind of keeping the cream of the crop for themselves, which is understandable, but at the same time frustrating if you're not in that click, you know, which most people are not, right? So, anyways, that's my weekly rant on all things about how Daario is an evil person. And uh I didn't really mean to go down that track, but apparently that's where I ended up again. I don't really hate Daario, but uh I don't really love the industry. All right. Anything else you want to say about this? >> Well, it just reminds me how much like the the it's so stacked against defenders. Like, >> yeah, >> you have to have a plan for this kind of stuff before it happens cuz it happens so fast. So you've really got to have a a mitigation plan in place to kind of shut something down that at machine speed. >> Yeah. >> Yeah. It's going to there's going to be some bumps over the next 24 months for sure. But I do think you know it's a large enough problem. There's enough money being thrown at it. It will get solved. It's just there's going to be some casualties in the meantime unfortunately. So anyways, on a more positive note, have you seen anything in the last week? >> Yeah, this one is about two and a half weeks old, but um I kind of started diving into Cloudflare's new um capabilities that's part of Cloudflare 1. Um and it's really just focused on MCP um traffic monitoring as well as applying um controls at different parts of that. So this is this uh Cloudflare 1 tool is expected to be used like in conjunction with others as a multi-layered approach. So there's already like an MCP server um portal and what that does is you can take multiple MCP servers basically throw it behind um just one single HTTP access point. Um and that kind of just streamlines access simplifies things a little bit. Um you can also apply um like you can customize which tools go can go to which portal or things like that. So you can make a couple of customizations there. Um you can also have some logging um on at that level as well um for individual requests and that observability will help with like DLP efforts as well. So these controls are intended to help admins um see if agents are following approved paths or if they're somehow circumventing them. So I'm going to talk about three places that you can apply MCP um controls. This is in addition to logging because logging tells you what happened but these are places where you can actually block something that you is against your policy or something like that. So first bot is inside the MCP client. Um so this is like a client hook that um it can work really well including for um even like local um servers as well that never cross the network. Um the limitation is that this can be kind of tricky to implement and keep it like standardized um and and maintained. But it does give um you a place to implement like an allow list of servers. Okay, you can talk to these MCP servers. Anything else is just a no. Right? So that's spot one. Spot number two is looking at the devices um network boundary where you have other controls. Um so this would be like a secure web gateway um or something similar to that. Um that just gives you observability which means that if you have like TLS traffic, you have to be able to decrypt it at that level to really see what's going on inside there. Um the limitation is that proxies can't see what's off network and um you also can't see like what is it std IO so like if it's just local you won't be able to see that because it's not going across the wire right >> um third place this is your this this third one is the last spot where you could logically like implement a control that would actually stop something that you didn't want to happen and that would be before the MCP server invokes a tool. Um, so with this you're the the cool thing is like by monitoring at the server level, the MCP server side, you've got really rich context because your user's already authenticated. Um, you've parsed the MCP message, things like that. You've got a lot of context that you can then use to make informed decisions. Um so the kind of tool we're talking about here is something like an agents handler um or similar middle middleware I guess um and here's where so what are the benefits of this you can apply uh rate limits you can apply um inspect the arguments as well as um authorize the caller for specific tools right you can put some logic in there or or lists so what Cloudflare uses is called right guard card. Um there there isn't a limitation with this as well. Once again, you're not going to see um like servers that didn't go across uh that that didn't create traffic to it. Um and really it's only going to protect the servers on which you're putting those controls on. Um so you have to apply it to each server that you want protected, right? Um, one of the challenges just speaking generally about like your MCP like security um, approach is that we've got different versions of the specs. And if you look at just the URL, it doesn't automatically tell you whether or not that's MCP traffic or not. It is getting better. So, there was a new spec um, released in July of this year that has a like it's a little bit more fleshed out. It really changed the schema there. Um, so it's getting easier, but like step one in like monitoring your MCP traffic is like kind of figuring out is this MCP traffic just because there are a couple different standards that are being followed. But yeah, anything you want to add so far, Bryce? >> Yeah, they did a major update to MCP like a couple months ago and um the protocol itself is significantly better now. Like so if you used MCPs in like a year ago versus like a month ago, if the MCP server is using the newer technology, they they will run a lot more token efficient and they they will that they're better. So essentially they move to a stateless design which is essentially how like the other model providers have been operating for a long time. Um, so it's kind of more in line with the way like OpenAI, Anthropic, and all the everybody else's APIs are operating already. So that's a plus. Uh, it's also like, yeah, a lot more token efficient. I I do think like with stateless designs, there are specific security risks associated with those. Um, albeit like I haven't really fleshed out in my mind if those would be applicable to the MCP environments. I kind of feel like most of those attacks are specific to like the model provider like the models that are being served up. So I'm not sure that really matters much. I also just want to say like yeah there there is a lot of different ways and even like deprecated ways for MCPs to operate. So like not all MCP servers or MCP services are like are equal, right? There's a this is very much like a changing technology like like yeah you can do MCPS over standard input like standard IO you can do MCPS over anyways like a bunch of different ways. So I I think it's great like that one that they're still making updates to it so they can stay relevant and two that providers like Cloudflare are like looking at those and saying like how can we actually get more visibility into this cuz when you have an AI agent it essentially has you know two points like if if you're going to say the model is a black box and we don't have the tools or capabilities to look inside that black box 100% then you know, communications to the model providers would be like a good way to hook in through like the light LM proxies or like open routers or things like that. But then uh you know, you also want to hook in where it talks to tools and having these like MCP proxies there give you like a lot of capabilities that I mean there there are providers that have had some of these capabilities in the past, but they're not they're not widespread. And I only know like, you know, a couple companies that have actually looked to implement it that the tool the tool interception type calls. And I think you want to do both if you want to have a secure AI agent environment. Anyways, there you go. I I don't know. >> Exactly. Yeah, you're getting ahead of me. Fantastic. >> Sorry. Sorry, I'm I'm wavelength. >> So basically just want to say like Cloudflare gateway. Um so step one, it has to identify the MCP traffic. So it's got a new detection heristic. There's just a boolean selector. I think it might still have like the experimental label on it or something, but it's been applied to everyone's traffics already for a couple weeks now. Um, so step one, figure out what's what's MCP traffic. Step two, we can start evaluating it, right? So, I want to talk about two like possible problems that you want to kind of be looking for when you're evaluating your MCP traffic um just generally from um this perspective. One is shadow MCP. That's when um someone inside of your organization is connecting to an unapproved server. Um this >> I'm sure none of us would ever do that, right? We need proof tools. Okay. But if you are a security team trying to implement controls here, um there are a couple concerns if that happens. Number one, uh you don't know what's being which tools are being exposed, right? Maybe that's supposed to be internal only and you don't know which what data is being sent out. Um, so how does this even happen? How does shadow MCP happen? You know, it's, you know, maybe someone just read something on the internet or got advice from someone that said connect to this server and then they just add that directly to their MCP client. Um, or maybe they were social engineered or whatever it was. Um, okay. So that's what it is. Let's talk about the solution. Um, CloudFare Gate Gateway is the solution when you're talking about like managed networks because that's a a spot where you can apply policy, right? You have an allow list. Um, the second concern with MCP is portal bypass. So maybe some people are listening to this and thinking, I don't even have a portal, so I don't care if people bypass the portal. But let me tell you what we can do with a portal. >> Okay. >> Um, so port portal bypass is exactly what it sounds like. Basically, you have an approved MCP server um in your organization somewhere, but someone is connecting to it upstream, right? You want them to go through the portal and instead they're going directly to the upstream URL um and plugging directly into that instead of through your designated portal. The concerns are if you're using your portal to its um the to its benefits um what you're getting there is access policy possibly a curated tool catalog um data loss prevention tools and tool level audit trail. So that's like that point where that all happens. So if you've got people going out the other way, you've you've got blind spots, right? Um so what's the solution for portal bypass? Um this one is a little tricky. It requires network control in addition to an origin that can like your server has to be able to reject direct requests and there are a couple different ways you can do that whether by source IP or whatever it is. So um so Cloudflare's new tool and capabilities they've they've made a dashboard which is nice because it gives you a path from discovery of like what's my traffic look like all the way to governance to being able to apply your policy and enforce that. Um so you can like find where the servers are. You can approve them, block them or um close pathways around um around those portals. Um currently there is a limitation where if you've got um if you've got a server that's on a private or internal only network that is not going to work behind Cloudflare's portal because it's not going to hit Cloudflare, right? So, um, so they are working currently to develop something that where you can have your portal be like the one-stop shop and even it can go to a uh like something that's on the web or to a server that's internal. Um, so that is currently in development. They're also working um on on making some additional like granular functionality and visibility. But um just wanted to throw out there a couple steps if you're trying to establish an MCP security program in your organization. Um overall generally you want to understand your traffic, how the MCPS are being used and establish your approved set of tools. And so number one, inspect your MCP traffic. Um, number two, move your approved servers behind those portals, those MCP portals. And then number three is enforce your boundary. So that's going to be closing off those um those alternate routes that aren't giving you any benefit and then creating and you do that by creating your gateway policies. So that is MCP from Shelby. >> Yeah, that's awesome. That's Yeah, that's great. I, you know, I know there's a space where there's a lot of updates rapidly, so it's good to, even if you've heard some of this before, get a refresher on it all. And I definitely learned some things when you're talking about it. I didn't know some of these capabilities existed. And yeah, I I also, you know, I'm I'm aware that Cloudflare has Cloudflare tunnels, um, which might be something that might help you like expose an internal AI agent to the Cloudflare ecosystem, but and and I do know like Cloudflare has like their own AI agent framework that allows you to run the agents in their cloud as well, but >> Oh, sweet. I don't know if like I that I haven't seen really anybody use that. So I I'm not sure that's got a lot of adoption. So but uh anyways I I'm a big fan of Cloudflare. I I I really like their services and I think they do like they're generally very security conscious. Um so I yeah and they're kind of forward thinking. seems like they're definitely have some smart people on their team and they're able to bring those that intelligence to the market, right, through like actually exposed services we could all use, which I always appreciate. >> Yeah, I feel like they're definitely a leader in the space, right? >> Yeah. >> How that looks. >> Yeah. Um, well, if you have any thoughts about the following, you should leave in the comments below, especially if you have additional questions about MCPS or any of the things we're talking about so that we know what you want to hear in future episodes. So, uh, with that being said, are you ready for my fun thing of the week, Shelby? >> Yes. Tell me. >> Uh, you're gonna be really impressed. Ready, set. >> Okay, >> go. Boom. >> Boom. I know everyone that's only on audio only can really enjoy this. But it's a Star Wars Lego uh vehicle. Uh, so yeah, it's got one of the bounty hunter guys and uh and uh the guy that you think is Boba Fett, but is not actually Boba Fett. He just found his armor in the in the sand. I remember that for the episode. That's as much as I remember. Uh but yeah. Anyway, so I got some extra time. Well, no, I didn't. I I was on some Zoom calls, let's be honest. And people couldn't see what I was doing below, even though I probably obvious clicking Legos. >> Click, click, click. >> Um, so I built that set while I was stuck on some Zoom call. just want to be on a call with you, a work meeting one day and just call you out for it. Be like, "Bryce, is that the click clacking of Legos?" >> I'll just say right now, not the worst thing I've ever done on a Zoom call. I've I've definitely like uh I'll just Yeah, I I I once got called out because of the like I had the monitor back here and they could see on the Zoom call the reflection on the monitor to like what I was doing, but I was like, you know, straight up just playing Nintendo Switch, right? Like on the call cuz it's like one of those calls where it's like they need me for like five minutes, but they're going to make me wait through like 50 minutes of everybody else talking and like >> Right. And that probably wasn't the right decision, but I don't know. I was like, >> it was a Friday and I was burned out. I was like, but uh anyways, >> I have some practical uh tips to share with our viewers as well. >> Yeah. >> Yesterday I decided I would like some relaxation. And so I pulled out my hammock and I set it up in my basement. >> Nice. and it was really nice for the first seven seconds and then it broke and I fell. So my advice to you today is if you're setting up a hammock, just put a pillow under your bum in case it breaks. That's all. >> What What was this a fix to? Did you just like put this on some books or something? Like what what was this set? one of those cool tall um beds. >> It's and so it's like but you sleep up here by the ceiling and then underneath it in a diagonal shape I put a hammock and it works fine usually but the hammock is getting old and like the where it attaches to itself to make a loop it just slipped out. >> Oh. >> So it wasn't my fault. I blame the hammock. >> Yeah. Well, I blame the hammock too in that scenario. That seems like you had reasonable controls there. you, me, and my tailbone are all aligned. >> Yeah, that's the worst part. You could be right, but then still be in pain. So, it's like doesn't matter that you're right. >> Seriously. >> All right. Well, that's >> that's all we have for today. >> Yeah, that's a wrap. Uh AI and cyber security move fast. We're here to keep you ahead of it. So, we'll see you next time. Bye.