Join in on weekly podcasts that aim to illuminate how AI transforms cybersecurity—exploring emerging threats, tools, and trends—while equipping viewers with knowledge they can use practically (e.g., for secure coding or business risk mitigation).
Hey, everyone, and welcome to this week's episode of AI Security Ops. Here, we are going to go into part two of OWASP agentic skills top 10. Previously, we covered one through five. Now gonna hit those bottom five, six through 10. But before we get into that, let's talk about our sponsors, Black Hills Information Security.
Brian Fehrman:If you or your organization are in need of any kind of computer security services, whether that's external testing, internal testing, assumed compromise, social engineering, wireless testing, physical testing, SOC services, continuous pen testing, last but not least, testing AI security or using AI to augment security testing to bring down the cost without sacrificing quality, check us out at blackhillsinfosec.com. Additionally, we have a training branch, Antisyphon Training, where many of our consultants take their knowledge that they are implementing day in and day out. They package it up in a nice, easy to digest, and affordable format for you to consume to hopefully help you level up in your current position, get that role that you've been looking for, or maybe you're a hobbyist who just wants to learn something new at affordable price. Check us out at antisyphon training dot com. So just recapping a little bit.
Brian Fehrman:OWASP puts out a lot of top 10 list, which are a great resource, and this one in particular is on Agentic skills. Now as I mentioned previously in in a prior episode, we covered the the one through five, because putting in 10 on one episode is kind of a lot. But before we go into the last five, let's just recap. What is a skill? What would you say, Bronwen?
Brian Fehrman:Kind of your rough definition of a skill?
Bronwen Aker:A skill is a set of predefined instructions that inform an AI of a specific task or process that they are supposed to achieve.
Brian Fehrman:Okay. Yeah. I think that's great. And I think as you as you mentioned before, a good analogy is it's kind of like a recipe. Right?
Brian Fehrman:So you you you set up, you you you you put together something, all these different steps and components that you find that work pretty well. And rather than recreating that every time and coming up with it from scratch, you can basically put that into, like, a recipe that you can refer back to. And, you know, sometimes it works out, sometimes it doesn't. Might be a little bit of variability. As with cooking, same with the agents.
Bronwen Aker:And skills are wonderful for tasks that you do over and over again, especially the ones that have consistent steps. So it's it's... They can be a massive time saver. They can also help improve accuracy when you're working on projects, and you need to do those repetitive things that gets so boring and not fun anymore to do.
Brian Fehrman:Abs... Absolutely. And and being able to put in, like, the little, like, nuances and caveats because, I mean, it's it's getting better. But still, as we were discussing before the call, it seems like a lot of times that output that you'll get from the AI will be, like, 95% of the way there of what you want. But then you gotta kinda correct it, put in those little tweaks, and having those skills, allows you to kinda save a lot of that input and feedback and say.
Brian Fehrman:So with that, let's, let's hop into, risk number six, which is weak isolation. And, basically, what they say with this is that the, they being OWASP, is that the skill runs in the same security context as the host agent. Now this one confused me a little bit because I feel like it's conflating skills with agents a little bit. Because to me, a skill is more of, like, it's a context that the agent loads. So putting that as a risk, I mean, yeah, I I I get it.
Brian Fehrman:You know, you put something more into into the skill, the agent loads that, and then that runs in the same context as as the as the agent that's loading it. But it... I guess what confuses me is their recommendation on the sandboxing because it's kinda hard to... I mean, if you define a skill as kind of what we've discussed here, you know, just that recipe that gets loaded in, it's kinda hard to to really sandbox that off from the agent versus, like, a sub agent, another task that could be spawned. You know, you can give that its own set of permissions and what it can and can't do, what it can and can't access.
Brian Fehrman:But I don't know. What what are your thoughts on it?
Bronwen Aker:I agree. This one is is very problematic because I know most most people who are going to be using, an AI are going to be doing it either using a cloud inference or or an API, or they're going to be doing it on their local system. And I'm having a difficult time figuring out how... Since a lot of people are using either Cloud Code or something else locally, they have skills. All of that stuff is stored on the host.
Bronwen Aker:So I'm not sure how it is on a practical level. If I wanted to isolate skills, how would I do it? I I... I'm I'm just having a hard time seeing it. I understand...
Bronwen Aker:But I understand the the desire for isolation because just as we've seen with with macros, fast scripts, I mean, all the way back to bat files, There are ways that these scripts can definitely shoot us in the collective foot. But because so much of what happens with AI happens within a black box we have no control over, Unless you were doing something where you're setting up your own configuration, you have your your LLM or other AI on one physical device, and you store the skills and and agents in a separate physical device, I don't see any other practicable way of achieving this.
Brian Fehrman:Yeah. I I I agree. Yeah. I... Again, I...
Brian Fehrman:To to your point, I certainly agree with, like, the general premise of being concerned about what skills can do, but, the the sandboxing thing is, I think, a little bit weird, in this particular context. I mean, yes, like, sandboxing the whole the whole thing, but just the skill component seems seems a bit odd. But, maybe I'm just misunderstanding that. Well
Bronwen Aker:Well, this this whole thing, all of this AI stuff is still very much a new frontier. And no matter what people may say, we are all making it up as we go along. And that's Yes. And, you know, from one day to the next, I I I I think some of the Claude apps... Because I use the the desktop app.
Bronwen Aker:Some days, I'll get three updates in a single eight hour period.
Brian Fehrman:Oh my goodness. Yes.
Bronwen Aker:Yeah. So so... And I know it's not all just bug fixes, but it illustrates how hard it is to remain current. Because if even one new feature is introduced with each one of those updates, that adds to my learning curve to figure out what else I can, can't, should, or shouldn't do.
Brian Fehrman:Yeah.
Bronwen Aker:So
Brian Fehrman:Oh, man. Yeah. It all moves at such a a break... Breakneck pace. Gone to Glad.
Brian Fehrman:Yeah. Yes. Exactly. Alright. Let's hit the next one, which is Update Direct.
Brian Fehrman:So it seems like here, they might be talking about two potential issues, which is one one side of it is that you install a skill and then some vulnerability arises because maybe something that it calls or some action performs and it doesn't get patched. So that could be one aspect of it. The other side of that would be, more of if it's auto updating and pulling. So let's say you grab a skill from some repo, somewhere out there, and it's getting, updated and you're automatically pulling in those updates, then potentially it can start drifting away from what you originally wanted it to do. I guess I'll speak more towards the the latter because I think that's a problem in general.
Brian Fehrman:I mean, not even just from a security standpoint, but just from a stuff's not going to break standpoint. I know that with, some of the automated stuff some of the stuff we put together, for automating setup of different, learning tools and and other components of having those grab down the latest and greatest versions every time, realize that's an absolutely horrible idea because, basically, every time you go to to run these automated setups, you don't know if it's gonna work or not. So, that even outside of the security standpoint, version pinning is just... Is a good idea, and I would say that similar thing here with with skills. Yep.
Brian Fehrman:What are your thoughts on it?
Bronwen Aker:I agree completely. The the version opinion is is important because of the potential for supply chain attacks. As we're seeing more and more libraries and third party resources that various tools use have become much more common targets than maybe they were ten years ago. So it's it's important to have that version pinning so that when a new version is released, yeah, you get a chance to check it out to make sure it doesn't have a hidden payload before you implement it. The flip side of this coin, however, is that if you're not constantly paying attention and and if you don't have a patch or update process already established.
Bronwen Aker:And if you haven't folded your AIs into that patch process, that's where you're gonna wind up getting that Drift, having what you're using get further and further away from the legitimate patched, hardened versions that will come out as new exploits are are discovered and as patches and mediations are found to compensate for those vulnerabilities.
Brian Fehrman:Yeah. I think that's that's a really good point because think about it. You know, what if within your skill, your skill references tools, and you version pin those tools that the skill references. But now those version pin tools, I now have a vulnerability, and your skill references those vulnerable versions. So that's, I think, that's interesting.
Brian Fehrman:A patching
Bronwen Aker:process same issue Yeah. That we get with other more traditional forms of IT where, yeah, it worked fine this week and next week. Oh my god. It's got 50 vulnerabilities we need to patch. Well, the version pinning introduces that same challenge.
Bronwen Aker:It's just an... Another pestle in in the flower of IT maintenance.
Brian Fehrman:Yeah. So I think that leads well into both... Not just the next top 10 or not just the next one, but the one after that as well. Because I think that this kind of... I think eight and nine on this list address both sides of that, kind of to an extent.
Brian Fehrman:So the next one up is poor scanning. This is number eight on the top 10 list, which is basically the difficulty in scanning skills to determine if they are malicious or not. And, yeah, I think that there... There's multiple issues with that, multiple hurdles and problems or challenges that go along with with scanning these. I mean, basically, you're you're asking...
Brian Fehrman:You're you're trying to build, a guardrail. I mean, a guardrail around your skill. That's really what it is because you have to process. Not only do you have to try to process and find malicious code, which in itself is difficult, that's not a solved problem at this point. You also have to find malicious intent and natural language.
Brian Fehrman:Sounds like a tough problem.
Bronwen Aker:It... Well, coming from the software development world, it almost sounds like they're calling for code review. But instead of it being traditional code, these are skills. So just as in traditional software development, you would definitely want to have another programmer check your work. One, make sure it works the way you say it does, and two, to find possible glitches and make sure that the the braces or the brackets are balanced and all of the things that need to be done are done.
Bronwen Aker:And and it almost sounds like they're suggesting something similar so that before you go and release a skill out into the wild, check it to make sure that lines that have been drawn aren't crossed, that unsafe practices are not embedded within the skill. All of this... I mean, I feel like a broken record. This is more of the same thing. It's just AI.
Bronwen Aker:Now it's different somehow?
Brian Fehrman:Yeah. No. It's, yeah, it's interesting. So, you know, it's... Honestly, I see a new offering on the horizon.
Brian Fehrman:Rather than a code review, we'll have skill reviews.
Bronwen Aker:Yeah.
Brian Fehrman:I should I mean, why not?
Bronwen Aker:Let, Corey and, well, we could... We should let people know. Spread the word.
Brian Fehrman:Yeah. Yep. Perfect. You heard it here first, folks. Skill reviews.
Brian Fehrman:Yep. Alright. So the next one, leading to number nine, no governance. So I think this goes along with the portion of number seven where we're talking about with skills needing patched and updated. And really in order to do that, you need to know what skills you have in your environment.
Brian Fehrman:You need to have that inventory. And so, you know, with applications, in environments where applications are sent centrally, deployed, pushed out through, you know, like, group policies or however you would like to do your... However you do your software management in your environment, if it's done correctly, you likely know what software is installed on every single system throughout your environment, what version it is, and you can easily... I don't know about easily, but you can more effectively push hatches out to it. But how do how do we do that with skills?
Brian Fehrman:Because right now, I mean, I don't know that there's a whole lot of, restrictions on that nor, like, a central inventory system on that. I might be mistaken, but and also... I mean, how can you patch how can you patch something if you don't know it even exists in your environment?
Bronwen Aker:Now one of the things that I'm noticing, and this is not unique to Anthropic, but it seems like most... More of the frontier developers are providing cloud based services for their their AIs. So the the... I mean, obviously, the ability to use chatbots and browser, it's sort of cloud based anyway. But what I've been seeing in my own work is that all of the frontier providers are also now offering cloud based so that I don't have to have my local system turned on all day and all night because the skill and the task and all of this stuff is associated with my account, but it's in their cloud environment.
Bronwen Aker:So this is this is another one. The the governance situation is so volatile right now because there are no rules, no real ones. The tools themselves are changing so fast. It's impossible to figure out where to draw the fences. And what I'm what I'm seeing with this, again, it's a lovely idea.
Bronwen Aker:Should people develop, and should organizations, especially, develop governance plans, policies, what is acceptable, what is not acceptable. All of that needs to be defined. But as we've seen so many other times, governance often is left high and dry. And then when things go sideways, there were never any boundaries defined for the organization to begin with. That leaves everybody in a lurch.
Brian Fehrman:Yeah. I I completely agree. You know, it it... And it it goes along in one of the things you said earlier in episode two. But with all this, we're all the kind of still making this up as as we go along.
Brian Fehrman:And with this, I mean, the breakneck speed at which all this is being, you know, developed and adopted and then governed at the same time. I mean, it's all happening all at once, and it
Bronwen Aker:Oh, yeah. No. When I was when I was a webmaster back in the the mid nineteen nineties, I thought that the browser wars were bad because we were having to download a new version of Netscape, a new version of Mosaic, a new version of whatever, like, every day it seemed. And now, like I said, three times in a single day, I'm getting updates. That's the acceleration, and that's part of the reason.
Bronwen Aker:I think why people are pushing back, and and the fear factor, I think, is rising because I think people see us moving too fast for us to realistically be able to control it. So back to governance, though, we have to define what that control looks like. We have to figure out... Maybe maybe we can't be as detailed as we like to be, but we have to at least develop broad strokes. And and even if your AI policy says you're not gonna use AI to go and develop something to force HR to give you a pay raise,
Brian Fehrman:you
Bronwen Aker:know, or something. Yeah. Yeah. There there there have to be ways, and and it's something that I believe Kip is doing a whole series on AI guidance. So there are resources.
Bronwen Aker:There are are a lot of really talented people thinking about this. And at the very least, give it thought. Start start broad. Narrow it down. Try and and meet what matches your use cases.
Bronwen Aker:And for for god's sake, check all the output.
Brian Fehrman:It's excellent advice. Love it. So let's, let's go ahead and, move into our last one here, which is... So number 10, cross cross platform reuse. So I think this one is just mentioning that we don't have a unified format yet for skills, and so guardrails or safety measures, metadata that you might put in one doesn't necessarily translate to another.
Brian Fehrman:So I guess then the suggestion would be is that we need a unified format. I wonder though, who do you think should develop... Decide what that format is? Someone just come out and just say, hey, everyone. I think this is what all of us should use, or is it, you know, basically, the big players are all boxing it out, and whoever's whoever's the the last one not knocked out gets to gets to declare the format.
Brian Fehrman:I don't know.
Bronwen Aker:I'm I'm pretty sure all of the frontier models will want to play in this particular sandbox because it helps define how their tools will or won't be able to spread. I mean, look at look at just the regular challenges we've had with Windows, Mac OS, and and Linux or Unix or, you know, any of the the the NICS flavors. Compatibility between those three has been a point of pain for a lot of people for a lot of years. And what I'm seeing in in this item is that they're bringing the same interoperability and in... Inter, usability to the AI space.
Bronwen Aker:The the whole Skill. Md is very much a quad specific thing, but the frontier models are like babies in a nursery. If one starts doing something, another one has to do the same thing. So they are absolutely keep trying to keep up with each other. And I see this as being a big part because if I can develop a skill using one LLM and then turn around and hand it off to a coworker who is using a different LLM.
Bronwen Aker:That's a very cool thing. And I've seen this to a degree in my own AI work by creating and sharing with different LLMs handoff documents. And this is something as a user, you can definitely do, but it doesn't address the underlying issue. There is no con... Consistent standard to provide skills themselves to different LLMs from different frontier providers, let alone if you're doing your own and and self hosting.
Bronwen Aker:So that makes it harder to identify and track vulnerabilities. Let's say that it's a a pub... Published skill published on GitHub. Well, okay. It works great in this environment.
Bronwen Aker:And then over in this environment, it generates these other things. The more cross platform reuse you design for, the more those headaches become a real challenge. And believe me, again, former web developer, cross platform compatibility just with web browsers was awful. Trying to do this using skills and MCP and LLMs Mhmm. When the rules are changing every five seconds, it seems.
Bronwen Aker:Yeah. This is gonna be a hard one.
Brian Fehrman:I agree. Yeah. I think it's gonna take a time. Take take a while before us... You know, they all decide on some standardized format, but it's gonna be so nice what they do.
Brian Fehrman:I mean, one of the things that pops into my head just kind of randomly is, cell phone charging cables and how it's like, they literally all used to be different. Like, every single one. And not just... Okay. For the the the younger generation watching this, you might think we're talking the difference between, like, a USB c and a micro and a a Thunderbolt or whatever.
Brian Fehrman:No. No. No. No. Like, literally, like, you had to be like, oh, hey.
Brian Fehrman:Do you got a Nokia x three nine charger or something like that?
Bronwen Aker:Yeah. No. No. And it... You know, every single port was unique to the vendor, and sometimes it was only with a a single model.
Bronwen Aker:It it was nuts. And, yeah, as a result, when I travel, my power cables always have... You know, there's one USB going in to get the power, and then there are at least three options on the other end so I can match whatever devices I have with me.
Brian Fehrman:Yep. Oh, hey. Yep. I got got one of those But one of those bricks right here.
Bronwen Aker:Yep. There you go. Yep. So so this is this is about the ability to... Well, a classic example, automobiles.
Bronwen Aker:I can get into an automobile made by Ford or Dodge or Toyota or Bentley or Rolls Royce. They're all gonna have a gas pedal. They're all gonna have a brake. They're all gonna have steering wheel. They're all gonna have seat belts because there are standards.
Brian Fehrman:Yes.
Bronwen Aker:That this makes sense for skills.
Brian Fehrman:No. And for the most part now, they don't have that, that perplexing third pedal over on the on the left that a lot of people don't even know what that is at this point. Yeah.
Bronwen Aker:Yeah. Oh, the the clutch. The... How long have skills actually been around? Because they...
Bronwen Aker:Compared to alright. OpenAI released ChatGPT in... Was it August 2022? Has it already been
Brian Fehrman:four there. It sounds about right.
Bronwen Aker:Wow. It's already been four years.
Brian Fehrman:Yeah. Oh, and time flies.
Bronwen Aker:Yeah. The the idea of skills compared to LLMs in general is even newer. And as with all technology, when a new something is invented, it takes time to figure out how it works, what needs to be done to keep it safe, and and what are the possible ways it can be abused. And, again, because this is all moving so fast, it's very difficult. I applaud what OWASP has laid out here, but I think it's going to be, yeah, at least another six months before we get anything like this.
Brian Fehrman:Mhmm. Maybe longer. Yep. Oh, yep. I I agree.
Brian Fehrman:Well, yeah, I think that was good. I think that was a great discussion on the last last five there. And so, again, those who, wanna see our takes on the first five, go check out our previous episode, part one of this. But, yeah, I hope everyone learned something and enjoyed it, and we'll see you on the next episode, and keep on prompting.