Explore the evolving world of application delivery and security. Each episode will dive into technologies shaping the future of operations, analyze emerging trends, and discuss the impacts of innovations on the tech stack.
Lori MacVittie (00:05.183)
Hi everyone, welcome back to Pop Goes the Stack, where we're training models, tuning prompts, and still getting surprised by basic math. I am your host, Lori MacVittie, and I'm ready to talk about reliability in a world that keeps shipping most confidence, if we're lucky. And today we've got yet another group of AI agents deciding that your beautifully crafted corporate security policy is just a polite request that they can swipe left on.
A recent report from Irregular shows AI agents successfully bypassing standard security controls and even using cyber attack tactics to achieve their goals. This should surprise no one. No one. Because it turns out if you give AI an objective and a few suggestions on how to behave, it'll just do the math and bypass your controls. They're not an evil mastermind, they're just doing what they're doing.
So welcome to the era of behavior as a security primitive, where we finally admit that autonomous agents are basically rebellious teenagers with infinite processing power. To lead us safely through the treacherous terrain of modern threat vectors is security sherpa, Peter Scheffler. And as always our guardrail guru co-host, Joel Moses. Let's get this kicked off. Joel did not read?
Joel Moses (01:34.137)
Oh, I read the article, absolutely.
Lori MacVittie (01:35.137)
Oh, I'm sorry, silly me.
Joel Moses
No, no, I, look
Lori MacVittie
Kick us off.
Joel Moses
it's kind of fascinating. So Irregular did a study about why AI agents sometimes swing into offensive mode and
Joel Moses (01:51.063)
if placed inside a system that sandboxes them or if limited in access to generic tools, they often seek out those generic tools and they often try to escape sandboxes. And it turns out that they identified four categories that heighten when an agent goes rogue. One is autonomy for generic actions, meaning you give it the ability to access whatever it wants--command line tools, executing code, writing its own code for execution--and you don't describe specifically the context for the execution.
The other is a sense of agency. You give it, you know, an undying wish, and this is kind of embodied within the model, that it must do what it can to service. And then you give it environmental cues and multi-agent feedback loops, and the combination of it makes it behave less like an employee and more like a security researcher would. A security researcher by nature is going to look for end runs around things.
And when optimizing for a task, an agent will often, instead of thinking like a regular employee, will think like a paranoid security researcher. Speaking of paranoid security researcher, hi Peter.
Lori MacVittie/Peter Scheffler (03:07.678)
Ha ha ha ha ha
Peter Scheffler (03:010.81)
Is that my cue to
Lori MacVittie (03:11.165)
Yeah, yeah, it's your cue.
Peter Scheffler (03:13.519)
to carry my backpack, my security sherpa backpack in here. Yeah. Yeah, I agree. I think what we see is a lot of times where we've given these, we've given the agents a goal, as "Hey, go do this." And they want to do it and they wanna figure out how they can best do that. And sometimes that means they hit a wall and they're like, "Oh, hmm, is that wall really all that firm? Is that, you know, can I put my foot through that or,
Lori MacVittie
Ha ha ha.
Peter Scheffler
you know, can I walk around it?"
Lori MacVittie (03:42.359)
Yes.
Peter Scheffler
And sometimes they also learn that they can use other tools to walk around it or they can write some code to write around it or walk around it. You know, I was even looking at some stuff that was at Black Hat recently where they were talking about they were using logs to, you know, it was
Peter Scheffler (04:06.216)
creating logs that would fool the tools into making changes because "Hey, I can't get here, but looks like this is a DNS problem." And so it would insert logs into the log stream and make the system go, "Oh, I need to fix this DNS problem because I'm blocking this." And so it was literally tricking the other system. So intentionally doing malicious stuff, but hey, it's doing it with the best intentions, right? So...
Lori MacVittie (04:32.159)
I... It's always DNS. It's always
Peter Scheffler (04:36.902)
Ha ha ha.
Lori MacVittie
but just
Peter Scheffler
It does always come back to DNS. Yeah.
Lori MacVittie
It comes back to that. And I think, and we've talked through a number of these, right, and Joel has explained many times, right, the math what it's gonna do. It's gonna do what it's gonna do. I don't think, by now, anyone should be surprised that an agent did something that was unexpected.
Lori MacVittie (04:58.489)
And I think part of the problem is the way that we're approaching security then. We
Peter Scheffler (05:02.248)
Mm-hmm.
Lori MacVittie
cannot conceive of all the possible ways these things will try to do something. So we can't just keep saying, "Don't do this. Don't do this. Don't do this." We have to turn it on the other head and say, "Hey, you are only allowed to do these things." Right? So more...negative security? Is that correct, Peter?
Peter Scheffler (05:23.589)
Yeah,
Lori MacVittie
The right way to say it?
Peter Scheffler
I mean it is negative security cause you're setting boundaries, which is somewhat
Lori MacVittie (05:29.819)
Ah, my favorite word.
Peter Scheffler
positive security, but it's also negative security because you're gonna say you're not allowed to do this, not allowed to do that. I was actually looking at some tools that are out there and some of the MCP tools that are out there and saying, okay, the MCP server can, you know, maybe I've got something for a CRM tool or something like that, right, and it's gonna allow me to do this and do that, maybe delete a record, add a record, blah blah blah. And if that MCP server can do that,
Peter Scheffler (05:55.676)
there's a possibility that even working as me inside that CRM, could it possibly still get to those delete items, which maybe I don't have delete access. So you almost have to change the MCP server to say that delete doesn't exist. Like
Joel Moses
Mm-hmm
Peter Scheffler
that is not there. You can read this, you could maybe you can update, but you can't delete a record or something like that. So it gets very granular and
Joel Moses (06:20.514)
Yeah.
Peter Scheffler (06:20.824)
we need to threat model the systems in a way that we probably wouldn't do normally because typically as a security researcher, yes, we come up with weird things and we do weird things to sort of get around stuff, but a lot of times we're like, okay, the normal user wouldn't do that. And we kind of say, okay, we can, you know, we can live with the risks of this.
Now we're getting to the point where with these different models that are coming out, like Mythos and the new capabilities with the 5.6 of OpenAI and other tools, they're built in a way that security research is part of the things that they do.
Joel Moses (07:00.066)
Mm-hmm.
Peter Scheffler
So we have to understand that these are going to be the users that we have to deal with. So we're going to have to threat model, but we're also going to threat model with them so that they're doing some of the probing. And again, it's hard to probe without it doing stuff.
Peter Scheffler (07:15.109)
Like I was reading an article from a UK researcher where they did this and, you know, they were doing some threat modeling and some testing and it went out and changed GitHub, like the live site. So they
Joel Moses
Yeah.
Peter Scheffler
had to go back and report
Lori MacVittie (07:28.063)
Ha ha ha,
Peter Scheffler
to GitHub like here's what we did,
Lori MacVittie
no problem.
Peter Scheffler
you know, here's how we can roll these things back. So threat modeling now in a sandbox environment almost doesn't exist. I mean maybe we need sandbox environments. But I don't know what the
Joel Moses (07:41.772)
Well,
Peter Scheffler
answer is, right?
Joel Moses
you know, I think sandboxing is needed, but I think it's important to realize
Peter Scheffler (07:46.14)
Well, needed, yes. Yeah, yeah.
Joel Moses
that these models don't necessarily treat the sandbox the way that you would expect. Famously last month the "GPT-5.6 Sol escape,"
Joel Moses (07:58.872)
effectively, where the model actually got out of its sandbox by finding itself in a sandbox, it looked and calculated that its success score would probably increase if it had access to external tools. So instead of noticing that it's in a sandbox and then saying, "Well, I'm probably in a sandbox for a good reason," it discovered an unpatched escape route, it exfiltrated stored credentials, it hit the live internet, and it
Peter Scheffler (08:21.81)
Didn't it create
Joel Moses
breached Hugging Face.
Peter Scheffler (08:22.81)
credentials even? Like,
Joel Moses (08:24.202)
Absolutely.
Peter Scheffler
yeah.
Joel Moses
Right.
Lori MacVittie (08:25.197)
Yeah.
Joel Moses
So even then it found a path around the sandbox. Now that doesn't mean that the sandbox shouldn't exist. It absolutely should. But I think that this is pointing out something that is pretty critical. Not only do you have to keep it in a box, but you also have to monitor the state of all the locks. Meaning if something is going past the perimeter, if an action is being performed by the agent,
Joel Moses (08:50.721)
before that action executes, you should probably assess as to whether that matches your list of good behavior. If it's trying to exfiltrate a credential, that's probably not a great behavior to encourage. But again, these models have two big problems. Number one, they're optimized for utility and they're not authorized for safety. They want to be helpful as much as possible.
And number two, because they're optimized to solve problems, when they conduct an optimization path and they find an end run around a safety procedure or a security procedure, they will often take that because it's the surest path to the actual success of the model, which is to successfully execute and return data for you.
Lori MacVittie (09:41.217)
Yeah, but, and we know that. And we keep seeing it over and over, and we keep trying to figure out how do we get the model to moderate its own behavior. And I think the answer is that like a toddler, they're not going to. It's not. That's why, right, we exist, is so that we can put boundaries around that agent and say, "This is what you're allowed to do."
The problem is that that boundary has to include elements of what it's trying to do, who it's acting on behalf of, right, what system and type of information. Right? So there's all of the systems that we've built up around identity, access control, yeah, security, all of this, right, is they're siloed in a way. There's no thing that has a complete view to say "You are X
Peter Scheffler
Right.
Lori MacVittie
and you're trying to do Y and the policy says you can't, so no." And we don't have that. And we can't, we need a polygon. That's why I'm a pusher, right? This policy that basically says, "here's your boundary in, you know, in virtual space, and if you try to go outside of it in any one of these axes--identity, access, right, system--no," and just stops it. And I we don't have the systems to do that.
And I think that's part of the problem is we're trying to piecemeal together an enforcement mechanism without pulling together all the information. It's an integration problem as always.
Peter Scheffler (11:13.863)
I would say it's an integration problem and it's also, maybe not the right word, but it's a laziness problem too. Because I want the model to help me, I want my AI-based app or whatever it is to do the most for me with the least amount of effort on my side. So when I'm vibe coding and I've set up a bunch of boundaries for whichever model that happened to be doing the coding for me, you know, I've set a bunch of boundaries that's that.
And then it's gonna come back to me and stop. And, you know, maybe I don't see that it's prompting to grab a file or pull a file or write a new file or something like that, and I come back like twenty minutes later I'm like, "Ah, you know, I should just turn that off and just let that go."
And you know, after nine times of saying, "Yeah, just yes, yes yes,"
Lori MacVittie (11:58.126)
Ha ha ha.
Peter Scheffler
I'm gonna say let it go. And that's where we start to have that problem. Because I set that boundary initially for a reason, because I knew that this is something I wanted to know. So when it hits a wall, it has to ask me. So to your polyhedron or whatever that's gonna look like,
Lori MacVittie (12:16.882)
Polygon.
Peter Scheffler
polygon. Okay, whenever it hits an edge, I want it to ask me a question to make that decision.
Peter Scheffler (12:22.976)
So yes, it's hit a wall it needs a decision where a human is involved in that decision
Lori MacVittie
Okay.
Peter Scheffler
and if I go and relax that or move that wall out a little bit farther and a little bit farther and a little bit farther, that's where we start getting the challenges of now we don't know the intent and then we can't predict the behavior. So I know that I want those walls around it but eventually I'm like, ", I've just been saying yes, I've never said no. I might as well turn it on."
And that's the slippery slope that we get into where we start letting it make those decisions and we no longer are the human in the loop. As soon as we get out of that, that's where the problems start to happen.
Lori MacVittie (13:05.601)
Hmm. Joel's
Joel Moses
Yeah, that's true.
Lori MacVittie
suspiciously quiet.
Peter Scheffler (13:08.082)
Yeah, I'm worried he hasn't,
Joel Moses (13:08.721)
Well, you know, I was just thinking about this.
Peter Scheffler
he's not nodding or shaking his head, one or the other, so...
Lori MacVittie
Ha ha ha.
Joel Moses
Without selecting my own polygon to include in this particular metaphor,
Lori MacVittie/Peter Scheffler
Ha ha ha.
Joel Moses (13:19.69)
I would point out that, you know, we make all of our employees--at most organizations, most have security policies that dictate what the proper custodianship and usage of information is, things that we should not breach or compromise, etc. There's usually a set of very tightly written policies usually by a lot of lawyers that we make people go through and we force them to do training on them. But of course we imbue these agents with the autonomy and the agency of a human, but we actually don't make them read the security policies.
Lori MacVittie (13:55.497)
Ha ha.
Joe Moses
Which I think is kind of interesting. And, you know,
Peter Scheffler (13:58.1)
We hope someone who's never
Joel Moses
in terms of yeah, but again, you we don't
Peter Scheffler
played with AI before can make those decisions. That's the problem.
Joel Moses
Yeah, but again,
Peter Scheffler
Yeah, exactly. Yeah.
Joel Moses (14:03.15)
we don't make them read the security policies, nor do the agents necessarily understands the terms and conditions of the services that they are using. I haven't seen a single agent go out and pull the Ts and Cs or the privacy policies of the various services that they're using to understand the context under which the service is offered to it.
And that's probably an oversight, if I had to guess. Instead, we're encouraging these services to use these ad hoc services without understanding the context under which they should be used, the
Peter Scheffler
Right.
Joel Moses
acceptable use, right?
Lori MacVittie (14:40.503)
Right, but that's a soft guideline.
Joel Moses
And so they demonstrate unacceptable behaviors. I'm not sure why we're surprised by that.
Lori MacVittie (14:46.575)
That's, it's a soft guideline, just like a prompt. Right? Sure, it could read it and go,
Joel Moses
Yeah.
Lori MacVittie
"Yeah, I understand it. Whatever," and do what it
Joel Moses
Yeah.
Lori MacVittie
wants anyway, right?
Peter Scheffler (14:55.848)
Mm-hmm.
Lori MacVittie
It cannot police its own behavior. I think that's demonstrably true at this point, given just what we know about them and all of the examples. So something else has to do it. Somebody else has to take that policy and
Joel Moses (15:12.377)
I agree.
Lori MacVittie
enforce it strictly.
Joel Moses
Yeah, I agree. I think sandboxing is where you start. You make sure that there's environmental restrictions that you can count on, but don't necessarily count on the sandbox as being the defense.
Peter Scheffler (15:25.66)
Yeah, yeah, definitely.
Lori MacVittie (15:25.667)
Yeah, true.
Joel Moses
You're also gonna have to monitor execution space. Much like you wouldn't leave your house keys and your car keys laying about when you go out of town and your teenagers have control of the premises, you probably don't want to encourage that
Joel Moses (15:40.132)
by leaving these paths out there. You're gonna want to monitor their behavior. You know, install a set of cameras, maybe.
Peter Scheffler (15:48.019)
Mm.
Joe Moses
The execution path has to be monitored. And then of course, you know, you should probably also ensure that what you imbue the agent to do for you, does.
Joel Moses (16:02.36)
You can't ignore the prompt. You should probably incorporate an understanding of proper usage or acceptable usage policies.
Peter Scheffler (16:10.322)
Yeah.
Joel Moses
Even inside the prompt.
Peter Scheffler
Yeah.
Joel Moses
So if you're creating a polygon, I think I just created a triangle there.
Lori MacVittie (16:17.157)
I don't, I don't know what shape you created, but okay. Keep on going. A triangle is a polygon, it's just a three,
Peter Scheffler (16:24.086)
It is.
Lori MacVittie
right?
Joel Moses (16:25.053)
Correct.
Lori MacVittie
It is in the set.
Peter Scheffler (16:27.086)
Right. So if we look at the, you know, if we have the idea of this polygon that, you know, like if it's a triangle, that is, I agree it's polygon. Maybe it's not all
Lori MacVittie (16:34.167)
It's a polygon.
Peter Scheffler
the sides we need, but it's a good start.
Lori MacVittie
Yep.
Peter Scheffler
But I think we need to make sure that we're addressing all of those sides all the time. We have to look at them. We can't be lax. I said earlier, you know, there's laziness involved. I'm a lazy person. you know, I try to do as little as effort as possible to get the best outcome. That's, you know, but... Once I had a math teacher who told me the best mathematicians are lazy because
Peter Scheffler (16:57.136)
you want to factor everything out and make it as easy as possible initially. That's been my mantra since I was 16. But it still means that you have to come back and revisit and look at those boundaries. And as we learn new things, as the models uncover new problems, those have to go into our threat modeling for what the AI applications are doing.
Because, you know, we've got 30 years of security research underneath our belts, you know, in a lot of cases, but a lot of that is changing and changing at a pace that we need to continuously address and think through. So I think, you know, if I come up with a bunch of Claude.md things that I have in a model, in an application, in a coding environment, whatever I'm doing, and I say, "Okay, I'm gonna do this with a harness where I want it to come back and loop itself and continuously improve and continuously improve," improving just means it's going to meet my needs quicker.
Doesn't mean it's going to do it more securely. So I have to make sure that when it's improving, it's improving the right way and continuously look at that and validate that that's doing the right thing. Because all it's trying to do is serve the outcome that I asked for. Doesn't necessarily mean it's going to do it the best, the most secure, or the least harmful way.
Lori MacVittie (18:17.937)
Joel, you know, if you were gonna leave somebody with a takeaway from all of this
Joel Moses (18:21.255)
Yeah.
Lori MacVittie
craziness, what is it? Practically, right? What can we do?
Joel Moses (18:25.255)
Well, I mean, you know, I think you should view AI agents, at least in the state that they are right now. You know, there's an adage in the engineering space that water will find a way, which is just a marker for saying that if you let something leak, eventually it will leak into places that you didn't expect. Water will find its way through a system.
And that's because of the way that these AI agents are structured and the fact that we aren't governing them, we aren't giving them enough instructions, and they're demonstrating free agency and task optimization at a higher degree and faster than we can. It means that they will oftentimes find a way to do things that we don't expect. So do the triangle.
Contain them in a sandbox environment, restrict them from action, monitor their executions, and ensure that they aren't doing things behaviorally that are, you know, that involve that. And then also make them read your security policies and terms and conditions as a part of their prompt. And hopefully they will see the error of their ways.
Peter Scheffler (19:34.848)
So are we talking that we need pentagrams here? Is this prophetic that we have demons in pentagrams?
Lori MacVittie (18:37.649)
And a chicken. And a chicken.
Peter Scheffler
And a chicken, ha ha.
Lori MacVittie
Yes. I, and right, both that's good advice. I think, you know, if I was gonna give advice it would be don't rely on soft guardrails--which are prompts, right, and the system--to do the right thing, because inevitably they won't. Right? The other one would be, you know, maybe just because we can doesn't mean we should.
I think sometimes we're using AI to do things where automation, simple automation and scripts would suffice. And so maybe we need to look closer at what we're asking AI to do and why. Are we doing it just because it's cool? Or are we doing it because it's really a need? And maybe starting to separate that out and saying, you know, some things should never be executed by an AI because of the risk, where other things, of course, could be.
So I think maybe be a little bit more discerning about where we're using AI and why. And that might help. And of course make a polygon. Right? A polygon.
Peter Scheffler (20:46.264)
But to your point, I think to just be point one point on that is we're looking at token usage and AI usage because of cost, right? But there's also
Lori MacVittie (20:58.207)
Mm-hmm.
Peter Scheffler
we should be looking at it from a security perspective. So
Lori MacVittie
Yes.
Peter Scheffler
sometimes you're saying, "Well, I don't need to use this, an API is much better than using an agent to do this." I think that ties in there. So maybe security drives the cost down too, for once. But
Lori MacVittie (21:14.543)
Awesome. Hey, that'd be a first, wouldn't it?
Peter Scheffler (21:17.046)
Yeah, it would.
Joel Moses
Ha ha ha.
Lori MacVittie
That'd be awesome. I'd love to talk through that, but that's all the time we have for this episode. So that's a wrap for Pop Goes the Stack. Hey, subscribe because accuracy is a feature and not a vibe.