Pop Goes the Stack

“Securing the AI model” sounds like a clean, fundable story. In practice, it’s often the wrong security target. In this episode of Pop Goes the Stack, F5's Lori MacVittie and Joel Moses are joined by Mark Menger to cut through the myth and focus on where the real risk lives: the inferencing server, the data sources it can reach, and the runtime environment that’s actually exposed to traffic.
 
They make the point plainly: a model file is usually just static weights, a heavy spreadsheet sitting at rest. If you trained a proprietary model, protecting that artifact matters. But most enterprises aren’t training producers; they’re training consumers using open-weight models, and obsessing over encrypting and isolating a freely downloadable file won’t stop the failures showing up in headlines.
 
The real battleground is everything around the model: what data the inference system can access, how RAG sources are protected, what APIs agents can invoke, what credentials get embedded in “skills” files, and how the system behaves under production-scale load. Mark frames it as an iceberg problem: the shiny GPU layer is above the waterline, but reliability, security, performance, and resilience are won or lost in the unglamorous infrastructure underneath.
 
A key architectural theme is loose coupling. Adding control points between clients, RAG, object stores, and inference services limits blast radius and prevents “pilot success” from turning into production Thanksgiving. The practical advice is to stop treating the model file as the center of gravity, build strong boundaries around the runtime, and stress test for real scale and real failure modes before rollout.

Creators and Guests

Host
Joel Moses
Distinguished Engineer and VP, Strategic Engineer at F5, Joel has over 30 years of industry experience in cybersecurity and networking fields. He holds several US patents related to encryption technique.
Host
Lori MacVittie
Distinguished Engineer and Chief Evangelist at F5, Lori has more than 25 years of industry experience spanning application development, IT architecture, and network and systems' operation. She co-authored the CADD profile for ANSI NCITS 320-1998 and is a prolific author with books spanning security, cloud, and enterprise architecture.
Guest
Mark Menger
F5 Solutions Architect
Producer
Tabitha R.R. Powell
Technical Thought Leadership Evangelist producing content that makes complex ideas clear and engaging.

What is Pop Goes the Stack?

Explore the evolving world of application delivery and security. Each episode will dive into technologies shaping the future of operations, analyze emerging trends, and discuss the impacts of innovations on the tech stack.

Lori MacVittie (00:06.382)
Welcome back to Pop Goes the Stack, the podcast where we gently peer under the hood of today's loudest tech trends and politely point out, yeah, that the engine is missing. I am Lori MacVittie here with co-host Joel Moses.

Joel Moses (00:21.93)
Hello, Lori.

Lori MacVittie
Hi. You know what we're gonna do today?

Joel Moses (00:26.396)
Oh, I hope we talk about something really entertaining.

Lori MacVittie (00:29.024)
Okay, alright. Well, just for you then, we're gonna tackle a marketing myth so pervasive, so aggressively funded that it has its own zip code in the valley

Joel Moses (00:38.918)
Oh boy.

Lori MacVittie
called securing the AI model. Yeah, yeah. We're hoping someone writes another venture back check to lock down a static file of mathematical weights and parameters sitting quietly on an S3 bucket.

Joel Moses
Yeah.

Lori MacVittie
We'll wait. Yeah. Well, spoiler alert.

Lori MacVittie (00:56.14)
And you know this, right? The model's basically a really heavy glorified Excel spreadsheet, or as Joel would say, a CSV.

Joel Moses (01:04.662)
CSV with

Lori MacVittie
Same thing.

Joel Moses
an attitude, Lori,

Lori MacVittie
There

Joel Moses
that's what it is.

Lori MacVittie
Okay, alright. Yes, but it's a static list of numbers and it doesn't do anything on its own. It's not plotting a coup, it's not even running. So if you want to find the real security battleground, you have to look at the thing actually sweating under the workload, and that's the inferencing server.

Lori MacVittie (01:23.884)
So today we're gonna talk about that. Our guest today is Mark Menger. Hi, Mark.

Mark Menger (01:29.506)
Hello.

Lori MacVittie
All right,

Mark Menger
Thanks for the invite.

Lori MacVittie
are you ready to help us? You're gonna help us cut through?

Mark Menger
I am certainly gonna give it a shot.

Lori MacVittie
Excellent. All right, because nobody recognizes a marketing hallucination like Mark, right?

Mark Menger/Joel Moses/Lori MacVittie (01:41.492)
Ha ha ha.

Lori MacVittie (01:44.015)
He doesn't know how to answer that.

Mark Menger (01:46.846)
No,

Lori MacVittie
All right. Well let's

Mark Menger
I will quietly let that go.

Lori MacVittie
You'll let it go, thank you.

Joel Moses
Yeah.

Lori MacVittie
Thank you. I want Joel to give us the rundown

Joel Moses (01:54.07)
Well

Lori MacVittie
because the premise here is

Joel Moses
Sure.

Lori MacVittie
it's just a file. So,

Joel Moses (01:54.07)
Well sure. Let's

Lori MacVittie
come on.

Joel Moses
let me turn this around on Mark. Mark, what's the obsession with the model file? I mean there are a lot of solutions out there that surround protecting the model file, encrypting it inside of S3 buckets, ensuring it's secure transport, that sort of thing.

Mark Menger (02:17.986)
Mm-hmm.

Joel Moses
When is that a good thing to do and when is it kind of, you know, maybe not as important?

Mark Menger (02:25.56)
Well, so first off I would s-. Okay, broadly I'll throw out a couple of things in this regard. One, I think that generally people, in the conversations I've had with customers both on site or at events, there seems to be this overall, and this kind of goes to Lori's point a moment ago, that there's an overarching focus on what's going on with just the AI part of the piece. You know? Oh, I've got to look at my GPUs.

Oh, I've got to focus, you know, minutely on all of these facets of delivering my AI application. And so I think that's part of it, is that we're not really looking at the bigger picture consistently. And I think part of that might be driven by the marketing messaging that Lori was mentioning earlier. The other part of it has to do, and I think part of the say broader trend has to do with the expansion of the AI pie.

You know, a few years ago, the pie was mostly filled up with the people who were really exploring doing the frontier work. And as a consequence, the pie was predominantly training focused in terms of the assets that were being created. And given what we know about what it takes to train a model, right, the people who are doing this are spending lots and lots of money. And so

Joel Moses
Right.

Mark Menger
obviously they would be interested in protecting it. But for the rest of the organizations, the enterprises that we serve, the customers that we talk to, they're not training producers, they're training consumers. And so then the landscape of, you know, to what degree I pay, you know, I'm interested in that protecting that model changes.

Joel Moses (04:08.96)
So I guess early on in the growth of models, people spent a lot of time and attention on actually training models and creating their own parameters and weights and things like that. And if you've spent a lot of money doing that, obviously the resulting file that contains your proprietary weights is something you probably want to protect a little bit more than the average thing. But for the most part it seems like enterprises and most of the world is shifting towards open weight models.

And those are just freely downloadable pieces of sheet music, if you get my meaning on that. So you should still protect a model file, but maybe only if it has significant value. But you don't need to spend a lot of time on that. What should you spend your time doing instead?

Lori MacVittie (04:59.01)
Securing other things, like all the holes that AI exposes. Like that's what I would do. I'm just saying, you know. I don't know.

Mark Menger (05:08.844)
Yeah, I agree entirely. I mean, I would say, you know, what do you need to protect? Look at the news.

Joel Moses (05:14.431)
Yeah.

Lori MacVittie (05:14.702)
Ha ha ha. Crap, ha ha.

Mark Menger
Right? It's, you know, there's your guidance on what you should be protecting, where things are going in a less than desirable direction. And that has little to nothing to do with what model was chosen, you know, what spread, you know, what CSV was loaded. It has more to do with the engine that you built

Mark Menger (05:38.925)
and the permissions and environment that you created around it, allowing it to operate, either simply from an inferencing perspective, what data does it have access to and what is it saying about that to whoever the user of the system is, or in more sophisticated use cases, agentic. And now all of a sudden it's not only what data are you providing it access to, but what APIs are you allowing it to invoke and with what permissions, what capabilities?

Lori MacVittie (06:08.046)
I would go even down to the system level, like, right, especially with agents. Right? We're putting agent skills, which are basically directions for how to access other systems--sometimes, of course, including credentials, because we haven't learned that lesson either. So you've got this

Mark Menger
Ha ha.

Lori MacVittie
file sitting here on a system. And I would be more worried about somebody gaining access to the system or somehow injecting something into that or extracting that information because it contains valuable information about the rest of your systems. I would be more worried about that than I would be about somebody going in and changing a weight in my model. I'm just saying. Like these are like two different things and one is infinitely more likely as agentic keeps growing and we keep relying on, you know, static files to define integrations.

Joel Moses
Mm-hmm.

Mark Menger (07:01.962)
Mm-hmm. Yeah, agreed. And to

Lori MacVittie (07:03.906)
Okay, we're done. There you go. Ha ha.

Joel Moses
Ha ha ha.

Mark Menger
Well and to reinf-

Lori MacVittie
Next, that's it.

Joel Moses (07:06.038)
That was an easy one.

Lori MacVittie
We agreed. Nice.

Mark Menger
Yeah to reinforce, reinforce that. I mean, I'm gonna say, you know, look at the news again. It's like who's actually--so distilling, for those who don't know, right, distilling a model is interacting with it in such a way that I figure out what how that model was trained, what how it behaves--and look at the news about who's complaining about distilling. It's not Enterprise ABC or XYZ.

Lori MacVittie
No.

Mark Menger (07:31.785)
It's OpenAI complaining about DeepSeek distilling their model. Right, so this is basically a form of the model theft. Right? How do you protect the model? And but it's occurring in the frontier model training space. For most people, right, coming back to the point, there are other more important things that you need to be paying attention to that you are paying attention to, because model theft distillation is not something of concern to you. And so broadly stated

Mark Menger (08:01.596)
when I've spoken to customers, this metaphor kind of resonates is that they should be looking at their AI estate as kind of an iceberg. I mentioned earlier, above the waterline, below the waterline. They really need to take that to heart. If the traditional AI infrastructure that is below the waterline--right, it's not glitzy, it's not shiny, it's not GPUs--but if that is not properly addressed, everything's gonna flip, it's not gonna operate properly, it's not gonna be secure, it's not gonna be performant, it's not gonna be resilient.

And they need to address that while simultaneously looking at above-the-waterline concerns. Right? So that's where your AI guardrails solutions come into place. That's where automated red teaming comes into play. Right. And you need to have all of that in concert. You need to have all of that together, which I think then addresses to--goes to your point, Lori--right, which is you really need to look at the entire estate. You need to look at the whole landscape, you need to look at how you're protecting your data because that's the traditional stuff.

You need to make sure that that's performant and resilient. And the thing that kind of changes about this landscape that we've seen with our customers, and the fact that it's informed some of their purchases from us is scale. AI operations at production scale look--well, a big change there is what in traditional AIT looks like the rare event, in AI scale is the normal day-to-day.

You need to prepare for that. If you're going to be spending your time preparing for production rollout, you need to be investing time preparing for that from a security perspective, from a performance perspective, from a reliability and resilience perspective.

Lori MacVittie (09:42.776)
Yeah, I like that. That's basically the, you know, if you leave the windows open when you lock the door, well, I

Mark Menger
Ha ha.

Lori MacVittie
you know, I don't know how to help you there. They're gonna get in one way or another. Right, you can't, or

Joel Moses (09:54.39)
Mm, my house needs the ventilation Lori, so, you know.

Lori MacVittie (09:59.307)
Oh, you got, you secured your ventilation? Like get the duct?

Joel Moses (10:01.404)
Sure, why not?

Lori MacVittie
Yeah. Ha ha.

Joel Moses
So let me decompose this if I could, Mark. So if, I think the instruction here is not that the model file is utterly unimportant. If you spend a lot of money creating a closed weight model of your own design, you should probably worry about

Mark Menger (10:16.208)
Oh, yeah.

Joel Moses
who has access to that particular file. But if you're using an open weight model, wasting your time ensuring that that is encrypted to memory and the whole nine yards is probably above what it really needs to be treated. Instead, what are some of the other focus areas? I can think of

Joel Moses (10:32.804)
ensuring that your RAG data sources, if you're augmenting the foundational model, ensuring that they are up to date and that they are useful, they contain data that is relevant to the task of the augmentation. I would probably spend time making sure that my API access to my vLLM or Triton Inference Server instances is governed properly, that only the people who, the services that need to have access are authorized to access that model and send traffic through it.

I would also invest in guardrail systems so that people who have access to the system aren't trying to misuse the model in some way. What other things should they also pay attention to?

Mark Menger (11:20.29)
So architecturally, and this is actually kind of an old model, but we've seen it resonate and actually provide help with customers, is the practice of loose coupling. So loose coupling, simply stated, for those who are unaware of the term, is reducing direct dependencies between a client solution and the service solution.

We've been doing a lot of work in this regard with storage, but it applies to just about any service that you can imagine, including an inference service, including a RAG service. And the reason for that is one, it provides a really good control point for the kind of use cases that you were talking about, but it also, and I think this is important in the context of AI scale, is that it provides a really good point for blast radius control.

So as one system component has hiccups or misbehaves, that misbehavior does not propagate throughout the whole system. And the reason that that is critical is, you know to go back to my point a moment ago, is that at AI scale the rare event becomes the norm. The high stress, what we considered previously high stress traffic volume, is now the normal traffic volume.

And so when things go sideways, they go sideways really fast. We've had customers, for example, in front of their RAG systems and in front of their object storage systems place components and services that traditionally are placed external facing and now they're internal. I'm thinking of DDoS in particular,

Joel Moses
Mm-hmm.

Mark Menger
because the client behavior of AI systems actually can induce DDoS behaviors.

Joel Moses (13:07.606)
Toward object stores, for example.

Lori MacVittie (13:10.051)
Yeah.

Mark Menger (13:10.156)
Toward object stores, for example. And so customers are actually buying and installing DDoS in front of those services at that loosely coupled boundary. So that that creates a means by both I mean frankly it's a bi-directional benefit, right? Which is misbehaving clients don't take out the object store and misbehaviors within the storage service don't propagate into negative results for the the clients.

And that is one, it's non-trivial. And two, going back to the point of like where should you be spending your time, this is non-trivial work that you really should be doing because you're investing a substantial amount of time and energy to build these systems. You don't want them to go down. So you want to be making the proper investments to enforce resilience and reliability in addition to performance and security.

Joel Moses (14:04.213)
Well,

Lori MacVittie (14:04.568)
I like that. Yeah.

Joel Moses
I'm glad you explained loose coupling. I thought for a moment I was getting relationship advice. But

Lori MacVittie
Ha ha ha.

Mark Menger (14:10.444)
Well there's that too.

Joel Moses
Okay, good.

Mark Menger (14:10.444)
Ha ha, exactly.

Lori MacVittie (14:14.326)
Wait, wait. No, nope, no relationship advice for Joel today. Let's keep it

Mark Menger (14:18.741)
No swiping left or right.

Joel Moses (14:20.405)
Mm.

Lori MacVittie
Yeah, no swiping left or right. I did like the architectural focus because I think a lot of that is lost on people when we start talking about things like RAG or augmenting with even tools, as that the very composition of the data center, right, the idea of three tiers, if you will, right, is kind of gone.

Lori MacVittie (14:40.3)
Right, things can talk in and out and like you said, you have to watch internally as well as externally for things like attacks, for you know, systems gone wrong. You know, the best way to describe that is there is no perimeter.

Joel Moses
Mm-hmm.

Lori MacVittie
There's not. Everything has to be kind of like paying more attention architecturally to like what's going on because things are moving. RAG has the data that you should be protecting, generally speaking, if you're augmenting

Mark Menger
Yeah.

Lori MacVittie
it with something like that. You need to be worried about you just kind of moved up a data source in your architecture closer to the, you know, client basically. How are you protecting it? Right. This isn't a database

Mark Menger
Mm-hmm.

Joel Moses
Yeah.

Lori MacVittie
deep in the data center anymore. So what do you do?

Joel Moses (15:23.241)
Interesting.

Lori MacVittie
what do you do?

Joel Moses
We should have named this topic When did my stack diagram turn into a hypercube? Yeah.

Mark Menger (15:31.734)
Ha ha ha.

Lori MacVittie (15:32.116)
Ooo, nice.

Joel Moses
Yeah.

Lori MacVittie
Nice yeah, yeah.

Joel Moses (15:35.574)
Yeah, that's an interesting observation, Mark, that oftentimes, you know, because of the way that the technology works, where it's directly accessing and choosing its own path to your object stores to augment the output, that a denial of service can be achieved simply by the prompts that you send through the system that result in the object store accesses. And you have to think about that. You have to think about patterns of access that your object stores have probably never seen before.

Mark Menger (16:04.96)
Mm-hmm. Yeah. And so part of this then goes back to like what are you doing in the pilot? What does the pilot teach you? What does it not teach you?

Joel Moses
Mm-hmm.

Mark Menger
And so direct coupling, like tight coupling, is something that is very expedient for your pilot, right? As opposed to, you know, taking the time to introduce a loose coupling layer like placing an ADC in front of there and configuring it properly with proper policy from a security

Joel Moses
Yep.

Mark Menger
and reliability perspective. You know, when you just want to start, you know, running inference and if you just throw in and start connecting agents to things, you're like, Oh, I don't want to do that. And in the pilot, you're probably not going to see any negative consequences. Why? Because the number of clients that you have can be measured on a couple of hands. Right? The volume of traffic that you have is modest. But then when you go to production, right, you're going to go from dozens to thousands or tens of thousands of simultaneous client connections.

And, right, so all of a sudden the whole scale environment catches you by surprise. There's, on that point, there's a kind of an apocryphal story by a financial analyst, or at least he used to be a trader, Nicholas Nassim Taleb, who wrote Black Swan, Antifragile, and he talks about the first thousand days of a turkey. And how the turkey learns that humankind is about making sure that the health and welfare of the turkey is ideal, right? This is your pilot. Your first thousand days of your application in its pilot phase is like, Oh my God, everything is fine. And then it's Thanksgiving.

Lori MacVittie (17:45.742)
I was gonna say is this,

Joel Moses (17:46.995)
Hey, maybe this could

Lori MacVittie
please air this on Thanksgiving.

Joel Moses
this could be our Thanksgiving episode.

Lori MacVittie
Thanksgiving episode. Ha ha ha.

Joel Moses
It's fantastic. Ha ha ha.

Mark Menger (17:53.269)
And so, there you go, right? So Thanksgiving for the turkey completely upends their view of humanity, right? And then and what you've learned from your pilot for the you know, the thousand days in your pilot phase, you go to production, and I'm sorry, it's like Thanksgiving. Everything that you learned, you discover is you learned things that might be valuable, but the most important lessons you should have learned,

Joel Moses
Yeah.

Mark Menger
you didn't.

Lori MacVittie (18:20.066)
But who does the turkey tell? I mean it's Thanksgiving. He's not telling anybody! I mean how does, so

Mark Menger (18:26.498)
Exactly, exactly.

Joel Moses (18:29.717)
That's great. Well, I

Lori MacVittie (18:31.022)
Wow, maybe, maybe yeah, we're, alright, we've run that into the ground.

Mark Menger (18:38.186)
Well

Lori MacVittie
So maybe let's give some practical advice other than

Joel Moses (18:41.247)
Yeah. I have

Lori MacVittie
don't be a turkey.

Joel Moses
learned don't be a turkey. Yeah, that's number one.

Mark Menger (18:46.592)
Mm-hmm.

Joel Moses (18:47.059)
But to avoid being a turkey, what you should do is avoid a single minded focus on securing things that probably don't need the security level that you think that they may require simply because you think that they're critical to the core of the system. If you're using an open weight model and you're spending a whole lot of time on protecting and isolating that file, it's a freely downloadable thing. Instead, pay attention to loose coupling architecture, pay attention to the front doors, and pay attention to things that the AI system may access and make those as bulletproof as possible. And that's the better way to work this.

Mark Menger (19:27.594)
And I agree. I mean to belabor the turkey for just one more moment. Humor me here.

Lori MacVittie
Ha ha ha, okay.

Mark Menger
Right, which is the question to ask, you know, to help the pilot not be a turkey is to ask, What was the worst day that we've had in our pilot phase? And if the answer is We haven't had any, okay, then you haven't learned all your lessons yet. And so you need to be really, you know, part of your pilot preparation for your production rollout needs to be stress testing. Stress testing security, stress testing resiliency, stress testing performance.

Lori MacVittie (20:02.67)
Nice, nice. I like that. I like that. Okay. Don't be a turkey and stress test.

Joel Moses (20:09.193)
Man, I'm hungry now.

Lori MacVittie
All right.

Mark Menger
Ha ha ha.

Lori MacVittie
I know. It's time for a, that's it. It's, that's a wrap.

Joel Moses
Time for a sandwich.

Lori MacVittie
It's time for a sandwich. So we're gonna wrap up and say, hey, can you subscribe? Because you could lock down a file all day, but if your runtime is wide open, you didn't secure AI, you secured your delusion.