Pop Goes the Stack

AI “epic fails” aren’t just funny headlines; they’re patterns you can design against. In this episode of Pop Goes the Stack, F5's Lori MacVittie, Joel Moses, and Buu Lam walk through why so many AI-powered features keep going off the rails, from chatbots inventing policies to agents deleting real infrastructure, and what those failures teach us about building safer systems.

Joel frames most incidents in two buckets. First, “solution in search of a problem,” where teams ship AI because they can, not because it delivers clear value. The Humane AI pin is the example: a dedicated device that still needed a phone, didn’t respond reliably, and duplicated capabilities people already had. Second, treating a statistical prediction engine like an authority. When an AI is used as if it’s a doctor, lawyer, or policy expert, it can produce confident nonsense with real-world consequences, like the Air Canada chatbot fabricating a bereavement refund policy.

Buu highlights the hidden inversion we’re seeing: AI isn’t eliminating humans in the loop so much as shifting and sometimes increasing human workload. Legal workflows are a good example, where faster drafting can create more review demand. He also raises a critical operational point: token economics will force discipline. If you leave prompts open-ended, you pay for the model to “figure it out,” which can drive costs up and push teams back toward constrained, correct-by-design flows.

The practical enterprise takeaway is permissions and agency. An agent “doing the thing” is still doing it as you. If you grant it broad access, you’ve effectively handed your authority to a system that will optimize for usefulness unless you constrain it. Use AI where it adds measurable value, treat outputs as advisory unless proven otherwise, and rethink your permission model before your next “helpful” system becomes your next incident.

Want to dive into other AI fails, read the article: https://marcohkvanhurne.medium.com/the-ten-biggest-ai-fails-of-2025-5d14fe876b2a

Creators and Guests

Host
Joel Moses
Distinguished Engineer and VP, Strategic Engineer at F5, Joel has over 30 years of industry experience in cybersecurity and networking fields. He holds several US patents related to encryption technique.
Host
Lori MacVittie
Distinguished Engineer and Chief Evangelist at F5, Lori has more than 25 years of industry experience spanning application development, IT architecture, and network and systems' operation. She co-authored the CADD profile for ANSI NCITS 320-1998 and is a prolific author with books spanning security, cloud, and enterprise architecture.
Guest
Buu Lam
Director Community Evangelism
Producer
Tabitha R.R. Powell
Technical Thought Leadership Evangelist producing content that makes complex ideas clear and engaging.

What is Pop Goes the Stack?

Explore the evolving world of application delivery and security. Each episode will dive into technologies shaping the future of operations, analyze emerging trends, and discuss the impacts of innovations on the tech stack.

Lori MacVittie (00:02.397)
Welcome back to Pop Goes the Stack, where AI powered means good luck debugging and the logs have decided their performance art. I'm Lori MacVittie and we're here to make the chaos explain itself. Cause it's too hard to explain it for us. It's really crazy. So this week we have, of course, our co-host Joel Moses.

Joel Moses (00:24.514)
Good to see you, Lori.

Lori MacVittie
Glad you're here. Cross fingers

Lori MacVittie (00:29.659)
nothing else goes, we've had some problems getting Joel connected. So

Joel Moses (00:34.328)
Gah, technical difficulties.

Lori MacVittie
we're just hoping. Yeah, we're gonna blame it on AI. And Buu Lam is here today. Buu, welcome.

Buu Lam (00:40.77)
Hello Lori. I'm glad I did not wear my My Little Pony shirt today.

Lori MacVittie (00:45.361)
You

Buu Lam
We would have clashed.

Lori MacVittie
No, they would, like,

Buu Lam (00:50.934)
You know

Lori MacVittie
Next

Buu Lam
who probably has one is Chase. I'm gonna throw

Lori MacVittie
Yeah, okay,

Buu Lam
that out there right now. So...

Lori MacVittie
Well alright. Okay, I know who to ask then and to prep so that we dress alike. So this could be fun, but today it's time to talk about AI Epic Fails or Epic AI fails, depending on, you know, what version of grammar you'd like to use.

Lori MacVittie (01:12.689)
Because nothing says cutting edge innovation quite like Google confidently hallucinating entirely new NASA histories or creepy little neck snitch gadgets that are marketed as your friend when really they're just surveilling you and nobody asked for it. So this episode we want to talk through different kinds of AI fails. There's many of them. We could spend three episodes talking about them, but let's just pick some of the favorites that give us some lesson that we can take away from it and apply to life, in general.

So, I don't know, Joel, you're always good at this. Why don't you kick us off? Pick one and let's just run with it.

Joel Moses (01:56.503)
Well, let's see. I, you know, I'm sure we're gonna link the article that all of us read that is kind of a dog's breakfast of AI failures,

Lori MacVittie (02:04.581)
It's crazy.

Joel Moses
and that's a good one. Of course there have been plenty of other AI failures outside of that, but they tend for me to fit into a couple of generic classes. One is the solution in search of a problem. Like, okay, we can do this, so let's just go ahead and do it. We won't think about

Joel Moses (02:24.194)
whether it has value, we won't think about anything else. We're just gonna try it. And things that aren't really, you know, intended to or you don't understand how they deliver value, they typically fail. That's not a surprise.

Lori MacVittie (02:36.775)
Ha ha ha.

Joel Moses
The other thing is that, you know, a lot of times people forget that a statistical prediction engine is not an authority. It's not a doctor.

Lori MacVittie
What?

Joel Moses
It's an advisor. It's just a thing that might help you orient your thoughts, but you still have to be

Joel Moses (02:52.748)
the one who decides whether those thoughts are correct and whether you should apply them or not. So to have an AI, you know, suggest that you should switch out the salt in your diet for something utterly poisonous, well, you know, maybe ask a doctor or some a nutritionist maybe, someone who actually has some specialized knowledge and is not just a generalized engine.

But the one that I kind of like to think about is the humane AI initiative, which was to try to create a pin that you could ask questions of. And that the pin also had a little camera in it so you could look down at it. I don't know what it your face would look like looking down at that pin,

Lori MacVittie
Ha ha ha.

Joel Moses
but somehow it would understand, you know, what you were wanting to do. And of course the pin had some, oh I don't know, shall we say flaws. Not being able to respond in time or taking five minutes to connect and initiate. And here's the funny part, the pin needed a phone to work, right?

Lori MacVittie (03:51.389)
Oh I remember that, yes.

Joel Moses (03:51.523)
But you could also do exactly

Lori MacVittie
Yeah. Mm-hmm.

Joel Moses
the same stuff you were asking the humane AI

Lori MacVittie
Just do it on your phone.

Joel Moses
pin to do on your phone. And nobody thought that that was a problem. And so

Lori MacVittie (04:02.301)
Ha ha ha.

Joel Moses
I fit that in that first category, which is, you know, it's a solution in search of a problem. You know, it's technical brilliance, even though I absolutely adore technical brilliance, is not equivalent to creating value. It's just technical brilliance for brilliance sake. And that doesn't connect to anybody.

Lori MacVittie (04:19.323)
Yeah. Yeah. What do you think, Buu? Would you buy it? You know, you need a little pin or a friend that hangs around your neck and listens to your conversations and everyone else's too?

Buu Lam (04:30.711)
I think that's what my dog's for and

Lori MacVittie (04:32.701)
Yeah, ha ha.

Buu Lam
it my dog doesn't judge me most of the time. Sometimes it gives me a funny look and I don't know if it's based on it actually knowing a little bit of English, but yeah, I've got one of those already here.

Lori MacVittie (04:47.139)
Alright.

Joel Moses (04:48.056)
Yeah.

Lori MacVittie
Yeah, that, I got one too. He's sleeping now, but yes, he also listens to everything. It's kinda creepy.

Buu Lam (04:55.847)
I think I do, you know, I do have an example, and I this is, follow me here, a bit of a stretch, but I was inspired by what I just read about the Taco Bell and their voice drive-thru handing out or trying to order 18,000 cups of water. And so I got my mind on remembering that if you guys have seen the documentary around the McDonald's ice cream, no, the

Joel Moses
Oh.

Buu Lam
McDonald's milkshake

Buu Lam (05:25.065)
thing, right. So McDonald's has this proprietary milkshake device and to repair it, it is highly locked down so there's no manuals or anything. You have to call some company that is the only company in the entire world that knows how to do this milkshake machine but through AI they were able to reverse engineer it and open source a manual on how to fix it and how to read the error codes on it.

And, lo and behold, the project gets shut down because that is not aligned to the value in that case. And so even though AI produced some sort of value and some sort of outcome, that was not aligned to the ultimate outcome. I'm not saying, you know, I don't have a tinfoil hat on right now, but I think that is a a little bit suspicious to me.

Joel Moses (06:20.472)
That's true.

Lori MacVittie (06:20.783)
Yes, yes. That's definitely an interesting,

Joel Moses (06:25.198)
Yeah.

Lori MacVittie
they can fail while not failing, right? I mean that's I think the interesting part. That comes back to hallucinations and bias and whatnot, but really, I mean AI can be working perfectly fine and still completely fail.

Joel Moses (06:39.725)
Yeah.

Lori MacVittie
As Joel said, like keep the salt in your diet, okay? Just saying.

Joel Moses (06:43.446)
Yeah.

Buu Lam (06:43.608)
A hu

Joel Moses
Of course.

Buu Lam (06:43.608)
A human jumped in the loop on that one.

Lori MacVittie (06:48.188)
Yeah.

Joel Moses
Yeah. I, you know, it's all about, you know, it's interesting, in the field of engineering we often optimize for correctness and performance. Like we want it to work reliably and we want it to work correctly. And I think that right now in AI engineering there is a tendency to optimize for usefulness

Joel Moses (07:07.682)
which is a different thing than

Lori MacVittie (07:09.648)
Mm-hmm.

Joel Moses
correctness. A search engine doesn't have to be perfectly correct to be useful. A legal opinion has to be correct to be useful. An insurance denial has to be correct to be useful. And a medical recommendation does as well. Right? So I think that part of this is the fact that we are number one, as humans, we're trying to trust something that isn't an authority. And number two, the people who are developing it are not thinking about correctness, they're thinking about utility.

Joel Moses (07:37.45)
And that's different. That they don't align with each other.

Buu Lam (07:41.814)
Well, you mentioned the case of legal review. And so, you know, when GenAI first becoming broadly available, they're saying, "Oh, this is gonna replace lawyers." But now as you read, it's actually created the need for more lawyers because you still need humans in the loop. And it has accelerated cases to the point where there's a pileup of cases because now you can get through that initial examination a little bit quicker and then need to involve more lawyers to be able to see it the whole way through.

I just think it that's kind of interesting where it's given the reverse effect of what they thought would happen.

Lori MacVittie (08:20.657)
Right, but we've s also seen, speaking of AI fails, cases where using AI instead and getting rid of the human in the loop results in very bad, right, outcomes like a 16X denial rate for insurance claims. Caused all sorts of right craziness because the AI is just like nope, no, no, no, no, just left and right,

Joel Moses (08:45.356)
Mm-hmm.

Lori MacVittie
right where, you know, when it's that much higher. So completely removing the humans from the loop, not a good idea either, but you know, there's gotta be some happy medium in there.

Buu Lam (08:56.695)
I'll pose this question, what if you put AI in charge of DNS?

Joel Moses
Ha ha ha.

Lori MacVittie (09:04.345)
Umm, can we not? Ha ha ha ha.

Buu Lam (09:07.285)
Would it

Lori MacVittie
Ha ha ha ha.

Buu Lam
Would it be good or bad?

Joel Moses (09:09.922)
Well, that's a good question. So one of the stories linked to on that article was the story about the Air Canada chatbot, where the Air Canada chatbot just out of thin air created a bereavement refund policy, right?

Lori MacVittie (09:21.96)
Yeah. Mm-hmm.

Joel Moses
I would be afraid that if you turned over DNS to the AI that you would have, out of thin air, creations about, you know, which sites route to what.

Lori MacVittie/Buu Lam (09:34.331)
Ha ha ha.

Joel Moses (09:34.883)
You know, the airline, it's interesting to note that the airline at the time argued that the chatbot doesn't represent the company, but let's be real, it does represent the company. It's an incredible lesson for an enterprise to learn that AI output speaking with authority and not as an advisor is indistinguishable to people from an employee that speaks on behalf of the company. And that's something you gotta think about when you're ensuring again correctness and not just simple utility, usefulness.

Lori MacVittie (10:10.309)
Right, right. And sometimes it's too much trust, sometimes it's wrong tool for the, you know, wrong job. It just, it's a mismatch. But I think in terms of epic AI fails to keep that trend going, I think 2026 is going to be the year of what database did AI delete today? Because there've been so many of them. Like you looking at the headlines, okay, who got it today? There's been, you know, Cursor. What was, Replit? What was the other another one? Replit AI.

Joel Moses
Yeah.

Lori MacVittie
That's like that's a remote code. But there have been, right, multiple ones in the headlines where it just kinda goes off the rails and it RMRFs and you're done.

Joel Moses (10:59.384)
Yeah. You know, one of the examples that was given in the article was about exam proctoring. And that one

Lori MacVittie (11:05.474)
Mm.

Joel Moses
So I'm thinking

Lori MacVittie
That was interesting.

Joel Moses
now about my two things about, you know, is it suited for purpose and or is it a good idea with no value whatsoever? And then trusting in something as an authority over an advisor. The exam proctor one is a little bit different. It's both simultaneously.

Joel Moses (11:24.142)
Students were flagged for doing strange behaviors like looking away or blinking too much or changing their

Lori MacVittie (11:30.065)
Breathing.

Joel Moses
posture or being in poor lighting. And

Lori MacVittie
Breathing.

Joel Moses
you know, the funny part is the model wasn't actually detecting cheating, it was detecting deviations from its training.

Buu Lam (11:40.631)
Mm.

Joel Moses
That's all. And if your fraud detector has never seen an unusual customer and can't tell the difference,

Buu Lam
Mm-hmm.

Joel Moses
congratulations, you've just invented customer harassment.

Buu Lam (11:51.137)
Mm-hmm. Well, so

Lori MacVittie (11:53.035)
Wow, ha ha ha.

Joel Moses (11:53.469)
Right.

Buu Lam
think of, you know, your eye movement. What if somebody had an eye operation

Joel Moses (12:00.452)
Absolutely.

Buu Lam
and they're wearing a patch? Or

Joel Moses
Yeah.

Buu Lam
their glasses were slightly lower. They were looking down, they're it was coming down the bridge of their nose and so it's covering their eyeballs.

Joel Moses (12:12.14)
Yeah, here's the simple truth. Humans are edge cases. We're almost always edge cases. And so if you're keying things based on a limited training set, those fraud detection, those exam proctoring systems are looking at a limited set and they think that they have a baseline, but they really don't. And that again, that's there are real risks to that. And again, it's because the people who put together that system, first of all, they're trying to apply the technology to something that actually, you know, just put an exam proctor in the room.

Lori MacVittie (12:45.949)
Well

Buu Lam
Mm-hmm

Lori MacVittie
but it goes both

Joel Moses (12:46.739)
Second

Lori MacVittie
ways.

Joel Moses
Yeah. Second, they're trusting it as an authority and not an advisor. And that's

Lori MacVittie (12:52.603)
Back to the authority. Yeah. Mm-hmm.

Joel Moses
again, it's an amalgam of both.

Buu Lam
Mm-hmm.

Lori MacVittie
Yeah, and I like that, but and this kind of fits in there too, you know, the did AI write this? Right. Oh well, you know, it will tell you yes or no. Sometimes it's right, sometimes it's wrong. It seems to be fifty-fifty here, you know, you're never really sure. But on the reverse side now you've got people who are absolutely certain that AI wrote it if there's an em dash in it; like this was just created by AI or something.

Joel Moses (13:24.204)
I write with the em dashes all the time. I'm sorry.

Lori MacVittie (13:26.417)
Oh, lord.

Joel Moses
I look, hey,

Lori MacVittie
You're

Joel Moses
I

Buu Lam (13:29.079)
Yeah.

Joel Moses
You know which style book I like, Lori. So come on.

Lori MacVittie
Yes, yes. Mm-hmm. Yes, yes.

Buu Lam
So Joel trained AI, I get it now. All right.

Lori MacVittie
That yes. It was trained exclusively on Joel's content over the years, so that explains everything now.

Buu Lam (13:43.383)
Prolific. Wow.

Lori MacVittie
It, yeah.

Joel Moses (13:45.24)
Yes.

Buu Lam
Yeah.

Lori MacVittie (13:45.628)
Yeah, it really is. I mean that's, right, that's happening back and forth, right? And that kind of plays to your, you know, don't treat it like an authority. Well, but it speaks with such authority. Well, they trained it to speak with authority. So you've got, you know, how much of this is train the AIs in a certain way and how much of it is train people

Buu Lam
Mm.

Lori MacVittie
to not do these kinds of things, right?

Joel Moses (14:08.077)
Yeah.

Lori MacVittie
I mean

Buu Lam (14:08.727)
You know, I think for the case of okay, you're trying to replace a human doing support, chat support, for Air Canada in that case. But I can guarantee you, I've worked in a call center in the early part of my career, you are hired off the street and you're given a basically a call flow tree to go through. And then if something strays outside of there and you can't figure it out yourself you call for escalation at that point.

Joel Moses (14:40.878)
Mm-hmm.

Buu Lam
You're not supposed to try to just randomly figure out something in the case of tech support that could delete a database. You're encouraged to, hey, here's more resources to bring another human into the loop at that point.

Joel Moses (14:57.442)
Yeah. And again, those scripts are optimizing for correctness, not for usefulness. Right, so the AI doesn't think about "Oh, I should be correct when I say something, and so if I need to escalate to someone else who knows more than I do, then I will." Instead, it really wants to bend over backwards to help you and in doing so will make things up like bereavement policies and refunds for those out of thin air.

Lori MacVittie (15:22.119)
But I think these

Joel Moses
So lot to learn from.

Lori MacVittie
Yeah, but I think people are learning because I'm starting to see more chatbots that allow me to type anything I want. But there are, right, it's very canned, right? It will only do on topic. If you try to stray off, it'll be like not

Buu Lam (15:37.791)
Not y-

Lori MacVittie
that doesn't count. I won't answer that. So they're

Buu Lam
Not her again.

Lori MacVittie
starting to build that in. Yeah, right, ha ha.

Joel Moses (15:42.072)
Mm-hmm, mm-hmm.

Lori MacVittie
Don't. You can't. Or you can only pick, right, certain questions, right? I've seen so basically a glorified FAQ.

Lori MacVittie (15:50.248)
But it looks like a chatbot, so I guess we're, I don't know, we're supposed to think it's AI even though it's not. But they are pulling back from the just here's a chatbot, talk to it, and getting more prescriptive so that they

Joel Moses (16:02.222)
Mm-hmm.

Lori MacVittie
can get to that correctness, I think. So...

Buu Lam (16:05.089)
Well, I think there's two things in play is that if you're leaving every single respon- or every single prompt to be open-ended, then think of the token spend of just saying, "Hey, here's a something and just go figure it out." That's gonna incur, it could incur a whole lot depending on what you would allow it. Hey, do you want it to actually start writing scripts and actually be able to execute something that goes and retrieves an editor and looks at different bereavement policies and compiles it together and gives that back to the person.

Or do you actually put in very prescriptive scenarios in there saying, "Okay, this person asked for something way outside of the FAQ. Let's not go beyond that and just respond and say, maybe you should call in to someone else." But we're seeing everybody got hooked on Claude Code subscription. Cool, I have these session limits and I can just do whatever I want within here. And then that's been a bit more of a cookie to get people on board. And then I think we already saw last month with Microsoft Copilot saying, "Oh, we're gonna start doing usage based on you

Lori MacVittie (17:17.725)
Mm-hmm.

Buu Lam
now and you're gonna actually gonna see how many tokens this all costs and this might change how you think about all the stuff you're vibe coding."

Joel Moses (17:25.102)
It's

Lori MacVittie (17:25.846)
It

Joel Moses
funny when people offer all you can eat subscriptions to things that have actual fixed costs

Buu Lam (17:30.229)
Ha ha, yeah.

Joel Moses
behind them.

Lori MacVittie
Ha ha ha.

Joel Moses
And then find

Buu Lam
Crazy.

Joel Moses
out later on, hey, you know, that maybe wasn't such a great idea. Yeah, that's this is gonna have to happen. And you know what? Perhaps the concern over token spend is the thing that will get people to not optimize for usefulness but optimize for correctness. Maybe that's exactly what needs to happen, but yeah, you're absolutely right.

Lori MacVittie (17:54.238)
Yeah, I think it that will also help with things like making, hey this friend that talks to an AI, right, that's gonna cost a lot. Are people really willing to pay that much for it? Right? It may pull back on some of those crazy ideas that people are having about how everything needs to be AI enabled, which is good. And that should be true in the enterprise too.

Not every application, not every process, not every thing needs to use or be augmented by or be replaced by AI. And I think that's kind of one of the lessons from all of these fails is like maybe you shouldn't have in the first place. Right? Just because we can doesn't mean we should. So what other lessons would you have people take away from all these fails as they're going about their busy enterprise days?

Joel Moses (18:45.902)
Well the takeaway that I have is first of all, these things are remaining true. We've got two years plus of experience with some of these AI technologies being deployed in mass, and they remain as vulnerable and possibly off the rails as ever. They're getting better, don't go get me wrong. But you know, when people fail to connect, "hey, we can do this," with "we should do this because it has value," you're going to make, you're gonna have errors, you're gonna have things that go off the rails.

When people optimize for the wrong thing and assume that an AI is an authority and not just a simple advisor, it's gonna go off the rails. And so, you know, we can control one, we can't really control the other. But we've got to do better by and large with how we behave around AI.

Lori MacVittie (19:40.899)
Absolutely. Buu, what advice would you leave our listeners with?

Buu Lam (19:45.621)
I think we are if folks haven't already need to examine their permissions models for how they interact with their enterprise resources. And so I think we're at a point where people are experimenting with agents and agents are something that you give agency to to do a thing. And right now I think in people's heads they might think, "Oh, the agent is doing it." Well, the agent's doing it on your behalf, so you have passed on your agency to the agent to execute on the thing.

And that is something that people have to accept and start to say, "Okay. Oh, now I have to think about what my agent is doing as me," as opposed to this other nebulous thing that it just goes off and there's GPUs and something happens and I get my bereavement policy back. Once that keys in, I think, people might start thinking about what they're doing a little bit differently.

Lori MacVittie (20:44.645)
I like that. You got kind of two sides here, right? Like Joel's, right, think about the value. Right? It's got to add value to make it worth it. And, right, be careful with that. And that's the, you know, with great power comes great responsibility. Right? I mean this is that's you doing it, somebody doing it for you. So I think those are both good sides because they're both true. Right, we have to get better about picking the right place to insert and inject AI.

And we have to get better about keeping it where it needs to be and not letting it go off the rails. So, well that is a wrap for this week's Pop Goes the Stack. Please subscribe before the next vendor promises observability with zero instrumentation and maximum confidence.