Explore the evolving world of application delivery and security. Each episode will dive into technologies shaping the future of operations, analyze emerging trends, and discuss the impacts of innovations on the tech stack.
Lori MacVittie (00:02.956)
Welcome back to Pop Goes the Stack, where every simple change arrives with bonus dependencies and a side of surprise billing. I'm Lori MacVittie. Let's open the box labeled No Risk. I'm all alone today because Joel has been yelling at AI so much that he's lost his voice. So everyone wish him a speedy recovery so he can come back and share his witticisms with us. In the meantime, today we're going to talk about AI and where it's gonna live.
Because for two breathless years, there have been headlines suggesting that every AI workload's gonna live in the cloud. Does that sound familiar? Because I've heard this one before. But reality has arrived and it's carrying a rack of GPUs and very large power bills. Dell's AI server revenue is up 757%. Yeah, that's a lot if you're doing the math.
HPE is actually sitting on a multi million, no, multi-billion dollar AI backload. So there is a lot of demand for all of this. So who could have predicted this? I think our guest, David Linthicum, might have been able to. I don't know. Maybe, you think you could have predicted this one, David?
David Linthicum (01:20.622)
No, I did predict it. Yeah,
Lori MacVittie
Oh, okay.
David Linthicum
there should be a shrine to me on every, HPE should have one, IBM should have one, Dell should have one. There should be just everybody worshiped at my altar as they come in and out of the building. 'Cause I called this a long time ago and took a lot of grief for it,
Lori MacVittie (01:36.822)
Mm-hmm.
David Linthicum
but here I am proved right.
Lori MacVittie
Oh yeah, for a while, just saying the word repatriation was like, you know, just sign right the severance, just you're gone. It was really kind of crazy. But I, you know, way back in 2014, you could look at the landscape and go, it's always gonna be multi-cloud and hybrid. People are not gonna leave the data center wholesale, they're not gonna, you know, leave the cloud wholesale.
There's gonna be workloads all over. And the question has always been which workload goes where? And really, I think that's the point that enterprises are at. And that's why they're now buying infrastructure to try to figure this out is not all workloads are good in the cloud, not all workloads are good on-prem. The question is, yeah, how do you decide?
And you've actually spent a lot of your career like helping organizations like figure out how to architect solutions, where to put things, write strategies. So I thought it'd be a good discussion today if we kinda had that. Like, you know, I've got a new workload, David, where do I put it? You know, what should I consider?
David Linthicum (02:49.005)
It's easy, where it can bring the most value back to the business. Who would have thunk that? I mean, that's
Lori MacVittie
What?
David Linthicum
conversations I have with everybody every day. And people what does that mean? Well, you just mentioned it, there's reasons to put things on the cloud, a public cloud provider, reasons to put things on a neocloud, reasons to put things on premises with you know something in your own equipment. And you have to look at the business value of doing each, which means you have to just kind of spend a little time, figure out, you know, why you're doing it, what are the core attributes of the business that you're looking to solve, what are your security parameters, governance parameters.
And the big thing, how much it's gonna cost? You know, putting things in the cloud, certainly AI in the cloud, that's gonna be ten times the amount of money to use the GPUs as a service from the public cloud providers than the on-prem version. Now there's a reason that on-prem may be more disad-, you know, more or less advantageous to your business problems, things like that. You have to have people around to maintain it.
But you have to look holistically as to what it is and where it should be optimized. And I always saw for a long time that if people are gonna put their AI workloads up in the cloud, they can't afford it. They don't have a big pile of money. They're not printing money in the back. And so if they are gonna run the AI loads that everybody says they're gonna run, then it's gonna be on their own equipment because that's the only stack they're gonna be able to afford. And they're gonna just get good at running it. Great thing about running things on premises now is it's not what it was twenty years ago.
You're going to use the co-location providers, maintenance service providers. You're gonna have leasing companies to lease your equipment, people to maintain it, things like that. And even with that overhead, it's going to be a fraction of what it's gonna cost if you put it up in the clouds. Now I'm getting, you know, token shock with my clients now. They're calling me
Lori MacVittie
Token shock.
David Linthicum
and, you know, they built a you know series of agent-, you know, swarms of agentic AI stuff. They're leveraging remote LLMs that are up on a hyperscaler and they're getting these hundred thousand dollar bills where they thought it was gonna be a thousand. And
Lori MacVittie (04:33.889)
Mmm, ha.
David Linthicum (04:34.049)
they're asking me for help. And I said, "Well, it's just a reality. They're not going to be able to give you a discount. That's even going to get worse. And so we have to figure out alternatives. If you want to run those systems, we have to do so in a more optimized way. And either we optimize your agents or optimize, you know, your monolithic AI system or you know, put it back on equipment that's going to be much cheaper to run, or even put it to downsize-based processors." Everything doesn't have to run on a GPU.
Lori MacVittie (04:59.297)
Ha ha ha. What?
David Linthicum
I don't know where that religion came from.
Lori MacVittie
What?
David Linthicum (05:02.291)
And so looking at, you know, looking at the ability to, you know, basically create systems that are not gonna break the bank and that's the reality of it. And I think many of the tech providers are seeing that reality now as people are moving back on prem.
Lori MacVittie (05:17.537)
Wow, yeah. It's not that much different than the cloud where people were like, wait, why am I getting these surprise giant bills? Like I didn't realize I left things running. I launched too much, I, you know, I, I, I, I, I, I didn't. Right? You can't forecast dynamic demand. And especially when you look at AI and you got, okay, how many tokens? Well, I don't know. You know, what is your customer gonna be typing in the chat bot?
You know, how much? How are they gonna use it? Right? You can't necessarily forecast that. I see a lot of people responding to that particular problem by constraining their chat bots, like you can only ask these questions. Like it already knows. It won't let you type anything. It's basically a glorified FAQ, right? That just acts like it, you know, can talk to you. And that's a valid, you know, way to deal with the billing surprises is well, we'll just constrain it somehow.
But for things like agents, like part of the benefits of that is that it is dynamic, that it's figuring out how to do things in ways maybe you didn't consider that might be more efficient and bring more value, but you have to let them run, which means you have to accept that it could be crazy costs. And maybe running those kinds of things on-prem where you control a little bit more might be a good idea. Even just to start. Right?
Why, you know, I'm gonna test this. Your test cost three million dollars and it failed. Oh, now what? You know, you're out that money. You know, try it somewhere else.
David Linthicum (06:56.311)
Yeah. And also it's commoditizing the processes that should be commoditized. There nothing that's magical about running a remote LLM, you know, a foundational model on a public cloud provider versus running it on-prem. And I built some, you know, quasi foundational models on old laptops that were 10 years old just to kind of prove out what I could actually do. And it wasn't that bad.
And so at the end of the day, if you're pretty good high end servers or leveraging a managed service provider that're providing you the servers, things like that, or even a neocloud, you're going to find a better alternative than the public cloud providers. Not for everything, because obviously the public cloud providers, you know, come with a huge ecosystem and everything and anything to do IT. It's all there, middleware, security, governance, all that kind of stuff.
However, if you're going to run these systems, you need to figure out where it's going to be optimized and how you're going to mix and match it. There's not a law that says you have to run everything up in the public cloud providers. The reality is we've been doing things in a distributed hybrid and multi-cloud way for a long period of time.And so you pick the best of breed system that's going to bring the most value and the one that's going to be the most optimized.
You gotta remember these architectures, you know, these are gonna be five factorial and permutations in terms of the kinds of equipment we can use to solve them, on-prem and off-prem, things like that. But there it's only a few that are gonna be optimized. So you have to get a good architect that's not only just looking at making something work. Guess what? I can make everything work. Given enough money and time, I can.
But the reality is there's a cost component in there that needs to be considered as well. And that is times 20 now with the AI stuff. Back in the cloud computing days, if we screwed something up, it was going to be three times, four times the cost, wouldn't really kill you. AI is not that. It's poison
Lori MacVittie (08:30.539)
Ha ha ha.
David Linthicum
as far as just taking money out of your account. And it's gonna get worse. They're, you know, using the LLM-based stuff as a lost leader now. So even though people are paying hundred thousand dollar token bills and balking at that, you ain't seen nothing yet.
David Linthicum (08:43.639)
They're losing five dollars for every dollar that they're billing you. And so eventually they're gonna have to be profitable and raise those prices up. So if you become addicted to them, you bind your applications to them, you're gonna find that you're gonna be in a cost problem. In other words, it's gonna drain your resources. You're not able to put the resources on the areas of the business that you need to focus on. And that's gonna be a catastrophe.
So much so that I think many of these businesses are gonna go out of business because of they're misaligning their AI spend with the strategic needs of the business and how many resources they can apply to it. And if they just take an architectural look, take a more objective look, take a best of breed look right now, understand that we're going to be dealing with heterogeneity. We're not going to deal with a single cloud provider.
Everything's not going to be on AWS, Microsoft, or Google. But ultimately it's going to get you in a better place. It's going to provide the scalability and the cost effectiveness to bring your architecture and use of AI, you know, into 2030. And I think that's what people are focused on right now: is how do I take our business to the next level? And where should I be making the investments?
Also, the investments you're making today are going to come along with lock-in. It's going to be very difficult to change platforms, you know, in three years, five years, very much like the repatriation we're seeing right now. And that's going to be technical debt that ultimately has to be fixed. But you may generate so much technical debt, you go bankrupt.
Lori MacVittie (09:56.334)
And that I think the lock-in and the tying your apps to it with just an LLM and using an API, most people who are providing some sort of AI service are adopting the OpenAI API, which is fairly simple if you've ever looked at it, right, from a development perspective. It's nothing magical, it's really easy. So if something else supports it, you can switch it out. That's good.
But when you get into agents, now we're looking at entire frameworks, right, entire ecosystems that you're suddenly buying into. That is going to be much more difficult to get out of in the future if you go, well, we're just going to do it all in this cloud using this framework because it makes it easy. And yeah, the easy is what traps you, right? It's also easy to walk into quicksand. Not so easy to get out.
So I think the agent side is really where we're going to see a lot of people getting into trouble and tying their futures to a particular cloud or provider and then not being able to extricate themselves later when new solutions come along. Because there are people looking at how do we make, you know, transformers more efficient? Or how do we do this in a way that's more efficient than transformers? And they're having some success.
So this is LLMs that we see today, this generation, are not the final generation. And it's going to change. And being able to take advantage of that, which may come at a much lower cost if it's more efficient, is gonna be incredibly difficult if you just go all in on, "hey, we're just grabbing this agent framework. That's it. We're going with this." Because those are processes.
And once you tie processes to a framework, you're pretty much, you know, back in the, "why do you still have a mainframe? Well, we have this process..." Right? I mean it's, this is how these things happen. You're stuck with it until it gets completely replaced. It's kinda scary.
David Linthicum (11:54.038)
It is scary and I think you kind of hit the nail on the head. I mean, at the end of the day too, we have to look at the applications and what's going to be bringing business value in. So what are the applications look like? I think everybody is trying to boil the ocean now, do the moon shots, and we're going to create, you know, an LLM for banking, things like that. The reality is the AI applications, the ones I'm teaching people to build right now in my AI architecture class, is going to be very narrowly focused, very tactically focused use cases for AI.
It's going to be inventory control, supply chain integration, really kind of things that are core to the business. Sales order entry, automation support, all those sorts of things, which are easy problems to solve, certainly if you're using AI. But at the end of the day, the use of AI of those systems is going to be fairly rudimentary. In some cases, we can use traditional machine learning, you know, versus throwing generative AI at it. And there's no reason to agent-fi everything just because it's cool, you know, like the cool tech bros are doing right now.
But use the technology that's going to be the minimum viable solution for your particular application. And that's where people are missing now. And so they want to use the best and brightest, the latest and greatest LLMs and the latest and greatest foundational models, things like that, the latest and greatest agentic frameworks. That's not what success looks like. What success looks like is I'm going to use the minimum viable technology, AI or not AI, that's going to solve the problem.
And I always, you know, in my architecture world and I talk about this in my speeches as well, it needs to be good enough for what it's for. It doesn't need to be good,
Lori MacVittie (13:23.713)
Mm-hmm.
David Linthicum
you know, good enough for, you know, something, you know, some kind of proposed future that we don't know yet. And I understand AI can be a force multiplier for your business. I built businesses around that. And, you know, my time at Deloitte automated lots of folks and got a lot of people going on AI and they're getting huge amount of benefits from it now.
David Linthicum (13:39.138)
But the reality is I took a same sort of a minimum viable approach to make it happen. I divided up their problem domain into particular application solutions able to bring to bear. You know, some using AI, some not. And I'm basically not reinventing everything with some nifty new technology. I'm looking at the minimum viable technology I'm able to bring to bear to make it happen. Sometimes that may be agentic, most times not. You know, that's we're kind
Lori MacVittie
Mm-hmm.
David Linthicum
of overusing that technology. Sometimes generative AI, sometimes not, sometimes AI, sometimes not. And so having that kind of discrimination, that doesn't seem to be out there right now because we're in the middle of the AI hype, is something that's kind of missing. We're not bringing the objective framework to making it happen. I mean, my role is a designated buzzkill. And so I'm going in and, you know, auditing projects and mentoring projects right now.
My question's always gonna be, "okay, why are we using that? What's that for? You know, what benefit does it bring us?" And you can't explain it, you know, in a couple of sentences, then probably shouldn't be using it. There's probably going to be some alternatives going to be, you know, a hundredth of the cost and much easier to operate that are going to get us to where we need to be. Now it's great that we have AI stuff. We're able to bring to bear some magical technology that's able to bring things to the next level.
Supply chain integration, fraud detection, I'd build those applications all the time. They love generative AI. They love aspects of LLMs. All that stuff is great. But I'm only going to use it if I need it. You got to remember there's 150 LLMs out there.
Lori MacVittie
Yeah.
David Linthicum
And what are they adding, you know, three or four every week. You know, as we're revising and doing that stuff, you don't want to get on that train. You want to get on the train
Lori MacVittie (15:09.665)
No.
David Linthicum
of what technology you're going to leverage, how you're going leverage it, and how you're going to find a purposeful use of it. And that needs to be the way that we're thinking right now. And by the way, when I talk about this, everybody's heads are nodding. I get it, you agree with me, but you have to kind of live up to it and have the political wherewithal within these organizations to start asking some of the tougher questions that we don't seem to be asking.
David Linthicum (15:28.769)
But we are seeing kind of a shadow IT thing occurring where people are buying with the on-premise stuff. They're not necessarily going to the cloud. Or many times are going to the cloud to do their prototypes, but they're deploying, you know, on-prem or on-prem analogs, managed service providers, co-location providers.
Lori MacVittie
Mm-hmm.
David Linthicum
So that tells me that people are thinking properly about how to do this and also they can't afford pushing everything into the cloud.
Lori MacVittie (15:50.626)
Yeah, I don't see the same, well, I see the same with AI as we did with cloud, in that there are mandates, right? Cloud first, right? AI first. You must use AI. But when we actually go out and you know, survey and look at what's really happening, they're using different models, a lot of open source models, especially on-prem, and they're using them for specific use cases.
Like, "well, I want to do automation of business stuff, so I want that to be on-prem, and that means I'm gonna use open source." Okay. And then we see people also going, "Yeah, I want the benefit of the newest and latest, greatest for these other use cases, so I'm gonna use the cloud." And we see that mix like already. They're using an average of seven different model families, but they're also using a mix of some AI as a service. We're using some cloud provider stuff, and then we're also on-prem with these open source models.
So I see people almost as if they've learned the lesson of cloud already and went, "yeah, we know it's gonna be everywhere." And now the trick really is, like you said, figuring out which ones go where based on value, architectural fit-- which should never be underestimated. Right, if everything else is on-prem, why are you using an AI somewhere in the cloud? Like I, that doesn't make sense necessarily, right?
Think about why you're using that particular model over a different one. What benefits does using, you know, the latest model of GPT or Gemini give you over Mistral, open source, right, whatever you're using on-prem? That's also a question. I think that's the harder one for people to answer is, "what do those models give you over something else?" And you kind of touched on that. And you said, right, "why are you using AI? Why are you using this model for this, you know, AI purpose?
I think that's another question people have to ask. Is there a difference? What is it? What are you getting? What's the value add?
David Linthicum (17:51.928)
Yeah, almost none. I mean, the reality is that the way you use like a remote LLM from is going to be, you know, plus or minus three percent, you know, versus ChatGPT 3.5, which came out a couple of years ago. And so the reason why we're typically using it for business purposes and the use of that system is gonna be rudimentary. And so also the most important thing for that model to understand is your own data.
And so if I'm able to do a good analog of what an LLM is, I'm not going to be able to run ChatGPT 5.4, you know, on my local server, but I'm going to be able to do so in such a good enough way--open source stuff, things like that--where I'm also able to consume my own data and have a model that's customized for my particular need. So, in essence, kind of an LLM for myself that's able to run on my equipment, which is gonna cost me no overhead. In other words, it cost me the cost of the equipment. There's capital expenses.
But as far as, you know, getting these huge cloud bills at the end of the month and having to worry about token consumption, that's not necessarily gonna be in the cards. And that's not only gonna be better than I think leveraging some sort of remote system, but it's gonna be customized for your particular business purposes. In other words, if you're in the automobile business, the retail business, healthcare business, things like that, we're able to build models that reflect what our business is.
We're not able to, you know, augment existing mega models, the LLMs up there, using our own data, which is a bit dangerous, start uploading information and have it read information directly from us. That may or may not be able to live up to the purpose. And also the horsepower that's up there isn't necessarily going to be needed. So we're using a very small percentage of what those AI systems provide, yet paying for the whole thing. Certainly with AI consumptions. If you ever look at an AI dialogue, there's lots of stuff going on, that's why the tokens are flying back and forth.
You have to re-enable state, you know, every time you
Lori MacVittie (19:39.245)
Ah, yes.
David Linthicum
communicate back with the AI and kind of figure out where it is. Where I'm not necessarily gonna have to do that. I don't care ev-, I'm able to maintain perfect state because I'm gonna own the system. I'm not going to be using somebody else's stuff. So it's just a consideration. Now now people come to me and they say, "well, Dave, that's gonna cause heterogeneity because I have, you know, different technologies and different systems that are there, different physical systems, as well as complexity." Yeah, duh. That's the way it's gonna be. In other words, they're gonna leverage best of breed technology.
David Linthicum (20:06.173)
Get good at leveraging complexity. You're already
Lori MacVittie
Yeah.
David Linthicum
going to be doing it anyway. Even if you outsource all your AI systems, you know, into the cloud, many of the systems are going to run across in a hybrid tripe fashion. We're going to have stuff in neoclouds, you know, public cloud providers, AWS, Microsoft, and Google, our own equipment, managed service provider, edge-based systems; it's gonna be all over the place. So we have to get good at managing the complexity and managing the heterogeneity.
And if we're able to do that, we're able to pick and choose the right systems that are going to be best of breed for its purpose. And that's going to lead you down the right path. I'm meeting companies right now that are spending between $100,000 and $100 million on basically doing the same technology. And the reason why, they're making bad decisions. And they have the cloud providers in there and they have the, you know, AI tech bros in there basically pushing them in the direction of these larger base systems. "Oh, everything's gonna be remote, everything's gonna be outsourced."
And they feel good about that because that's what they're getting at the conferences. But the reality is I can do it at a fraction of the cost. And guess what? Somebody like me is going to show up in a couple of years and tell your leadership that you guys spent $100 million when you should have spent $100,000. And that'll be a good conversation you're going to have. And very much like we had with the cloud. But we're not seeing the multipliers
Lori MacVittie
Ha ha ha.
David Linthicum
with this kind of waste that's going on. So use your head. This is a business game. This is not a technology game. You know, it's not he who burns the most tokens wins. That token matching stuff was, you know, industrial string stupid. So let's not follow stupid with stupid. Let's think about using this technology in a way that's going to be optimized for its particular purposes. Not devaluing AI, not devaluing the clouds. Those are going to have a purpose.
We're going to have a reason to use those sorts of things. I'm just saying right application use cases for the right platforms, the right purpose, and spending the right amount of money, or at least get as close as you can. I think we can't do the waste thing like we did back in cloud computing. I'm seeing it reinventing itself now. That's just going to get companies uber in trouble. They're going to run out of money.
They're not gonna be able to participate in AI because they don't have any resources to move it 'cause they blew it all with a few prototype systems that they built back in 2006.
Lori MacVittie (22:07.757)
Absolutely. And I, we're out of time, but I think you just summed up the takeaway: right tool for the right job in the right place. Like that's what I hear is that same theme. Like, right, pay attention to what you're doing. Make right decisions based on data, on locality, on cost, on value, and match these together and you will be on a better path than if you just flock to the latest, greatest. Right? Don't over provision your AI. Use the right one.
David Linthicum (22:38.541)
Who would
Lori MacVittie
So this is
David Linthicum
have thought in two thousand, you know, in 2026 we'd be having this discussion. I figured people would have figured it out by now. So it's different decade, same argument.
Lori MacVittie (22:45.677)
Same argument. There are new people that need to hear it. So, you know, it's good to say it. And it's good to hear it from you because, right, you've seen this, you know this. Keep saying it because people obviously need to hear it. So thank you for coming on and sharing your wisdom. Really appreciate it. But that's a wrap for Pop Goes the Stack. So please hit subscribe because the next episode might be triggered by a minor patch and major regret.