Explore the evolving world of application delivery and security. Each episode will dive into technologies shaping the future of operations, analyze emerging trends, and discuss the impacts of innovations on the tech stack.
Lori MacVittie (00:02.501)
Welcome back to Pop Goes the Stack, where model routing promises the best brains for every task and delivers a new layer of "it depends" to debug at 2 a.m. I'm Lori MacVittie, joined as always by Joel Moses. Welcome, Joel.
Joel Moses (00:18.154)
Good to be here.
Lori MacVittie
Awesome. Today we're going to talk about choosing models without letting the chooser become the problem. Now, this is actually important because on paper, model routing
Lori MacVittie (00:30.309)
looks a little bit like traditional load balancing wearing a brand new coat. But under the hood, we aren't just steering packets or requests anymore. We've replaced simple network metrics with a whole new beast of variables: token budget, semantic intent, and model cost benefit parameters in real time.
So grab your coffee, hold on to your routing tables, and let's discuss why you can't just throw around Robin at a modern LLM cluster without everything falling apart.
Joel Moses
Ha ha ha.
Lori MacVittie
To help us navigate this, ha ha, did you see what I did there? It's ni-
Joel Moses (01:08.286)
Route it.
Lori MacVittie
It's kinda like routing. Yeah,
Joel Moses
Mm-hmm. Yeah.
Lori MacVittie
routed.
Joel Moses
Got it.
Lori MacVittie
Yeah. Okay, I tried. I tried. is Patrick Roughan. Patrick Roughan, welcome.
Patrick Roughan
Yep.
Lori MacVittie (01:19.227)
He's like, yeah, I'm here, but I'm not sure I want to be anymore after that intro, so...
Joel Moses/Patrick Roughan (01:23.701)
Ha ha ha.
Lori MacVittie
Let's start. I mean, model routing, it is a form of load balancing, but it's not the load balancing I grew up with, correct?
Patrick Roughan (01:36.451)
Yeah, that's correct. Yeah.
Lori MacVittie
It's something else.
Patrick Roughan
Yeah. So traditional load balancing, which so many people know it's been around for such a long time, was very much focused on traditional kind of requests, you know, load and it didn't really care about the context of those requests. It was just simple things like least load or round robin or different kind of forms of routing. And this all came out of even networking devices and it's grown ever since then when you're talking about your application load balancers in AWS or whatever that might be.
But the problem with a lot of this now in today's world, especially around large language models, is context matters. Is what is that actual request and how fast--large language models traditionally are not very fast at processing things. They can take time. And if you can make it as efficient as possible and distribute the load as much as possible to try to get as fast as response as you can is very important.
And that's why you know context, as I said, really does matter and being able to understand not just how many requests there is, but what's in the request can actually help do a lot of that.
Joel Moses (02:53.118)
And so what do we call that? Is that like a semantic routing characteristic?
Patrick Roughan (02:57.376)
It's kind of
Joel Moses
What are we doing there?
Patrick Roughan
So there is different types that a lot of different companies are working on. We here in AI security in F5 have our own version that we've built ourselves, but you know Google have one with Gateway API that they're trying things out. And a lot of obviously companies, large language model companies, of the world are all working on this. Is what we're currently doing is one of the biggest beneficiaries of performance on a large language model is KV cache.
Joel Moses (03:30.376)
Mmm.
Patrick Roughan
When you get a request into a large language model, it has to compute that request. It has to go by token by token and it has to, you know, tokenize each part of the prompt and it stores it in a cache and then it uses and refers back to that cache consistently in order to do it. If you're sending similar things, it doesn't have to redo that prompt again and again. So one of the approaches we take is KV cache aware routing.
So when you get a request in, you can look at that request and you can look at the basically the hash of that request and with that look at indexes of where have I sent this cache before. Sorry, this type of prompt--how similar is it to a previous prompt--and you can then with that send it to a GPU that has already done that work. So then you don't have to redo that work.
And with that then, one, you're reducing the load. So if you're having to do this on a traditional load balancer--traditional load balancer Round Robin will go hit GPU by GPU--you're constantly getting a cold cache. You're constantly,
Joel Moses (04:40.307)
Mm-hmm.
Patrick Roughan
it's constantly having to redo that work every single time because you're constantly sending in, it's just constantly going around. Or least load; it's very difficult to understand least load in terms of GPUs because GPUs are very spiky. You know, they can jump to 100% for a second, drop to
Joel Moses
Right.
Patrick Roughan
zero, jump to a hundred percent again. And it's very hard to know well, is it actually busy? A lot of large language models depend on bandwidth, you know, and memory bandwidth is the bottleneck. And as such, again, it's very hard to understand when it comes to load balancing, is it busy, is it not busy? And that's where a lot of these new kind of solutions like KV cache aware routing or other types of systems are coming online to try to understand how do we route to large language models in a more efficient manner.
Lori MacVittie (05:33.307)
Yeah, and Joel, I mean when we were kind of pre-discussing this, like what is this topic about, you mentioned, right, architectural nightmares in
Joel Moses (05:41.868)
Oh, sure.
Lori MacVittie
this because what I'm hearing is that I mean the load balancer is still necessary, but it's not going to do the job of model routing. That's gonna be something
Lori MacVittie (05:50.875)
closer to the KV cache, to the GPUs, to the factories where it has more awareness, as you said, right? Aware of what's going on because the load balancer doesn't have that information. So we're seeing another architectural layer, right?
Joel Moses (06:06.558)
Yeah, you know, this is a really interesting engineering puzzle. In fact, in some cases it's almost an engineering paradox where, you know, in order to analyze something to get it to the most efficient place, some of the best ways or more the most accurate ways to do that are actually far more expensive and time consuming and costly than the model itself. And you can't do that. Like if you're gonna be a layer that does model routing, you cannot be more expensive than what you sit in front of.
You can't be slower than what you sit in front of. And so the, you know, the gut feeling is that you might be able to use something like sampling of prompt context and send that through an LLM of its own. And then you can have a way to route inside that model or to different models. And there's also different elements of model routing like token semantics, KV caching is another one, prompt context, model capabilities. But again, some of these are efficient and some of these are woefully inefficient. And so it's a difficult engineering challenge. Can you describe for us some of the puzzles that you've had to untangle Patrick?
Patrick Roughan (07:21.378)
Yeah, and so when we look at it, so we obviously were very early trying to figure this out. We moved to large language model guardrails nearly over a year ago, a year and a half ago. And it was that problem of once we go past one GPU, how do we do this? And you know, originally we were using Kubernetes services to do it and true Kubernetes services are just a probabilistic distribution that it tries to equal the load
Joel Moses
Yep.
Patrick Roughan
across multiple nodes. And we suddenly found one GPU was doing all the work and the other three GPUs are doing very little. And as such, it wasn't really solving the problem. We tried traditional AWS ALBs to do the same thing, you know, tried different strategies on the ALBs. We talked to Amazon to try to figure out how can we do this. And we came across a concept around KV cache routing
Joel Moses
Yeah.
Patrick Roughan
that we then got into to look at going, okay, how can we, how does this work? And it is a concept of basically, as I think Lori said, is the load balancer suddenly getting the context. Right? The load balancer doing a bit of extra work. So in our case, the load balancer hashes every prompt coming in and on the first part of the prompt and then has a lookup table. Basically it has a lookup table going, I've seen this hash before, I sent it over to this GPU-1. I seen this hash, I sent it to GPU-2.
So when a similar request comes in, it can go, Oh, I've seen this before, I sent it to GPU-2, that's most likely will have a hot cache and as such, I will send it there again in order to process it faster. Now there's a bit more logic because at some stage you will overload a GPU and you still need to maybe have a second GPU that has a cold cache. But now we'll have a warm cache for a certain, for instance, guardrail that we send through, because a lot of our guardrails are now designed to be KV cache compatible.
So we'll front-load a lot of tokens and stuff in the templating and stuff. So we know we'll get a very high
Joel Moses (09:33.841)
Mm-hmm.
Patric Roughan
sorry, a very large cache hit rate. So like at times we can get a cache hit rate of 50-60% on a lot of our prompts. So we know that's 50-60% of the work we do not need to do.
Patrick Roughan (09:44.841)
And we make sure that there is offlets. As I mentioned, is at some stage you will send too much to a GPU, so you do need the model router or the load balancer to understand I actually now have to send it to a different one. I have sent even though this is my match, this is where I've been said I should ideally send it to GPU-1, GPU-1 has too many active requests, I need to now send it to a new one. Might have a cold cache, but now suddenly you have two warm caches.
And it can start distributing the load better. And from there, from the original what I said is we'd have four or five GPUs, one would do all the work, the other three or four would be doing nothing. From very large performance tests, load tests we've run, we now see nearly equal load. It works so well that we see a distribution across every GPU when it comes to the actual performance of the GPUs, that all of them are nearly close, obviously not exact,
but they're quite close to doing all equal work.
So you're having this perfect distribution and load balancing across all the instances.
Lori MacVittie (10:50.298)
And that, one of the other things we're seeing, I mean this is, right, KV cache, right, the request and all of that makes sense because a lot of requests are very similar. There's just a few words that might be different, but they're essentially semantically the same. But one of the things that we're seeing is, and we started to see this early on, is model specialization.
Organizations were using things like we're gonna use Mistral or, you know, some other small model that we've chosen for operational automation. We're going to use, you know, OpenAI or Azure, you know, for our chatbots. Right? So there's some kind of specialization starting to be going on with different models. And model routing can be part of this. If you're taking in a bunch of prompts, you're like, "Oh, this is a request for X or Y," and model routing has to be able to discern that as well, not just look at the the KV cache, right?
Patrick Roughan (11:46.051)
Yeah, so yeah, there's lots of things going on right now around cost, around different capabilities, even within
Lori MacVittie (11:54.318)
Yeah.
Patrick Roughan
the models. Like a lot of the reasoning models, they have a form of routing in them that once you send a request, they're a combination, some of them are a combination of models. In which it'll go, Oh, that's a kind of a maths problem, I'm gonna send it over to this model over here, because it's gonna give me the best result.
Patrick Roughan (12:11.638)
Or that is some form of, you know, physics problem, I have a lovely physics model. And that's built into a lot of these large language models themselves, that they can be a combin-, or services, they can be a combination of models in the background. And they are doing
Lori MacVittie (12:25.636)
Interesting.
Patrick Roughan
already a form of model routing that they know our traditional large language model might not be 100% on maths, but if they have a perfect maths ML model that they can actually send off to, that the large language model is still doing that form of
Patrick Roughan (12:41.218)
verification and routing, as in it's still kind of the one in the middle. But yeah, a lot of even providers are using this kind of form of function routing in a way of I'm going to give it to the best. It's kind of like mixture of experts is the perfect explanation of it. And a lot of the reasoning models are this mixture of experts models in which it decides I have a ton of different expertise in different forms and fashions and I'm gonna give it to the best one that knows for maths or physics or
Joel Moses
Yeah.
Patrick Roughan
chemistry or whatever it might be, that it's gonna give me a more accurate answer than a traditional large language model, which is basically a text model. And it would and this is what reduces things like hallucinations and stuff like that. Because you're gonna get less hallucinations when you actually give it to a particular model that's the expert in that type of field.
Lori MacVittie (13:33.711)
So I'm trying to, I want to make sure I understand what you just said because it's really cool. But what I'm hearing is that so the inference server that like handles all of these models or this app actually has its own internal routing, much in the same way, you know, you would do API-based routing
Joel Moses
Mm-hmm.
Lori MacVittie
inside an application where, you know, there's actually a router function that determines where a particular API, right, gets sent internally. And you're saying this is kind of what they're doing with these different models. Like it's almost like a supermodel that's like, Yeah, I'm gonna, you know, delegate and then pull it back, correct?
Patrick Roughan (14:12.864)
In a way, yeah, they're actually called
Lori MacVittie (14:14.434)
Okay.
Patrick Roughan
mixture of experts models.
Lori MacVittie
Okay. Mm-hmm.
Joel Moses
Mm-hmm.
Patrick Roughan
So there's lots of them out there. And a lot of models are mixture of experts models at this stage. And that's what it is, is they determined that you know large language models were really good at generalization. But when it comes to specificity of a certain topic, you may need some other model that is particularly trained for that kind of conversation, if it's maths or whatever.
Patrick Roughan (14:40.684)
And you kind of saw that from a year or two years ago when large language models used to be terrible at doing your accounts.
Lori MacVittie
Ha ha ha, yeah.
Joel Moses (14:46.976)
Yeah.
Patrick Roughan
If you try to get a "Hey, well, how much tax would I'd pay?" And they give you a number, you're like, "Oh, my God. That's terrible."
Lori MacVittie (14:53.71)
Ha ha ha.
Patrick Roughan
And then you'd start looking at the numbers going, "Wait, how much am I getting paid? That doesn't make any sense."
Joel Moses (14:58.228)
Yeah.
Lori MacVittie
Ha ha ha.
Patrick Roughan
And in the last year or so, you suddenly realize, no, they're really good at maths. How have they suddenly become really good at maths? And some of this is this whole mixture of experts topic, is which
Patrick Roughan (15:10.156)
you know, trying to not be such a generalist in everything, but have specific things that you know, and then route to those specific things in a way. But that's kind of within models itself. That's it's not
Joel Moses (15:21.91)
Yeah.
Patrick Roughan
something external to the model. It's not something
Lori MacVittie (15:24.451)
Okay. Mm-hmm.
Patrick Roughan
that you kind of break out. It's kind of in the model architecture that they build these days.
Joel Moses (15:30.272)
So Patrick, let's talk about something that might overlap the areas of traditional load balancing and model routing. What role does model routing in particular have to play in the area of sovereign AI? That's kind of become a
Lori MacVittie
Oooo.
Joel Moses
hot topic as of late.
Patrick Roughan (15:46.191)
Yeah. So sovereign AI is actually a very kind of important topic. And we from an early day from kind of the AI security side of it, we saw this as something that was very important, you know, with large customers, not just governments but banks, defense contractors, they do not trust anybody in a way. And as such, they want this solution self-hosted. And when I mean self-hosted, this can be bare metal, their own sovereign cloud, it could be hyper cloud somewhere but in that jurisdiction.
And as such, we our solution is fully built. We give our own SaaS solution that people can use, but we also get all the functionality that feature by feature you can install entirely in your own sovereign cloud and do what you need to do with security, but with your own data. You can pump through it and fully know that it is, is all the data resides where it lives, in your where you want it to be. And it isn't a burden.
And that those feature sets of so all of that model routing and load balancing and all that smartness we do around AI security, we package all of that and we distribute it out. That customers can download it, it installs, it runs wherever they want it to run, and it can do everything we do there, we do in our SaaS, so it's like for like.
Joel Moses (17:12.032)
So model
Lori MacVittie (17:12.366)
That's yeah.
Joel Moses
routing can also be used to ensure the locality remains solid
Patrick Roughan (17:17.794)
Yeah.
Joel Moses
for sovereign AI.
Patrick Roughan
So for sovereign AI, yeah.
Lori MacVittie (17:20.506)
Nice.
Patrick Roughan
It's part of these types of things around as well as, yeah, it's where you want your data. And part of Guardrails is you can do items like that. You know we secure for things like the EU AI Act and stuff like that to verify things are not. And with that customers can kind of bundle using Guardrails to make decisions on where data is sent. So you can kind of bundle
Patrick Roughan (17:47.171)
the model routers which are saying its jurisdictions or where you want to send the data, you can bundle that with types of things like guardrails that if the model router is not sure, like "I'm not sure this context, the content in this data, should this be distributed outside the EU, I'm not sure. Send it to guardrails." Guardrails comes back, go, "No, actually this violates XY or
Joel Moses
Interesting.
Lori MacVittie
Mmm.
Patrick Roughan
whatever guardrails, do not send that all forward." And some of those types of things are very important for customers to be able to make decisions dynamically on where data is going or what data should be sent to a large language model. It's like an OpenAI, is it a world or an Anthropic, or actually it shouldn't, it should be sent to their own local model that they have running in their own area. So some of the these things can be very important when it comes to trying to make decisions on. And again it comes back to that content, you know, not what the requests are, but what's in the request, you know, what the data is and so on.
Joel Moses
Right.
Lori MacVittie (18:50.552)
Yeah, that's a change too from traditional, you know, load balancing where generally we would collect a bunch of statistics and then sample it and like make decisions about what's fastest, what's available. And what you're saying is this is actually a coordination between multiple systems going, "Is this okay? Is it not?" It's architectural. It's not a here's a box and it works. You know, it's very different.
Patrick Roughan (19:14.606)
Well, one of the simplest examples to kind of understand is like one of our guardrails is PII. It looks for personal identifiable information and we do it from multiple different types of languages and countries. And one of those things can be that is a company is using a SaaS LLM, whatever your your favorite is, and they have some form of model router that's sending this data out or sending it to different ones, but they have Guardrails connected.
And they're sending it to guardrails before they send it out. And in that, now suddenly you've uploaded accidentally, somebody in accounts, I'm not gonna blame accounts, has asked about Joe's tax returns for the year and has put in personal information. You do not want that going out to a SaaS provider who now has it, is training on it, understands that context, and now suddenly your data is no longer in your hands.
And that's where things like guardrails and stuff can come in. That model router sends it to Guardrails, Guardrails comes back going, "Wait, there's PII in that. That is not allowed.
Joel Moses
Yeah.
Patrick Roughan
Do not forward that request." And aspects of that when it, it's part of the load balancing as well in a way is understanding what traffic not to send on.
Lori MacVittie (20:30.648)
Yeah,
Joel Moses (20:31.359)
Yeah.
Lori MacVittie
that is a good point. I would love to keep deep diving 'cause this is my like favorite topic, but
Joel Moses (20:38.325)
Yeah.
Lori MacVittie
we are getting, you know, to the end. So I kinda wanted to, you know, especially ask Joel, like, what are the takeaways here? What should an enterprise be thinking about doing? You know, what does this matter to them?
Joel Moses (20:49.536)
Yeah. Well I mean it's very clear you don't want to ask a traditional load balancer to do semantic routing. It's gonna have a core dump trying to network address "Translate a metaphor" if you do.
Lori MacVittie
Ha ha Ha.
Joel Moses
So yeah, but there are definitely ways that you can layer these technologies to work well together. I mean traditional load balancer is not an unknown quantity, nor is it a valueless piece of equipment. But the combination of the two things, traditional load balancer approach for large scaling and then much more precise controls in model routing, especially as it relates to management of KV cache and making sure that you are efficiently optimizing that path. That's important.
You know, I kind of, it puts me in the mindset of, you know, a traditional load balancer is a little bit like a massive postal office where you have mailbags and you're trying to sort them and put them in trucks that are least loaded. But a model routing function is more like a triage unit in a hospital, where you're trying to optimize a patient's treatment by sending them to the doctor that can perform the procedure that they need most effectively, and you're trying to isolate minor scratches from, you know, much more complex operations requiring, you know, a lot more treatment.
So yeah, I mean it strikes me that the AI race won't be the people who pursue the biggest, most complex models. They're gonna be the ones that have the smartest triage units in their hospital.
Lori MacVittie (22:23.844)
That was, that was an interesting mix of metaphors there. We went from traffic
Joel Moses (22:28.831)
Yep.
Lori MacVittie
to triage, but they both start with T, so alliteration for the win. That's what I'm going with. Patrick, what would you tell enterprises about model routing, load balancing? Like how do they need to think about it? What should they be, you know, planning for, doing, that kind of thing?
Patrick Roughan (22:47.0)
Biggest thing is trying to do it smarter. Everybody and it doesn't matter if you're the Googles of the world to the small startup of the world, GPUs are a constraint. Everybody's struggling to get them. And as such, it doesn't matter how much you want to grow, you are stuck with this constraint. And no matter what, you can't just keep on throwing more and more GPUs at the problem. You need to think smarter. You need to figure out how do I optimize for this? How do I reduce my cost? And how can I be more efficient with what I have and what I can get?
Lori MacVittie (23:24.506)
Excellent. Now it sounds like the old, right, saying, you know, you can't just keep throwing more hardware at the problem. You can't just keep throwing more GPUs. And I'm sure we'll hear that again in the future with technology Y, whatever that might be.
Well, that is all we have time for. So that's a wrap for this episode of Pop Goes the Stack. Please subscribe because if your router sends the request to the wrong model, you don't get AI magic, you get incorrect answers at premium speed.