RRE POV

In this episode of RRE POV, Raju Rishi and Will Porteous unpack two major forces shaping the future of AI: model distillation and the debate between open-source and proprietary models. They break down what distillation means, why it matters for model defensibility, and how open systems have historically reshaped major technology markets. From Linux and Android to the open internet, the conversation explores what past platform shifts can teach us about where AI is headed, and why founders and enterprises should focus less on model lock-in and more on durable workflows, data, and customer value.

Highlights:
(02:04) Explaining AI Distillation 
(03:10) How Distillation Copies AI Models
(06:47) LLMs Trained on Everyone Else's Data First
(08:26) The Sunk Cost Problem in AI
(14:30) 85% of AI Queries Are Repeats
(19:47) Open vs. Proprietary Models Explained
(21:28) What Windows vs. Linux Teaches Us About AI
(26:33) AOL's Walled Garden Warning
(34:34) Lock-In on Apps, Not AI Models

What is RRE POV?

Demystifying the conversations we're already having here at RRE and with our portfolio companies. In each episode, your hosts, Will Porteous and Raju Rishi, will dive deeply into topics that are shaping the future, from satellite technology to digital health to venture investing and much more.

Raju: Hello, listeners and viewers. Welcome to another episode of RRE POV. I’m Raju Rishi, and I’m joined by my partner, Will Porteous. And today we’re going to dive into a really important set of topics in the world of AI. Two of the most important current topics, and I think—and my partner, Will, also believes—they are going to define how the shape of AI pans out. The two that we’re going to discuss, we’re going to start with a topic called distillation, and then we’re going to end in a topic which is highly relevant to distillation, which is open-sourcing versus proprietary.

Will: I would just advise our audience that we feel like we’re at a moment in the AI landscape where these questions are—the answers are perhaps murky, but if you look to history, we may find some guidance for what happens next.

Raju: Yeah, yeah. And you know, just for our listeners’ purposes, you know, some folks on our calls and our video, they’re technical. They understand distillation, open-source versus proprietary. The ones that don’t, we’re going to explain it. We’re going to get very, very simplistic about what these things are because they are so important that these topics be understood, and I think it’s going to ultimately determine which models are going to succeed and which ones are going to fail. Okay?

So, humor us, give us a few minutes to sort of go through the topics and give you guys some level of grounding, and then we’re going to get into some topics, and we’re going to create parallels from the past that you guys will absolutely understand. Okay, so we’re going to start with distillation. And candidly, I’m a little annoyed that the word distillation has made its way into the center of AI, you know because I wish we could isolate it strictly for chemistry and alcohol production.

Will: [laugh]. Yeah, is that a still, actually, in the back behind you, Raju, that curly looking thing?

Raju: Yeah, no that’s—

Will: Are you distilling something?

Raju: Oh, you know what, that is a bottle of tequila, and I bought it because it did remind me of a little bit of a distillation tool. You know, look, whiskey, vodka, rum, gin, tequila, brandy, all distilled. We actually do need to do another podcast one day on the OG use of distillation, so we can get some taste testing. I think it’ll be incredibly—

Will: It involves firing some hickory in the backwoods somewhere, and—

Raju: Oh, dude. Moonshine, you know, like—that’s an open-source version of distillation, right there. Open-source.

Will: Very open source.

Raju: Macallan? Pretty proprietary [laugh].

Will: There you go. Viewers, listeners, you have the essence of what we’re going to talk about.

Raju: I just boiled it down. Okay… anyway, we have distilled LLMs now that do not taste nearly as good as the other things. But for those that don’t know, distillation is a method that’s commonly used by AI labs to train a large language model using the outputs of another, okay? So, you basically take a series of prompts and you submit them to the larger teacher model, and its responses are used to train the smaller student model. And the things to understand about this is knowledge is not actually transferred directly, but the output is mimicked.

So, you don’t actually have a model that has all of the data inside of it, but it really learns well, and it can do things really effectively. Now, this is highly acceptable when companies use it for their own models, and that’s called, like, authorized distillation. So, what does it do? Like, if you have a large language model and you want to create smaller ones for specific tasks, well, you don’t want to retrain the new model, the smaller model; you’ll just compress it, and you’ll focus on questions for specific tasks. And so, the small language model, or medium-sized language model, performs faster, costs less to develop and operate, and widely accepted.

So, if I’m Anthropic or OpenAI and I want to train a smaller model focused on particular things, distillation is an absolutely useful tool. The large language models that are proprietary basically say that unauthorized distribution puts their businesses at risk. And that’s when a company comes in, basically queries the crap out of the model, and builds their large language model upon it. And it’s widely suspected that this is what DeepSeek did to train their model for very inexpensive levels. So, I’d love to hear your thoughts on this, Will. You know, should the world kind of be, sort of like, really, like, [charred 00:05:12] over this, or is this, kind of, bound to happen?

Will: Well, I think the question is, is it a means to an end, and is the end accurate? Because to me, the logic of distillation is such that you’re going to end up with a trained model that doesn’t really have the internal logic or familiarity with the core data set to reliably get to accurate correct answers as variables change over time. So, you know, I kind of land on the thought that maybe this is a moment in time thing as we try to get more accurate results or more accurate results inside of the boundaries of an enterprise that need to rely on these junior trained models because you don’t want to expose proprietary data of the enterprise to the larger model. So, perhaps this is a moment in time. As you know, I’m fiendishly optimistic right now that we’re going to enter a new era of AI model efficiency, we’ll come to that a little later on, I’m sure, but this feels like a little bit of a bridge to that era.

Raju: Let’s imagine that there wasn’t any anomalous results, you got exactly the same output, and you could just basically leverage a big model to create sort of a new one. I got, like, a couple questions. Didn’t the LLMs initially train on OPD: Other People’s Data?

Will: Yeah, right. [laugh] sure.

Raju: And so, like, you know, like, you scour the internet, you basically use everything that’s out there, you look at reading, you look at writing, you look at graphics, you look at how artists create things. You know, and so to some extent they’ve built their whole business around ingestion of large quantities of data, and now they’re saying, “Hey, well, now you can’t take mine because we’ve spent billions of dollars on it.” So, there is a little bit of hypocrisy there. So, let’s just say, okay, yeah, we got to cut it out, right? We got to basically—like, it’s a bad thing, and we shouldn’t allow it. Can you really stop it?

Will: Oh, I don’t think you can. I mean, you know, a copy of a copy can beget a million, a billion, an infinite number of copies. So, in this context I don’t think it’s stoppable from this moment or for any recent moment because this pattern exists now.

Raju: So, this is—you know, we, I talk about this, you, I know we talk about it, but you got these trillion-dollar valuation companies that are built off of, like, hey we’ve spent all this money and time training. If it’s so easy to, sort of, distill, and you can’t kind of copy it, you know, what—you can’t really prevent the copying of it, right? You can’t—and so you know, do they lose value instantaneously? And so, are they worth the trillion-dollar valuation? Can they sustain—this is a very open question, and I’m not trying to imply anything. I’m not trying to apply that they’re not—

Will: No, but this does pick up on one of the big ideas I’ve heard you discuss, and our listeners and viewers have heard you discuss in the past, which is the overall sunk cost problem in AI, which is, you know, once you can make a copy that will produce as good a result as the master model that was trained on billions and billions, you know, what’s to protect the value that’s been invested?

Raju: So, open question. We’re not going to answer the question, but it is something for people to, sort of, think about and say, you know, a lot of money has been spent training the models; can you effectively lock it down? I think there are areas where you should lock it down, Will. Like, there are guardrails that are built into all of these LLMs that—I mean, the ones in the US, at least—that are, like, hey, you can’t ask it how to build a weapon, you know, for instance. Or you can’t ask it to do anything, you know, super nefarious.

You could eliminate those guardrails by copying it, you know, which is dangerous. So, I do think that we have to have some level of control over this, but I’m not sure how much we’re going to actually have. Distillation is a real thing. People are using it inside their own companies, the language models, and people are using it inside their corporations so they don’t have to spend so many tokens to retrain something. It makes it cheaper, and I can get a model that’s a little bit more—and then it can take a life of its own, but you know, it has all of the guts of the prior model very cheaply done.

So, I don’t know… I’m going to go—I’m going to take us back a little while, and I’m going to ask you some questions, and you’re going to know all of these things. Isn’t technological advancement always built on prior work?

Will: Oh, indeed. Yeah [laugh].

Raju: Isn’t it always? And so, you know the question is, like, you know, should we stop it? Yes, I agree, we should stop it for certain things, and yes, people have spent billions of dollars, but you know the question I’ll ask you is that, you know, should these LLMs actually be private companies or should they be government run, right, as a service?

Will: Yeah, or should they essentially become a public good, which they may indeed become over time. I mean, you’re picking up on one of my favorite topics, which is sort of the idea that anything is truly proprietary over time. I mean, you know, innovation is built upon layers [laugh].

Raju: Exactly.

Will: And other people laid down [crosstalk 00:10:53].

Raju: I’m going to walk us back a little bit, just, we’re going to play—like, I love doing this because we’re both relatively the same age, we’ve experienced this, but like, think about GUI. I mean, Xerox actually invented the Graphical User Interface, and you know, who leveraged it? Apple, Microsoft, and you know everything out there. The first smartwatch, do you remember? Ish? I don’t know if it was the first, first, first, but do you remember who [unintelligible 00:11:19]—

Will: Yeah… I can’t put my finger on this one.

Raju: Casio Infraceptor.

Will: Nice [laugh].

Raju: [laugh]. And obviously, you know, Apple and Samsung.

Will: Yeah.

Raju: This one you’re going to love. Short-form video? Do you remember our portfolio company, Vine?

Will: Of course.

Raju: Vine was basically the OG TikTok, and the OG Reel. And it was Vine, it was Periscope, it was Meerkat, which were like, you know, hey, like, create a short video, UCG, you know, user-generated, UGC, and now we have TikTok and Reels. And Reels basically copied TikTok to some degree. Even with these big ones, Social Network, Friendster, Myspace, now we have Facebook and LinkedIn, and all of these other ones—and Snap—there wasn’t any kind of, sort of like, resistance pattern, right? Like, you could technically take some of the essence of these tech technologies and build another one. So, how do you create defensibility?

Will: [laugh]. Standards or market share—

Raju: Yeah.

Will: —as I think Bill Gates famously said. You could take our audience into a history lesson on networks and networking technologies and network protocols because we both remember an era of proprietary network protocols, a handful of which still exist out there. And yet our internet today and all the value that’s been created, really is based around open protocols. And you need open technologies that also proliferate to get network effects that create a lot of value over time. And, you have, you know, any number of examples, I know, even in your personal past, of these things that went proprietary—were proprietary at one time, that people used to pay for, that got eclipsed by open innovation.

Raju: Yeah. I mean, this is so early innings. I love how we’re coining trillion-dollar winners of an industry that is so on the… very early innings. You know, think about this, right? Like, think about the fact that you can basically, like, use diffusion to copy a language model that people have spent billions of dollars building.

And it might be okay to do it, right? Like, I mean, like, definitely don’t do it without guardrails of, like, you know, social ill will or whatever, but there’s ways, right? Like, if you look at some of these things that I talked about, like, TikTok and Vine and everything, defensibility was kind of owning user behavior. If you could own user behavior, you—and so, like, where would I be pouring money into an LLM is definitely trying to create user behavior and get people locked in. You know, owning user data, right, like, prior queries. And you know your history bar is an area where—and getting to know you as an individual and sort of ingraining that but not necessarily the route [answers 00:14:30] and by the way this is a crazy stat. 85% of agentic queries are repeats. 85%, Will. If you think about that in the scheme of how many tokens are being used today—

Will: [noises of incredulity].

Raju: —imagine if you could cache the answers. Imagine if you could cache the answers. You would decrease token utilization overnight by two orders of magnitude, and that is not something that can be done if you create proprietary locked in one to one marriages, but if it’s an open-source thing, you could Akamai a query.

Will: The core principle behind caching is that you already know what people want, you already know what the world is going to demand, and so if you can stage the answers out there just the way we used to stage web pages, and still do, frankly, you already know what’s on people’s mind. You already know what they’re wrestling with.

Raju: I will say, I’m going to just, like, kind of reiterate—and the reason I’m going to reiterate the defensibility models is because I want people to understand, like, the LLM that does this is going to be winners, but also our companies, our application companies, if you can own user behavior, if you can own user data and sort of intelligence of, like, who they are and what they like to do, and you can integrate into their workflow and other data sets, you create lock-in.

Will: Because you have relied on the—you’ve learned to rely on the open models as a public utility, as a public utility that you only need so much of and it’s a consumable.

Raju: Yeah, I don’t know. I think the governments are trying to, you know, restrict this. That’s what patents were for, right? Like, I mean, if you think about the patent infrastructure, like, we created patents so that somebody couldn’t spend all this money doing something, but I think intelligence is hard to patent. Do you know what I mean? And so, it’s like—

Will: Well, intelligence that’s, by its very nature, derivative.

Raju: Exactly.

Will: If you look at the open models, at the large language models that have essentially devoured all of the public data, it’s essentially, these are interpreted results, these are derivative by their very nature from other intellectual property. And so, you know the fact that we can draw a line from what they consume to the insights that they provide, and they provide those insights with a great deal more utility, and by the way, as you pointed out the questions being asked are the same over and over again, it kind of begins to feel like we’re at a moment in time.

Raju: Yeah. I mean, just, like, reflect on this for a second. 85% of agenda queries are repeats. We don’t have enough data centers. We don’t have data centers, and the reason we don’t have enough data centers is because of all these queries. 85% of them are repeats. You wouldn’t need that many data centers if you would cache this stuff. I mean, just the ramifications are pretty profound.

Will: Yep.

Raju: So, I think, like, you know, understanding this distillation piece is going to be an important piece, but what I really want to move toward is open-source versus closed, open-source versus proprietary. And I’d love to move to that portion of this discussion because it is—distillation is one piece of that, right? Distillation is a mechanism to effectively open-source a model, even if it doesn’t want to be open-sourced. But again the largest proprietary models that we see—and this is going to crack everybody up—the first one is OpenAI. It starts with ‘open.’ It’s proprietary. It is proprietary. It’s, like, crazy that it’s called OpenAI, but the model’s proprietary.

Google Gemini? Proprietary. Anthropic Claude? Proprietary. Microsoft Copilot? Proprietary. The largest open ones actually fall into two categories, right? One is ones with a large number of parameters, so like the Trillion-Parameter Club, which is Kimi, DeepSeek, and Mistral. And then there’s Dense-Model Club, and Llama is Dense Model, and Nvidia has an open-source one called Nemotron. And it’s interesting and not surprising that the model that Nvidia is creating is open-source because they’re arms merchants to the world, right?

They say, every model is going to need processing power. We make the chips, so like, why would we close things off? We’d like everybody to have their own model. And Meta did, like, a basically 180 a long time ago, and said we’re going to create an open model. So, they’re the open ones, and the ones that you know are trillion-dollar going public soon, they’re closed. And so, let’s talk about the core issues associated with open versus closed. Like, what is better and why is it better?

So, the first is, like, code weight and access. So, open LLMs allow public people to download and inspect. You can download and inspect everything. It’s sort of like a glass box. You know how the model makes decisions, you know why it makes decisions, and there’s a lot of value in knowing that, right? And the propriety ones, everything’s hidden. You just get an answer, and it doesn’t tell you how it got there or what data it uses.

A second core, sort of, difference is data privacy and security. Turns out the open LLMs can be hosted on-prem, okay? So, that means if you have proprietary data, you’re going to use an open LLM because you can host it on-prem, you download it, you leverage it for your purposes, your data never leaves the firewall. Whereas you know, the proprietary ones, it has to go to their servers, and their servers are elsewhere. And so, there’s regulatory risk, you know, for things like healthcare and banking.

Can they actually use, you know, proprietary models where the data leaves? And you know, I know that they’re working on proprietary models that are relegated to your data, and your data is isolated to them, but it’s not quite the same as a proprietary model. And the last piece is customization. Proprietary offers limited customization; open, infinite customization. So, you get, if you really just want out-of-the-box, you know, you might go proprietary. If you really have regulatory concerns, your data can’t leave the firewall, you might have to use open-source. If you want to know exactly how the decision is made, you might need open-source.

And it’s interesting to see the ramifications of this because I want to go back to history again one more time and just play this out in terms of, like, the models that existed in other sectors and which ones won. So, Windows versus Linux. You remember that—

Will: Sure.

Raju: —battle? And so, when you look at that one, Will, like, Windows won the consumer, but Linux won the back-end.

Will: Oh gosh, yeah. Yeah, and Red Hat won the Linux game, building a great business on top of open-source in a way that, at the time, seemed sort of a little bit screwy, and then everyone realized that they needed to pay for service and support and someone to help them navigate the open-source journey. So, it’s kind of staggering. Today, you also have to look at all the capital that was invested in operating systems at the then time. And this is a direct analogy, I think, to the moment that we’re in.

Raju: Absolutely. Absolutely, please go forward with this because this is a direct analogy.

Will: The operating system wars were the AI model wars of their day, and you had enormous amounts of capital being poured into the development of operating systems, and yet you had open alternatives that were scaled in terms of performance and security, and lots of other layered features, like Linux, a contribution of the community. But they needed—they were good at certain things because they were open-source, but they were, by definition, also not productized in some cases. And so, this tension of how the open models become products that enterprises can reliably trust is perhaps one of the profound opportunities that we should probably talk about.

Also that notion I remember the dialog we were having about operating system lock-in and vendor lock-in and how that would lead to architectural constraints that our engineers hated at the time because we would be forced into that vendor’s development environment, we would be forced into their database, which we hated. Because databases were there, another venue where the open-source alternatives provided far more freedom and customization and better performance and security, and all these things. So, we may live it again, here.

Raju: We may live it again. We may live again. Okay, another one: iOS versus Android, right? IOS, look, I mean, there’s something to be said that the phone works, it works, it works, it works. Kind of relegated to the Apple ecosystem, you know? You can’t, like, get all the stuff that you know if you want it on a larger screen, you got to buy their iPad. You can’t use, like, a different iPad. And you know, I think, like, iOS kind of won in the US. Like, people get pissed off if you have a green, like, in their text messages.

Will: Sure.

Raju: I mean, they get really annoyed. You know, I’ll tell you my funny story. My son was the only one in our household who had an Android device. He wanted the customization. And you know, his girlfriend and him broke up, and he was starting to date again, and he’s like, I need an iPhone [laugh] because of dating.

Will: The reason they won is ultimately because of network effects, whether it’s everybody being on iMessage and being able to do the same things, or being in the app ecosystem, or having better access. And I’m not saying that Android didn’t catch up. Android provided—it’s interesting because Android provided us with an open hardware architecture parallel that we could easily relate to from server wars and other wars.

Raju: Absolutely.

Will: It was open-ish [laugh] because of Google, and the way that they ultimately controlled it.

Raju: Yeah. You’re right, it is open-ish. They’re both kind of a little bit proprietary. But I will tell you, Android is 70% of the international market.

Will: Yeah.

Raju: 70%. So, it’s hard to say iOS won. They won in the US, for sure, but, like, 70% of a global market? I mean, geez, that’s huge. So—

Will: Well, you know what? Free is popular [laugh]. It’s always been popular, and it still feels like free if it’s ad supported.

Raju: Yeah. And I will tell you, that is absolutely why the LLM wars are not over. Free is popular. Free is popular. Free has always been popular, and so you know, data costs are not zero, token costs are high and going to get higher, and data standards are not—we need more of them. We’d probably need a whole lot less if we cached the queries, as I mentioned, but, like, okay, so this is going to be an interesting debate. Okay, AOL versus the internet.

Will: [laugh]. Right. I mean, it uh… was a wonderful world inside that walled garden for a period of time. And in hindsight, it looks like such a wonderfully and amazingly simple business model of, you just charge people to get in, and you determine what happens inside that world. And yet the vast abundance of the open internet would outstrip it in no time at all, in terms of creativity and opportunity. All of it available through open standards-based platforms, beginning with the web browsers.

Raju: Exactly.

Will: And even the proprietary web browsers had to be free, from the [future 00:27:18].

Raju: Exactly.

Will: That to me is, sort of, staggering. I mean, I suppose we paid for Internet Explorer at some point as part of a bundle, but in truth, they were all free.

Raju: It’s interesting. It’s interesting because what ha—AOL was great. Walled garden, you got the great, you know, experience, the iOS, you know, experience, you know, for data, and you can find things and store things, but they couldn’t get all the information inside. I mean, so much terabytes and terabytes of information is not in AOL, so like, yeah, there’s no way you can win.

Will: Well, and also it just, it captures the creative, kind of, paradox of being in the media and entertainment business, which is, you only have so many resources, and if everything that’s going to go into the walled garden is something that you have to pay someone to create, you’re never going to catch up with the abundance that’s out there. And that holds true in the open media world of today, which is ultimately what’s powering today’s large language models.

Raju: Yeah, exactly. All right, so this is the last one that I’ll do a history lesson on, but I’m going to bring us back to the parallel because I think this is—as funny as this may be, we talk a lot about AI, and you know, we always paint a vision of, like, look, it’s early innings, and so you can’t get lock-in. Lock-in you know, people who think there’s lock-in. And we’ve explained why from a variety of—and we explained, like, how to play the game by getting into some applications that are more nuanced, that have, like, you know, workflow integration, and they’re very sticky, but this is the first podcast where we’re getting into, sort of, the deepest—well, two of the deeper issues that LLMs are wrestling with, and we’re coming to the same questions. We’re coming to the same questions.

But here’s a really important analogy, Oracle as a database versus Postgres SQL and MySQL, which were open, and Oracle was closed. If you look—and I’m not going to wait for you to answer because I’m going to tell you the answer—like, Oracle won in the corporation, right, but every modern architecture that is being built today is being built on open SQL servers.

Will: Yeah.

Raju: Postgres and MySQL. Think about that. The modern architectures. And so, you know, yeah, it won because it had needs that more met for the enterprise, the proprietary nature of that whole—but there’s a massive shift away from it. And you know, internet is definitely, you know, not AOL, which was a proprietary version, and Android is 70%, and Linux has won the back-end, which is core.

And I think that the key thing to, you know, sort of remind ourselves is, as this is playing out, you know, this LLM, we’re going to take advantage of it all. But you said something very important, like, free is popular. And people are willing to give stuff away for free, you know? And I think that, you know, the ecosystem that’s building around proprietary without some versioning of open, you know, may miss out.

Will: That I think is one likely conclusion. I also think you have to think through the motives of the sponsors of the open models, though, too. I mean, you know, training a model costs a lot of money, right? You’ve got to buy GPUs, you’ve got to have hardware, you’ve got to have global routing capability, like, having a great open model isn’t just a matter of motivating a collective of engineers, you’ve got to have somebody to underwrite it. And so, there’s an end game that has to be played for these open models, just the way Google played the end game with Android and the ad business. I mean, perhaps ultimately that’s where these open model sponsors have to end up, or we need global public infrastructure that can support it.

Raju: Yeah. And I don’t know if we’re ready for global public, but you know, interestingly enough, OpenAI was built under the cornerstone of, like, hey, this is a service.

Will: Right.

Raju: It’s not going to be a corporation. It’s going to be something that’s more of a, you know, public service and a public good. And we are quickly rushing to build a bunch of proprietary models that the world doesn’t have enough data centers to support, and it’s going to take a long time before we have enough capability to support them. But more importantly, you know, like, we could short circuit some of this stuff. This is really going to be a balancing act of, like, you know, like, if you look at Claude, and why it’s deployed so heavily, they bet on coders, right? They said we’re going to make coding incredibly effective. We’re going to try to get to, you know, the singularity in coding, if you will, where we have, you know, super intelligence.

Turns out 80% of the coders are in businesses, and so they got this drag-along effect into businesses. But the sticking point—the sticking point—is a lot of the businesses like healthcare and finance cannot—they need to have—they have regulatory frameworks where the data needs to—can never seep. It can’t seep, and so, do they push the boundaries and go open-source for those kind of businesses? And how is Anthropic going to respond?

Will: So, to harken back to one of our prior episodes, I think data governance rules the day on that decision, ultimately, and I think that we will surface in the next few quarters, some horror stories about data leakage that are already starting to come out around public company information and that sort of thing that are going to force big companies to tighten their data governance policies, and that’s going to be deterministic in terms of this question you’re posing, and in a way that’s favorable to the open models, and to those who build businesses that support them.

Raju: Exactly. And I think the one person not to discount here is Nvidia, right? Like, the fact that they are building a dense open model, they are open-sourcing it. Like, Nematron is open.

Will: Yeah.

Raju: Yes, they got super wealthy off the back of, like, we need more data centers, and you know, they also got wealthy off the back of, we need more crypto because they were mining crypto, and they got super wealthy off of the gaming, you know? So, they had three inflection points upon which they built the backs, you know, of their company. But you know, like, the fact that they are open-sourcing some of this stuff, it will be absolutely a competitor, like, an option, a different option for how things get done. So, you know, I just kind of want to play this back for folks, so they understand, like, you know, distillation is a tool where you can literally copy an LLM. Not an exact replica, but you can mimic it almost to perfection.

You know, it is a tool that exists. There’s a bunch of Chinese companies that we know are using it. I know the big LLMs that are proprietary are, you know, complaining, and they’re, you know, sort of like foreign policy issues that are taking place, but that is a reality. And the debate is, like, how locked down it should be, and you know how to play that out.

And the second is, like, should they be proprietary? Like, is there a world where open-source wins, or is there a world where proprietary wins, or do both sit alive? And given these couple of open questions in AI, if you’re an end-user or one of our listeners, you don’t sit there and want to lock in on models. You don’t want to lock in on LLMs because there’s a lot of shakeout that’s going to happen. What you can lock in on is applications and tools that have value, and it kind of supports the same supposition that we had in a different way.

Will: I think that was a great recap. For our listeners and viewers, I also want you to hear Raju asking the question about lock-in and the implied question about how you create defensibility in the businesses of the large language models over time, and whether their current market position, and all the billions that have been invested in some ways erodes in defensibility over time, in terms of ability to extract revenue that somehow erodes towards a public good or that is eclipsed by the open models over time.

Raju: Thank you, Will. And given that it’s late in the day, I might actually do a little bit more distillation research.

Will: [laugh]. We ran just long enough today for the distillation run to happen in the background, and I believe there’s a sample there for you to—

Raju: [laugh].

Will: Partake in. So, thanks as always for a great conversation.

Raju: Yeah, likewise, Will. I love talking to you because we both have history.

Will: Indeed.

Raju: Thank you.