From frontier labs and enterprise platforms to emerging startups reshaping entire industries, The Deep View: Conversations podcast interviews the brightest minds and the most influential leaders in AI.
Jeff Morgan: This is the stuff that's powering data centers that's coming onto your laptop. And what's so impressive is, these are models that are optimized for high throughput. You're able to get thousands of tokens of seconds processed and processing. So you're getting data center speeds to all that happening on this three to $5,000 machine. And what's great about that is that it's a fixed cost and it's private. And so you don't have to go figure out, hey, where's the compute? Is it in a secure location? Is it in the right country? It's right there with you.
Jason Hiner: In this episode, I talked to Jeff Morgan, CEO and co-founder of Ollama. The software that enables companies and developers to run open models instead of renting their AI from the big frontier labs. Because open models are having such a huge moment right now, especially in the US, I wanted to have Jeff on the podcast to talk about the momentum Ollama is seen around new startups in the US building open models, enterprises adopting open models, and hardware makers like Apple and NVIDIA and AMD, making hardware that developers and AI enthusiasts can use to run AI faster and save a ton of money. We talked about how Ollama is seen businesses and developers put AI to use since over 80% of the Fortune 500 uses the Ollama software today. We talked about why companies would want to customize open source models or even build their own models as more and more of them look to own the intelligence layer in their company. Today, only about 20% of AI tokens run through open models according to the estimation of Jeff and his team. But this episode talks about why that's likely to dramatically flip in the years to come. All right, so here it is, our conversation with Jeff Morgan of Ollama. All right, so Jeff, you and I have talked before and when we talked last time, there's just, because there's so much happening in AI right now around many of these topics, enterprises, developers, open models, I said, I really want to have you come on the podcast because we want to talk more about this stuff. We want other people to, I really want other people to get the advantage of kind of your insights, perspectives, wisdom. So, but for those who aren't familiar, tell us a little bit about, tell the audience a little bit about what Ollama does in your role at the company.
Jeff Morgan: Yeah, well, first off, thanks for having me, Jason. It's great to be on the show. Ollama is the easiest way to get up and running with open models. And so, whether that's models that you can run as a developer on your existing hardware or laptop, all the way to the largest open models in the cloud. So, think the, you know, GLM, DeepSeek, Nemotron, Ultra from NVIDIA. Ollama just makes it easy and specifically what it does is, you know, as you know, there's been this huge surge of usage and tokens from coding agents. And our goal is really to make sure that open models can power this next generation of software. That's really driving, you know, the AI industry. And that, you know, you can take an open model that really starts with weights on the internet, but actually make it plug into your, your favorite coding agents and be very productive with that while staying safe and secure. And so, that's kind of been how, you know, what we're thinking about as, you know, what's the next step for the company and the project? It's really around powering these coding agents.
Jason Hiner: Very good. And your role with the company too?
Jeff Morgan: Yeah, I'm this co-founder and CEO. And I've been here since the very beginning.
Jason Hiner: Yeah, very good. And then you and your team too, or some of your co-founders came from this startup that you did before Ollama, which has some connection to that you sold to Docker. Tell us a little bit about that. What were you doing before Ollama?
Jeff Morgan: So I started this company in Ollama with the same co-founder as my previous company, Michael Hoa, I met in college. And, you know, as we were exiting college as best friends, we said, hey, let's go start a company. And we, you know, having really been students and engineers, we want to spread drone itch, which was a developer tool product because we were engineers. And so we saw this thing Docker take off, but we realized it was really hard for developers to actually get up and running with Docker to get their first container running on their Mac or their Windows computer. So we ended up launching what became Docker desktop, which is now used by over 10 million developers a day, generating your hundreds of millions of revenue for Docker. And the lesson, the takeaway there was really, I mean, first off, work with awesome people that you've already known trust along your journey. But importantly, it's that if you can really solve, make things easy for developers, it's incredible what can happen. And the level of creativity that gets unleashed for that, both for an individual developer, just starting to build their projects or next app, all the way to entire enterprise is being able to leverage this to move their business forward, just the power of a great developer experience. And so when we were starting with llama, you know, that was very much our DNA and how we thought about how we could bring value to this world of AI and open models.
Jason Hiner: What was the product call that you sold at the became Docker desktop, just out of curiosity?
Jeff Morgan: Yeah, it was called Kitematic and which kind of became part of the Docker product line. And then ultimately became Docker desktop, Docker for Mac, Docker for Windows.
Jason Hiner: Yeah, very good. Okay, how about the team that you have now, how big is the team, when did Ollama launch its first product?
Jeff Morgan: Team's 15 people, the first project we launched, actually, Ollama was a pivot out of a security tool business that we started in 2021. And so, you know, for two years, we worked on securing infrastructure. And look, I think security will have an incredible, important place as we know, security and safety with open models. But we were really drawn three years ago to, you know, it was three years ago, almost two, you know, the day a month ago. And we saw the launch of these open models come out and we couldn't help but ask ourselves and a bunch of our team, we had worked with a Docker and a bunch of them had worked at VMware. So we really knew how to build these run times and we kind of scrapped, like, you're scratching ahead saying, well, why is there no really easy to use runtime for these open models? And that's really kind of what started the Ollama project. And we wrote the first version in two weeks, we shipped it, happened to ship the same day as Ollama too. And the rest is history, it's just grown incredibly fast since then.
Jason Hiner: What interesting quits it is too, that you, one that you developed it and shipped it so fast, but then at the same time, as Ollama, which was this meta's open model, which was really kind of one of the first big open models and amazing timing, but also amazing sort of branding, you know, alignment there. That almost did some of the marketing for you, did it? Is that true? Because when Ollama came out, it generated a bunch of interest and then here it comes like, you all have this product that can help you sort of make it a lot easier to just run in your enterprise locally.
Jeff Morgan: Yeah, if you look at what drove Ollama's growth from the first thousand developers when we launched it to over nine million developers a month right now, it was, it was never about Ollama. It was about how could we get out of the way and help power the launches of these incredible open models? And more recently, power the, you know, the coding agents and harnesses that really enabled them to do spectacular things. And so it's never been about Ollama as much as how do we, you know, get out of the way for the developer and help these open models shine the best because they're, you know, they've gone from a model that, you know, the first Ollama model was a completion model so it could auto complete, you know, given some text. And these latest open models are incredibly powerful agentic models. They have multi modality. They're able to do long form tool calling. They have a million token context window. And so it's always been for us to power those launches and to help developers get access to the next wave of open model capability. And so each model launch drove a ton of growth for Ollama, whether it was the Llama models, the Gemma models, Mistral was a very big model launch back in September of 2023. In 2024, we saw the rise of multi modality 2025. We saw GPT OSS in deep-seek, which is just an incredible moment, both of them broken models. And that's been the flywheel. That's really, you know, grown Ollama incredibly fast.
Jason Hiner: Yeah, I want to get to that because 2025 and 2026, like the world has moved in your direction in a big way. And so I want to talk about that, you know, for sure. One question I would love to unpack, though, with you all is, you know, this, I love the mission from you all of, like, our job is to get out of the developer's way, to get out of people's way, and make the technology just work better. You know, there's something refreshing about that at a time when so many places are trying to amplify their brand and they sort of are creating a lot of energy, you know, around that. And so I would love to know a little bit of where that comes from and some of the sort of thinking and vision around that, which is you're all's job, as you think about it, you're all's job is to not make Ollama necessarily stick in people's minds, but to be a tool that lets them do some really powerful things and not have to pay that much attention to what Ollama is or to know the brand or the story that well.
Jeff Morgan: Absolutely. I mean, the product, it by design, it sits in the background and powers your favorite apps. And like the beating heart of Ollama, it's really, you know, these partnerships we have, whether it's the model providers and the model launches, that's one, there's the hardware and inference providers, you know, we're partners with NVIDIA, AMD, Intel, Qualcomm. Then there's the third component, which is the developers application or harness is a very, you know, more popular way to call it now. And the, you know, we saw DeepSeek launched a harness this week. It's the mix of these three things working together that is really what we try to nail. And that just means getting out of the way so the developer can enjoy their harness with their favorite model on the hardware they love or if it's in the cloud, they can get access to secure US and Europe hosting that they wouldn't usually be able to access because of, you know, extreme demand and compute.
Jason Hiner: It reminds me, you guys are like the really good friend who's always introducing you to cool people. That's sort of like the, that's what the bad thing is.
Jeff Morgan: Never really in the way, exactly.
Jason Hiner: Yeah, they're just sort of like, you know, always always doing those kinds of things. Those are like your most valuable friends, by the way, you know, the ones that are always introducing you to great people that you didn't know about.
Jeff Morgan: Yeah, it's really the staple of a good developer, a good data tool or a good infrastructure product is generally that they fade into the background, but they're always there to support you on the things you really don't want to build as a developer. You want to focus on your app, your harness, the task at hand. You don't want to focus on how do you make sure that, you know, the models, caching properly, that it's able to batch different requests. You know, these like low level problems, they're very interesting, but if you're a developer trying to build the next application of harness for yourself or your business, that's where, you know, that stuff should get out of the way.
Jason Hiner: Yeah, very good. All right, let's talk a little bit about this moment that we're in, you know, with open source, OpenAI, sorry. Yeah, open models, OpenAI models, OpenAI themselves, because as you mentioned, they've released GPT-OSS. Very little known, by the way, I find that very few people have a sense, even people who work in the industry, who follow news and all of that have a very little sense that OpenAI does actually have an open model out there, and I've talked to them, and I do have a feature store I wrote about why they did it, you know, on the deep view, but we'll talk a little bit more about some of the providers of open models, but there is this moment that's happening in the US that all of a sudden open models have gained an enormous amount of airtime and attention, and the hope is, is this a moment that turns into a movement or is it just a moment, which we'll see, but the momentum that has emerged around this has been amplified, of course, by Satya Nadella and Jensen Wong, you know, in the last few weeks, really releasing this open letter assigned by a lot of folks, including Ollama, supporting the US government not getting too restrictive about open models, about distillation, that's one of the ways that you sort of, if you look at somebody else's model and kind of learn from it, and not getting too, you know, regulatory about open models, because there was this fear, because most of the open models are Chinese that the US might just say, okay, we're gonna sort of ban open models or do something too heavy-handed. What, but the opposite sort of, you know, seems to now be happening, which is like the rise in interest among US businesses and developers in open models. Are you seeing that? What's your perspective? I mean, clearly, your whole mission is to help people use open models in really, you know, smart ways. So, but what have you seen from this happening, just this momentum over the last few weeks?
Jeff Morgan: Yeah, and it's been an incredible month. In a transition that really started with Satya's tweet was amplified by Jensen, and now it's becoming really the resounding approach from every single business we speak to, which is that, you know, open models which started as a research project three years ago, and for the longest time, you know, we're perennially behind the frontier closed models, you know, have now reached parity, right? And if you asked me three years ago, I would not have predicted this to happen so quickly. Specifically, like in a month, right? To have such a strong shift. I think it's largely driven by customers. And, you know, if you look at the companies that are adopting open models, which is effectively every business in every particle we speak to, largely because just like open source software, open models have a few key advantages. One is that, you know, I say three things, it's cost, it's privacy, and it's control. With cost, you know, as you saw in the news earlier this year, many businesses are just completely topped out of the coding, they're token budgets, and it's only getting worse because they want to drive their business forward with more and more of these really token heavy AI models. Whether they're running coding agents on the developer teams, or they're running kind of personal assistance, which we saw the rise of open claw in Hermes as two example frameworks, that tier token demand has led many businesses to say, well, how can I get 10x more usage? And open models are a great answer to that. First off, because you have this huge ecosystem of businesses optimizing them. And then secondly, you can run them on your own hardware, which you've already paid to acquire for your business. So that cost is the first big one. But, and that was really there at the beginning of the year as well, like earlier and later in the spring, early summer. But what we've seen in the last month is it's not just about costs, it's about owning your data, and then owning your competitive advantage. And if you think about it, there's a key loop that happens, and Satya talked about this, the learning loop. When you adopt a model, you start using it, that outputs data and important traces that think it's fed back into improving your AI strategy and your AI stack, whether that's training a model, it's customizing your own harness in the business. That loop has started to form and in every business we speak of. And what they're specifically doing is they're taking some off-the-shelf AI tools, whether it's a coding agent, it's a personal assistant like Open Claw, and then they're customizing it into their own AI furnaces, that's unique to their business, that has a unique task and unique outcome and goal. And now they're able to develop that. And that becomes this asset that they wouldn't trade for anything. And the challenge that these customers have had with existing model providers, the closed model providers, is they don't have control over that harness. Instead, they're being told to use an off-the-shelf one. And for a business where, you know, competitive advantage is everything, that's just, it's an easy choice. And so, whether it's the individual developer teams who are responsible for building these harnesses or it's the CIOs of these businesses, they're all asking how do I adopt open models? And it's just an incredible moment in time. Like I said, you know, a year ago, maybe almost none of the tokens being used by, you know, enterprises in the US-Europe were open models. And now it's probably about 20%. And I think even in one to two years, it'll be a super majority of tokens, right? You can see the line of sight into like 50 to 80% or something, you know, over the next few years. Yeah. Absolutely. So it's an exciting moment.
Jason Hiner: Yeah. For those who aren't familiar, when we talk about harness, hopefully you guys were talking about Claude Code and Codex, you know, and other, these coding agents that can go, they take the models, but then they can do incredible things with them. You know, these sort of long running tasks, they can do stuff that is agentic, of course these, you know, a harness is essentially an agent. It's a piece of software that uses the model in really smart ways. You know, Jeff, you've talked already about that being the fact that you started really focused on models themselves, open models, but a lot of what your team is thinking about and focusing on now is these harnesses. And one of the things that a lot of people don't, maybe don't always know either, is that open models harness Codex has been, open AIs, you know, harness has always been open. It's been an open source harness from the beginning. Claude Code is not open, but the newer harnesses, a lot of these harnesses are open as well. So you can not only customize the models, these open models, these open source models, but you can actually customize the software that runs them and decides what to do with them and does kind of the smartest, most powerful, you know, things with them. What does that look like for your team now? That's got to be a pretty big part of the, you know, 2026 momentum.
Jeff Morgan: Absolutely. We have not seen a growth trajectory at the rate, you know, we're experiencing in the last three to six months, largely because open models can finally power these harnesses. And, you know, I think the introductory one was really Claude Code over the holiday season back in December, January, there was the Claude Code moment. And that really unlocked this idea to your point of like, there's the model and then there's an outcome. And then there's something, there's a bunch of stuff in the middle to get that to work. And the way we see it as, you know, Jensen classically uses a five layer cake analogy for the entire AI stack, where you've got the application or the harness, you've got the model, you've got the infrastructure, you've got the chips and you've got the energy. And with Open, the problem is that top of the stack, and the bottom of the stack is pretty similar, right? You've got these big data center build outs, you've got, you know, NVIDIA and other GPUs, you've got kind of the infrastructure layer. But then as you get into the model and the harness layer, everything changes with Open, where all of a sudden you have the model by itself, you've got a task and a harness to run, and all the stuff you have to do in the middle. And that's what we're focused on. And so for our team to answer your question, it's largely been, how do we, as the harnesses evolve and get more sophisticated, how do we provide the technology that you need to power these harnesses that the models don't come with? So a few examples of this, for example, like tool calling, for a model to take an action and go do work, that's a really hard thing to get right, because the tool call needs to be effective and needs to be safe and secure. Teams we talk to need to monitor the harnesses, that's something you get out of the box with some of the closed frontier labs you don't get with open models. They need execution environments. I saw a great, great video about, from some of the anthropic team, about what they call the API platform. And it came down to three problems, knowledge, execution, and coordination. And we've seen the rise of a bunch of router tools. And I think that's a great example of the coordination layer, of how do you know what model to use? How do you make sure that a model if it needs to spawn other models to go do some tasks, it can do that reliably? These are the kind of problems we're focused on. Some agents are a great example of that. And it's only the beginning, for these long running agents, you can imagine where they run the same agent three times, they're delegating to subagents, you have a coordinator model, that's really a really big model, and you have a bunch of small models that are working in tandem. So we're starting to see the emergence of this. And for open models, that's something that developers need a tool to go help with. And that's how we're thinking about the roadmap of Ollama from here.
Jason Hiner: Yeah, okay, so for a lot of the folks in the audience who are here in this conversation, if they aren't using open models yet, I'm thinking about a way to sort of frame this, is that when you pay for a Claude Code, when you pay for codex, when even you pay for, you might pay Microsoft or Google, to do some of the things that they do for you. One of the things they're doing is sort of the connective tissue in the middle of how to take the models and then go do smart enterprise things for you in safe, smart, sort of private ways. And what you're talking about is the fact that there are a lot of open models that a company can just download and then use in their environment. And that's what you, Ollama, helps them do. It's like you don't need to go and pay a big subscription fee to any of these folks. You can download that model. You can then your team of developers, IT, your team can customize it to the things that you need to your own data, all of those pieces. But what happens is in the middle, there's some things that anthropic and Microsoft and OpenAI and others do that helps this work. You mentioned tool calling, so that's like being able to use tools, being able to use applications, tap into the things that you use in your organization. And what Ollama is doing is it's sort of taking care of some of that stuff in the middle. Is that free? Jeff, talk about how that works, because essentially you're helping companies if they can do this themselves, if they can let their developers and their IT team do it, they can save a lot of money and potentially get better or customized more private outcomes than if they're sort of renting from some of the big players.
Jeff Morgan: Absolutely, I don't think it's just free, it's also open source. If you think about the modern AI stack with open models, most of the stack is open source. To your point, the Codex harness is open source, the model is open source. So what we call kind of this runtime layer, it absolutely needs to be open source, which provides developers and teams a way to customize it all the way down to their business. And so it's absolutely open source. And if I think about Ollama, how we think about the business, it's much more, you know, they may not have the hardware. So we'll be able to provide that as a cloud service for teams that don't have, you know, the capability to run it themselves. But this runtime layer will get more and more to sophisticated over time. There'll be many, you know, new problems to solve around safety, security, coordination, plug it into your data that, you know, will become a critical component of the stack. Again, it's something you get generally out of the box from, you know, OpenAI and anthropic, but for open models, developer teams are ultimately needing to build themselves or find a tool that they're able to plug in.
Jason Hiner: Very good. So that's the business model for you all, is like, they can use your tool, they can go do it all themselves, but then they also have to host the models on their own infrastructure.
Jeff Morgan: Exactly.
Jason Hiner: That has, you have to buy GPUs, which will are very expensive and hard to come by. There's a lot that you have to do. Even if you sort of have a colo or some, you know, cloud capacity, it's still you have to manage it. And what your business model is that you're going to offer the ability to do some of that, that sort of hosting part for you, but then the company is completely in control. They can, their data is private, they can customize the models, they can do all that. And so that's where Ollama comes in in terms of how you're building sort of the future of the business.
Jeff Morgan: I'd say the other thing that's really important is, how do you actually get this plug in the model into Codex, right? Into, in Codex comes with Ollama built in for open models as an example. But going that extra mile to make sure that, you know, you can take a harness that you have or a coding agent application, you can plug in models, it should be one button. And I think that extra level of, you know, back to the Docker desktop story, making it right there in your environment, making it seamless to plug in. So there's a layer of developer experience and design as well that I think will unlock so many more developers to use open models. So this kind of a platform level way to look at it, which is, hey, there's just a series of needs that every developer and team is going to have. And then there's also how do you deliver that in a way that's well designed and easy to use. And that's where we've seen, you know, we're putting a lot of our energy is trying to make it so that if you have Codex and you have an open model that just came out, you know, this morning, the Qwen, the new Qwen 2.8, 27B model came out, state of the art model for its size. How do you plug that into Codex? There's a lot to get right there.
Jason Hiner: Okay. So one of the things that your team is doing is working with all these model providers, making sure that their models work with Ollama because this is what a lot of, you know, businesses and developers are using to run open models. And then you're also working with hardware providers as well, making sure Ollama runs in their environment is optimized, you know, for it as well. We'll talk about more of them in the second, you know, because one of the ways that I first learned about it was on the Apple's MacBook Pro M5 Max running Ollama on that because all of a sudden now you could run like these huge models that you previously would need at a data center. Now you're running on your laptop. It's pretty crazy. So we'll talk more about that in a second, but you, because you have this perspective of working with vendors on both sides of hardware and software, you also have a vision of a view into sort of what's happening in the ecosystem. So one of the questions I have for you is, you know, with this moment that we have with open models gaining so much interest in the U.S., knowing that most of the open models, the leading open models, like the ones you've mentioned, Qwen from, you know, Alibaba, Kimmy, DeepSeek, all of these, all of the open models that essentially can compete with the Frontier models, those are mostly in China and then the Frontier models are in the U.S. So the U.S. is sort of saying people in the U.S. in the AI ecosystem and beyond in the government and other places are saying, look, we need open U.S.-based open models, you know, here. We need companies working on it. And I will say when I was at AI for event a couple weeks ago, I did have a couple of companies that are coming up to me and saying like, hey, we're an open model, you know, we're building open models that are here. So I sort of have seen a little bit of it, but clearly those companies and those teams working on that stuff, they have to come to teams like you and you don't have to say the names, but I guess I'm wondering, are you already seeing after this momentum over the last couple months, are you seeing new maybe U.S. and European-based open models arising? And it's not a, I think it's not just a U.S. versus China thing. It's more that, you know, they're gonna understand the whether it's in Canada, whether it's in the U.S., or this Europe, that if people are making open models, they're gonna understand the needs, right, of some of those local GOs and be able to work more collaboratively, you know, with them and that. Are you seeing that? Are you already seeing a sort of a surge in interest of people building open models sort of over here?
Jeff Morgan: Absolutely. I mean, Monday with meta-superintelligence labs launching Use Glimmer and then committing as well to open a version of Muse Spark is a huge step forward. I mean, these are state-of-the-art models. And the way I think about it is this kind of two different, you know, approaches here. For these frontier models where, again, they're looking for superintelligence and seeking the frontier, it requires a tremendous amount of resources. And I think what's so amazing about Monday was that one of the biggest labs, which, you know, is meta-superintelligence labs, one of the most resource labs, committed to open. And that's a huge moment for this market. And again, again, this is the team that started it all with the original llama models. So what an incredible moment. And then we also meet with, you know, some of the new neolabs, which are just working on some incredible models, specifically around enabling these different use cases. So some of them may be focused on coding, some of them may be focused on cowork, some may be focused on model customization. And they're also well-resourced. And that's one of the special things about Silicon Valley. It's just the amount of resources that are available here, whether it's talent, obviously funding. And so we're seeing that and that being focused on open. Were you asked me a year ago where the next neolabs and the top five labs in the U.S. focused on open models, no. But with this change that's happened in the last month, a good portion of them are. And many of them supported Jensen's letter in Satya. And I think right at Pivotal Moment for that. And then I do think there are startups that are approaching it on how can they build very differentiated small models? Were small models don't require as many resources as the larger models? Because generally at your point earlier, they can be distilled. And we see a great set of startups that are distilling models around ones that can run on the phone or some that can run on your laptop. And I think that's, I think it's incredible moment for all tiers of models that all pistons are firing right now. And so it's incredibly exciting. But a year ago that definitely wasn't the case. And what I can hope for, and I think will happen because customers are so importantly focused on owning their AI that I think the idea from now we'll be looking back and saying, wow, that was just the beginning of the open model wave, that there's an entire ecosystem that's built. That effectively every model lab is releasing open models. And I think we could easily get there in the year.
Jason Hiner: So you see the moment becoming a movement as the direction it's going right now?
Jeff Morgan: It's been three years since we've been launching models with Ollama, but I think it's really only in your point, the last three weeks that we've seen a pivotal change in the market. And it really starts with ultimately the largest players in the space. So when the labs online with the talent and the resources come online, and then you've got the hardware support from NVIDIA to go open, I think you've got everything you need to jumpstart an ecosystem. That's unlike anything we've seen the last three years. And open models have already been deployed in enterprises, Ollama's used by 85% of the Fortune 500, but ultimately I think it's just at the beginning. If you look at the token rate, how many tokens are actually going through open models, that's still quite small. And I think we're about to see an incredible inversion where open models are the super majority of tokens.
Jason Hiner: Okay, you said 20% of models today globally are going through open models.
Jeff Morgan: 20% of the token volume is the team that I have. And just out of token volume, right? Because if you think about, if you're going to build the next business on top of AI and AI platform, that really wasn't available to you until the last 30 days because of the models, because of the backing from the industry, because the model and because of the hardware partner support. And so I think that's all come together and that's going to completely change that 20% into an 80 to 90%. You see a world easily where for everyday tasks, everything but the very frontier research, open models are the best choice. And at that point, you have a focal point where most of the tokens in the world are just processed through everyday open models. And that's great for businesses, right? Because the costs get lower, they have more control over customization. It unlocks a whole new wave of innovation. It's not locked up in one group of people.
Jason Hiner: Yeah. All right, I want to talk for a moment about developers because obviously it's a core audience for you all. And I saw this movement happen over the last few months. This is maybe three or four months ago when people really started to feel the pain of token costs and inference. And one of the things, we had businesses like Uber saying, they ran through their whole token budget in four months for the first four months of the year. This is really driven by what we talked about earlier by agents, these agents, they eat up a lot of tokens but they also do a lot. They can do a lot more all of a sudden. AI has gotten enormously more capable in 2026 just because of the agents, the agent harnesses alone. One of the consequences of that is that developers are spending sometimes crazy amounts of money. I had a CTO tell me that the leads on their development team were spending 1.5x their total compensation and tokens essentially every month. And so we've seen, this is another one of the areas that you all play in. It was how we first met. I wanted to ask you about this M5 MacBook Pro M5 Macs came out and all of a sudden you were able to run these larger models, 30 billion parameter models, 70 billion parameter models. For those who aren't familiar, that's typically you would need a data center to do that before you could only run maybe a two billion parameter model on a really powerful laptop. Well, now you could run this on your laptop. And so all of a sudden you had all of these benefits of you had a number of benefits. You were saving a lot in token costs. Your token costs might be zero, might be whatever the cost of the hardware is. Even if that laptop cost five grand, if you were spending, if it saved you in that month, it saved you more than that in one month in token costs. But then you also had a benefit of it was faster, it was more private, and you also had the ability to run it when you were offline. You could be, there were these tweets that went live of like, I'm doing all my work on a laptop at 20,000 feet, you know, I'm connected to this local model. And so I'm interested to hear your all's perspective on that because all of a sudden this year, the hardware has gotten really interesting and good. You all work with the hardware providers, including Apple, AMD, Nvidia. Are you seeing that? Is this sort of just a flashy thing? We see in some tweets or are you also seeing, you know, some real momentum around developers and people wanting to run local models in that way.
Jeff Morgan: Yeah, it's a, there's an inflection point occurring in the, you know, workstation hardware market. And that started with the Apple Silicon laptops. And what are the key things that have changed? The first one is that you've got an extreme increase in memory, you know, just the sheer amount of memory that comes with your computer. Whereas, you know, a fixed graphics card for the last, for the gaming and workstation market of the past might come with eight, 10, 16, maybe 32 gigabytes of VRAM. You can get a laptop for a couple thousand dollars. That is 128 gigabytes. And that's the first step of this unified memory because you're able to pool the system and the graphics memory into one place. So that's the first thing that changed. The second thing, which is really new at the M5s is these new hardware instructions. So the graphics cards of the past were really focused on rendering graphics, you know, they focused on 3D graphics. And that has the, you know, because of technology like CUDA, you know, which NVIDIA built for many years leading up to the first neural network, you know, that enabled the flexibility to enable these neural networks, but that doesn't mean the hardware was optimized for it. And now we're seeing a case with M5 and the new DGX Spark from NVIDIA and the RTX Spark where seeing these unified memory be met with these hardware instructions that are specific for accelerating neural networks. The neural accelerator is an Apple. And then of course you've got the Blackwell and, you know, very Ruben chips coming from NVIDIA that are purposefully focused on accelerating neural networks.
Jason Hiner: Yeah. And it comes with the... Just running AI workloads. Like it's like a separate processor just to run AI workloads.
Jeff Morgan: Exactly. You know, just to handle attention, which is a very expensive operation. And then lastly, you've got this suite of software that's being released to accelerate, to take that and then make it extremely accessible for tools like Ollama and other developers to then integrate that into the developer workflow, right? And so with Apple, there's the MLX project which Ollama leverages to be state-of-the-art speeds on Mac. And then of course there's the new CUDA stack and Nemotron stack of the application layer from NVIDIA. And so you take these three things and you bring them together and it's less about like, hey, we're seeing desktops get good enough. It's like, this is the stuff that's powering data centers that's coming onto your laptop. And that's extremely powerful.
Jason Hiner: At a price point, which is not much more expensive than the laptops of the last 10 years, right? It's a little bit more, but it's not order a magnitude more. And that creates a limit. It creates a limit. You can get a really good one, three to five K.
Jeff Morgan: Absolutely.
Jason Hiner: When you're spending that on five or 10 K on token costs, like you can be like, great, I spend that. And now I'm saving a ton of token costs.
Jeff Morgan: Exactly. And what's so impressive is for some of the models, you take the Gemma for a 26 billion parameter model and the Qwen 35B model. And these are models that are optimized for high throughput. And so on a M5 Mac or an RTX Spark or DGX Spark, you're able to get thousands of tokens of seconds processed and prompt processing and over 100 to 200 tokens a second in output. So you're getting data center speeds to all that happening on this three to $5,000 machine. And so, and what's great about that is that it's a fixed cost and it's private. And so you don't have to go figure out, hey, where's the compute? Is it in a secure location? Is it in the right country? It's right there with you. And so the way we think about this, is this the personal moment for a personal computer moment for AI, it's a PC moment. Because all the things that went from a mainframe back, IBM mainframe back in the day to the PC, you probably remember when you got your first PC. It was an incredible experience, right? It wasn't the most powerful machine compared to what you could, you know what, I'm sure what banks were using for mainframes, but it was yours. And it was your data and it was your applications. And you didn't have to worry about metered costs. And that same moment's happening right now. And it's going to completely transform developers, consumers as well, but most importantly, you know, businesses.
Jason Hiner: So a lot of the, we're in this sort of compute crisis still, right? Like the cost of memory is so high, it's rare. GPUs are hard to get a hold of. But one of the things, that is one of the things that I see as one of the movements in the second half of the trends for the second half of 2026 is this local processing. Both sort of on the sort of desktop, laptop, and you know, even kind of locally. All of that has the potential to ease up some of this burden, this data center burden that, you know, everyone's running, building data centers, because they're worried that, you know, AI is growing so quickly and we won't have the power to do it. And these, so open models and plus this hardware has the opportunity to help ease some of that burden. Yeah.
Jeff Morgan: Exactly. And I think it's not all of it to your point, but there's a set of use cases where it makes so much sense to run it locally. And as we stand today, and I hope this changes and I believe it will in the future, the hardest coding tasks, you still got to use a really large model. You saw the Kimi K3 models, the 2.8 trillion parameter model, the Qwen Max model, future models were rumored to be very large in the trillions of parameters, the open models. And so, and those are there to rival these large frontier closed models, and they really shine with these really hard coding problems. However, for a lot of the cowork trying to get work done, trying to do document processing, that's already absolutely possible with the 2,200 billion parameter models that run great on this kind of hardware we're talking about. And so I see that for the run-of-the-mill workloads like this, it will absolutely be local first. And you'll be in a world just like with iCloud, for example, where some of it will happen on your machine, and some of it you'll offload to the cloud. And we'll just need to get really good at making that seamless as the hardware catches up. And so you can imagine a world where exactly, this is where routing comes in. And routing's a huge topic right now from a cost standpoint, but also from a latency and just customer experience standpoint, routing has so many potential applications.
Jason Hiner: Very good. Now, are you working with other providers, too? I'm looking and I'll hold one up. This is the AMD Ryzen AI Halo. And when AMD sort of put this in my hands, they said, OK, try our version of this. This is essentially their version of the one you mentioned earlier, the Nvidia's box. You mentioned Nvidia and Apple. They're two heavyweights in this industry. But there are others doing, are you working with other hardware? Is there more coming? You don't have to say the names, but do you have a line of sight into more hardware being able to do the same kind of things like we talked about with Nvidia and Apple?
Jeff Morgan: Absolutely. Just mention the two, maybe call it sorted by market cap. But we're seeing this transformation in every major hardware vendor. Because the applications are shifting to these neural networks, these models. And this unified architecture is becoming the most powerful way to support these models. And we're seeing both AMD and Intel are two example partners of ours. And we're so excited to collaborate on. Because to your point about it costing 3,000 to 5,000, well, what if it wasn't even that expensive? And I think when you have a collaboration of hardware vendors come together to solve this problem, whether it's an introductory laptop or a really powerful high-end machine, there'll be options across the board. And so we're just really excited across the board from both AMD Intel as well as Nvidia and Apple. And partnering with them and making sure that when somebody buys one of these machines, they can open it up, get up and running with a model and start doing work with it. But those are two examples of more partners. We're also partners with Qualcomm and extremely excited about what they're doing as well.
Jason Hiner: Okay. Those are pretty much all the big chip and hardware makers in the space. Then there's some laptop makers as well. So that's a really interesting development. We're seeing the AI ecosystem sort of, there was a few names that dominated maybe a year ago and kind of to your point, as we've been going through the conversation today, you can see that a lot, there's gonna be a lot more participation. And in a sense, you're gonna need that because there are gonna be more distinct use cases, there are gonna be more, there's gonna be some more freedom and options for people to customize and do some of their own things. One of the last sort of questions I had for now, Jeff, for all of the ecosystem stuff is, I recently ran into the CTO of Thomson Reuters and he was saying that they built their own model and they had this view that intelligence was gonna become the most important part of what they do. Now, we think that Thomson Reuters Reuters is like a news organization, they have a bigger business. One of the things they do is like law, they have a lot of like their own kind of law services that they offer. And their conclusion was, if we're gonna really make these models a key part of what we do, we have to own the intelligence, like we can't outsource that to somebody else. And that sort of broadened my perspective to be like, that is sort of the logical conclusion. You can see a lot more, a lot more companies come into the conclusion, like, okay, this is doing important, we're putting important stuff on this. Like we also, we wanna be really careful about who we wanna send that to, right? That becomes our maybe most important, most valuable IP. And so because you work with 80% of the Fortune 500, where does that, I mean, making our own model is only maybe a step further from like taking an open model and customizing it for what we do. But I just wanted to get your perspective on that. Are you seeing more and more companies basically coming to that conclusion, like, oh, we need to own the intelligence layer. We can't just outsource that to a provider. That would be silly to take our most important part of our business and do that, no matter what it was.
Jeff Morgan: Yeah, I'm seeing three phases from these very large businesses who are very much wanna focus on owning their AI. And customizing it. The first is, you know, the first phase is how do you, how do you switch to open models? So you basically switch to a stack where you control it. And the step two is then how do you customize it? And then what's so exciting is the step three that I was whiteboarding earlier this week with a major financial firm, which is then how do you give it, how do you do that with your partners and how do you share that in a way that benefits the ecosystem and your partner ecosystem? And this will all come back kind of to the first point. But, you know, step one is just making sure that you're building on intelligence that you can ultimately run it yourself. You can take control of it. You know, the incentives are aligned for you and the providers of that intelligence to, you know, move your business forward. And if you look back to, you know, the largest cloud providers AWS Google, they're built on open source and that provides an incredible amount of agency to these businesses. So that's step one. Step two then is how do you customize that loop we spoke about with Satya, right? With Satya's tweet. And that involves then how do you take your data? Like, contributions have so much data, right? Around all the press releases, all this legal data for that part of the division. And how do you customize the models, the harnesses to really make sure that you're able to solve your customer's problem in a level you're not before. But this third step of giving back is what excites me the most, which is this major financial force time. But well, what if we open source the model ourselves? Because we could build our own AI, deliver that to customers, but then our customers are now in the same shoes as us with the previous closed model providers. And so it's this interesting, I love that insight because what it meant was instead of taking the intelligence and keeping it to yourself and then, you know, and then not providing to your customers, how do you actually give back? And this idea that, well, if you take it, if you're able to build your own model and then actually open source that, the whole cycle continues. And now your customers can say, hey, I can work with this financial firm, for example, use their models and I can trust that they've got my back. And then that financial firm has a great way to service these customers with a ton of infrastructure and value added services around that model. And this is really what's special about open models is you're always paying it forward to your customers because everyone's a customer of somebody. And if you can have open models all the way down the stack from the foundation models, all the way to the vertically customized models, all the way to the end, you know, provider and retail, you know, and provider to the consumer. And then guess who benefits the most? The consumer because they've got incentives along all the way down the stack and then the consumer ultimately owns their intelligence. And that thinking was so powerful to me. And I thought that was something that, you know, every business will embrace.
Jason Hiner: Yeah, very good. Love that, love that altruistic note in terms of where the ecosystem is going. And the conversation in general, like thinking about open models, there's a lot more empowerment there. There's a lot less top down, a lot more bottom up and all of that, you know, are good things. Okay, Jeff, I'm gonna let us sort of wrap with the two questions I ask everybody, you know, at the end of the podcast, which is the first one is, what's the AI tool that you're using, you know, today that's surprising you that maybe not as many people know about that you'd love for people to take a look at?
Jeff Morgan: I'm really excited. You know, I'm neck deep in the developer world. I'm just excited by this new layer of software that's being formed. And I'd say there's kind of two that we're spending a lot of time with, within Ollama. One is the MLX software from Apple, which many people don't know is from the founding PyTorch team, which is kind of the software that runs all the inference in the world right now. And they went to Apple to start this new project called MLX. And what's so awesome a lot of people don't know is it's what comes across as another way to use GPUs to run models is actually this flexible tool for you to do training, for you to do inference to kind of build anything you want. And I think that's so powerful because it can be done on a consumer laptop. And so I'm super excited by that. With any spare time I have, I'm just reading about it, scratching the edge. And I think very much within video we're seeing the same thing with the DGX Spark. And I'd say if there's a number two, it's the DGX Spark. We were so lucky to be one of the first people to receive one. And what started as kind of a product maybe focused on researchers and developers, I think has become a little bit of a sensation because it's so powerful and you can run almost any kind of AI workload on it. And I think both whether it's an Apple Silicon laptop with MLX or it's a DGX Spark with the new CUDA stack for it, I think they're just powerful tools. That's the answer for developers. I think I'll say the one that I'm excited about is my Gmail search box finally has AI. I'm so excited for that. It doesn't have much to do with open models, but I'm just glad I can get through my email much more effectively now.
Jason Hiner: Very good. Those are AI tools. Well, thank you. We always try to get the tools that people in the industry are using the AI tools. So those are two great ones. Yeah, from the sort of developer standpoint, all the way down to the stuff that everybody can use. I also, yeah, very glad to have the ability even just to search and find emails in Gmail. It's a big upgrade. Okay, the last one, last question that I ask everyone is, you know, there was always this promise in AI that when AI sort of advanced, that it would do a lot of the mundane tasks for us, and we would have more time to sort of do other things in life, and I've found over the last couple years more of the opposite issue. Most of the people, the leaders I know, they have less time and not more at the moment. Maybe because they see the opportunity and are very excited by the opportunity that they can do so much more, and AI can do a lot of the things for them. But because of that, everyone is more and more conscious about time, and as leaders often phrase it, they talk about how do I get maximum leverage for my time, and so I love to always ask leaders, you know, what's your best tip? What are you doing right now that's helping you get maximum leverage for your time? You know, what's the tip that you would share for other leaders?
Jeff Morgan: I think everyone, you know, we talked about a loop at a business level, but I think everyone individually has their own loop that they go through in the day. And what I'm encouraging myself and my team to do is to say, how do you make that loop as automatic as possible, and obviously use AI for that? And, you know, and it's hard because you got to get the right connectors and model and runtime, and you got to find the right harness. But I think ultimately the biggest tip I have is if you can automate that away, if you can take these parts of your day-to-day life that are not special, you know, running a startup's really hard. I always ask myself, what are things I can do to automate so I can spend more face time with my team so I can focus on, you know, my personal life and spend more time with my wife. I think that's where encouraging everyone to build their loop and actually optimize it and figure that out. The best example I have for this, actually, is one of Ollama's lead investors, his name is Tamash Tungus. He, while he's a well-known investor in AI and data, he's automated so much of his life with a coding agent named Pi. And everything from research to emails to documents to presentations, he does it all, he lives in his terminal. And I think that's the future. And I think, you know, it may not be in a terminal, maybe in the Codex or claw or an open source alternative. I think that world is gonna be so exciting. To the point where like, you just don't even have to be in the loop anymore, you can start your day and end your day and you have a whole best friend coworker that's done most of the boring work of your life. So that's my tip is to strive for that loop because I think it's not just businesses doing it, I think it starts with the individuals at the business.
Jason Hiner: Very good. Well, Jeff, thank you for the time. Great conversation. I'm sure we're gonna talk again because there's so much happening in this space, but what a moment to be able to do this and really appreciate your insights on all the things that we talked about, especially sort of where open models go from here. It's a great discussion.
Jeff Morgan: Thank you so much for having me.
Jason Hiner: All right.