From frontier labs and enterprise platforms to emerging startups reshaping entire industries, The Deep View: Conversations podcast interviews the brightest minds and the most influential leaders in AI.
Caroline Ingeborn: The thing that has never changed throughout this period is that we believe that multi-modality is the way towards AGI. It's the way towards the kind of intelligence that can operate and generate alongside us as humans. And so the way forward is to build unified models. If you are building intelligence that have only really seen one modality, then you end up with video models over here, image models over here, language models over here, 3D models over here, and audio models over here. And that is sort of saying that, okay, we're going to take the most advanced technology that we've ever seen as human beings by becoming them all together.
Jason Hiner: In this episode, senior reporter Nat Rubio-Licht sat down with Caroline Ingeborn, the COO of Luma AI. A lab that's building multimodal intelligence that transcends language alone. While most of the AI conversation is still dominated by the latest LLMs, Luma is betting on something broader, systems that can understand and generate across text, images, and video in a unified way. Ingeborn argues that this reflects how humans actually process the world, not in a single modality but across many signals at once. Nat and Caroline talk about why that shift matters and how world models could change the path toward AGI. They also explore where this technology is already being applied today, especially in creative industries such as entertainment, advertising, and marketing, and why Luma sees AI as a creative amplifier rather than a replacement for human work. Ingeborn also walks through the bigger implications, from how multimodal systems may reshape workflows to what physical AI could mean for labor markets and centralized power. So here it is, our conversation with Caroline Ingeborn of Luma AI.
Nat Rubio-Licht: Caroline, thank you so much for joining us today. I want to start by talking a little bit about Luma's roots. You guys were originally founded in 2021, which was a very nascent time for AI. It was before the, it was like a round or before the whole Will Smith eating spaghetti of it all. How has your company's place in the AI market shifted since then?
Caroline Ingeborn: A lot, I think is the right word. So we started, our founders met when they were working on Vision Pro at Apple. And so the foundation, it's sort of, the foundation was always research. That was always the foundation of Luma. The initial focus was 3D. When the technology was becoming good enough to see that like video will be possible. And then Sora released like early, early sort of examples of what video could do. At that point, Luma had already sort of behind the scenes started working on video. And then in June, 2024, when we released our first video model, that was like that was a moment that the world had never seen. And so I joined like two weeks before we launched that video model, because there was no one that was focusing on launching a video model. And that seemed like a good thing. I don't want to focus on. But so we then we developed from there to build sort of video models and image model and singularities. And the thing that has never changed throughout this period is that we believe that multimodality is the way towards AGI. It's the way towards kind of intelligence that can operate and generate alongside us as humans. And so it might sound like, oh, there was 3D and there was this and there was this. But like the mission of the company has never really changed. And so we we just spent time building video models and image models. And around 12 to 18 months ago, we realized that the way forward is to build unified models. And that is models that have not only seen one modality, because if you have, if you are building intelligence that have only really seen one modality, then you end up with video models over here, image models over here, language models over here, 3D models over here, and audio models over here, and God knows what other models. And then you need to plumb them together. Because those models have only seen that type of data. And that is sort of saying that, OK, we're going to take the most advanced technology that we've ever seen as human beings, and we're going to succumb it to plumbing, by plumbing them all together. Or we're never going to do that. And we're going to just like build the kind of intelligence that is really, really good as long as we stay in one modality, which is what we're seeing now in language, which is, OK, you get really, really good at language tasks such as research encoding, which is awesome. But I don't think that is ever going to take us to the next level. And my non-engineering way of thinking about this is that I don't know what happens in my brain when I sleep or when I dream, but I know that I've never dreamt in images or text alone. And so that's kind of weaving together different kind of modalities. That's what our brain has done for a myriad of times. And so we have evolved a lot as a company, but the things that have always sort of been at the core of Luma is our focus on multimodality towards AGI and research at its core. Like we are, first and foremost, a research lab that has then evolved to bring this intelligence to market.
Nat Rubio-Licht: I definitely want to talk to you more about the AGI of it all, the multimodal models, the world models conversation, which is something that you and I have talked about before. But before I jump into that, I want to talk a little bit about where Luma is fitting into the creative industry, because I know that you guys offer creative agents as part of your suite. As someone that lives in LA, I'm very aware that AI in the creative industries is a pretty contentious topic. So where do you guys see yourself fitting into the creative workflow?
Caroline Ingeborn: We see ourselves because of our agents working with creative professionals with end-to-end creative work. And so that differs depending on what your creative work looks like. If you work in a studio, then it's a lot of... There's a lot of early ideation, pre-vis, etc. But AI has been pretty good for that for quite some time. And where we're seeing a huge uptick is on the production side. And what has sort of really driven that is the use of modified video or video-to-video, where you can use real-life actors and you record them in whatever you want. And then you use images of environments and situations that you want to transfer them to. And so that means that you can make productions that really wasn't possible. And so for all the fear that I hear and see in your hometown that I understand and I empathize with and that I think is very human. I'm an optimist. I'm an optimist, especially when it comes to human and technology. And I see a Hollywood that have stopped greenlighting any sort of projects with risks. And I do think that AI is a path towards greenlighting projects that otherwise would not be greenlighted or definitely not produced in LA. They would be all made in a country far away where it's super cheap. And so I think this is a good thing for filmmaking and it's a great thing for LA over time. Are there going to be bumps? Of course. But that's sort of very LA focus. What we're seeing in marketing and in advertising is content production at scale. The sort of conversation that has been for the past, I don't know for how long in marketing of personalized content is becoming true. And that's really, really, really exciting. Both agencies and brands, I'm happy to go into specific cases, but both agencies and brands are really grasping and jumping at this opportunity to not so much see it as a cost cutting measure, but rather saying if I continue to spend as much as I did before, can I serve my customers better? Can I reach new customers? And through this, can I build a much bigger business?
Nat Rubio-Licht: Now it's time for a word from this week's sponsor, Microsoft. What does it take to go from an AI idea to your first paying customer? Most teams are experimenting with AI, but experimentation doesn't generate revenue. Shipping does. Microsoft AI Envisioning Day is a free video series built for developers and software companies ready to turn ideas into real products. You'll learn how to take an AI idea to a working MVP that customers will pay for using practical patterns and guidance. So that you can build something that actually generates revenue. No fluff, just clear frameworks and steps you can start using right away. If you're building with AI and want to start closing deals, not just shipping demos, this is how you get there. You can get started today at aka.ms slash theDVU. That's aka.ms slash theDVU. And we thank Microsoft for their support of theDVU. Now back to the show. It's so interesting because I've heard this from a lot of people that I've talked to in the AI video space. This idea that the tech can sort of democratize the ability to create, like you said, these more risky sort of ideas or films or ads or anything in the creative space. Do you think however, are there any ethical red lines or places that AI maybe doesn't belong or do you think that's sort of a moving target that the industry is still figuring out?
Caroline Ingeborn: I still think people are figuring it out. I mean, I think that some of the same sort of rules that we've always lived by should apply here as well. You should not be able to do an output of a very well known IP. And if you use the technology and you bend it backward and forward and you're able to do that or you introduced said IPs into this technology, you're acting in bad faith. We should not take responsibility for that, but bad actors should be notified. That's true in the non-AI world and that should be true in the AI world as well. I mean, this is not an ethical thing, but I think a lot about taste and what's exciting in 2026, especially with our Luma agents is that the technology has become so good and the products that are built on top of the intelligence is built with creative professionals, not just in mind, but we built it together with them. 25% of the people that works at Luma are creative professionals. And so what's exciting is to see people that have worked in the creative industry for decades and that know they know what it should look like and they are using this technology and therefore we're seeing much better results every month. And so while it's true that a lot of jobs are going to change, what is not going to change is if you're very good at what you do and you learn this technology, your outputs are going to be better than anyone else and people will pay you for that. They might pay you even more for that today than they are willing to pay you if you did not do it with AI. And I think that's exciting. That's exciting for anyone who's watching any sort of content honestly. And that's exciting for anyone who's really good at a craft.
Nat Rubio-Licht: I want to pivot to talk a little bit about this idea of world models and physical AI and visual AI generally. It's a huge topic of conversation. It's something that we talked a lot about at HumanX where we first met. Why do you think the industry has started to pivot its attention to this area of AI?
Caroline Ingeborn: I think there's really three things. One is the technology is becoming good enough. The second one is that if you solve some of the problems that you need to solve in robotics, you also solve it for everything that we're doing on the creative side as an example. Because that means that you have models that are understanding the physical space. And that has a certain limit when you have models that do not with anything visual. The third thing is just addresses a fundamental mean of just production globally and labor shortages. So I think that there is a huge demand for this and opportunity here. Not just financially but for all of us in terms of what we do and how we spend our time, etc. And so I think all of these are like, you know, which is the biggest reason, I don't know. But there's a big push for it and timing is right.
Nat Rubio-Licht: I know that there are different ideas about how to achieve this, how to achieve a world model, I guess there are different kinds of world models. And I know that some companies believe that these can be achieved like world models, human level intelligence by just scaling video models. And you guys believe in this omnimodal approach. And you talked a little bit about this at the start, but why did you guys decide to take that approach to building world models, building towards AGI?
Caroline Ingeborn: Staying in video land, like it caps out. It caps out. And it comes back to what I said in the beginning about plumbing. And so you need to sort of settle for less or like settle for a very good video model. But it's hard to see how a very good video model can become that, just as like it's hard to see how a good language model can become that. And so I think, you know, that's based on what I'm hearing about what other prominent research labs are doing. This is becoming more and more of the shared sort of conclusion about how to get there. So I think anyone who made a different bet than we did 12 months ago is in a much worse position. That doesn't mean that we automatically win. But it does mean that we are in a very good position in the market for this.
Nat Rubio-Licht: What would you say is the end goal of Lumas omnimodal models? And what would you say is your most recent milestone?
Caroline Ingeborn: Our most recent milestone is Uni-1, which is our image thinking model, which is a text and image model. So it understands both language and an image. It's a very powerful model. And we are able to compete with both Google and open AI on this front. And that's a huge feat for a team that is so much smaller and resourced in a different way than what those companies are. And I think we do that because we have fantastic people, but also because we, this is all we focus on. This is all we do when we do something. We give it our full attention. I think that's also why AI is exciting at the moment.
Nat Rubio-Licht: And what is the, what do you think is the end goal?
Caroline Ingeborn: I think the end goal for the end goal for us is general intelligence that can operate and reason alongside us humans.
Nat Rubio-Licht: What do you think are the most promising use cases for that kind of power or for world models and omnimodal intelligence generally?
Caroline Ingeborn: I mean, there's a lot. You know, we talked a little bit about the, we talked a little bit about the robotics case, which is, you know, addresses the, that addresses the global labor shortages. It makes it possible for us to build and produce and service in a way that we've never been able to do before. But what, what, what world models also does is that it helps or it enables everyone to visually communicate. And not just to communicate, but to create. And so if you think about how, how blocked most of us are in terms of not being able to visually communicate what we need or want or visually communicate, like our main sources speaking like you and I do now or writing, but 70 to 80% of us are visual thinkers. So my ideas never come in text. My I like my reasoning almost never comes in texts. And so that's why we assume I'm like whiteboards, but our main communication tool on whiteboards still turns out to be texts because we're at least me, I'm very bad at drawing. Now if I could easily visualize for you the kind of world that I want to create or the kind of product I want to create and what that vision looks like, you would understand me a lot better. And that would go for any potential customers, any potential partners, any potential, any relationship that you have. And so I know that I sound a little bit like, and then we will have world peace because we all understand each other so well. I think we're pretty good at like, not, not, I don't think it's the one all solution. But I do think that we would be able to solve really, really hard problems a lot better with this and solving hard problems is important because there's a lot of hard problems that we as humans need to solve if we want to continue to coexist the way we do at the moment.
Nat Rubio-Licht: I'd love to talk a little bit about your approach to general intelligence and this idea of generalization. You guys recently launched the open physical AI lab, which is a collaborative effort to solve this issue. Why did Luma decide to take that approach to solving generalization, this sort of like open project collaborative approach?
Caroline Ingeborn: I mean, that simply because we don't believe it should be closed. That kind of technology should be open and available to everyone. And so this was our way of setting up the lab is our way of saying, this is how important this is and give it, give it our full attention. And the second thing is to make it open simply because it's, it's the right thing to do. So it wasn't a very once you sort of like had that clearly in front of you that sort of these two things came together.
Nat Rubio-Licht: What are the dangers of keeping it proprietary or keeping this kind of research closed?
Caroline Ingeborn: If it is closed, then a few centralized monopolies will control the world's physical infrastructure. And so just think about that for a second. Everything that is around us or could be around us could be controlled by a handful of people because many of these companies are not just owned by everyone and run by everyone. It's, I mean, it's controlled by a handful of people. That seems super scary to me. Super scary, especially because a lot of those people doesn't seem to be great at handling the powers entrusted into them already. And so if we add on the world's physical infrastructure into that, like I would like there to be many places, but at least a place where there is a foundational lab that says, hey, this is an open science lab. We are going to share this. We want to co-create and make this accessible to everyone to build on. So that it is not controlled by a handful of people.
Nat Rubio-Licht: Do you think the way that LLM stand right now, we are sort of facing that centralization problem already?
Caroline Ingeborn: Yes and no. I mean, yes. And to some extent, because there are three companies that are very dominant in the Western hemisphere. No, because I'm seeing this sort of push towards very good and reliable open-weight models because the technology has become good enough. And so every day now you read about, and it's not just like, oh, five guys building something in a garage. It's like the largest companies in the world saying, hey, we don't want to pay this much or only rely on this one player when it comes to this. We want to use open-weight models instead. And so I think that that is going to shift this power balance that we've seen a lot. And I also don't think that anyone in the LLM side has a real moat at the moment. And that makes my answer to your question, which a few months ago probably would have been more yes. And the firm yes to be like, nah, I'm seeing a complete shift on the horizon here. That is shifting this power balance just as we speak.
Nat Rubio-Licht: I am curious as it relates to LLMs and world models, because there is this growing movement towards world models being, and world models, omnimodal models, visual intelligence being sort of the key to achieving generalization, AGI. Where does that leave all of the work that has been done for LLMs? I guess the idea that LLM scaling laws will get us to AGI.
Caroline Ingeborn: I don't know, but I don't think it's really a problem. Because I think that the encoding is having its moment, right? And so there is this real push to continue to improve this flywheel as much as possible. I don't know how much the people that are doing that are like, oh, is that leading us to AGI or not leading us to AGI? I'm sure there's a handful of people that are focusing on it. But they are, I mean, they're also just trying to win that race and build that business.
Nat Rubio-Licht: I guess who would you say are the other than LLM, of course, the labs or the, I don't know, researchers, the organizations that are doing exciting work in the world model space. Yeah, who, yeah, what else is exciting you in world models right now?
Caroline Ingeborn: All the large labs and I mean, on the robotics side, there is also a bunch of smaller labs and we're going to see more and more of them coming. And there is a, there's different views of like how to get to what, but that's always the exciting part of research. So I think it's a, I'm excited about what I'm seeing at with all of them. It's just, it's too, it's too early. Yeah, it's too early to tap winners.
Nat Rubio-Licht: I guess last question I've got for you. And this is kind of a broad one. So take it however you want to. What is your AI hot take?
Caroline Ingeborn: My AI hot take is that, I mean, I don't know how hot these are, but I don't think anyone has a moat in AI and anyone who claims that they do are either lying to themselves or to others. There's just not. I wish there were, but there's not. And the second one is that if anything, I think the real moat is deployment. So on the Luma side, what we start doing very early, and, and which is why we're having such a, I think warm welcome both among, you know, C-levels at enterprises, but primarily among the creatives inside of enterprises, because every CMO, every studio head, everyone also like, this sounds like great, but how do I get my creative professionals to, to, to work this way? I mean, the answer is that you do not, but our forward deploy creatives do. And they come in and they, they are creative professionals who started using AI early, who is very good in their crafts and they come in and they work alongside you. I mean, that is, is that something that only Luma can do? No, it's not, but doing that really, really well is a big differentiator. Same thing on the forward deploy engineering side. Deployment is the differentiator. And so for us to coin this sort of title of forward deploy creatives, and that being synonymous with Luma has been very good for creative professionals inside large organizations because it ensures that this is not a technology that is dumped on them and then they need to swim or sink. It's technology that is being deployed together with them by people in their field. It makes it seem less scary and daunting.
Nat Rubio-Licht: Yeah.
Caroline Ingeborn: And you get, like, and you get to see, like, exactly how your workflow changes and what the technology can do here for you now.
Nat Rubio-Licht: All right. Well, thank you so much, Caroline, for joining us today. I really appreciate your time and it was a pleasure to have you.
Caroline Ingeborn: Yeah. Thank you so much.