The Deep View: Conversations

Nearly every smartphone launched in the past year features agentic AI capabilities, offering users an early look at what a fully agentic smartphone future could do for them. Of course, the tech powering it is driven by the chipsets.

In this episode of The Deep View Conversations, we talked with Vinesh Sukumar, Qualcomm's VP of AI at the Snapdragon Summit, the company's annual conference where it launches its latest processors. This year, the launch included the mobile platforms Snapdragon 8 Elite Extreme Gen 6 for phones and Snapdragon Sound Elite Gen 2 for wearables.

Vinesh discussed how the chipsets came to be, including the special considerations made during their design such as improving connectivity, on-device support for large models, longer battery life, and other features crucial to smoothly running agentic AI applications. We also discussed what the future of a truly agentic AI phone looks like and what's been holding it back.

Topics covered:
• What an ideal agentic smartphone experience would look like 
• The demands agentic AI models make of mobile chipsets
• The obstacles to agentic solutions becoming a game changer 
• The crawl, run, walk phases of agentic solutions, and where we are now
• The role of other smart devices in creating agentic experiences 
• How support for a 30 billion MoE on-device model was made possible 
• Qualcomm's role in working with partners to bring AI experiences to life

If you want to learn more about how the latest chipsets will change the future of Android flagship devices in the next year, including new AI experiences, this conversation will give you a clear idea.

📺 Watch on YouTube: https://youtu.be/YXvw9Gw41Xg 
🎧 Listen in your favorite podcast player: 

Subscribe to Deep View Conversations for interviews with the leaders shaping the future of AI, business, and technology.

And don't forget to sign up for The Deep View daily newsletter. We don’t just cover AI, we decode it. In a world flooded with hype, we deliver sharp, no-nonsense insights to keep you ahead of the curve and help you put AI to work every day: subscribe.thedeepview.com

Disclaimer: Sabrina Ortiz's travel to Snapdragon Summit was paid for by Qualcomm. The Deep View's coverage is editorially independent from the companies we cover.

Creators and Guests

Host
Sabrina Ortiz
Senior Reporter at The Deep View

What is The Deep View: Conversations?

From frontier labs and enterprise platforms to emerging startups reshaping entire industries, The Deep View: Conversations podcast interviews the brightest minds and the most influential leaders in AI.

Sabrina Ortiz: Well, thank you, Vinesh, so much for joining us on this episode of the podcast. We are having a special episode today. We're live from Hawaii at Snapdragon Summit. We've had a lot of announcements in the past few days. And last time we spoke was a year ago. We talked about the future of mobile. That's been a big topic here. And I just want to kick off the podcast asking you how you think the role of the smartphone has evolved within the last year, particularly because of AI advancements in agentic AI?

Vinesh Sukumar: First of all, thanks for having me on your podcast. It's been a while. I think the last time we met was the Samsung event about 12 months ago. So nice to catch up in a nice place of Maui. Now, coming back to your question, I think we've kind of come to an interesting phase of edge AI computing, starting with a mobile device. Historically, when you look at mobile devices, there was a lot of focus around perception, related use cases like detection, classification, segmentation, around single modalities, which is still useful, still happening. But now the transition is now moving towards an appless environment, where we believe the voice is going to be a primary interface. And that is going to be driving not simple action, but compound actions. Whatever the thing that is, you don't use your phone to say, you know, switch on the light, turn off the phone kind of stuff. But you know, trying to see if it could be a lot more productive, can augment you as your digital twin, plan your vacations. If I'm late picking up kids, come up with a nice crispy answer, I suppose, to my wife. So those kind of personalized agentic experiences is what is going to be the next big thing. And we at Qualcomm are putting a lot of emphasis to really make sure, you know, we are successful in that front.

Sabrina Ortiz: So this is why I wanted to ask you this question to kick it off, because agentic experiences require a lot more from the smartphone than what it has typically had to do, right? It has to be running, basically have the compute to run for much longer periods of time to be able to power these kind of always on experiences. And it's to be connected more than it ever has been. Ideally be able to process more on device. You're not always experiencing the lag or of sending it to the cloud. And all of this ultimately on a smartphone comes down to the chip that could power it. And that's where Qualcomm comes in, right? So we'd love to hear a bit more about you, about the two big chipset releases that we saw yesterday, specifically for mobile, and then how building these chipsets and rolling them out were impacted by everything we just talked about, the new demands that AI poses on the hardware.

Vinesh Sukumar: Great question. So, as I was saying before, as you transition towards agentic experiences, about taking a lot more compound use cases, you know, context becomes critical. As always, keep having fun with my colleagues and with my family, these days, agents lack common sense. And if you want to make it a lot more personalized, context makes a big deal. And when you're trying to gather context, context can come from your browsing history, your responses, your preferences, your choices, all kind of stuff. So which means, you know, it's constantly gathering information behind the scenes. So which means that the most important element has the least amount of impact of the battery life. So the two platforms we mentioned yesterday, the 8 Elite Extreme Gen 6 and the 8 Elite Gen 6 is all focused on that, is all about token per joule or inference per second per watt, you know, are you able to really push the days of use. The second thing is, once you create these knowledge graphs, is it safe? Is it secure? That's very critical. And, you know, you don't want this information to be given out to anything, but to anybody out there. So we really try to make sure this information is secure, encrypted, and only on users' permission can be made accessible to the applications. So that way the user is fully in control of the private information about him or her. For that's another big thought process that we push into. And last but not the least is, you know, user experience matters. For any kind of request, I'm able to finish a certain request within a certain quality of service. And for that, you definitely have to invest in compute, you have to invest in bandwidth, and you have to really make sure that you're able to do several requests parallely at the same time. And that's where, you know, we have a Hexagon NPU that we presented yesterday, and we introduced the element-wise accelerator for the first time to really push for longer context length of about 32K. That is about 50 pages of dense text to really get this through. So I think, you know, in totality, as you transition more towards agentic AI, it's just not one thing. You've got to do a lot of things together to really make sure we can accomplish that. And that's what we're trying to do that here.

Sabrina Ortiz: One moment during the keynote yesterday where I literally bopped my head up from just, you know, taking notes on my laptop was when I heard the 30 billion parameter mix of experts model that can now run locally on phones. And for context, our audience probably already knows, but at the moment, the most we've seen running on like the most advanced phones are basically four billion parameter dense models. So we'd love to hear from you a bit about the advancements necessary to get the number so high and what exactly the mixture of experts model architecture means because it's a bit different than just your standard, just dense model.

Vinesh Sukumar: All right. So when you look at the model architecture, historically, for the last couple of years, we've been focusing on dense models. So which means that you have one model and then one model is, you know, very general in nature and is able to get one specific domain accomplished. In this case, let's assume a domain is a productivity domain. And then we have a lot of low ranking adapters, or what we call as LoRA adapters, which are fine to do certain tasks. So which means that, you know, you want to go with a suggestive reply for a certain text. I bring in a certain LoRA adapter. If I want to do any kind of summarizations of my messages, of my text, I bring in a certain adapter. So, you know, for specific tasks within a certain domain, we have these adapters. Now, in the last, you know, commercial release phones that we have done with Samsung and other partners, we have pushed these adapters to be around 25, 30 adapters. And at the same time, you know, these models can go maximum, let's say, three or four different languages. It's just not English. You want to speak Korean, you want to speak Mandarin, you want to speak French, it can do that on device. But there's now a lot more ecosystem pull to extend the number of domains, just not on productivity, but also on content creation, on content consumption, on video editing, expand beyond four languages to 12 or 13 languages. Now, we kind of hit a brick wall where in the four billion parameters, dense, can no longer sustain or support this stuff. So, we have to kind of think outside the box. And there's also been an evolution and deep learning architectures that people are using mixture of experts, where in, you know, just for your audience, if it's a 30 billion mixture of experts at any given point of time, we have certain experts which are always be active. There could be five experts, eight experts, 10 experts kind of stuff. Those experts are fine tuned to a specific task for a certain specific domain. And depending upon the task, you're able to get that specific expert pulled in and execute function. And what you've seen is, you know, the 30 billion MoE models now provide you the infrastructure and foundation to enable a lot more user experiences across multiple domains, across multiple tasks. And that's what we're pushing for as a next evolution of mobile computing.

Jason Hiner: Now, it's time for a word from this week's sponsor, Nooks. Most sales teams run outbound, inbound and deal management in three separate tools. A lead comes in Friday night. Nobody calls till Monday. A rep hangs up from a great call. The recap goes out on Thursday. Deals slip and nobody can say where. Nooks puts every motion on one execution layer. AI agents work alongside your reps. They get calls answered, get emails delivered, send the follow up the same hour and prep to secure the next step. While the reps do the selling. Teams on Nooks book three times more meetings, follow up 20% faster and sales reps get back 10 hours a week. Over 1800 revenue teams already run their go to market on Nooks. See what revenue agents can do at Nooks.ai slash Deep View. That's Nooks.ai slash Deep View. And we thank Nooks for their support of the Deep View. And now back to the show.

Sabrina Ortiz: As you mentioned, AI for different tasks that also reminds me of something that I always love being reminded at specifically at Snapdragon Summit or during any of Qualcomm's announcements, it's that the chipset and even AI goes far beyond just your agent experiences, which although that is currently the most cutting edge and at the frontier, AI is also used, or at least you spoke about yesterday, to do things like improving your sound experiences, like isolating ambient noise and which helps just with capture, but also your calls. It helps with camera enhancing the quality of your captures and then also how it reads the information so that you could do more comprehensive edits, all of that. Curious again, how you're able to build these experiences. Not only are you thinking about how to improve AI, like agentic experiences, but also how to use AI to kind of transform the entire user experience throughout the phone too.

Vinesh Sukumar: Yeah. So when at Qualcomm, we have been enabling AI for the last 15 years. And our journey in AI started off with perception use cases I mentioned before around cam. It actually started with audio and then transition to camera and into video. And when you start looking at audio, you have a lot of these interactions with our ecosystem partners, with our OEMs, with our developers, with our consumers, and we tried to understand exactly what are some of the pain points. And we often found out that it's about echo cancellation, background noise suppression, or if you're trying to look at any kind of language translation, but all that is all doable once you understand the problem statement. So that's what we put a lot more focus to really understand is once we understand the pain points, how do we look at system innovations? How do we look at hardware architecture innovations to really get to that point? Now, once we have it, then there's a lot of system constraints. Hey, you know, you cannot impact days of use. Can I process these functions with a very high accuracy kind of stuff? So that's when we put a lot more energy into our algorithms, into our software and get this going. And it's just on the audio piece. We do something very similar on the camera front, on the video front as well, because they're very similar. And one of the things which is still is in the research phase, not yet resolved is I don't know if you've seen this, but for me, you know, I use AI to take, you know, transcription, you know, especially in the meetings. And often in my meetings, I have six people, seven people talk over each other. And my meeting transcription does not really recognize when you have seven people talk over each other. How can AI isolate these seven different speakers and make sure that, you know, it is the transcription notes is mapped to that specific speaker? So we still have some problems to be solved. So this is one such example. We'll see. That's what, you know, it's a challenge and that's keep the worlds of AI more exciting, I suppose.

Sabrina Ortiz: You're also building this platform, but the execution of what features are brought to the new hardware devices are ultimately up to, I guess, how the customer, which the customer can include companies like Samsung or many other OEMs, how they decide to, I guess, use what the powers and the capabilities that you have enabled, but then also build on them. How do you work with your clients to make sure that they're making the most out of all of these new productivity or optimizations, power enhancements that you've made enabled to fit in the new chip?

Vinesh Sukumar: It's a great question. So when you look at AI, one of the biggest challenges has always been as what is to be expected next, you know, what is the next big thing you want to resolve? And the way we have kind of looked at this is we have a very strong partnership with our ecosystem players, either could be on the frontier labs or making models with our developer community, with independent software vendors or with OEMs. And especially with quite savvy OEMs like Samsung or Vivo or Xiaomi or Honor, we have a very strong partnership where in the upfront kind of share a vision statement on what is to be expected for the next 12 months. And as part of it, we put a lot of joint effort in making sure we get to that point. Now, it is not to say that if you talk about 10 different experiences, all 10 go into production. It's quite possible. Only two or three of them go in, but we put in the work to really get that done. So over the years, you have maintained a very strong relationship to understand, you know, what is to be expected. Now, obviously, there's something new that comes up and there's always going to be, hey, can you get this done in the next 30 days? That's always going to be scenarios like this. But we try to avoid them as much as possible. That's only possible if you continue to maintain this deep collaboration with all of our partners to understand what problems you're trying to solve. How can we improve on UX using AI?

Sabrina Ortiz: Got it. And with that, I have to ask a bit of a, you know, harder question, but we've been hearing about agentic experiences in smartphones. We've been, for a couple of years now, I've been reviewing basically every phone that's come out since, you know, the agentic era has started. And while we've seen glimpses of it, some really interesting features, which actually I'm sure you've worked with partners to enable, like, um, uh, Samsung phones or Google's phones have like the, um, task automation features, which is somewhat like, you know, the most futuristic agentic experiences I've seen yet on a phone. I've yet to reach a point where I use my phone every day and I feel like there is this really new way that I can interact with it to get things done and that I could just, you know, talk to it and it has this context of me and I barely have to type anything and it just understands what I'm requesting. And I think that's what people want when they hear agentic experiences. What's the holdup? What's the obstacle? Cause every year again, I always sit here and I'm like, that sounds awesome. And yet it doesn't seem to come to my phone just yet. So.

Vinesh Sukumar: I think, you know, um, it's a, it's a fair observation. Usually I would put it this way. It usually takes a village to get to that epitome of success. And when I mentioned by saying that as a, uh, do we happen to have these small language models which are capable enough, uh, to do these functions on device? So that's one thing. We have seen an evolution of small language models, which are going better and better, you know, a year over year. The second thing is, uh, agentic experiences by definition has to involve cloud. It cannot be completely done on device. So there has to be this hybrid transition path, but hybrid transition path has been very static in nature. I mean, it's predefined. I mean, for example, you asked something on psychology, I suppose it goes to the cloud, but if you want to do some simple tasks, it gets done on device, but you want to have these harness layer, as they call it, wherein it can understand the intent of the user. And depending on the context is able to make a switch automatically between the cloud and the edge that is still in research phase. It is not taken off to an extent where people can say, man, this harness has a brain of its own. It's able to decide, you know, where these models need to run. So I think we have some homework to do there. Then I talked about the elements of personalization because without context, most of these, uh, you know, agentic applications have these multi-turn conversations, which never take off. And, uh, we're kind of building that infrastructure layer to really get to that point of establishment. And, uh, last but not the least is, um, you know, user data across applications is very private. How can you make applications from one domain interact with other? It could be either an OEM based applications, could be, uh, Google based applications, could be Meta based applications, could be Microsoft based applications, are they willing to share information across domains? Um, and that, you know, kind of gets into user privacy, kind of gets into, uh, um, you know, elements where you need to have certain number of negotiations across players that this information is safe and secure, is interchangeable. So we have some ways to go, but I think you're making progress. Uh, every day is a good thing, uh, because you're seeing a lot of traction, uh, from just not Qualcomm, but players around Qualcomm to really get to that next phase. So again, um, you know, um, planning for success. I'm hoping this year, you know, you're going to see a lot more movement come in that of last year.

Jason Hiner: Hey everybody, thanks for listening to this episode. Quick note, and then we'll get you back to the conversation. We love bringing you this content every week. We're always trying to figure out how we can deliver you the most value to help you understand how AI is transforming business and transforming the world. If you're enjoying the show, there's an easy way for you to give a little value back and help others learn about the show as well. If you're on Apple Podcasts or Spotify, drop us a rating and leave us a review. And if you're on YouTube, hit like, subscribe, or leave us a comment. That's it. It only takes a minute, but it's a huge help. So thanks in advance for pitching in. And now back to the show.

Sabrina Ortiz: When do you think that vision of that truly agentic experience where you just seamlessly get to talk to your phone will be truly possible both from the infrastructure part and also just from the actual, I guess, research and model developments that have to be like layered on top of that to make this a reality.

Vinesh Sukumar: I think when you look at agent experience, by definition, I will always say there is a crawl walk run phase. So we are transitioning from crawl to walk, I would suppose. And as you start to deploy a lot more into commercial production with consumers with enterprise, we learn a lot of new things. And as you start putting in commercial production, these feedback from our consumers from our enterprises give us new challenges. So I expect every year we'll make progress. And I'm hoping we get into the run phase ASAP, but that's my blue sky vision. Hopefully we'll get there as soon as possible.

Sabrina Ortiz: So we're still from what you're observing, we're still in the crawl to walk transition.

Vinesh Sukumar: That is correct.

Sabrina Ortiz: Okay, got it. That's interesting. With that, I think part of the bigger vision, if you pan out even a bit further, and I actually really enjoyed seeing this in some of the more conceptual demos during the keynotes, there'll be a world in where your phone and your smart glasses and your watch, maybe your ring, who knows, maybe pins actually become a thing that are more viable than they are now, kind of interact seamlessly. So again, you could spend less time on screen, more time just talking and observing and world around you and asking your devices to handle tasks for you. What's the holdup currently preventing that from being more seamless than it is now? And I have multiple smart glasses, I love them, but again, it's not quite there yet, as what we were seeing in these demos of what could happen one day.

Vinesh Sukumar: Yeah, so the thing is, for us, the vision is clear. Qualcomm is one of the only players that participates from doorbells to data centers, which means from milliwatt to gigawatt. So we have a breadth of form factors and devices we totally participate in, and our expectation is each of the devices has a certain context about the user, you know, a digital twin of the user. Can I get all of them connected? Now, to really make this happen is, hey, do we happen to have the inter-device communication infrastructure set up? We at Qualcomm are getting that done. So that's one major rock. The second question kind of becomes, okay, now that you are able to share context, where do I host them? Does it get hosted on the cloud, or does it get hosted on a certain device that you think is your master? And if that, so there's always this debate on if that master device gets stolen or lost, your entire context is lost. For sure, I'll be having a backup on the cloud. So there's always been this debate on, you know, can I get that stuff done? And the third most important thing is, you know, when I transition outside Snapdragon Garden of Products, I don't imagine why people will not use Snapdragon, but any other devices, how do I look at, you know, non-Qualcomm, non-Snapdragon devices? Will that context be shareable if it's within the user's, you know, ownership? That needs to be resolved. So I suspect, again, these are, you know, as I mentioned, it's always like Swiss cheese with a lot of holes. You've got to fill one hole at a time. And then, you know, you enjoy the ride because these are good challenges to have, and eventually we'll get to that goalpost.

Sabrina Ortiz: And with that in mind too, and I don't know, our listeners might not even know this, but Qualcomm, specifically with, you know, powering the AI wearable segment, is doing a pretty darn good job. Meta Ray-Bans, which are the, you know, highest selling smart glasses currently on the market, are powered by Snapdragon chipsets. The most cutting edge XR and AR wearables, for instance, last week I was at the Snap Specs launch, and they're here today. Those are powered by Snapdragon too, XREAL powered by Snapdragon. So you're positioned in a pretty good, you're in a good spot right now for at least that segment. Why do you think Qualcomm was particularly ready to kind of hop into a new product kind of space and lead it, essentially? It is almost like you've captured that market at the moment.

Vinesh Sukumar: I think our CEO Cristiano has laid out a strong vision. He truly believes that edge AI and edge AI computing is the next big milestone we want to really accomplish. We have been in the space of AI for quite some time, especially from an inferencing standpoint, and doing inferencing using mobile as your primary foundation. And when you participate in the mobile space, you have to look at multi-modality. You have to look at concurrency. You have to look at power efficiency, because all these are, you know, quite critical to our mobile partners. And once you happen to have the kind of experience, it becomes quite more easier to kind of transition to hearables, to wearables, to PC platforms, because you're able to gain that knowledge, and then you're trying to adapt to that form factor, and then, you know, try to supplement that with more use cases for that specific streamline. So that's been the foundation, you know, so far. It's been spectacular success, as you mentioned. But, you know, there's a lot more things to be solved, especially in the software side, on cross-device orchestration standpoint, and we'll get there.

Sabrina Ortiz: And you're still also diversifying the portfolio further. It's worth mentioning that today's announcement actually was a new chipset that is powering still AI wearables, but less so things like Meta Ray-Bans with a camera. Those are still the chips that's been being used. But this is more to power, like, audio experiences, giving context to your wearables. Was that also, I found it super timely, because, you know, people are, we've been talking about privacy a bit, particularly with the smart glasses. A lot of people are, you know, hesitant about surveillance, nonconsensual, you know, captures. So I think there'll be a sharp rise in smart glasses and other wearables that either don't have cameras or have cameras just for context and not for content capture. Today's chipset was perfectly almost timed with that. Yeah, love to hear more about that release and if that was also a role, just kind of taking consumer sentiment into account for a bit there too.

Vinesh Sukumar: Yeah, absolutely. As you're transitioning towards an agentic era, one of the earlier points I mentioned is all about context. You know, are you able to see things around you, hear things around you, and then can I make an intelligent graph about that? And it's just not going to be based on one device. You know, it'll be based off your glasses, based off your hearables around it. And you also want these devices to be a little bit more intelligent. For example, if you happen to have your buds, I need to be touching on the buds, or if you happen to have an app that's connected to your phone, you have to let them know that you're in a certain environment. But in today's keynote, you know, there was a mention wherein the buds are now going to come with cameras, wherein it can understand that you are, you know, in a noisy environment, or you're working, you're walking on a pedestrian footpath, so you have to be cautious around the environment around you. So it automatically adjusts things based off your surroundings. So you don't, it doesn't have to be human in the loop at all times to make the decision for you. So that's what you're transitioning towards. And it's one of those stepping stones wherein we continue to build on context, and this is one such great example.

Sabrina Ortiz: We're getting close to the end of our time here. So before we wrap up, I did want to ask you a funner forward-looking question. We talked about when agentic AI hits that run phase between your phone and the ecosystem of your devices, what is one thing that you are the most excited to actually become a reality, and that you think will be able to materialize? And let's just say, let's keep it a little bit more short-sighted. In the next two to five years, something that would really be a game changer for you and you think will be actually possible.

Vinesh Sukumar: You're making me think now. I think one of the things I passionately want to get this stuff done, and I think it'll happen, is can I have reinforcement learning done on device? What I'm saying that is if you happen to have an agent, agent is giving you a response. You don't like the response. Maybe after 24 hours, about a certain period of time, you ask the agent the same question. It better learn from the previous interaction and improve on itself. You know, that is difficult. That is hard. But at Qualcomm, we're putting all these investments to really make sure we can get this stuff done. Today, most of the reinforcement learning happens on the cloud with a lot of feedback from the user. The data is labeled and it self-learns every 24 hours to get to that point. Nobody has really accomplished that on endpoint devices. I totally feel we can get it done because we got the talent, we got the technical acumen, we got the infrastructure, and definitely a hard-pushing CEO. So I'm pretty sure we'll get there sooner than two years.

Sabrina Ortiz: All right. Well, with that super ambitious end-note, I think that's a great place to stop the podcast. Thank you so much for your time today, Vinesh.

Vinesh Sukumar: Absolutely.

Sabrina Ortiz: I hope we get to talk a year again, a year out from now, and there'll be even more to talk about, more advancements in the world of mobile. Who knows? Maybe they'll look entirely different by then.

Vinesh Sukumar: Absolutely. The world of AI is changing so rapidly.

Sabrina Ortiz: Exactly. All right. Thank you so much.

Vinesh Sukumar: Thank you.