Speak to an Agent explores what it really takes to bring AI into large organizations, why AI can't fix a problem you haven't already solved, why real understanding matters more than a polished voice, and why the domain expert who knows how to use AI is becoming the most valuable person in the business.
Featuring the people and ideas shaping the future of CX AI, the innovators, the thought leaders, and those making the decisions. They get candid about what's working, what isn't, and what changes now that machines can reason, plan, and act at a human level, and sometimes beyond it. If you're responsible for putting AI to work inside your organization, this is a conversation you won't want to miss.
In the next six months, there are gonna be real time speech LMs that are gonna be much smarter and are gonna solve a lot of these problems.
Neeraj:Welcome to Speak to an Agent. Words used to mean one thing, get me a human. Now the agent people actually want is AI. It's faster, sharper, and never puts you on hold. If that surprises you, you're exactly who this show is for.
Neeraj:The builders, the decision makers, the implementers. Comparing notes while the next generation of technology gets built. You're not just listening. You're in the room. I'm your host, Neeraj Verma.
Neeraj:Joining me today is Dylan Fox, founder and CEO of Assembly AI, the speech AI company powering transcription and voice intelligence for thousands of developers. Let's get into it. Voice was sort of destined to die. Right? Five years ago, six years ago when I was in this industry, it was all around text chat agents.
Neeraj:You remember this. Right? Yeah. And it was it was like, voice is kind of this old thing that's reserved for human agents, and now it's completely reversed. Right?
Neeraj:AI agents are they're out. They're reliable. Voice technology is growing. How have you seen that trend over the last three to four years?
Dylan:Yeah. I mean, it's really, you know, it feels like there's been this overnight shift, but in reality, it's been, you know, just like years of the technology getting better and better. But I can tell you now, I mean, we all feel it, right? As a consumer of technology and as a user of technology, as soon as you have this experience with an agent that's actually helpful and really fast, that's actually like what you prefer to use. Right.
Dylan:And I think about even over text, a couple of years ago, if you would talk to a support agent, it's just frustratingly bad. You're like, oh my God, I need to like escape this matrix like immediately. But now I look for, you know, if I'm ever looking at a new tool or something, look for the, Hey, let me talk to the AI agent here because it's gonna give me good answers. It's gonna be super fast. I can interrogate it, you know, on any question I have.
Dylan:And it's going to be super helpful. And that's what I default to and what I use. And I've even found this, like, you know, we have an API platform and we now have this like ask AI feature over our API documentation. And it can write you code. It can switch the Python examples to C sharp if you want.
Dylan:And people prefer that because it's a tool. It's helping them. They prefer that over our You know, we have a support engineering team and the support engineering team, you know, we find like when people are doing long sessions with our agent, it's because they're wanting like hands on help with how they're writing code or how they're like understanding the product. And so when I think about voice, you know, as a input to an agent only recently has gotten, you know, fast enough, accurate enough, reliable enough where, okay, you can build those habits as a consumer of like, I wanna use this thing. And it's just the beginning.
Dylan:You know, I last night, I was driving around San Francisco with my kids looking for a restaurant. And I call a restaurant and I talking to a voice agent. And I know it's a voice agent because, you know, I know it's a voice agent. And I'm like, hey, do you have a kids menu? And then my kid says something in the background and I have to turn around and ask my son to be quiet for a second.
Dylan:And the voice agent just totally breaks. Right? And I know why and I hang up and I'm like, okay, we can't go there. But that's an example of like, there's still so much room for improvement. We've crossed this threshold where, yeah, it's working.
Dylan:And for a lot of use cases and a lot of, for a lot of end users, it's a great experience.
Neeraj:You know, it's interesting. I think that the use case that you mentioned previously around the bottom of documentation website, it feels almost like people are used to the Cloud AI and sort of ChatGPT experience. Yeah. But it's a long tail experience. Right?
Neeraj:You're expecting when you type something in, it's gonna do its research, and then it starts typing something out. It's a long tail experience. I think voice is such a nuanced case. Right? That the average person only has two seconds.
Neeraj:Right? The average agent Yeah. AI agent has two seconds to listen to your voice, to transcribe your voice, to understand your voice, to call tools, and then to write it out and do TTS. It's a it's a really weird space. I think maturity is coming.
Dylan:Yeah. But it's clear to
Neeraj:me as I talk to more and more voice agents that it's close, but just not, not quite there yet.
Dylan:Yeah. There's some things around the voice UX that still are clunky, right? Like you don't have to do all the reasoning in that loop. You can, like when you're talking to a human and there's a long running tool call that the human has to go make, it will still talk to you while it's running that tool call. It will, I realize I'm describing human as it.
Dylan:The person you're talking to will start asking you like, how's your day going? And or hold on, let me put you on hold for two minutes while I go look this up. And then, you know, okay, something's happening. I need to just wait and be patient. And we're still building those experiences into voice agents.
Dylan:And I think like there's there's still some UX clunkiness that needs to be solved, but like, yeah, it will it will get there pretty quickly.
Neeraj:It's, you know, it's interesting because you get this concept of uncanny valley. Yeah. I'm okay with you asking me how the day is going, but if Claude asks me how my day is going, I get a
Dylan:little creeped out. I'm like, I'm
Neeraj:not really interested in that. Right.
Dylan:Because you still know, like, is This is not a human. This is a machine. Right? I'm like in command of this machine. And this, you know, the yeah.
Dylan:There's a dynamic here. Right. And so I think it's maybe not how's your day going, but it's okay. Hey. Hang on
Neeraj:for a few minutes.
Dylan:Right. I'm looking that up right now. It might take a few minutes. Give me a second here. And then if you ask something back like, hey, are you still there?
Dylan:It's like, yep. I'm just still waiting on this verification. There's some clunkiness there where, you know, because the voice agent stack is today mostly built from like all these independent technologies. Right? You have voice activity detection, you have noise cancellation, you have speech to text, maybe you have a turn detection model, you have an LM, you have a tech like it's and none of them talk to each other and are wired together.
Dylan:So you have this like very like Frankenstein almost thing that you're trying to trying to make sound human. And that's hard today still. So I think as parts of those technologies start getting kind of merged together, it'll become easier. But yeah, those are some of the Voice UX challenges that I think once you run into those, you're like, okay, I wanna bail out of this. And those will be solved, you know, quickly.
Neeraj:So so you talked about sort of and I agree by the way that that enterprise voice orchestration is a conglomeration of technologies Right. That have latencies between them that add up over time. Right? Right. It's clear.
Neeraj:So when you think about all those conglomerations of technologies, how do you feel about voice to voice models essentially compressing the entire stack into one thing? Know Google is a really big advocate of sort of using their new Gemini flash 3.1, which is the flagship voice to voice model. And it sounds amazing.
Dylan:Yeah. So our our opinion is like there are gonna be stages of this. So yeah, so speech to speech model that collapses all these tasks into one model, you solve a lot of these problems that we're talking about, but you introduce new ones. And in the meantime, you know, as we've talked about, even with as clunky as a Cascade and architecture is right now, you can package that up and go deliver amazing customer value. I mean, you guys are seeing this, right, up We're talking about a lot of the pitfalls of a voice agent right now, but you can have a resolution rate of, you know, 50% with a voice agent in the real world.
Dylan:And that's amazing. I mean, it could even be even higher for just simple, like the resolution rate could be even higher for simple use cases. And so there's real value there with the existing technology. And I think where it's difficult is like, how do you timing's really important. If you're too near term invested, you struggle.
Dylan:And if you're too long term invested, you struggle. Right? And so, like if you guys were to just say, hey, we're gonna just go all in on a voice to voice model right now. The voice agents that you're building would probably struggle in the market. Whereas if you're still on a fully cascaded solution six to nine months from now, you're also gonna struggle.
Dylan:So that timing is really important. You know, what I think about right now is like where you'll definitely see things collapse is like in the front end. And so what I mean by that is right now you have, you know, usually like noise cancellation, a voice activity detection model, speech to text, turn detection. Turn detection, yeah. Right?
Dylan:And that's a lot to figure out like, is someone talking? Right. You sometimes have voice focusing models so you can ignore the background speaker or a TV on in the background. So you don't respond to that. You just respond to the primary speaker.
Dylan:You know, real time speech LLM models can can do a lot of those tasks at the front end. And so imagine just a stream of tokens like, person started talking. This is what they said. Child in the background started talking. This is what they said.
Dylan:TV's on. This is what it said. You have this real time stream of like intelligence at the front end. That that will happen. I think that will help a lot
Neeraj:with this It's more like a speech intelligence model than this
Dylan:speech recognition model. Because speech to text models are still just doing, you know, like a data processing tasks. They're kind of dumb. But a speech LM, a real time speech LM model is smarter. It can know, it can understand, okay, that's a TV and this is what the TV is saying versus like, this is the person, the primary speaker that, you know, this is the primary speaker.
Dylan:And then you as the agent builder can leverage that information on like, okay, I'm gonna ignore everything that's not the primary speaker. You have all And that that's gonna help a lot with these like, you know, a lot of these clunky experiences are still around like interruptions, turn taking. The agent thinks it's being interrupted, but really it's just someone in the background. That's really where there's like clunkiness. So how do you feel about Actually, feel like people, and you tell me what you think, but I feel like users are a bit more accommodating of overall turn latency than they are of that front end interruption.
Dylan:That's like, oh my god.
Neeraj:I I I We find that on average, if your response in activity goes beyond two seconds Yeah. Users get very frustrated.
Dylan:Yeah. Two two seconds for sure. Over two seconds is a
Neeraj:long time. But you know, it's like two seconds is nothing. Right? You and I are talking and it's I'm sure I'm pausing for more than two seconds. Right?
Neeraj:But so how do you feel about Whisperesque models? Right? Whisperer is a LLM. Right? That's obviously a by the way, I used to write speech to text code for a living.
Neeraj:Was a research scientist way back in the day, seventeen years ago.
Dylan:Yeah. So Whisper is, you know, a sequence to sequence model, but still the decoder is not like, when I say a speech LLM, I mean like, do you have a LLM decoder? Lot like a lot like think of like Lama or something. Right? And you teach it, audio tasks.
Dylan:You train it to be really good at audio tasks and how to be really good at that modality. And you kinda get rid of some of the intelligence you don't need anymore. I don't need this thing to summarize anymore. Right. I don't need this thing to write markdown files.
Dylan:I need it to, I'm going to focus all the learning power on these audio tasks. And you kind of get this like audio general intelligence type model. And so for example, with our universal three pro model, you can give it context on like, Hey, this is a, like it, you can prompt the model, you can give it context. You can say, Hey, this is a doctor's office, dentist's office. And so when someone calls in and say, Hey, my, I've, I've some pain in my crown, you know, or whatever, I need a new crown or something.
Dylan:If the model is confused, like, did they say crayon or a crown? Of course. Right? That context helps the model. Oh, it's it's a this is a dental office.
Dylan:So person's probably probably saying crown, not crayon. So you can already do that today with the models that we're making. And they can also tag audio descriptions. And now we're working on making those like work in really low latency real time. Interesting.
Dylan:So all that at the front end That's
Neeraj:super valuable.
Dylan:Super valuable.
Neeraj:Super valuable.
Dylan:What we're also making these models do is be able to back channel with the TTS models. So every STT today has no clue what the TTS said, right? So if your voice agent says, hey, do you want to reschedule or cancel? And then the person says, you know, in a muffled way, you know, cancel or something. If you know that the agent just said, do you want to reschedule or cancel?
Dylan:That's a strong prior that can help you at inference time understanding that single short word in a noisy environment or something. I do ultimately think a real time voice to voice model as a front end, you know, is is definitely like where things will go. But you'll also wanna have a, you know, a real
Neeraj:TTS at the end.
Dylan:Stronger reasoning model that that can work with that. Yeah. So you basically think about it as like two agents. Interesting. And you have a voice to voice model that's really just like a front end.
Dylan:But, you know, I'll give you an example. Like our daughter had had a stomach bug and she had to go to the ER. And because she's she's three and, you know, she like got super dehydrated. We got a bill in the mail that the ER visit wasn't covered. We had like a $20,000 bill in the mail.
Dylan:And so we call into the insurance company and they're like, hold on, let me check that. Let me check into that. Like a couple minutes later they come back and they're like, okay, actually that was a billing error. You're all good. You know?
Dylan:Okay, bye. You there's a lot of reasoning that happened there. They had to look into policy information. Right?
Neeraj:Human orchestration.
Dylan:Yeah. A lot of reasoning that happened happened there. So it should be fine for a voice to voice model to kick out to a bigger, smarter model to say, hey, this person's calling this is what they're calling in about. I need you to go look into this. I'm gonna keep chatting with a human while you're doing that and then let me know.
Dylan:And then, okay, I, you know, that, that, that
Neeraj:reasoning's That's almost like this A to A concept, right? We always throw around. Right?
Dylan:A to A?
Neeraj:Yeah. Yeah. Agent agent concept. Right? This concept of agent agent.
Neeraj:Right?
Dylan:Right. Right. And we're already doing this with some of our support agents. Like we have a, you know, agent that has a lot of reasoning time and access to systems. Then we have an agent that can talk to customers that doesn't, but can go talk to that other smarter agent.
Dylan:So this is where I think you'll see That's really interesting. The concept
Neeraj:is really interesting. I think, you know, we we've been using Whisper a lot internally, obviously. Right? I'm sure everybody else has been too. Yep.
Neeraj:And we find this sort of prompting concept interesting. Right? We use it all the time and real time to do prompting. And Whisper what is Whisper? Whisper system GPT decoder with GPT two on top.
Neeraj:Works good enough. Are you finding one of the big problems with Whisper like models is that out of vocabulary phrases have to be essentially prompted in. Yeah. And they're not as reliable as a traditional sort of trans you know, traditional older style transformer model. Yeah.
Neeraj:How do you guys deal with that?
Dylan:So out of vocabulary is definitely less of an issue with the latest state of the art speech LM type architecture. The So it's definitely less of a problem. Where there are still problems are like very rare terms. So like think about what's rare even to GPT-five. And and that's like, okay, what's the frequency of of of that word just in the training data of a model, even a big one?
Dylan:Also rare words that require context. Right? So like our names require context. Right. There's no way you're gonna get our names right.
Dylan:I mean, what if my name was spelled not d l d y l a n, but d y l o n? I mean, some people spell their name or d I l l o n. Like, you have to have that prior context. And so I think that, you know, sometimes you have, like, you know the words ahead of time as context, but sometimes you don't. So you need models that can take in like hints of context.
Dylan:This is what we're building. And that then helps them with those paths. So like imagine, right? Like an LLM that was just trained knows my name and AssemblyAI, right? And so if you say, oh, spell the founder of AssemblyAI's name, it's going to be able to do that correctly.
Dylan:Maybe even, you know, like, like, maybe even some someone that just based on their LinkedIn profile or something, right, if it has that information. And so I think that's where you need to get to, like, smarter models that you can give context to that are gonna be able to reason a bit more about what they're doing and not just be like these these like classical like data. Because I still view Whisper as like a classical ML model where like it's definitely smarter than a, you know, RNNT model or something, but it's not, you know, like an LM there. You know? It's not yeah.
Neeraj:It's a tiny little SLM. It's not definitely not smart, but you still see hallucinations, right? Yeah, exactly. I think that's the bigger problem.
Dylan:That's a big issue. Yeah, totally. And so, you know, when you when you move to a model like that, even, you know, and this is something that, you know, we've had to deal with like these these speech LMs, they always want to predict more tokens, even like a, like a sequence sequence model, like whisper. And so you really have to focus on that during training. We spent a lot of time trying to train the models to not hallucinate when there's silence or background noise or Mhmm.
Dylan:If they're uncertain about something and and there's a lot of work in doing that.
Neeraj:And the silence hallucination has been just ridiculous. So Yeah. So when you look at the market and the industry you're in, it's clear that you guys have a lot of competition from Azure, Google, the typical hyperscalers. Right? I I'd argue the hyperscaler ASRs are still generation behind at best, but you've got other sort of newer age vendors like eleven Labs that even have TTS.
Neeraj:Are you guys thinking about how you fill the entire sort of agent stack?
Dylan:Yeah. Yeah. So where we're really focused is infrastructure. You know, so we're like a 100% at the infrastructure layer. And that's where we're excited about building and where we think there's a lot of value to be building, solving problems like scalability, security guardrails, all the primitives that you need to go have really strong voice capabilities in products or in workflows.
Dylan:And for us, that means, you know, offering all the technology stack you need at that voice infrastructure layer. We're real, we've been really focused on the listening part of the stack because all these problems that we talked about are parts at the listening problems at the listening part of the stack. There's a lot of edge cases with TTS on like, okay, acronyms, whatever dollar amounts, hallucinations. Like hallucinations and a TTS are so
Neeraj:They're funny too though.
Dylan:Okay. They're funny. Oh my god. Yeah. I was yeah.
Dylan:They're they're everyone has those funny examples. But a lot of it too is like vibe. Right? Like what what do you want the voice agent to sound like when you're calling into a restaurant? It's pretty different than a in like a media application.
Dylan:Right. So there's a lot around like vibe, I think, that's really important at the voice front in addition to the reliability and ability for it to match the vibe of the conversation and not be like unaware of the sentiment of the of the of the user is talking to.
Neeraj:Do you talk to Grok?
Dylan:A little bit. Yeah.
Neeraj:So so you know, it's interesting. My brother-in-law actually has a has a Tesla and you can talk to Grok on the Tesla. It will map to your not only to your voice, it'll map directly into your sentiment and how you're talking. It's amazing. Mhmm.
Neeraj:And it's super creepy. Super creepy. You can see the tone of the conversation change over time. It's not exactly what I want. Yeah.
Neeraj:But have you also seen the voice interpreter feature on Teams? No. If you want creepy, holy that is the creepiest thing you'll ever see. It is so the voice interpreter is like, and we were trying this earlier, you know, I've got a bunch of colleagues in Germany and I I run up that division now, and we were trying to translate between German and English. Not only will it translate between the languages, in that tiny little one second voice sample of me saying, hey, my name is Neeraj, it creates a TTS of your voice.
Neeraj:Yeah. And it sounds incredible in the foreign language. Yeah. And I I'm I'm dumbfounded at how amazing that is.
Dylan:Yeah. Yeah. The voice cloning is getting good. I mean, also think there's these like it's an interesting moment, right, where, you know, something I think I think about sometimes is there's this scene in the the new Dune movie, the first one, where Paul is learning about Arrakis on this, like, computer thing. And it's a projector, and it's talking, and it sounds very robotic.
Dylan:But in this, like, futuristic way.
Neeraj:That's it for this episode of Speak to an Agent. Big thanks to Dylan for giving us the real story on Voice AI. If you found the conversation valuable, subscribe wherever you listen and leave us a rating or review. It helps more listeners discover the show. We'll be back soon with more of the people making the AI agent the one you actually wanna talk to.
Neeraj:Thanks for listening.