1
00:00:03,000 --> 00:00:05,920
Evan: Welcome to Daily Paper Cast.

2
00:00:05,976 --> 00:00:14,896
Ashley: Today, we're discussing a paper from the Hugging Face daily paper list of September 16, 2026, which received 21 upvotes.

3
00:00:14,952 --> 00:00:19,812
Evan: The paper is titled 'StepAudio 3 Realtime Technical Report'.

4
00:00:19,872 --> 00:00:27,552
Ashley: The first two authors are Bin Lin and Bo Zhao, and the corresponding author is Boyang Zhang from the StepFun-Audio Team.

5
00:00:27,600 --> 00:00:30,240
Evan: Let's dive into the introduction of the paper.

6
00:00:30,288 --> 00:00:36,278
Ashley: Natural spoken interaction requires a system to follow the user while managing its own response.

7
00:00:36,278 --> 00:00:43,088
This interaction can be quite complex due to the need to handle pauses, acknowledgments, and interruptions fluidly.

8
00:00:43,152 --> 00:00:52,142
Evan: Right, and this complexity is compounded when dealing with multi-turn requests, where the model must reason carefully while keeping the conversation responsive.

9
00:00:52,142 --> 00:00:59,312
Tool use introduces an additional challenge since tasks initiated by spoken exchanges might outlast the dialogue itself.

10
00:00:59,426 --> 00:01:07,046
Ashley: Advances in speech recognition have leveraged acoustic representations combined with the linguistic knowledge of large language models.

11
00:01:07,046 --> 00:01:12,636
Audio-language models now support broader acoustic understanding and direct speech generation.

12
00:01:12,696 --> 00:01:20,566
Evan: Indeed, streaming and full-duplex systems have become more sophisticated, allowing listening and speaking to overlap.

13
00:01:20,566 --> 00:01:26,396
These capabilities consider linguistic content, vocal delivery, and conversational timing.

14
00:01:26,448 --> 00:01:32,278
Ashley: StepAudio 3 Realtime builds on the StepAudio series’ shared audio-language foundation.

15
00:01:32,278 --> 00:01:40,548
The focus of StepAudio 3 Realtime is on the seamless coordination of perception, reasoning, and action as a conversation unfolds.

16
00:01:40,608 --> 00:01:41,878
Evan: Interesting.

17
00:01:41,878 --> 00:01:44,508
How do they organize these core functions?

18
00:01:44,568 --> 00:01:50,578
Ashley: They organize them into a continuous loop of listening, conversing, thinking, and acting.

19
00:01:50,578 --> 00:01:55,468
Deep Perception is designed to capture linguistic and nonverbal acoustic evidence.

20
00:01:55,468 --> 00:02:01,208
Seamless Duplex utilizes both user and model speech to manage the conversational floor.

21
00:02:01,322 --> 00:02:10,552
Evan: Think-While-Speaking, as mentioned in the abstract, coordinates reasoning with spoken delivery, supported by Adaptive Thinking and multi-token prediction.

22
00:02:10,552 --> 00:02:13,132
How does the voice agent fit into this mix?

23
00:02:13,176 --> 00:02:17,416
Ashley: The Voice Agent extends the interaction to tasks that require tools.

24
00:02:17,416 --> 00:02:25,606
It processes the request and required arguments before execution, then incorporates returned evidence into the ongoing conversation.

25
00:02:25,606 --> 00:02:30,336
This means users can continue speaking while a task progresses in the background.

26
00:02:30,384 --> 00:02:32,764
Evan: That sounds incredibly efficient.

27
00:02:32,764 --> 00:02:37,344
So, what kind of performance has StepAudio 3 Realtime achieved?

28
00:02:37,392 --> 00:02:49,312
Ashley: StepAudio 3 Realtime leads the reported baselines on four out of eight audio-understanding benchmarks and achieves the highest reported Overall score on the Artificial Analysis Full-Duplex Bench.

29
00:02:49,312 --> 00:02:57,052
The results identify some gaps in multi-turn constraint following and retail tool-use tasks, which are discussed in later sections.

30
00:02:57,146 --> 00:03:05,336
Evan: It seems like the introduction sets a strong foundation for understanding the model's capabilities and its real-time interaction potential.

31
00:03:05,400 --> 00:03:08,520
Ashley: This concludes the introduction section of the paper.

32
00:03:09,830 --> 00:03:14,830
Evan: Alright Ashley, let's dive into the methods used in StepAudio 3 Realtime.

33
00:03:14,912 --> 00:03:15,722
Ashley: Evan.

34
00:03:15,722 --> 00:03:21,642
The methods are organized into several key components focusing on real-time spoken interaction.

35
00:03:21,698 --> 00:03:23,618
Evan: That sounds comprehensive.

36
00:03:23,618 --> 00:03:25,678
What's the first major component?

37
00:03:25,730 --> 00:03:28,570
Ashley: The first component is the system architecture.

38
00:03:28,570 --> 00:03:39,960
StepAudio 3 Realtime uses a mixture-of-experts architecture, which consists of approximately 196 billion total parameters with 11 billion active parameters per token.

39
00:03:39,960 --> 00:03:48,850
The language backbone builds on Step 3.7 Flash, while the audio frontend utilizes the Audio Transformer encoder from Qwen3-Omni.

40
00:03:48,974 --> 00:03:53,854
Evan: So, how does the full-duplex input path work within this architecture?

41
00:03:53,906 --> 00:03:58,696
Ashley: The full-duplex input path integrates user and model audio streams.

42
00:03:58,696 --> 00:04:06,326
Audio representations pass through the encoder and adapter to the language model decoder, which also receives text tokens separately.

43
00:04:06,326 --> 00:04:14,126
This setup allows joint conditioning on acoustic and textual context, enabling seamless integration between listening and speaking.

44
00:04:14,246 --> 00:04:18,386
Evan: It's fascinating how they manage such complex interactions.

45
00:04:18,386 --> 00:04:20,366
How about training these models?

46
00:04:20,426 --> 00:04:27,116
Ashley: Training is divided into three stages: modality alignment, multimodal mixed training, and cooldown.

47
00:04:27,116 --> 00:04:33,166
The modality-alignment stage establishes the interface between acoustic representations and the language model.

48
00:04:33,166 --> 00:04:37,566
Multimodal mixed training develops joint audio-text modeling at scale.

49
00:04:37,566 --> 00:04:42,986
Finally, the cooldown stage emphasizes high-quality data to refine the foundation.

50
00:04:43,034 --> 00:04:47,234
Evan: What kind of data are involved in this pretraining process?

51
00:04:47,282 --> 00:05:00,132
Ashley: Pretraining involves a mix of curated audio data processed through an automated pipeline for sound event detection, voice activity detection, metadata assignment, transcription, and language identification.

52
00:05:00,132 --> 00:05:05,142
This ensures high-quality, semantically complete samples suitable for training.

53
00:05:05,186 --> 00:05:09,246
Evan: And after pretraining comes midtraining, right?

54
00:05:09,290 --> 00:05:10,310
Ashley: Exactly.

55
00:05:10,310 --> 00:05:20,310
Midtraining expands context length to accommodate longer dialogue histories and integrates deeper perception, voice-agent interaction, and synthetic conversational data.

56
00:05:20,310 --> 00:05:26,370
This stage significantly increases the proportion of audio-understanding and agent-interaction data.

57
00:05:26,426 --> 00:05:29,206
Evan: You've mentioned Deep Perception earlier.

58
00:05:29,206 --> 00:05:31,946
How does that fit into the fine-tuning process?

59
00:05:31,994 --> 00:05:39,334
Ashley: Deep Perception covers lexical understanding and cues about the speaker, delivery, acoustic events, and temporal structure.

60
00:05:39,334 --> 00:05:44,244
This involves StepAudio 3 ASR Max, which is specialized for transcription.

61
00:05:44,244 --> 00:05:53,794
The ASR-specialized model is fine-tuned with sequence packed examples, applying time-frequency masking for augmentation and producing normalized transcripts.

62
00:05:53,918 --> 00:05:57,938
Evan: Are there any specific techniques to handle rare terminology?

63
00:05:57,986 --> 00:05:59,216
Ashley: Yes, indeed.

64
00:05:59,216 --> 00:06:03,966
Rare names and technical terms are augmented with targeted synthetic examples.

65
00:06:03,966 --> 00:06:13,366
An LLM expands a taxonomy to identify homophones, uncommon characters, abbreviations, and places these terms in natural carrier sentences.

66
00:06:13,366 --> 00:06:17,246
This helps the model differentiate between similar-sounding words.

67
00:06:17,366 --> 00:06:20,786
Evan: And how is audio understanding evaluated?

68
00:06:20,834 --> 00:06:27,934
Ashley: Audio understanding is evaluated through benchmarks like Big Bench Audio, MMSU, MMAU, and more.

69
00:06:27,934 --> 00:06:36,254
These benchmarks help assess audio-grounded reasoning, fine-grained perception, nonverbal acoustic cues, and multi-turn understanding.

70
00:06:36,314 --> 00:06:39,014
Evan: This brings us to Seamless Duplex.

71
00:06:39,014 --> 00:06:42,174
What's its role in conversational floor management?

72
00:06:42,218 --> 00:06:47,358
Ashley: Seamless Duplex is critical for managing conversational timing and reasoning progress.

73
00:06:47,358 --> 00:06:51,918
It handles pauses, backchannels, interruptions, and other speech phenomena.

74
00:06:51,918 --> 00:07:03,218
The system uses a time-interleaved representation, segmenting audio into 320 ms blocks followed by state or text tokens to continually track interaction state.

75
00:07:03,266 --> 00:07:08,306
Evan: So, how does it decide when to listen, speak, or yield the floor?

76
00:07:08,354 --> 00:07:17,704
Ashley: The decision to continue listening, respond, continue speaking, or yield the floor is guided by acoustic evidence and semantic context.

77
00:07:17,704 --> 00:07:29,634
The model interprets overlapping user utterances against its ongoing speech for contextual role determination, distinguishing pauses from turn completions and backchannels from interruptions.

78
00:07:29,690 --> 00:07:31,990
Evan: It seems very nuanced.

79
00:07:31,990 --> 00:07:34,490
How is conversational timing trained?

80
00:07:34,538 --> 00:07:44,448
Ashley: Midtraining adapts the model to full-duplex interaction through supervised training for streaming ASR, voice activity detection, and prediction of utterance completeness.

81
00:07:44,448 --> 00:07:54,338
Post-training refines these behaviors using high-quality interaction data covering turn taking, user backchannels, interruption handling, and background speech rejection.

82
00:07:54,386 --> 00:07:58,186
Evan: And full-duplex interaction needs robust evaluation.

83
00:07:58,186 --> 00:08:00,106
Which benchmarks are used here?

84
00:08:00,170 --> 00:08:14,130
Ashley: StepAudio 3 Realtime ranks first in the Artificial Analysis subset of Full Duplex Bench v1 and v1.5, scoring highly in pause handling, turn taking, user interruption handling, and backchannel handling.

85
00:08:14,216 --> 00:08:18,766
Evan: Now, moving to conversational intelligence and realtime reasoning.

86
00:08:18,766 --> 00:08:21,926
How does StepAudio 3 handle these aspects?

87
00:08:21,986 --> 00:08:31,776
Ashley: StepAudio 3 Realtime combines dialogue and reasoning training with Think-While-Speaking, which coordinates spoken responses with ongoing private reasoning.

88
00:08:31,776 --> 00:08:38,806
Think-While-Speaking uses concurrent calls to the audio model, generating private reasoning while managing streaming output.

89
00:08:38,928 --> 00:08:43,318
Evan: Adaptive Thinking sounds like a key component for reasoning efficiency.

90
00:08:43,318 --> 00:08:44,738
What does it involve?

91
00:08:44,786 --> 00:08:48,896
Ashley: Adaptive Thinking determines whether a response needs explicit reasoning.

92
00:08:48,896 --> 00:08:54,836
It constructs supervision at the assistant-turn level, labeled by whether reasoning changes answer quality.

93
00:08:54,836 --> 00:08:58,026
This helps allocate reasoning effort more efficiently.

94
00:08:58,172 --> 00:09:02,542
Evan: But how is private reasoning accelerated within this framework?

95
00:09:02,594 --> 00:09:13,204
Ashley: Private reasoning is accelerated using Multi-Token Prediction or MTP, which speeds up the decoding process through speculative and multi-token decoding heads.

96
00:09:13,204 --> 00:09:17,614
This reduces sequential target model steps and improves efficiency.

97
00:09:17,726 --> 00:09:22,086
Evan: How effective is Model Merging for capability integration?

98
00:09:22,130 --> 00:09:29,040
Ashley: Model Merging integrates specialized teacher models, leveraging their complementary strengths by parameter averaging.

99
00:09:29,040 --> 00:09:33,970
This balances overall capability without retraining on the unified data mix.

100
00:09:34,034 --> 00:09:39,514
Evan: Now, the full-duplex voice agent brings asynchronous tool execution into play.

101
00:09:39,514 --> 00:09:42,774
How does the model decide when to delegate tasks?

102
00:09:42,818 --> 00:09:53,638
Ashley: Requests are routed based on context: direct responses for stable knowledge, lightweight tools for up-to-date public information, and backend services for complex or prolonged tasks.

103
00:09:53,638 --> 00:10:01,058
The model gathers missing information via user clarification or context retrieval before committing to an external action.

104
00:10:01,106 --> 00:10:05,266
Evan: Can users interact with backend tasks during execution?

105
00:10:05,330 --> 00:10:12,260
Ashley: Users can request updates, provide additional requirements, or shift topics while backend tasks progress.

106
00:10:12,260 --> 00:10:18,370
Returned tool results update the conversational context, guiding subsequent responses and reasoning.

107
00:10:18,514 --> 00:10:22,494
Evan: What does training for conversational tool use involve?

108
00:10:22,538 --> 00:10:36,238
Ashley: Training combines targeted dialogues covering request routing, clarification, private-context retrieval, and execution-time updates with multi-step trajectories focusing on tool-call structure and argument consistency.

109
00:10:36,290 --> 00:10:39,710
Evan: And how is agentic task completion evaluated?

110
00:10:39,770 --> 00:10:48,550
Ashley: Evaluation uses the Artificial Analysis implementation of τ-Voice to assess tool-grounded customer-service tasks under challenging conditions.

111
00:10:48,550 --> 00:10:53,930
This measures the model's ability to track requests and coordinate interaction with tool use.

112
00:10:53,978 --> 00:10:56,618
Evan: Alright, that wraps up the methods section.

113
00:10:56,618 --> 00:11:02,978
It's fascinating how all these components integrate to handle real-time spoken interactions so effectively.

114
00:11:03,026 --> 00:11:06,926
Ashley: Indeed, it's the end of the Method section of the paper.

115
00:11:08,187 --> 00:11:10,987
Evan: We've covered the introduction and methods.

116
00:11:10,987 --> 00:11:14,607
Now, let's move on to the experiments and results.

117
00:11:14,697 --> 00:11:15,477
Ashley: Evan.

118
00:11:15,477 --> 00:11:30,827
StepAudio 3 Realtime underwent extensive evaluation across several key domains: speech recognition, audio understanding, dialogue and reasoning, full-duplex interaction, agentic task completion, and general text capabilities.

119
00:11:30,891 --> 00:11:33,341
Evan: That's a wide array of benchmarks.

120
00:11:33,341 --> 00:11:36,151
How did the model perform in speech recognition?

121
00:11:36,195 --> 00:11:44,435
Ashley: StepAudio 3 ASR Max, the specialized transcription model, showed excellent performance on multiple ASR benchmarks.

122
00:11:44,435 --> 00:11:52,025
It achieved a word error rate of 1.18% on LibriSpeech clean and 2.28% on LibriSpeech other.

123
00:11:52,025 --> 00:11:58,325
It also led AISHELL-1 with a character error rate of 0.49%.

124
00:11:58,325 --> 00:12:02,035
These results highlight its robust transcription accuracy.

125
00:12:02,161 --> 00:12:03,471
Evan: Impressive.

126
00:12:03,471 --> 00:12:06,371
How about its audio understanding performance?

127
00:12:06,435 --> 00:12:10,585
Ashley: StepAudio 3 Realtime excelled in audio understanding tasks.

128
00:12:10,585 --> 00:12:18,395
For instance, it scored 90.6 on the MMSU benchmark, which assesses multi-task spoken language understanding and reasoning.

129
00:12:18,395 --> 00:12:27,315
It also achieved 86.5 on the MMAR benchmark, showing strong performance in deep reasoning across speech, audio, and music domains.

130
00:12:27,423 --> 00:12:31,343
Evan: What other benchmarks were considered for audio understanding?

131
00:12:31,425 --> 00:12:38,715
Ashley: Other benchmarks included AudioMultiChallenge, MMAU, WildSpeech, Step-Caption, and MTalk-Bench.

132
00:12:38,715 --> 00:12:47,165
StepAudio 3 Realtime led four out of these eight benchmarks, with particularly notable gains in areas like MMSU and MMAR.

133
00:12:47,165 --> 00:12:55,595
These results underline the model's robust audio-grounded reasoning, fine-grained perception, and ability to handle nonverbal acoustic cues.

134
00:12:55,699 --> 00:13:00,159
Evan: How did the model perform on dialogue and reasoning benchmarks?

135
00:13:00,219 --> 00:13:08,879
Ashley: On the StepAudioChat dialogue benchmark, StepAudio 3 Realtime reached a macro average of 70.4 in its interactive mode.

136
00:13:08,879 --> 00:13:18,909
It showed competitive results in instruction following, faithfulness, reasoning, memory, knowledge, conversational pragmatics, and persona consistency.

137
00:13:18,909 --> 00:13:23,159
It closely matched dedicated reasoning models while speaking in real-time.

138
00:13:23,211 --> 00:13:29,651
Evan: Could you give an example of a specific dimension within dialogue and reasoning where the model excelled?

139
00:13:29,715 --> 00:13:30,515
Ashley: Of course.

140
00:13:30,515 --> 00:13:41,305
In reasoning, for example, the model scored 73.6, which is a close second to the best-performing model on this dimension, Kimi K3, that scored 81.9.

141
00:13:41,305 --> 00:13:44,395
This indicates strong aggregate reasoning abilities.

142
00:13:44,451 --> 00:13:48,711
Evan: And how did it fare in managing full-duplex interactions?

143
00:13:48,771 --> 00:13:56,651
Ashley: StepAudio 3 Realtime ranked first on the Artificial Analysis subset of Full Duplex Bench v1 and v1.5.

144
00:13:56,651 --> 00:14:03,161
It achieved an overall score of 98.9, with perfect scores for turn taking and user interruption handling.

145
00:14:03,161 --> 00:14:12,031
This demonstrated balanced conversational control, responding appropriately to pauses, interruptions, and backchannels while maintaining conversational flow.

146
00:14:12,075 --> 00:14:15,605
Evan: It sounds like it handled real-time interaction quite well.

147
00:14:15,605 --> 00:14:18,095
What about agentic task completion?

148
00:14:18,147 --> 00:14:29,037
Ashley: For agentic task completion, StepAudio 3 Realtime was assessed using the τ-Voice benchmark, which involves customer-service tasks in domains like airline, retail, and telecom.

149
00:14:29,037 --> 00:14:38,377
It achieved a macro task-success rate of 56.0%, close to the highest score of 56.5 achieved by Grok Voice Think Fast 2.0.

150
00:14:38,377 --> 00:14:44,187
Its performance was particularly strong in the telecom domain, with a score of 70.2%.

151
00:14:44,235 --> 00:14:48,575
Evan: And finally, how did it perform on general text capabilities?

152
00:14:48,627 --> 00:14:53,797
Ashley: StepAudio 3 Realtime also showed excellent results in general text benchmarks.

153
00:14:53,797 --> 00:15:09,277
It achieved an accuracy of 86.8% on the HMMT February 2026 benchmark, which focuses on competition mathematics, and scored 83.0% on GPQA Diamond, a graduate-level question answering benchmark.

154
00:15:09,277 --> 00:15:14,647
These results demonstrate solid performance in text-based reasoning and knowledge tasks.

155
00:15:14,751 --> 00:15:21,081
Evan: It's clear that StepAudio 3 Realtime has robust performance across multiple benchmarks and tasks.

156
00:15:21,081 --> 00:15:24,111
Do these results indicate any areas for improvement?

157
00:15:24,171 --> 00:15:30,661
Ashley: Yes, the results identified some gaps in multi-turn constraint following and retail tool-use tasks.

158
00:15:30,661 --> 00:15:34,271
These are areas where the model has room for further enhancement.

159
00:15:34,323 --> 00:15:35,633
Evan: Thank you, Ashley.

160
00:15:35,633 --> 00:15:38,543
That's the end of the Experiment and Results section.

161
00:15:39,797 --> 00:15:50,217
Evan: Now that we've explored the introduction, methods, and experimental results, let's dive into the related work that influenced the development of StepAudio 3 Realtime.

162
00:15:50,261 --> 00:15:51,441
Ashley: Sure, Evan.

163
00:15:51,441 --> 00:16:04,381
The related work in the paper primarily revolves around notable advances in speech recognition, audio-language models, and full-duplex systems, as well as methodologies incorporated into StepAudio 3 Realtime.

164
00:16:04,445 --> 00:16:08,465
Evan: What are some foundational works cited in this research?

165
00:16:08,525 --> 00:16:18,405
Ashley: The paper references important contributions to robust speech recognition using large-scale weak supervision, such as the work by Radford et al., 2023.

166
00:16:18,405 --> 00:16:24,265
This research integrates acoustic representations with the linguistic knowledge of large language models.

167
00:16:24,317 --> 00:16:26,017
Evan: That sounds crucial.

168
00:16:26,017 --> 00:16:29,417
Were there any specific models or techniques highlighted?

169
00:16:29,477 --> 00:16:30,147
Ashley: Yes.

170
00:16:30,147 --> 00:16:39,317
The work by Bai et al., 2024, on understanding diverse speech and contexts with LLM-based speech recognition is particularly noteworthy.

171
00:16:39,317 --> 00:16:52,717
Lin et al., 2026, further explored multi-token prediction for fast and long-context speech recognition, which laid the groundwork for advances in transcription accuracy in models like StepAudio 3 ASR Max.

172
00:16:52,781 --> 00:16:57,851
Evan: Speaking of advances, audio-language models have seen significant progress.

173
00:16:57,851 --> 00:17:00,801
Could you elaborate on how these models contributed?

174
00:17:00,845 --> 00:17:01,885
Ashley: Certainly.

175
00:17:01,885 --> 00:17:12,565
Borsos et al., 2023, introduced AudioLM, a language modeling approach to audio generation that supports broader acoustic understanding and direct speech generation.

176
00:17:12,565 --> 00:17:20,925
Models like Salmonn and Paralinguistics-aware LLMs have demonstrated the capabilities for natural and nuanced spoken exchanges.

177
00:17:21,041 --> 00:17:26,551
Evan: And full-duplex systems seem to play a vital role in managing conversational dynamics.

178
00:17:26,551 --> 00:17:28,981
What examples can you share from this realm?

179
00:17:29,045 --> 00:17:38,275
Ashley: Full-duplex models like the ones discussed by Wang et al., 2024, and Defossez et al., 2024, are integral to streaming audio systems.

180
00:17:38,275 --> 00:17:44,475
These models enable concurrent listening and speaking, thereby allowing for fluid and responsive dialogue management.

181
00:17:44,475 --> 00:17:55,305
Wu et al., 2026, highlighted chronological thinking in full-duplex spoken dialogue, which plays a crucial part in StepAudio 3 Realtime's Seamless Duplex component.

182
00:17:55,419 --> 00:17:58,249
Evan: It's fascinating how these contributions integrate.

183
00:17:58,249 --> 00:18:04,109
What methodologies were particularly influential in the development of StepAudio 3 Realtime?

184
00:18:04,157 --> 00:18:12,237
Ashley: Methodologies like Mind-Paced Speaking, by Wu et al., 2025, influenced the Think-While-Speaking capability.

185
00:18:12,237 --> 00:18:19,087
It employs a dual-brain approach, with parallel processes for generating private reasoning and managing streaming output.

186
00:18:19,087 --> 00:18:24,977
This methodology supports fluent conversational responses without compromising reasoning depth.

187
00:18:25,037 --> 00:18:29,697
Evan: Were there any specific training techniques used for enhancing these models?

188
00:18:29,741 --> 00:18:38,701
Ashley: Yes, SpecAugment introduced by Park et al., 2019, for data augmentation in automatic speech recognition played a significant role.

189
00:18:38,701 --> 00:18:51,481
Additionally, Recognizer Output Voting Error Reduction (ROVER) by Fiscus, 1997, is utilized to improve transcription accuracy by fusing hypotheses from multiple recognition systems.

190
00:18:51,593 --> 00:18:57,573
Evan: It's remarkable how these methodologies converge to enhance the StepAudio 3 Realtime model.

191
00:18:57,573 --> 00:19:03,333
Let's touch upon the ways these advances impact real-time interaction and tool use in the system.

192
00:19:03,389 --> 00:19:13,689
Ashley: Tool-grounded language modeling approaches like Toolformer by Schick et al., 2023, and Gorilla by Patil et al., 2024, have been significant.

193
00:19:13,689 --> 00:19:25,389
These models demonstrate how language models can effectively integrate external tools for deeper interaction and accurate execution of tasks, which is essential for StepAudio 3 Realtime's Voice Agent.

194
00:19:25,505 --> 00:19:30,725
Evan: Indeed, integrating tools seamlessly into dialogue systems is crucial.

195
00:19:30,725 --> 00:19:34,545
How does this relate to end-to-end task completion benchmarks?

196
00:19:34,619 --> 00:19:46,379
Ashley: End-to-end task completion benchmarks, such as those defined in τ-Voice by Artificial Analysis, 2026, assess full-duplex spoken interaction under challenging conditions.

197
00:19:46,379 --> 00:19:55,509
These benchmarks ensure that models like StepAudio 3 Realtime can track evolving user requests and coordinate spoken interaction with tool use effectively.

198
00:19:55,625 --> 00:20:04,105
Evan: It's impressive how all these elements come together to support robust, real-time spoken interactions in StepAudio 3 Realtime.

199
00:20:04,157 --> 00:20:07,017
Ashley: This concludes the Related Work section.

200
00:20:08,262 --> 00:20:15,582
Evan: We've walked through the introduction, methods, experimental results, and related work for StepAudio 3 Realtime.

201
00:20:15,582 --> 00:20:19,642
Now, let's summarize the key contributions and takeaways.

202
00:20:19,686 --> 00:20:25,456
Ashley: StepAudio 3 Realtime presents a comprehensive approach to real-time spoken interaction.

203
00:20:25,456 --> 00:20:31,746
It coordinates perception, reasoning, and action through a continuous listen-converse-think-act loop.

204
00:20:31,806 --> 00:20:41,346
Evan: One standout feature is Deep Perception, which captures rich acoustic cues to interpret user intent, enabling nuanced spoken exchanges.

205
00:20:41,436 --> 00:20:42,506
Ashley: Precisely.

206
00:20:42,506 --> 00:20:49,456
Seamless Duplex manages synchronized audio streams to fluidly handle pauses, backchannels, and interruptions.

207
00:20:49,456 --> 00:20:52,386
This ensures balanced conversational control.

208
00:20:52,506 --> 00:20:57,856
Evan: And Think-While-Speaking is key to reducing latency while maintaining deep reasoning.

209
00:20:57,856 --> 00:21:02,646
It allows the model to deliberate while producing spoken output in real time.

210
00:21:02,694 --> 00:21:03,764
Ashley: Indeed.

211
00:21:03,764 --> 00:21:13,294
The integration of a Voice Agent extends interactions to tasks requiring tools and asynchronous execution, making the system versatile and efficient.

212
00:21:13,410 --> 00:21:24,620
Evan: In terms of performance, StepAudio 3 Realtime leads in various benchmarks, including audio understanding, full-duplex interaction, dialogue, and agentic task completion.

213
00:21:24,620 --> 00:21:30,570
Particularly notable are its scores on the MMSU, MMAR, and τ-Voice benchmarks.

214
00:21:30,630 --> 00:21:38,600
Ashley: These achievements highlight the model's robust ability to handle multi-turn interactions, complex requests, and dynamic tool use.

215
00:21:38,600 --> 00:21:47,090
While there are areas for improvement, such as multi-turn constraint following and retail tool-use tasks, the overall performance is impressive.

216
00:21:47,142 --> 00:21:58,762
Evan: StepAudio 3 Realtime exemplifies advancements in real-time spoken interaction, showing that seamless coordination between listening, reasoning, and acting is achievable.

217
00:21:58,806 --> 00:21:59,906
Ashley: That's right, Evan.

218
00:21:59,906 --> 00:22:07,646
It's exciting to see how far we've come with audio-language models and their potential to enhance real-time dialogues and interactions.

219
00:22:07,710 --> 00:22:11,010
Evan: Thank you for tuning in to Daily Paper Cast.

220
00:22:11,010 --> 00:22:14,070
We hope you found today's discussion insightful.

221
00:22:14,118 --> 00:22:19,538
Ashley: Be sure to join us for future episodes where we'll continue to explore the latest in AI research.

222
00:22:19,538 --> 00:22:21,738
Until next time, take care!