1
00:00:03,030 --> 00:00:05,400
Evan: Welcome to Daily Paper Cast.

2
00:00:05,448 --> 00:00:13,128
Ashley: Today's paper is from the Hugging Face daily paper list of September 30, 2026, with 52 upvotes.

3
00:00:13,176 --> 00:00:27,236
Evan: It's titled 'VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models', written by Yang Xiao and Vidhyasaharan Sethu, with corresponding author Eun-Jung Holden from the University of Melbourne.

4
00:00:27,288 --> 00:00:27,988
Ashley: Great.

5
00:00:27,988 --> 00:00:30,388
Let's dive into the Introduction section.

6
00:00:30,432 --> 00:00:37,962
Evan: Large audio language models, or LALMs, have rapidly advanced the field of natural spoken interaction.

7
00:00:37,962 --> 00:00:44,272
These models enable dialogue systems to manage extended, multi-session conversations more effectively.

8
00:00:44,358 --> 00:00:45,288
Ashley: Evan.

9
00:00:45,288 --> 00:00:51,028
And for these systems to be truly effective, they must possess long-term memory capabilities.

10
00:00:51,028 --> 00:00:59,088
This means being able to accumulate information across past interactions and use that information to ground present responses.

11
00:00:59,136 --> 00:01:12,576
Evan: While simply expanding the context window can allow models to process longer dialogue histories directly, it doesn't necessarily guarantee that the model can reliably retrieve and use the information contained within those histories.

12
00:01:12,624 --> 00:01:13,554
Ashley: That's correct.

13
00:01:13,554 --> 00:01:21,084
As interaction histories grow, the critical challenge becomes whether information can be accurately retrieved and used across sessions.

14
00:01:21,084 --> 00:01:26,704
Evaluating spoken conversational memory and guiding future memory systems is essential.

15
00:01:26,860 --> 00:01:32,720
Evan: Existing benchmarks fall short because they generally lack a principled taxonomy of audio memory.

16
00:01:32,720 --> 00:01:41,320
They typically cover only a narrow range of spoken-memory capabilities, focusing predominantly on linguistic content and selected acoustic cues.

17
00:01:41,376 --> 00:01:42,136
Ashley: Right.

18
00:01:42,136 --> 00:01:49,026
And this results in evaluations that mainly assess memory for what was said, rather than how it was conveyed via audio.

19
00:01:49,026 --> 00:01:57,316
Such an approach makes it challenging to understand whether failures arise from particular evidence types, memory operations, or their interaction.

20
00:01:57,360 --> 00:02:08,160
Evan: VoxMem addresses these gaps by introducing a benchmark for evaluating spoken conversational memory across two complementary axes: acoustic evidence and memory operation.

21
00:02:08,208 --> 00:02:25,848
Ashley: Specifically, VoxMem defines four types of acoustic evidence: speech semantics, which is what was said, speaker identity, which is who said it, paralinguistic cues, which is how it was said, and environmental sound, covering what was heard around the speaker.

22
00:02:25,936 --> 00:02:39,326
Evan: And to capture how this information must be used, VoxMem defines four memory operations: information extraction, multi-session reasoning, temporal evolution tracking, and answer refusal.

23
00:02:39,326 --> 00:02:50,456
This helps in evaluating whether models can retrieve a single fact, integrate evidence across sessions, track changes over time, or refuse to answer when there's insufficient evidence.

24
00:02:50,520 --> 00:02:51,560
Ashley: Exactly.

25
00:02:51,560 --> 00:02:58,210
VoxMem evaluates each of these operations and evidence types across multi-session conversational histories.

26
00:02:58,210 --> 00:03:11,080
The benchmark consists of 3,196 evaluation instances spread over 34,743 spoken sessions, amounting to 177 hours of audio content.

27
00:03:11,136 --> 00:03:23,416
Evan: To enable controlled evaluation as context length scales, VoxMem spans four context budgets: 8K, 16K, 32K, and 64K tokens.

28
00:03:23,416 --> 00:03:33,136
Each setup ensures the question, answer, and evidence remain identical across different lengths, so any performance drop can be attributed to the growing context.

29
00:03:33,192 --> 00:03:45,742
Ashley: Testing across these various budgets revealed a significant finding: among 15 evaluated LALMs, none exceeded 40% overall accuracy at the 32K token budget.

30
00:03:45,742 --> 00:03:54,412
Models were particularly good at retaining 'what was said', but struggled with 'who said it', 'how it was said', and 'what was audible'.

31
00:03:54,456 --> 00:03:56,186
Evan: That's a substantial gap.

32
00:03:56,186 --> 00:04:04,736
Moreover, accuracy declines as the history grows across all evidence types, and error modes differ qualitatively across categories.

33
00:04:04,736 --> 00:04:11,676
VoxMem essentially aims to provide a foundation for measuring and driving progress in spoken conversational memory.

34
00:04:11,736 --> 00:04:14,956
Ashley: This brings us to the end of the Introduction section.

35
00:04:16,262 --> 00:04:25,422
Evan: Ashley, can you explain how VoxMem categorizes and evaluates the multimodal memory in large audio language models?

36
00:04:25,466 --> 00:04:32,756
Ashley: The VoxMem benchmark establishes a two-dimensional framework to evaluate spoken conversational memory systematically.

37
00:04:32,756 --> 00:04:41,126
On one axis, we have the types of acoustic evidence that must be retained, and on the other, the types of memory operations applied to this evidence.

38
00:04:41,186 --> 00:04:41,866
Evan: Got it.

39
00:04:41,866 --> 00:04:45,506
And how did the authors define these acoustic evidence types?

40
00:04:45,554 --> 00:04:48,984
Ashley: The benchmark targets four types of acoustic evidence.

41
00:04:48,984 --> 00:04:56,384
Firstly, speech semantics which is information recoverable from transcripts such as facts, plans, and preferences.

42
00:04:56,384 --> 00:04:59,294
This serves as the baseline for semantic memory.

43
00:04:59,354 --> 00:05:01,254
Evan: And what about the other types?

44
00:05:01,298 --> 00:05:08,568
Ashley: Secondly, we have speaker identity which requires the model to associate information with the voice of the speaker who produced it.

45
00:05:08,568 --> 00:05:18,008
Thirdly, paralinguistic cues cover how something was said, including vocal states and delivery characteristics like hesitation, emphasis, and laughter.

46
00:05:18,008 --> 00:05:29,678
Finally, environmental sound captures non-speech sound events such as traffic, alarms, and household noises—elements which are entirely audio-native and absent from transcripts.

47
00:05:29,738 --> 00:05:31,528
Evan: That’s quite comprehensive.

48
00:05:31,528 --> 00:05:34,338
Now, how are the memory operations defined?

49
00:05:34,424 --> 00:05:37,884
Ashley: VoxMem evaluates four distinct memory operations.

50
00:05:37,884 --> 00:05:43,864
Information extraction, multi-session reasoning, temporal evolution tracking, and answer refusal.

51
00:05:43,864 --> 00:05:49,634
Information extraction tests if the model can retrieve data from a specific session using acoustic evidence.

52
00:05:49,634 --> 00:05:54,084
Multi-session reasoning checks if the model can integrate evidence across sessions.

53
00:05:54,084 --> 00:05:59,234
Temporal evolution tracking evaluates the model’s ability to follow changes across sessions.

54
00:05:59,234 --> 00:06:04,234
Answer refusal checks if the model can abstain from answering when there’s insufficient evidence.

55
00:06:04,298 --> 00:06:05,658
Evan: Interesting.

56
00:06:05,658 --> 00:06:11,178
And how did the authors ensure naturalism and controlled scenarios in their benchmark construction?

57
00:06:11,234 --> 00:06:17,744
Ashley: VoxMem was constructed meticulously through multiple stages to ensure both naturalism and control.

58
00:06:17,744 --> 00:06:27,004
Each item was planned with a structured specification detailing the required evidence and valid sessions, before rendering them into user-assistant dialogues.

59
00:06:27,004 --> 00:06:33,794
During dialogue writing, user turns were authored independently from assistant turns to prevent answer leakage.

60
00:06:33,842 --> 00:06:36,902
Evan: But how did they translate this into spoken audio?

61
00:06:36,992 --> 00:06:44,402
Ashley: All sessions were synthesized using Higgs-TTS-3 with fixed VCTK voices for speakers.

62
00:06:44,402 --> 00:06:53,262
Paralinguistic cues were introduced via style controls, and environmental sounds were mixed in using assets from ESC-50.

63
00:06:53,306 --> 00:06:57,546
Evan: So, both the dialogue text and audio were carefully crafted.

64
00:06:57,546 --> 00:07:00,226
What about ensuring the context length scales?

65
00:07:00,290 --> 00:07:10,250
Ashley: To manage context scales, VoxMem includes histories at four token budgets: 8K, 16K, 32K, and 64K.

66
00:07:10,250 --> 00:07:16,280
These histories grow progressively while preserving the content, context, and session structure.

67
00:07:16,280 --> 00:07:25,390
Each history length strictly extends the shorter ones to prevent changes in the question itself, ensuring we’re genuinely measuring the impact of context length.

68
00:07:25,442 --> 00:07:28,952
Evan: Right, consistency is key in such evaluations.

69
00:07:28,952 --> 00:07:33,362
Can you elaborate on the evaluation protocol adopted for testing these models?

70
00:07:33,410 --> 00:07:34,270
Ashley: Sure.

71
00:07:34,270 --> 00:07:41,920
The evaluation aimed to measure models’ accuracy across all memory operations and evidence types at each context length.

72
00:07:41,920 --> 00:07:45,720
Models were given the interaction history followed by a query.

73
00:07:45,720 --> 00:07:54,250
Each response was judged for correctness using a Gemini model to avoid biases, particularly as Gemini models were among the evaluated ones.

74
00:07:54,354 --> 00:07:58,254
Evan: And how did they address potential biases in judging the responses?

75
00:07:58,298 --> 00:08:11,118
Ashley: Responses were scored using a rule-based agreement test between Gemini and GPT-5.6-Luna models, achieving a Cohen’s Kappa agreement of 0.95—that’s quite high.

76
00:08:11,162 --> 00:08:12,492
Evan: That's impressive.

77
00:08:12,492 --> 00:08:16,022
Any specific metrics they focused on during their analysis?

78
00:08:16,082 --> 00:08:17,232
Ashley: Indeed.

79
00:08:17,232 --> 00:08:21,942
They measured raw accuracy and retention relative to the 8K baseline.

80
00:08:21,942 --> 00:08:29,842
Accuracy declines were evident as history grew, which reinforces the complexity of maintaining reliable memory over longer histories.

81
00:08:29,976 --> 00:08:38,466
Evan: So essentially, while models retain semantic information better, they struggle with audio-native information as the context scales.

82
00:08:38,522 --> 00:08:39,452
Ashley: Precisely.

83
00:08:39,452 --> 00:08:50,152
Errors from these models also varied by operation type; for instance, temporal evolution tracking mostly failed due to localization breakdowns before any reasoning occurred.

84
00:08:50,152 --> 00:08:58,042
They also analyzed distinct error modes across different evidence types, showing varied underlying reasons for memory failures.

85
00:08:58,106 --> 00:09:03,156
Evan: This really shows the depth of insight VoxMem offers for improving LALMs.

86
00:09:03,156 --> 00:09:05,166
We’ve covered a lot about the method.

87
00:09:05,210 --> 00:09:07,810
Ashley: Yes, that wraps up the Method section.

88
00:09:09,115 --> 00:09:14,325
Evan: Ashley, let’s move on to the Experimental Setup and Results section of VoxMem.

89
00:09:14,325 --> 00:09:18,315
How did the authors design their experiments to evaluate these models?

90
00:09:18,363 --> 00:09:21,203
Ashley: The experimental setup was quite comprehensive.

91
00:09:21,203 --> 00:09:28,623
The authors evaluated 15 large audio language models, including ten open-weight and five proprietary systems.

92
00:09:28,623 --> 00:09:37,533
They ensured that each model was tested at four history lengths: 8K, 16K, 32K, and 64K tokens.

93
00:09:37,533 --> 00:09:42,763
This provides a controlled scale to see how performance changes as context length grows.

94
00:09:42,879 --> 00:09:47,279
Evan: And each model was evaluated through their native audio interfaces, right?

95
00:09:47,331 --> 00:09:48,531
Ashley: Exactly.

96
00:09:48,531 --> 00:09:53,161
Each model received user turns as audio and assistant turns as text.

97
00:09:53,161 --> 00:10:02,071
The final part of the interaction included an audio query, and models had to generate a single response based on the entire conversation history.

98
00:10:02,115 --> 00:10:05,095
Evan: How were the responses judged for correctness?

99
00:10:05,139 --> 00:10:14,419
Ashley: Responses were scored using a combination of exact string matching and a large language model judge, specifically the Gemini-3.7-Flash.

100
00:10:14,419 --> 00:10:21,019
This judge marked responses as correct or incorrect based on predefined answer types and rules.

101
00:10:21,075 --> 00:10:22,935
Evan: That sounds meticulous.

102
00:10:22,935 --> 00:10:25,855
So what kind of accuracy figures did they find?

103
00:10:25,899 --> 00:10:31,709
Ashley: Overall, no model exceeded 40% accuracy at the 32K token budget.

104
00:10:31,709 --> 00:10:39,189
The proprietary models averaged about 33%, while the open-weight models averaged around 21.9%.

105
00:10:39,189 --> 00:10:46,459
Across all evaluated models, performance remained far from what we could call reliable spoken conversational memory.

106
00:10:46,515 --> 00:10:52,475
Evan: Were there specific types of information that were easier or harder for the models to remember?

107
00:10:52,539 --> 00:10:57,959
Ashley: Yes, there was a marked difference between lexical and non-lexical information.

108
00:10:57,959 --> 00:11:08,279
For speech semantics, which involves the content of what was said, the models were relatively better, achieving an average of 55.6% accuracy.

109
00:11:08,279 --> 00:11:17,979
In contrast, speaker identity, paralinguistic cues, and environmental sound had much lower accuracies, often below 20%.

110
00:11:18,087 --> 00:11:24,207
Evan: So the models are much better at remembering what was said rather than who said it or how it was said.

111
00:11:24,267 --> 00:11:25,187
Ashley: Exactly.

112
00:11:25,187 --> 00:11:30,677
This gap highlights the difficulty LALMs have with non-lexical acoustic information.

113
00:11:30,677 --> 00:11:34,397
Even within non-lexical categories, there was variance.

114
00:11:34,397 --> 00:11:41,947
For instance, models performed a bit better with speaker identity compared to paralinguistic cues and environmental sounds.

115
00:11:42,003 --> 00:11:43,083
Evan: Interesting.

116
00:11:43,083 --> 00:11:47,603
Were there any insights about how these models handle different memory operations?

117
00:11:47,667 --> 00:11:53,447
Ashley: Yes, the researchers found that memory operations also influenced performance significantly.

118
00:11:53,447 --> 00:12:04,987
Multi-session reasoning was relatively robust, especially for speech semantics and speaker identity, indicating that models could handle combining evidence across sessions better than other operations.

119
00:12:05,043 --> 00:12:08,843
Evan: What about temporal evolution tracking and answer refusal?

120
00:12:08,907 --> 00:12:14,677
Ashley: Temporal evolution tracking, especially for non-lexical evidence, showed quite poor performance.

121
00:12:14,677 --> 00:12:23,397
For example, tracking paralinguistic cues or environmental sounds across sessions had very low accuracy rates, often collapsing altogether.

122
00:12:23,397 --> 00:12:34,787
On the other hand, answer refusal accuracy was higher for paralinguistic and environmental questions, possibly reflecting that models often abstain when they can't process acoustic information effectively.

123
00:12:34,851 --> 00:12:39,371
Evan: It appears the complexity of these tasks really challenges current models.

124
00:12:39,371 --> 00:12:42,331
Were there any additional error analyses conducted?

125
00:12:42,387 --> 00:12:47,607
Ashley: The authors performed fine-grained error analysis to understand why models failed.

126
00:12:47,607 --> 00:12:58,177
They categorized errors into several types, such as evidence recognition failures, binding errors, operation execution failures, and unsupported answers.

127
00:12:58,177 --> 00:13:06,307
For instance, many errors in information extraction tasks were due to models failing to correctly associate facts with the speakers who said them.

128
00:13:06,393 --> 00:13:11,023
Evan: And was there a pattern in how the models' performance degraded with longer histories?

129
00:13:11,067 --> 00:13:16,667
Ashley: Yes, for all evidence types, models' accuracy declined as history length increased.

130
00:13:16,667 --> 00:13:25,927
Speech semantics showed the highest retention rates, maintaining about 70.3% of their 8K token accuracy even at 64K tokens.

131
00:13:25,927 --> 00:13:32,307
In contrast, speaker identity and environmental sound had the most significant drop-offs in performance.

132
00:13:32,355 --> 00:13:37,335
Evan: So increasing the context window isn't a silver bullet for better performance.

133
00:13:37,395 --> 00:13:46,145
Ashley: Correct, simply expanding the context window doesn't address the inherent difficulties LALMs face with non-lexical acoustic information.

134
00:13:46,145 --> 00:13:52,795
The results suggest that more sophisticated mechanisms are needed for models to effectively use longer histories.

135
00:13:52,851 --> 00:13:59,951
Evan: This really highlights the gap between current capabilities and the goal of reliable long-term conversational memory.

136
00:14:00,003 --> 00:14:00,873
Ashley: Indeed.

137
00:14:00,873 --> 00:14:09,503
The experiments underline the need for improvement in how models deal with non-lexical acoustic evidence and handle various memory operations.

138
00:14:09,555 --> 00:14:13,315
Evan: This brings us to the end of the Experiment and Results section.

139
00:14:14,653 --> 00:14:17,213
Evan: Let's delve into the Related Work section.

140
00:14:17,213 --> 00:14:25,573
Ashley, can you give us an overview of how VoxMem positions itself among existing memory benchmarks for spoken dialogues?

141
00:14:25,667 --> 00:14:26,937
Ashley: Of course, Evan.

142
00:14:26,937 --> 00:14:36,217
VoxMem builds on various dimensions of memory evaluation for spoken dialogues that previous works have touched upon but never fully covered comprehensively.

143
00:14:36,269 --> 00:14:37,339
Evan: Interesting.

144
00:14:37,339 --> 00:14:40,269
How does it differentiate from earlier benchmarks?

145
00:14:40,325 --> 00:14:53,425
Ashley: Well, earlier benchmarks like SpokenWOZ and ContextDialog focus mainly on the lexical content of what was said, often neglecting the richness of the audio signal that captures how it was spoken or who said it.

146
00:14:53,477 --> 00:14:59,997
Evan: So, these earlier benchmarks miss out on evaluating the full scope of spoken conversational memory?

147
00:15:00,053 --> 00:15:01,373
Ashley: Exactly.

148
00:15:01,373 --> 00:15:14,373
Others, like AudioMarathon and LongSpeech, primarily target single long-form recordings to assess understanding of extended audio, rather than retaining spoken memory across distinct conversational sessions.

149
00:15:14,429 --> 00:15:15,539
Evan: That makes sense.

150
00:15:15,539 --> 00:15:19,409
And how does VoxMem tackle the issue of cross-session memory?

151
00:15:19,469 --> 00:15:23,579
Ashley: Cross-session memory is a crucial area that VoxMem addresses.

152
00:15:23,579 --> 00:15:28,729
Real spoken interactions are rarely confined to single, unbroken recordings.

153
00:15:28,729 --> 00:15:41,809
Instead, they unfold over multiple sessions, often with varying topics and contexts, making it essential to evaluate models on their ability to retrieve and integrate information across these separated instances.

154
00:15:41,861 --> 00:15:47,681
Evan: So, VoxMem ensures that this multi-session aspect is a core part of its evaluation?

155
00:15:47,741 --> 00:15:48,901
Ashley: Precisely.

156
00:15:48,901 --> 00:16:01,401
VoxMem stands out as the first benchmark built from the ground up to evaluate multi-session spoken conversational memory, with a diverse set of topics and conditions that mimic real-world interactions.

157
00:16:01,445 --> 00:16:06,845
Evan: Can you elaborate on the memory operations and how previous benchmarks have dealt with them?

158
00:16:06,893 --> 00:16:11,023
Ashley: Memory operations are about how remembered information must be used.

159
00:16:11,023 --> 00:16:14,803
Existing benchmarks rarely go beyond basic retrieval tasks.

160
00:16:14,803 --> 00:16:32,933
For instance, most focus on whether a model can retrieve specific information from a past session, but they don't evaluate more complex operations like integrating evidence from multiple sessions, tracking how information evolves, or recognizing when there isn't enough information to provide a coherent answer.

161
00:16:33,041 --> 00:16:36,001
Evan: And how does VoxMem enhance this aspect?

162
00:16:36,053 --> 00:16:48,613
Ashley: VoxMem enhances this by defining four distinct memory operations, as we touched on earlier: information extraction, multi-session reasoning, temporal evolution tracking, and answer refusal.

163
00:16:48,613 --> 00:16:59,393
By covering these, VoxMem ensures that models are tested not just on retrieving information, but on performing a variety of tasks that reflect the true complexity of human memory.

164
00:16:59,513 --> 00:17:05,613
Evan: Did the authors mention any qualitative failures in previous benchmarks that VoxMem aims to address?

165
00:17:05,669 --> 00:17:14,249
Ashley: Yes, the authors pointed out that earlier benchmarks often don’t account for structured failures that can tell us a lot about where and why a model is going wrong.

166
00:17:14,249 --> 00:17:27,669
Without a framework that characterizes both the types of memory operations and the kinds of acoustic evidence to be remembered, it's challenging to pinpoint whether a failure is due to the type of evidence, the operation applied, or their interaction.

167
00:17:27,785 --> 00:17:33,245
Evan: So VoxMem's comprehensive taxonomy helps in analyzing these structured failures?

168
00:17:33,293 --> 00:17:34,293
Ashley: Exactly.

169
00:17:34,293 --> 00:17:43,963
The detailed taxonomy in VoxMem allows for fine-grained error analysis, enabling researchers to identify specific failure modes and their underlying causes.

170
00:17:43,963 --> 00:17:47,633
This helps in developing targeted improvements for LALMs.

171
00:17:47,723 --> 00:17:56,973
Evan: It sounds like VoxMem not only fills gaps in earlier research but also sets a new standard for comprehensive and detailed evaluation in this field.

172
00:17:57,029 --> 00:17:58,019
Ashley: That's right.

173
00:17:58,019 --> 00:18:10,429
By addressing the limitations of previous benchmarks and introducing a detailed taxonomy, VoxMem offers a more holistic and rigorous evaluation framework for spoken conversational memory models.

174
00:18:10,493 --> 00:18:14,013
Evan: This brings us to the end of the Related Work section.

175
00:18:15,330 --> 00:18:20,270
Evan: Alright, Ashley, let's recap the key contributions and takeaways from this paper.

176
00:18:20,334 --> 00:18:21,684
Ashley: Certainly, Evan.

177
00:18:21,684 --> 00:18:34,274
VoxMem introduces a principled taxonomy for evaluating spoken conversational memory, covering both the types of acoustic evidence that need to be retained and the memory operations applied to this evidence.

178
00:18:34,326 --> 00:18:47,586
Evan: It focuses on four types of acoustic evidence: speech semantics, speaker identity, paralinguistic cues, and environmental sounds, ensuring a holistic evaluation that goes beyond just textual content.

179
00:18:47,646 --> 00:18:58,826
Ashley: Yes, and crucially, it evaluates four distinct memory operations: information extraction, multi-session reasoning, temporal evolution tracking, and answer refusal.

180
00:18:58,826 --> 00:19:05,346
This allows for a comprehensive assessment of how well models perform various tasks related to spoken memory.

181
00:19:05,406 --> 00:19:15,906
Evan: The benchmark is extensive, comprising 3,196 evaluation instances over 34,743 spoken sessions.

182
00:19:15,906 --> 00:19:29,566
The design ensures consistent testing across different context lengths—8K, 16K, 32K, and 64K tokens—which helps in observing how performance scales with increasing history.

183
00:19:29,622 --> 00:19:45,482
Ashley: Findings showed that none of the 15 evaluated models exceeded 40% accuracy at 32K tokens, highlighting significant challenges in retaining and using non-lexical acoustic information, particularly as context length grows.

184
00:19:45,534 --> 00:19:57,074
Evan: The error analysis provided valuable insights into the varied failure modes across different evidence types and memory operations, emphasizing areas where LALMs need improvement.

185
00:19:57,126 --> 00:20:07,946
Ashley: Overall, VoxMem sets a new standard for benchmarking the full scope of spoken conversational memory, providing a solid foundation for future research and development in this field.

186
00:20:07,998 --> 00:20:10,278
Evan: That's a wrap on today’s episode.

187
00:20:10,278 --> 00:20:12,818
Thanks for tuning in to Daily Paper Cast.

188
00:20:12,870 --> 00:20:15,550
Ashley: We hope you found this discussion insightful.

189
00:20:15,550 --> 00:20:23,430
Be sure to join us next time as we dive into more papers from the world of AI, NLP, CV, and beyond.

190
00:20:23,558 --> 00:20:29,818
Evan: Until then, keep exploring, stay curious, and we'll see you in the next episode.