1
00:00:00,045 --> 00:00:00,735
Frankcx: Hey, everyone.

2
00:00:01,015 --> 00:00:06,305
Last week, I said that AI conversations
were moving from cost to control.

3
00:00:07,025 --> 00:00:08,605
Well, cost just crashed too.

4
00:00:09,645 --> 00:00:14,235
The best AI is getting dramatically
cheaper, and the AI you can

5
00:00:14,235 --> 00:00:17,895
run on your own device is also
getting dramatically better.

6
00:00:18,575 --> 00:00:24,745
this week's question is simple: If being
smart gets cheap, are the big AI companies

7
00:00:24,745 --> 00:00:26,485
still gonna charge a premium for?

8
00:00:27,325 --> 00:00:30,115
From cost to control to what's left.

9
00:00:30,855 --> 00:00:31,325
find out

10
00:00:47,191 --> 00:00:49,261
Welcome back to The Local Host.

11
00:00:49,351 --> 00:00:52,881
I'm Frank, and this week,
I am your local host again.

12
00:00:53,811 --> 00:00:58,311
This is show colon 0002 port speak.

13
00:00:59,331 --> 00:01:02,921
hey, let's do a little bit of around the
horn with, Jacob, Chauncey, and Robert.

14
00:01:02,961 --> 00:01:06,851
And we know Neal is, on a
important call, so he will join us.

15
00:01:06,951 --> 00:01:10,351
You know, customers come first, but
he'll, join us as soon as he's available.

16
00:01:10,371 --> 00:01:12,211
So hey, Jacob, how are

17
00:01:12,279 --> 00:01:13,779
Jacob: Hi, I'm doing great, Frank.

18
00:01:13,809 --> 00:01:15,059
Yeah, I'm Jacob.

19
00:01:15,169 --> 00:01:17,629
I work for Microsoft, so
that's an important disclosure.

20
00:01:17,879 --> 00:01:20,929
I love telling stories about
on-device productivity, including AI.

21
00:01:20,989 --> 00:01:21,569
Excited to be here.

22
00:01:23,009 --> 00:01:23,609
Frankcx: Awesome.

23
00:01:23,940 --> 00:01:24,240
Chauncey: Awesome.

24
00:01:24,450 --> 00:01:24,810
Yep.

25
00:01:24,880 --> 00:01:27,940
Chauncey Larson, tech nerd,
also worked for Microsoft

26
00:01:28,253 --> 00:01:31,153
Frankcx: And our hockey fan from Florida.

27
00:01:31,859 --> 00:01:33,569
Robert: It almost seems like an oxymoron.

28
00:01:34,329 --> 00:01:35,679
I am a little bit of both.

29
00:01:36,019 --> 00:01:41,109
And I'm Robert, and, I too am, employed
by Microsoft under full disclosure,

30
00:01:41,929 --> 00:01:46,509
I'm gonna do my best to match Frank's,
electric energy on today's show

31
00:01:47,846 --> 00:01:50,326
Chauncey: You are bringing the best
hat of the crew though, that's for sure

32
00:01:51,015 --> 00:01:53,505
Frankcx: yeah, that, the
Copilot, "I am your father."

33
00:01:54,018 --> 00:01:54,388
Chauncey: Love it

34
00:01:54,745 --> 00:01:55,835
Frankcx: Hey, let's get into it.

35
00:01:55,885 --> 00:01:57,305
It's been a wild week.

36
00:01:57,335 --> 00:02:02,635
I swear every day I look at the news and I
see something else that's coming in around

37
00:02:02,665 --> 00:02:06,065
AI and it's just an ever-evolving, piece.

38
00:02:06,065 --> 00:02:08,755
But we have this section of the
show we call the rundown, which

39
00:02:08,755 --> 00:02:10,195
is kind of news of the week.

40
00:02:10,245 --> 00:02:13,825
Hopefully you're listening to this and
you understood that, as we're recording

41
00:02:13,825 --> 00:02:18,205
this, it's been reported that NVIDIA,
is thinking of buying Hugging Face.

42
00:02:18,555 --> 00:02:19,475
What's Hugging Face?

43
00:02:19,475 --> 00:02:23,725
Well, Hugging Face is essentially where
all the AI models are stored, from

44
00:02:23,725 --> 00:02:27,475
cloud-based models that can be deployed
up to a trillion parameters, all the way

45
00:02:27,475 --> 00:02:29,875
down to what can run on a Copilot+ PC.

46
00:02:30,175 --> 00:02:33,615
If you want to find it, there's a little
slider bar that says, "I have this much

47
00:02:33,615 --> 00:02:37,275
memory," or, "This is the kind of device
I have," and it'll tell you essentially

48
00:02:37,275 --> 00:02:38,615
the model that you can download.

49
00:02:38,995 --> 00:02:42,255
It's literally where… It's
the App Store of, models.

50
00:02:42,295 --> 00:02:47,465
And it's super interesting that NVIDIA,
I always say is essentially the company

51
00:02:47,465 --> 00:02:51,475
that's selling the shovels for you to
go out and mine AI, is also the company

52
00:02:51,475 --> 00:02:56,115
that's now becoming the claim where you
come in and tell, the AI where it is that

53
00:02:56,115 --> 00:02:58,695
you've, found AI where you can go get it.

54
00:02:58,755 --> 00:02:59,225
hey, Neil.

55
00:02:59,455 --> 00:03:02,425
We were just talking about,
NVIDIA buying Hugging Face.

56
00:03:02,695 --> 00:03:06,085
Let's, see if anybody has any
opinions towards this announcement

57
00:03:06,917 --> 00:03:10,007
Jacob: The discourse online has
been pretty rampant about this.

58
00:03:10,107 --> 00:03:11,717
Kind of felt like it was a matter of time.

59
00:03:13,707 --> 00:03:17,527
It seems like, people are
generally positive to Nvidia.

60
00:03:17,527 --> 00:03:22,007
I think it's a better option than
some of the other model providers, and

61
00:03:22,007 --> 00:03:25,547
model creators because Nvidia has been
moving in an open source direction

62
00:03:25,547 --> 00:03:30,177
already with some of their in-house
models, Nemotron and things like that.

63
00:03:30,187 --> 00:03:32,207
So I think it's par for the course.

64
00:03:32,227 --> 00:03:33,997
Probably a good thing, but that's my take.

65
00:03:35,395 --> 00:03:37,165
Robert: Does it dilute their focus though?

66
00:03:37,295 --> 00:03:39,575
I mean, I don't think
Hugging Face makes money,

67
00:03:39,885 --> 00:03:40,335
Chauncey: Yet.

68
00:03:40,515 --> 00:03:41,705
It doesn't make money yet

69
00:03:42,064 --> 00:03:44,314
Jacob: I heard they
actually do make money.

70
00:03:44,644 --> 00:03:47,224
They do cover their costs, but they
have a ton of investment from a

71
00:03:47,224 --> 00:03:50,074
whole bunch of different areas, so
that's not their focus, I don't think

72
00:03:50,601 --> 00:03:50,738
Chauncey: do they just make money?

73
00:03:50,738 --> 00:03:52,455
'Cause that might be also
the other thing, right?

74
00:03:57,583 --> 00:03:59,493
Frankcx: I think it's
a double down, right?

75
00:03:59,543 --> 00:04:02,673
I mean, I think it's a double down into
the idea of what an open weight model

76
00:04:02,693 --> 00:04:04,663
is and why it's important to NVIDIA.

77
00:04:04,753 --> 00:04:09,143
I mean, NVIDIA is known as the company
that essentially builds cloud-based

78
00:04:09,203 --> 00:04:14,113
AI, but they're also, with their
announcements and deployments of DGX

79
00:04:14,113 --> 00:04:18,833
Spark and RTX Spark coming, you know,
the idea would be is that now you have

80
00:04:18,833 --> 00:04:22,973
models that can run wherever you want
intelligence to be, whether that's on

81
00:04:22,973 --> 00:04:25,113
your device or in your data center.

82
00:04:25,133 --> 00:04:29,223
And Hugging Face is the tool that
deploys those different solutions.

83
00:04:29,233 --> 00:04:33,293
So to me, it aligns really
well with what NVIDIA is doing.

84
00:04:33,743 --> 00:04:36,434
But, you know, I think, time
will tell if they change

85
00:04:36,734 --> 00:04:36,924
Jacob: Well I

86
00:04:36,924 --> 00:04:39,464
mean, there's just a risk they
lose their users, like the

87
00:04:39,554 --> 00:04:41,354
community forks and like…

88
00:04:41,911 --> 00:04:42,261
Chauncey: Yeah

89
00:04:42,364 --> 00:04:44,784
Jacob: I think they've just
got the users and the models,

90
00:04:44,884 --> 00:04:46,314
and so that's like the market.

91
00:04:47,554 --> 00:04:48,854
Neil, you were gonna say something?

92
00:04:48,910 --> 00:04:53,820
Neil: I saw an analogy
that NVIDIA's chips, right?

93
00:04:53,830 --> 00:04:59,200
These are the ovens, and Hugging Face,
they have the library of recipes,

94
00:04:59,400 --> 00:05:02,700
and if you control the oven and
the recipes, you can start, really

95
00:05:02,700 --> 00:05:04,810
delivering outsized value to the market.

96
00:05:05,100 --> 00:05:09,890
I think it's also important to note that
this isn't, NVIDIA's first attempt, to,

97
00:05:10,303 --> 00:05:10,683
Chauncey: Hmm

98
00:05:10,820 --> 00:05:15,810
Neil: they actually invested, in
Hugging Face back in 2023, alongside

99
00:05:15,810 --> 00:05:18,160
Salesforce, Google, Amazon, and IBM.

100
00:05:18,510 --> 00:05:22,380
Last year, NVIDIA wanted to follow
on with another 500 million.

101
00:05:22,410 --> 00:05:25,280
Hugging Face said, "No, we don't wanna
give outsized control to a single

102
00:05:25,280 --> 00:05:31,110
investor." Potential hot take here would
be, did the, breach of Hugging Face,

103
00:05:31,140 --> 00:05:32,760
lead potentially to this acquisition?

104
00:05:33,060 --> 00:05:33,770
Frankcx: Oh

105
00:05:33,809 --> 00:05:35,539
Neil: need, we do need support from

106
00:05:35,974 --> 00:05:36,244
Hmm

107
00:05:36,369 --> 00:05:36,389
to,

108
00:05:37,284 --> 00:05:37,924
Chauncey: Interesting

109
00:05:38,279 --> 00:05:41,259
Neil: some of the security and,
and maximize, the, the value

110
00:05:41,259 --> 00:05:42,429
of this entity we've created"?

111
00:05:42,449 --> 00:05:44,129
So couple talking points there.

112
00:05:44,129 --> 00:05:45,409
This isn't, out of the blue.

113
00:05:45,479 --> 00:05:49,679
NVIDIA actually has, quite a extensive
history, including being on Hugging

114
00:05:49,879 --> 00:05:51,279
Face's cap table, three years ago

115
00:05:52,080 --> 00:05:55,340
Robert: Does that become a
hotspot for the regulators if they

116
00:05:55,340 --> 00:05:57,020
own the recipes and the ovens?

117
00:06:00,360 --> 00:06:02,930
Chauncey: I mean there's still--
Yeah, but there's still a decent

118
00:06:02,930 --> 00:06:03,910
amount of competition here.

119
00:06:03,910 --> 00:06:05,700
Maybe not great competition, right?

120
00:06:05,700 --> 00:06:07,540
But there are still other
platforms out there.

121
00:06:07,540 --> 00:06:11,270
So what grounds do the hawks
have to stand on in this case?

122
00:06:11,367 --> 00:06:13,747
Jacob: there's nothing preventing
those users from going to another

123
00:06:13,864 --> 00:06:14,284
Chauncey: Yeah

124
00:06:15,417 --> 00:06:16,157
Jacob: And that's why I think,

125
00:06:16,157 --> 00:06:17,027
Chauncey: Or even starting one

126
00:06:17,912 --> 00:06:21,452
Jacob: especially for things
like, uncensored models that

127
00:06:21,462 --> 00:06:23,302
strip away some of the guardrails.

128
00:06:23,602 --> 00:06:28,502
Like for that portion of the local AI
community, I think they will go elsewhere,

129
00:06:29,382 --> 00:06:30,872
away from the corporate entities

130
00:06:32,763 --> 00:06:35,783
Frankcx: Well, leaning into what Neil
was saying around they have a history

131
00:06:35,783 --> 00:06:38,663
of doing this, I used to work for
this company called Silicon Graphics

132
00:06:38,663 --> 00:06:42,533
that was making big time, you know,
graphics engines, and then NVIDIA

133
00:06:42,593 --> 00:06:45,443
comes along and literally scales it
so that they give it to everybody.

134
00:06:45,533 --> 00:06:49,533
So it's like they have a history
of finding out ways to scale their

135
00:06:49,533 --> 00:06:53,713
business to bring everybody into
the market, and not just somebody

136
00:06:53,713 --> 00:06:54,873
that can own a data center.

137
00:06:54,963 --> 00:06:59,003
So I think that that's a big play for
an NVIDIA here is like their customer

138
00:06:59,023 --> 00:07:02,943
isn't just the big giant data centers,
it's literally you and me, and, you know,

139
00:07:02,993 --> 00:07:08,803
anybody that can afford a 3090 card that's
used or buy a new RTX Spark device, you

140
00:07:08,803 --> 00:07:10,063
know, that's who they're playing for.

141
00:07:10,103 --> 00:07:11,693
So I like it personally.

142
00:07:11,763 --> 00:07:13,693
I think it helps our story, for local AI,

143
00:07:13,993 --> 00:07:14,423
Robert: hopefully it'll

144
00:07:14,505 --> 00:07:14,895
Chauncey: fingers

145
00:07:14,903 --> 00:07:16,303
Robert: the prices of RAM down

146
00:07:17,399 --> 00:07:17,949
Frankcx: Right.

147
00:07:18,329 --> 00:07:18,839
Robert: Maybe they can

148
00:07:18,975 --> 00:07:19,515
Chauncey: Mm-hmm.

149
00:07:21,257 --> 00:07:21,797
Frankcx: You're right.

150
00:07:21,877 --> 00:07:25,617
Yeah, I mean, that is an interesting…
Maybe that's a show coming up of like, is

151
00:07:25,617 --> 00:07:29,147
local AI gonna take some of the heat and
pressure off a data center deployment?

152
00:07:29,267 --> 00:07:31,077
You know, like, how do we balance that?

153
00:07:31,197 --> 00:07:36,767
But there will always be, in my mind,
cloud AI, but, what level of it balances

154
00:07:36,767 --> 00:07:41,267
in the future is up to be seen, So moving
on to the next topic of news that I'm

155
00:07:41,267 --> 00:07:46,057
bringing to the table today is around
Qwen, which is our favorite Chinese model.

156
00:07:46,057 --> 00:07:47,037
At least it's one of mine.

157
00:07:47,037 --> 00:07:48,277
It's pretty capable.

158
00:07:48,537 --> 00:07:50,137
It's got amazing intelligence.

159
00:07:50,407 --> 00:07:54,387
It scales or quantizes,
if I'm saying that right.

160
00:07:54,657 --> 00:07:57,947
But the interesting thing that
Qwen, it had released a model

161
00:07:57,947 --> 00:07:59,677
3.8 a little over a week ago.

162
00:07:59,677 --> 00:08:00,517
We benchmarked it.

163
00:08:00,807 --> 00:08:04,427
But it also released last week or
this week this interesting model

164
00:08:04,457 --> 00:08:10,197
called Qwen3.8-Flash.NEXT, which is
actually a preview towards Qwen4.

165
00:08:10,537 --> 00:08:14,557
But it has some really interesting memory
capabilities, and I know Jacob, you've

166
00:08:14,557 --> 00:08:17,367
played around a little bit with this, so
maybe you can kind of tell us what you

167
00:08:17,424 --> 00:08:20,774
Jacob: I'm very, very passionate and
interested in just like experimentation,

168
00:08:20,794 --> 00:08:24,144
'cause at this point I'm just learning,
trying to be a sponge as much as possible.

169
00:08:24,144 --> 00:08:29,994
So the moment that it dropped, I had my
agents working on the best way to figure

170
00:08:29,994 --> 00:08:34,594
out how to install it, 'cause I knew
that it was 170, 180 billion parameters,

171
00:08:34,594 --> 00:08:38,454
and I was like, that seems like too much
for what I can fit on a consumer card.

172
00:08:39,284 --> 00:08:40,664
it seems like it's just too big.

173
00:08:40,904 --> 00:08:44,104
I've got that external 3090 that I
was thinking about running it on.

174
00:08:44,934 --> 00:08:48,964
And, I realized that there's a lot that
I'm still learning, about this lookup

175
00:08:48,994 --> 00:08:53,914
table, these Ngram tables that are
essentially like available for being

176
00:08:53,914 --> 00:08:57,734
able to… Well, they just don't need
the same availability as the rest of the

177
00:08:57,734 --> 00:08:59,654
model does, is the way I think of it.

178
00:08:59,944 --> 00:09:05,294
They can use slower memory, they
can use system RAM, and I heard some

179
00:09:05,294 --> 00:09:09,154
folks in the community talking about
being able to use SSD storage for it.

180
00:09:09,984 --> 00:09:13,744
so that was what was cool for me, and
I had some success with running the

181
00:09:13,744 --> 00:09:19,194
bulk of those model weights on SSD in
the Ngram tables so that it could look

182
00:09:19,204 --> 00:09:20,824
up and then a small amount on VRAM

183
00:09:22,429 --> 00:09:24,319
Robert: Now, is this an MOE model, Frank?

184
00:09:26,209 --> 00:09:27,939
Is it a mixture of experts model

185
00:09:28,078 --> 00:09:30,118
Jacob: it is an MoE model technically.

186
00:09:30,118 --> 00:09:32,998
I think there are 6 billion
parameters that are active at a time.

187
00:09:33,098 --> 00:09:33,988
I'll have to double-check.

188
00:09:34,068 --> 00:09:39,968
And then there's 51 billion parameters
that are in that n-gram table, and then

189
00:09:39,968 --> 00:09:46,508
there's another 120 are available that
should be on some fast RAM as well,

190
00:09:46,518 --> 00:09:51,238
but it doesn't need to be VRAM because
they're not active all the time being MoE

191
00:09:52,861 --> 00:09:55,761
Frankcx: Yeah, it feels like even though
you look at the size of the model and

192
00:09:55,761 --> 00:09:59,371
you're like, "That can't run on my
device," the way it loads the portions

193
00:09:59,371 --> 00:10:02,181
of the model that you need are kind
of in the chunks that you're asking

194
00:10:02,181 --> 00:10:04,011
for, not like everything all at once.

195
00:10:04,061 --> 00:10:04,779
So yeah, it's a,

196
00:10:05,079 --> 00:10:08,069
Robert: Which is the whole
premise of a MoE model, right?

197
00:10:08,069 --> 00:10:11,057
It's a combination of a bunch
of smaller models together,

198
00:10:11,357 --> 00:10:11,367
Chauncey: Mmhmm

199
00:10:11,479 --> 00:10:13,899
Robert: and it could pull
from any one of them.

200
00:10:13,909 --> 00:10:16,789
there's a component that lives in
there that says, "Oh, let's use this

201
00:10:16,789 --> 00:10:19,899
model for this task and this model
for that one," which is awesome.

202
00:10:19,909 --> 00:10:20,829
It's super cool

203
00:10:21,958 --> 00:10:23,138
Jacob: Yeah, I think
there's something like

204
00:10:23,201 --> 00:10:23,451
Frankcx: that,

205
00:10:23,751 --> 00:10:28,049
Jacob: 512 experts, and each one of those
experts is a certain size in this model.

206
00:10:28,349 --> 00:10:28,689
Frankcx: Right

207
00:10:28,888 --> 00:10:31,138
Jacob: I'm still learning,
but it's very interesting.

208
00:10:31,831 --> 00:10:35,281
Frankcx: I tried to run it on the CPU
on my device, which is a Xeon-based

209
00:10:35,281 --> 00:10:38,621
device, and it was getting like four
tokens per second, which is like

210
00:10:38,901 --> 00:10:41,891
barely… Like, if you think of a
token as like three quarters of a word,

211
00:10:42,191 --> 00:10:45,274
that's like somebody is slurring their
words, like they're having trouble

212
00:10:45,524 --> 00:10:50,834
ran it on the 3090, it got up to
50 tokens per second, but because

213
00:10:50,834 --> 00:10:54,484
it was just chunking along, it
had trouble finishing its tasks.

214
00:10:54,484 --> 00:10:57,204
So, there's definitely some
tuning that has to be done.

215
00:10:57,514 --> 00:11:00,554
I think in the end, what I found
is it's not built for the 3090.

216
00:11:00,594 --> 00:11:06,534
It has this weird lookup table that's
26 gigs and, a 3090 only has 24 gigs

217
00:11:06,534 --> 00:11:09,234
of VRAM, so it just doesn't fit right.

218
00:11:09,614 --> 00:11:15,374
It'll be interesting when, RTX Spark or
DGX Spark or, other devices that ha- Like,

219
00:11:15,614 --> 00:11:19,884
Neil, I'd be interested if you tested
it on a Mac that has unified memory, how

220
00:11:19,894 --> 00:11:23,624
that would run, because it seems like it's
more geared towards that type of hardware

221
00:11:23,954 --> 00:11:26,604
than it is something like a 3090 card.

222
00:11:29,299 --> 00:11:30,999
Neil: Got my homework for next week, so

223
00:11:31,756 --> 00:11:32,196
Frankcx: All right

224
00:11:32,496 --> 00:11:35,479
Neil: received some feedback folks
liked the, our ability to translate

225
00:11:35,489 --> 00:11:37,339
technical topics into business terms.

226
00:11:37,359 --> 00:11:40,919
I think pausing for a moment on Mixture
of Experts, all the rage right now.

227
00:11:41,199 --> 00:11:43,169
Maybe this analogy fits, maybe it doesn't.

228
00:11:43,409 --> 00:11:48,039
I like to view Mixture of Experts as,
analogous to the Encyclopedia Britannica

229
00:11:48,169 --> 00:11:50,209
bookshelf at my grandparents' house.

230
00:11:50,529 --> 00:11:50,799
You've got

231
00:11:50,992 --> 00:11:51,662
Frankcx: Oh, right

232
00:11:51,909 --> 00:11:52,679
Neil: of all

233
00:11:52,753 --> 00:11:53,263
Chauncey: All right

234
00:11:53,509 --> 00:11:54,149
Neil: topics.

235
00:11:54,589 --> 00:11:55,309
If you know

236
00:11:55,692 --> 00:11:57,902
Frankcx: B through C-A

237
00:11:57,959 --> 00:12:00,399
Neil: You don't need to open all
of the books at the same time.

238
00:12:00,639 --> 00:12:03,602
So, just, again, wanna pause for a
moment, define Mixture of Experts.

239
00:12:03,902 --> 00:12:04,592
Frankcx: love it

240
00:12:04,679 --> 00:12:09,339
Neil: this is going to allow us to
leverage larger parameter models

241
00:12:09,589 --> 00:12:12,749
without, running into some of the
memory bandwidth constraints that we've

242
00:12:12,749 --> 00:12:16,699
seen with, previous, models before
Mixture of Experts became, top of mind

243
00:12:17,616 --> 00:12:19,406
Robert: And I'll share
some stuff later on,

244
00:12:19,416 --> 00:12:19,726
Chauncey: Oh, nice

245
00:12:19,934 --> 00:12:21,594
Robert: this week with MoE models

246
00:12:22,828 --> 00:12:23,608
Frankcx: Oh, nice.

247
00:12:23,848 --> 00:12:27,618
Neil, given that you're talking about
unified memory and kind of memory

248
00:12:27,618 --> 00:12:30,678
management, I think you've been
doing some stuff with Apple as well.

249
00:12:30,678 --> 00:12:34,608
So what can you tell us about
some of the Mac stuff or the news

250
00:12:34,608 --> 00:12:36,048
that you're excited about for Mac?

251
00:12:36,807 --> 00:12:41,097
Neil: So big announcements
this week from the Apple world.

252
00:12:41,147 --> 00:12:47,917
They announced M5, M6 processors, Mac
Minis, which became all the rage with the

253
00:12:47,947 --> 00:12:50,857
OpenCL moment earlier this calendar year.

254
00:12:51,127 --> 00:12:57,867
We're continuing to see Apple
double down on this local AI moment.

255
00:12:57,887 --> 00:13:01,077
The twenty twenty-six was
supposed to be the year of agents.

256
00:13:01,387 --> 00:13:05,197
I think one can make an argument that it's
becoming the year of local and hybrid AI.

257
00:13:05,197 --> 00:13:05,702
Um, Apple continues to, um, innovate
at the silicon layer, uh, dropping

258
00:13:05,702 --> 00:13:09,772
down to a two you look at some of the
models like, Kimi V3, right, and three

259
00:13:09,772 --> 00:13:13,792
hundred and fifty gigs, you see almost
a direct parallel between the, the

260
00:13:13,802 --> 00:13:17,522
type of silicon and, and these types of
MoE models that we're seeing released.

261
00:13:17,772 --> 00:13:22,662
Ran some tests on an M3 Max,
versus the thirty ninety.

262
00:13:22,662 --> 00:13:24,702
So Frank, thanks for sharing
those benchmark tests.

263
00:13:24,992 --> 00:13:27,592
I, I think it boils down to
what are you trying to get

264
00:13:27,612 --> 00:13:29,542
out of these local AI models?

265
00:13:29,772 --> 00:13:35,062
Ran those five scenarios, and the
consistent takeaway was anything

266
00:13:35,062 --> 00:13:40,362
that's bandwidth bound, memory bound,
Mac is outpacing the competition.

267
00:13:40,672 --> 00:13:45,122
When it comes to raw compute, NVIDIA's,
you know, taking the cake, right?

268
00:13:45,372 --> 00:13:47,412
Apple is two to six X slower.

269
00:13:47,412 --> 00:13:52,202
So again, that's a, a couple generations
back, but you're seeing Apple continue

270
00:13:52,202 --> 00:13:56,372
to double down on these, bandwidth,
memory bound constraints, and you're

271
00:13:56,372 --> 00:14:00,422
seeing NVIDIA to, you know, continuing
to outpace the competition when

272
00:14:00,422 --> 00:14:02,102
it comes to, to raw compute power.

273
00:14:02,102 --> 00:14:06,142
So it'll be interesting to see the
shift in paradigm and, again, lots of,

274
00:14:06,142 --> 00:14:08,882
excitement for Mac enthusiasts, this week.

275
00:14:09,192 --> 00:14:12,812
Curious any other takes from the crew
here on, Apple's recent announcements?

276
00:14:13,431 --> 00:14:16,261
Chauncey: I mean, I would say as a,
fellow hardware nerd and just tracking

277
00:14:16,261 --> 00:14:19,461
these other guys, like, this is only just
beneficial to the entire ecosystem, right?

278
00:14:19,461 --> 00:14:21,881
If Apple's going this way,
Windows has to follow.

279
00:14:21,911 --> 00:14:25,711
Like, as an ecosystem, of course, if
you think about Surface and Dell and

280
00:14:25,711 --> 00:14:28,871
Lenovo and everything else that's out
there, like, it's so critical that

281
00:14:28,871 --> 00:14:30,771
we all continue to match this space.

282
00:14:30,781 --> 00:14:34,321
So I'm always excited when new things
come out because that means that there's

283
00:14:34,321 --> 00:14:37,541
gonna be new things on the horizon for
everyone else and we get some, I mean,

284
00:14:37,541 --> 00:14:38,961
it's just gonna continue to get better.

285
00:14:39,091 --> 00:14:42,481
Just imagine the amount of power we
have now and what we're able to achieve,

286
00:14:42,491 --> 00:14:45,391
which all you guys are talking about
so far of, like, testing locally.

287
00:14:45,531 --> 00:14:47,771
What we'll be able to do in
just a year or two years it's

288
00:14:47,771 --> 00:14:49,221
just exponential from here out.

289
00:14:50,100 --> 00:14:50,150
Robert: I

290
00:14:50,150 --> 00:14:52,830
wouldn't exactly say
Windows is following, right?

291
00:14:52,830 --> 00:14:55,910
But just to be clear, I mean, there's
some pretty amazing things that are

292
00:14:55,910 --> 00:15:00,533
happening in the Windows ecosystem Look,
Neil and I have had these conversations.

293
00:15:00,583 --> 00:15:05,373
It was just cloud, then it was cloud and
device, and now it's cloud, edge, and

294
00:15:05,373 --> 00:15:07,593
device in sort of a three-tiered model.

295
00:15:08,283 --> 00:15:10,713
once you get to the phone, you

296
00:15:10,713 --> 00:15:11,913
know, bets are off

297
00:15:12,811 --> 00:15:15,561
Frankcx: I think the thing that I see
is, look at what happened this past week.

298
00:15:15,581 --> 00:15:19,341
Two things: NVIDIA looking at Hugging
Face, that's a double down on local

299
00:15:19,341 --> 00:15:24,461
AI and open-weight models, and
then Apple, another multi-trillion

300
00:15:24,461 --> 00:15:28,371
dollar company, doubles down on
local AI themselves by releasing

301
00:15:28,371 --> 00:15:30,061
these Macs that are in that space.

302
00:15:30,101 --> 00:15:34,971
So I would argue, and I'm not with
Microsoft anymore, but I would argue

303
00:15:34,971 --> 00:15:40,041
that the governance that Microsoft is
gonna bring to these local AI spaces, I

304
00:15:40,041 --> 00:15:43,121
think that's why a lot of people should
be excited about what RTX Spark could

305
00:15:43,121 --> 00:15:47,881
bring, 'cause you're bringing this
capability of what CUDA and speeds are,

306
00:15:48,131 --> 00:15:51,701
but you're bringing the governance of
what you need, because many enterprise

307
00:15:51,721 --> 00:15:53,533
companies are scared to death of running,

308
00:15:53,833 --> 00:15:55,183
Robert: And they should be, yeah.

309
00:15:55,193 --> 00:15:56,053
Well, not Linux, but

310
00:15:56,063 --> 00:15:57,363
Of running agents for sure

311
00:15:57,873 --> 00:15:58,223
Chauncey: Yeah

312
00:15:58,666 --> 00:15:59,406
Jacob: I was talking to a customer

313
00:15:59,571 --> 00:16:02,341
Frankcx: being able to have that
governance that Microsoft can bring over

314
00:16:02,341 --> 00:16:07,061
top of what local AI can, bring from an
intelligence standpoint, there's some

315
00:16:07,061 --> 00:16:09,001
exciting things just around the corner

316
00:16:09,070 --> 00:16:14,310
Jacob: I was talking to a company in Japan
this week, and they were pretty frank

317
00:16:14,310 --> 00:16:18,840
with me, like: "Hey, right now we can't
run agents at all in our organization."

318
00:16:19,060 --> 00:16:20,990
And they are very AI forward.

319
00:16:20,990 --> 00:16:23,860
So AI forward, they actually have
a partnership now to run local

320
00:16:23,860 --> 00:16:25,960
models on all of their laptops.

321
00:16:26,030 --> 00:16:29,380
So they're running local models
everywhere, but those local models

322
00:16:29,420 --> 00:16:33,100
are not oper-- They're-- it's all
just connecting to the cloud or to

323
00:16:33,100 --> 00:16:35,060
specific outputs and not agented tasks.

324
00:16:35,060 --> 00:16:40,950
So they're like: "We want to run like
Scout or things like OpenClaw or, you

325
00:16:40,950 --> 00:16:46,080
know, a whole host of other agents."
And, I think there's a continued story

326
00:16:46,080 --> 00:16:50,040
that, Windows will continue to add value
with Microsoft Execution Containers,

327
00:16:50,280 --> 00:16:52,740
Agent 365, all of that goodness.

328
00:16:54,041 --> 00:16:55,671
Frankcx: Well, you guys will
get a kick out of this story.

329
00:16:55,671 --> 00:16:58,281
You may not know this, but I've been
playing around with doing a little bit

330
00:16:58,281 --> 00:17:00,719
of Uber driving myself, just on like a,

331
00:17:01,019 --> 00:17:02,059
Robert: In the sprinter?

332
00:17:02,081 --> 00:17:05,571
Frankcx: I want, I wanna get down
into Cap Hill and see what the

333
00:17:05,571 --> 00:17:09,261
cool kids are doing, so I'm driving
them around, dad in his Volvo.

334
00:17:09,651 --> 00:17:13,741
I picked up this guy, and we had this
conversation, and I'm like: "Hey, what

335
00:17:13,741 --> 00:17:17,651
do you do?" And he's like: "Oh I run
an aerospace company." I'm like: "Oh,

336
00:17:17,651 --> 00:17:21,221
that's pretty cool." I was like: "Are
you guys using much AI?" He's like:

337
00:17:21,481 --> 00:17:25,211
"I use it to write emails, but I can't
use it for engineering because we're

338
00:17:25,211 --> 00:17:29,231
doing contracts with Boeing and the
defense side that I can't share the IP

339
00:17:29,231 --> 00:17:30,881
of what we're doing with the cloud."

340
00:17:30,961 --> 00:17:34,311
And I was like: "Well, do you know much
about local AI?" And he's like: "Well,

341
00:17:34,311 --> 00:17:37,931
tell me." And I'm like, here's this
Uber driver telling him about local AI.

342
00:17:38,311 --> 00:17:41,581
He used to be in this band called
Fifth Angel, and he was the lead

343
00:17:41,581 --> 00:17:44,861
singer for it, and it's this Seattle
hair metal band that came out

344
00:17:44,891 --> 00:17:46,821
Robert: may- maybe his next
band will be called Air Gap

345
00:17:47,679 --> 00:17:48,699
Chauncey: Ooh, nice.

346
00:17:48,749 --> 00:17:49,449
That's good.

347
00:17:49,519 --> 00:17:50,039
Robert: Hair, you got a

348
00:17:50,369 --> 00:17:51,019
Chauncey: That's good.

349
00:17:51,229 --> 00:17:51,419
Robert: gap

350
00:17:51,549 --> 00:17:51,999
Frankcx: gap.

351
00:17:52,569 --> 00:17:54,439
Chauncey: Robert coming up
with the band names, man

352
00:17:55,199 --> 00:17:55,519
Frankcx: Yeah.

353
00:17:56,179 --> 00:17:58,239
Does anyone else have any
news they wanna share?

354
00:17:59,739 --> 00:18:02,509
All right, if not, it feels
like we should do a baseline.

355
00:18:02,559 --> 00:18:06,009
Baseline to me is like, Chauncey, I know
you're interested in making sure we're all

356
00:18:06,009 --> 00:18:08,019
talking speak that everybody understands.

357
00:18:08,019 --> 00:18:09,489
Maybe you can give us
a little bit of a clue

358
00:18:09,547 --> 00:18:12,347
Chauncey: This is also referred to
as the section of the podcast where

359
00:18:12,347 --> 00:18:14,097
it's like, I have no idea what you
guys are actually talking about.

360
00:18:14,097 --> 00:18:14,907
Can you please explain it?

361
00:18:14,957 --> 00:18:16,607
So can you do a couple things?

362
00:18:16,667 --> 00:18:17,807
One, I know some of it.

363
00:18:18,077 --> 00:18:20,467
Like what's the difference
when we're-- Well, first of

364
00:18:20,467 --> 00:18:21,967
all, like what is quantization?

365
00:18:21,977 --> 00:18:23,717
So you kind of mentioned
quantizing earlier.

366
00:18:23,717 --> 00:18:27,167
Like, I think it's probably really
worth diving a little bit into the

367
00:18:27,167 --> 00:18:28,967
details of what quantization means.

368
00:18:29,227 --> 00:18:32,207
And then I would actually love
to know your thoughts on the

369
00:18:32,207 --> 00:18:35,317
difference in cost per token and
how we're thinking about that too.

370
00:18:37,595 --> 00:18:41,275
Frankcx: Yeah, I mean, I hear a lot of
people talking about, the $25 token or the

371
00:18:41,275 --> 00:18:43,785
$6 token or, like, you know, what is that?

372
00:18:43,795 --> 00:18:46,225
Nobody's paying $6 for a token.

373
00:18:46,235 --> 00:18:49,285
Well, it comes down to
it's $6 per million tokens.

374
00:18:49,315 --> 00:18:54,745
And what the funny part is that this past
week, you know, Grok, which I think is

375
00:18:54,745 --> 00:18:58,965
Neil's favorite AI, like it's, I called
it last week your drunk uncle, but your

376
00:18:58,965 --> 00:19:03,955
drunk uncle sobered up and became a pretty
capable, model at $6 per million tokens,

377
00:19:03,955 --> 00:19:09,335
which honestly is like a fifth of the
cost of what, OpenAI and Claude charge

378
00:19:09,345 --> 00:19:12,315
for their, frontier models at $25 a token.

379
00:19:12,935 --> 00:19:16,345
But I think the point is, is that
because of the ability to start

380
00:19:16,345 --> 00:19:20,875
running these models locally, it's
starting to drive down, you know…

381
00:19:20,875 --> 00:19:25,495
And when Grok comes out with a $6 per
million, you know, token cost, that

382
00:19:25,495 --> 00:19:28,665
really makes everybody else look at
it like, "Hey, what are you doing

383
00:19:28,665 --> 00:19:32,395
there? And how come you're not charging
as much as we are?" Well, you know,

384
00:19:32,395 --> 00:19:35,815
they're not the premier, you know,
AI, but they are showing themselves

385
00:19:35,815 --> 00:19:37,345
to have just as much capability.

386
00:19:37,345 --> 00:19:40,965
And Neil, maybe you can talk a little bit
about Grok, but it'd be like, you know,

387
00:19:40,965 --> 00:19:45,005
it feels like there are models out there,
because they're all sharing algorithms,

388
00:19:45,005 --> 00:19:47,015
they're all becoming kind of equally smart

389
00:19:47,736 --> 00:19:51,706
Neil: Yeah, it's becoming more of
a commodity value-based discussion.

390
00:19:51,736 --> 00:19:54,286
The days of unlimited
tokens are over, right?

391
00:19:54,326 --> 00:20:00,046
As we mentioned, currently work at
Microsoft and we have set token limits.

392
00:20:00,076 --> 00:20:06,596
And so continuing to route every prompt
through frontier models, Opus 5, GPT

393
00:20:06,596 --> 00:20:09,156
Sol 5.6 no longer becomes practical.

394
00:20:09,166 --> 00:20:12,946
So, shifting some of the
development needs to Groq 4.6

395
00:20:13,166 --> 00:20:15,126
stretches my token budget further.

396
00:20:15,336 --> 00:20:19,466
And if it can do, as good or close to
as good of a job, if I can get more

397
00:20:19,466 --> 00:20:23,746
done with my existing token budget,
that's why, Groq, the drunken uncle,

398
00:20:23,746 --> 00:20:27,246
that sobered up, is now becoming, one
of my favorite models to work with.

399
00:20:27,306 --> 00:20:31,306
not all tokens are equal, dependent
upon which models you choose.

400
00:20:31,306 --> 00:20:34,776
And I think Groq offers, great
value at this moment in time

401
00:20:35,237 --> 00:20:37,857
Chauncey: I'm really interested
to see where we go with, model

402
00:20:37,857 --> 00:20:39,147
choice and making it easier.

403
00:20:39,147 --> 00:20:42,537
I mean, like the average person-- Your
developers, obviously, they're gonna

404
00:20:42,537 --> 00:20:45,287
know, like, "Oh, I should probably
choose something cheaper," and whatnot.

405
00:20:45,297 --> 00:20:48,697
But like the average person's just gonna
go, "Oh, this says it works better.

406
00:20:48,707 --> 00:20:52,397
I'm gonna click on work better."
Right? And, but ultimately, like I

407
00:20:52,397 --> 00:20:56,487
would say companies would value either
Microsoft or whoever making this kind

408
00:20:56,487 --> 00:21:00,397
of orchestration layer that says, "Hey,
for sure, I'm building a PowerPoint,

409
00:21:00,427 --> 00:21:02,457
I'm gonna have this level of model.

410
00:21:02,607 --> 00:21:05,207
But if I'm just answering a simple
question, I should be able to like

411
00:21:05,247 --> 00:21:06,887
dynamically switch to another model."

412
00:21:06,887 --> 00:21:09,687
I think that's where I'm really
excited to see some of this happen.

413
00:21:09,687 --> 00:21:12,547
And I know we have some auto
settings in, Cowork and in

414
00:21:12,547 --> 00:21:14,377
Scout, but not quite there yet.

415
00:21:14,387 --> 00:21:18,307
It's still a pre-selected handful
of models versus if it can really

416
00:21:18,307 --> 00:21:21,707
start dynamically going across
whatever your organization says.

417
00:21:21,717 --> 00:21:24,197
Like, I want only these
six models to ever run.

418
00:21:24,197 --> 00:21:26,691
I think that's when we're gonna start
seeing some pretty cool things happen.

419
00:21:26,991 --> 00:21:29,471
Robert: it goes down a whole
other layer that general

420
00:21:29,471 --> 00:21:30,821
users don't understand, right?

421
00:21:30,871 --> 00:21:31,431
Chauncey: Well, yeah.

422
00:21:31,672 --> 00:21:31,911
Yeah

423
00:21:31,972 --> 00:21:36,001
Robert: know, sonnet or whatever,
then there's like medium, max,

424
00:21:36,301 --> 00:21:36,761
Chauncey: Yep.

425
00:21:37,061 --> 00:21:39,052
Robert: extra large, extra max,

426
00:21:39,190 --> 00:21:39,450
Frankcx: Right

427
00:21:39,682 --> 00:21:40,782
Robert: big max,

428
00:21:40,862 --> 00:21:41,302
Chauncey: Plus

429
00:21:41,360 --> 00:21:44,030
Frankcx: wondering, to bring it
back to Chauncey's original question

430
00:21:44,040 --> 00:21:47,020
around quantization, I'm wondering
who can take a shot at that.

431
00:21:47,030 --> 00:21:49,460
And when you quantize
something, does it make it less

432
00:21:49,845 --> 00:21:53,785
Jacob: Well I'll do my first pass, and
then I'll get corrected by Robert or

433
00:21:53,785 --> 00:21:57,195
Neil, is probably gonna be my choice here.

434
00:21:57,495 --> 00:22:00,903
The way I think of quants is I think
of it like any other compression.

435
00:22:01,465 --> 00:22:04,905
Like you're making the model and
the weights take up less space,

436
00:22:05,775 --> 00:22:11,165
a result, you are potentially
cutting some of the detail.

437
00:22:11,945 --> 00:22:16,505
Just like when you have a JPEG
image that's compressed, you may

438
00:22:16,505 --> 00:22:18,015
be cutting out some of the detail.

439
00:22:18,015 --> 00:22:20,645
It's doing some calculations of,
"Hey, we're grouping these pixels

440
00:22:20,645 --> 00:22:24,255
together. They look the same."
It's saving you space on disk.

441
00:22:24,255 --> 00:22:25,705
It's saving you space in RAM.

442
00:22:26,085 --> 00:22:32,015
And so if a model like was a seven billion
parameter model, like its full weights

443
00:22:32,185 --> 00:22:35,305
would probably be like double that.

444
00:22:35,365 --> 00:22:37,075
They would probably be like 14 gigs.

445
00:22:38,075 --> 00:22:44,305
if you, compress it down to Q8,
then it might be seven gigs.

446
00:22:44,315 --> 00:22:47,415
So it might cut the space in half, that's
kinda how I think about it, is like you

447
00:22:47,415 --> 00:22:51,035
may be trading off some precision, you may
not be able to trust the outputs in the

448
00:22:51,035 --> 00:22:55,955
same way, in a world where you have models
where you're running them to do a task,

449
00:22:55,995 --> 00:22:59,975
and then you have another model checking
its work, then maybe it makes sense.

450
00:23:00,045 --> 00:23:00,285
Chauncey: So

451
00:23:00,381 --> 00:23:00,621
Robert: Yep,

452
00:23:00,645 --> 00:23:00,895
Chauncey: it

453
00:23:01,341 --> 00:23:02,841
Robert: The trade-off is accuracy.

454
00:23:03,251 --> 00:23:03,571
How do you

455
00:23:03,675 --> 00:23:04,035
Chauncey: Hmm

456
00:23:04,811 --> 00:23:08,341
Robert: more, you know, tighter
without losing the accuracy?

457
00:23:08,341 --> 00:23:09,741
And that's the trick that goes with it.

458
00:23:10,001 --> 00:23:13,201
And look, we're even seeing
some models get down to one-bit

459
00:23:13,301 --> 00:23:15,131
quantization, which is insane.

460
00:23:15,481 --> 00:23:17,401
But that's where the technology is going.

461
00:23:17,791 --> 00:23:20,981
But that, that's always, as I understand
it, that's the challenge is how do you

462
00:23:20,981 --> 00:23:27,823
make it and use less, but still maintain
the accuracy and the credibility of the,

463
00:23:28,123 --> 00:23:28,163
Chauncey: S-

464
00:23:28,171 --> 00:23:29,131
Robert: the answers you get?

465
00:23:29,303 --> 00:23:30,313
Chauncey: how do you choose then?

466
00:23:30,453 --> 00:23:34,923
Like, if I have this option of a frontier
version that's gonna be obviously the most

467
00:23:34,923 --> 00:23:38,733
powerful, the most intelligent, and like,
that's kind of like what I would think I

468
00:23:38,733 --> 00:23:42,433
would want all the time versus something
that is more efficient, that is gonna be

469
00:23:42,433 --> 00:23:44,253
smaller, that I can maybe run locally.

470
00:23:44,533 --> 00:23:47,353
Like, if there is kind of that
more gut reaction of just like,

471
00:23:47,363 --> 00:23:48,323
well, then how do I choose?

472
00:23:48,343 --> 00:23:49,743
Like, what would be your
kind of answer then?

473
00:23:50,882 --> 00:23:55,492
Jacob: I would say the highest quant
that you can the space that you have.

474
00:23:55,792 --> 00:24:00,142
So like the benchmarks for like 3.8, 27
billion parameters that we talked about

475
00:24:00,142 --> 00:24:03,432
last week, and, I think Robert, you might
have some more to share about this week.

476
00:24:03,722 --> 00:24:08,492
Like those benchmarks were generally ran
at full precision, like no quant at all.

477
00:24:08,892 --> 00:24:12,852
So you want it to be as precise
as possible, but there's a

478
00:24:12,852 --> 00:24:16,402
trade-off in speed and efficiency
and what can physically fit.

479
00:24:17,612 --> 00:24:18,222
That's my take.

480
00:24:19,797 --> 00:24:19,827
Frankcx: Yep.

481
00:24:21,607 --> 00:24:26,357
Well, getting to that, I think it's time
to go into the tension of the week, and

482
00:24:26,357 --> 00:24:29,847
I know that Robert, you said you'd be
willing to be on the hot seat, so I'm

483
00:24:29,955 --> 00:24:30,445
Robert: Why not, eh?

484
00:24:30,607 --> 00:24:31,289
Give it a go.

485
00:24:31,589 --> 00:24:31,689
Frankcx: it

486
00:24:31,877 --> 00:24:34,027
Robert: I'm always anxious
to get into the fight.

487
00:24:34,447 --> 00:24:38,227
So listen, we've been talking
about AI running it locally.

488
00:24:38,267 --> 00:24:39,907
Now we can run it ourselves.

489
00:24:40,217 --> 00:24:41,757
The models keep getting better.

490
00:24:42,047 --> 00:24:45,787
it, just obviously saves money than
using the ones… But the big cloud

491
00:24:45,787 --> 00:24:47,537
models keep getting cheaper as well.

492
00:24:48,337 --> 00:24:51,647
so the big question is: where
are the AI companies eventually

493
00:24:51,647 --> 00:24:52,617
gonna get their money?

494
00:24:52,637 --> 00:24:53,927
What are they gonna charge for?

495
00:24:54,337 --> 00:24:57,397
Wendell at Level 1 Tech
says the answer is speed.

496
00:24:57,487 --> 00:25:00,907
Not smarter answers, just
the same answer, but faster.

497
00:25:01,307 --> 00:25:04,987
the example I used, earlier off air
was, you know, I'm gonna order something

498
00:25:04,997 --> 00:25:07,367
from Amazon, it'll come in two days.

499
00:25:07,447 --> 00:25:10,527
If I want it in a day, I'm gonna
pay, two ninety-nine to get it.

500
00:25:10,537 --> 00:25:14,477
So we're already starting
to see it, in the AI world.

501
00:25:14,537 --> 00:25:16,377
If you wanna wait, you pay less.

502
00:25:16,387 --> 00:25:18,097
If you want it right now, you pay more.

503
00:25:18,447 --> 00:25:20,297
Same AI, just a different clock.

504
00:25:20,377 --> 00:25:23,897
And so, you know, the twenty-five
dollar, token is dead, right?

505
00:25:23,897 --> 00:25:28,227
I think within 18 months, the only
thing that frontier labs are-- really

506
00:25:28,227 --> 00:25:31,857
gonna be able to sell is the speed,
and not so much of the intelligence.

507
00:25:31,877 --> 00:25:35,267
I experienced some of this this
week, in loading, a thirty billion

508
00:25:35,267 --> 00:25:39,907
parameter model, onto one of my
devices, and I was blown away, at the

509
00:25:39,917 --> 00:25:44,157
speed and the accuracy and, just the
capability of it and how far it's come.

510
00:25:44,177 --> 00:25:48,317
So, in plain English, soon being
smarter will not be enough to

511
00:25:48,317 --> 00:25:51,507
charge more for, you're gonna have
to be faster, and that's my take.

512
00:25:52,969 --> 00:25:53,299
Frankcx: Yeah.

513
00:25:54,529 --> 00:25:57,779
I mean, you know, for me, I think
of it as like Neil's… He's Mr.

514
00:25:57,779 --> 00:26:01,839
Analogy, so I'll bring in the analogy
of like, to me, like a cloud AI

515
00:26:01,839 --> 00:26:03,569
is like hiring a consulting firm.

516
00:26:03,589 --> 00:26:05,079
They can bring in a lot of people.

517
00:26:05,359 --> 00:26:06,299
They can do a lot of work.

518
00:26:06,319 --> 00:26:07,619
They can scale very well.

519
00:26:07,639 --> 00:26:08,779
They're expensive.

520
00:26:09,259 --> 00:26:11,909
They don't know you very well, so
they have to take the time to learn

521
00:26:11,909 --> 00:26:13,419
you, you know, what you're doing.

522
00:26:14,139 --> 00:26:17,139
if you think about local AI,
could literally go out and hire

523
00:26:17,139 --> 00:26:20,919
an MBA from business school and
bring them in and train them.

524
00:26:21,119 --> 00:26:21,719
They're smart.

525
00:26:21,759 --> 00:26:24,689
They're just as smart as who you could
bring in from a consulting group,

526
00:26:25,059 --> 00:26:26,489
but they just can't do as much work.

527
00:26:26,539 --> 00:26:28,529
They don't have the same
volume that they can do.

528
00:26:28,539 --> 00:26:30,759
You could hire more and more of
them, but then you have to take

529
00:26:30,759 --> 00:26:34,419
on the responsibility of managing
them and taking care of them.

530
00:26:34,769 --> 00:26:38,979
for me, it's this, type of analogy
that the AI is getting smart enough

531
00:26:38,999 --> 00:26:42,199
that you can run it locally, and
then you're able to control privacy.

532
00:26:42,249 --> 00:26:44,569
That person isn't gonna take that
information and share it with

533
00:26:44,569 --> 00:26:46,009
some other consulting agency.

534
00:26:46,379 --> 00:26:48,979
It's sovereign, meaning that,
it's controlled as to where

535
00:26:48,999 --> 00:26:50,409
it goes and where it lives.

536
00:26:50,789 --> 00:26:55,079
It's local, so if something happened,
like, there was a storm and, things got

537
00:26:55,079 --> 00:26:58,339
knocked out, that consulting company
that, flies in every week would not be

538
00:26:58,339 --> 00:27:01,689
able to get to you, but that business
person can come to work that day.

539
00:27:02,099 --> 00:27:04,699
So there's a lot of
advantages of having local.

540
00:27:04,999 --> 00:27:07,259
But that doesn't mean that the
cloud's ever gonna go away.

541
00:27:07,279 --> 00:27:10,719
It's gonna be there for when you
need it and the tough problems.

542
00:27:11,039 --> 00:27:14,509
But I would argue that AI
being balanced is the way we're

543
00:27:14,509 --> 00:27:15,699
looking at AI in the future

544
00:27:19,475 --> 00:27:20,535
Chauncey: To your point about, the
analogy too, going back to speed,

545
00:27:20,975 --> 00:27:24,795
a bigger team of individuals may or
may not come to that answer faster.

546
00:27:24,825 --> 00:27:28,955
We can maybe throw more people at
that project, so you get that ability

547
00:27:28,965 --> 00:27:33,325
to get either a more cohesive answer
or a faster answer, because of

548
00:27:33,325 --> 00:27:34,425
that ability to throw more at it.

549
00:27:34,435 --> 00:27:35,745
Is that kinda the mindset here?

550
00:27:37,843 --> 00:27:38,085
Maybe?

551
00:27:38,143 --> 00:27:38,823
Frankcx: Yeah, I would agree

552
00:27:38,924 --> 00:27:43,064
Jacob: I don't know if I believe the
assertion fully, Robert,  because

553
00:27:43,089 --> 00:27:44,529
Robert: not say it with enough conviction?

554
00:27:45,459 --> 00:27:46,799
Chauncey: Ding, ding, let's get it on

555
00:27:46,946 --> 00:27:51,076
Jacob: I feel like everyone's just trying
to acquire users, and so it's a race

556
00:27:51,076 --> 00:27:54,596
to the bottom temporarily while they're
trying to prove that their models or

557
00:27:54,596 --> 00:27:57,366
their values, is where the user should be.

558
00:27:57,576 --> 00:28:01,646
it's to the point a couple weeks ago,
there's this Ox Alpha, like kind of

559
00:28:01,646 --> 00:28:05,296
like a stealth model on OpenRouter,
so you can choose if you're, running

560
00:28:05,296 --> 00:28:09,726
a cloud model via OpenRouter, you can
choose this Ox Alpha and it gives you a

561
00:28:09,766 --> 00:28:14,266
anonymous model that you don't know what
it is, but it's free to run and use.

562
00:28:14,286 --> 00:28:17,276
So you're, using a model, you don't know
what it is, you don't have any trust.

563
00:28:17,276 --> 00:28:20,206
It's in the cloud, it's not
local, but, it's free to use.

564
00:28:20,466 --> 00:28:22,286
And so they're trying to acquire

565
00:28:23,092 --> 00:28:24,932
Robert: Model roulette
is what that's called.

566
00:28:25,316 --> 00:28:25,926
Jacob: What's that?

567
00:28:26,226 --> 00:28:27,696
Robert: That's called model roulette.

568
00:28:27,996 --> 00:28:28,806
Jacob: Exactly.

569
00:28:28,842 --> 00:28:29,742
Robert: Good Lord

570
00:28:29,956 --> 00:28:34,406
Jacob: but everyone wants users, so, like,
they're gonna continue lowering the cost,

571
00:28:34,416 --> 00:28:36,376
but I don't think that that's sustainable.

572
00:28:36,856 --> 00:28:39,896
I'm not into the economics of it with
these companies, but, like, it just

573
00:28:39,896 --> 00:28:40,976
doesn't feel like that's sustainable.

574
00:28:40,976 --> 00:28:44,246
At some point, they're gonna, like,
entrench that they've got their

575
00:28:44,246 --> 00:28:46,566
users and the prices go back up

576
00:28:48,445 --> 00:28:51,835
Neil: I'll offer a
perspective in terms of speed.

577
00:28:52,225 --> 00:28:57,145
Robert, the argument that speed is
going to be what these frontier labs

578
00:28:57,225 --> 00:29:03,525
offer, how does that shift when we
get into asynchronous workflows?

579
00:29:03,525 --> 00:29:09,915
If I can craft my prompt to build the
code I need overnight, why would I

580
00:29:09,935 --> 00:29:12,035
pay a premium for a frontier model?

581
00:29:12,045 --> 00:29:15,645
Why would I pay the outsource
consultant, leveraging the example?

582
00:29:15,925 --> 00:29:21,585
Yes, maybe my MBAs take, much longer
to execute that task, but if I can set

583
00:29:21,585 --> 00:29:26,405
them up proactively and achieve the
desired result, certainly under some

584
00:29:26,405 --> 00:29:30,375
time constraint pressures, yeah, maybe I
need to outsource and execute a sprint.

585
00:29:30,775 --> 00:29:33,925
Does speed become less of a factor
when we start talking about,

586
00:29:34,205 --> 00:29:39,610
orchestration and autopilot workflows
that can accomplish work when

587
00:29:39,910 --> 00:29:40,230
Robert: But is,

588
00:29:40,425 --> 00:29:41,135
Neil: lid is shut?

589
00:29:41,400 --> 00:29:44,430
Robert: is that a builder
perspective or is that a general

590
00:29:44,430 --> 00:29:46,000
everyday user perspective?

591
00:29:47,561 --> 00:29:49,961
Neil: I would say right now
it's a builder perspective.

592
00:29:50,051 --> 00:29:53,671
I'm seeing AI from a knowledge
work perspective, a more one in

593
00:29:53,671 --> 00:29:57,121
one out, a more sophisticated Ask
Jeeves, Google search, if you will.

594
00:29:57,351 --> 00:30:00,961
Once we start talking asynchronous
workflows, which developers are, are

595
00:30:01,126 --> 00:30:01,416
Robert: Yeah.

596
00:30:01,536 --> 00:30:01,856
Yep

597
00:30:02,180 --> 00:30:02,540
Jacob: Yeah

598
00:30:02,641 --> 00:30:03,481
Neil: Right to that,

599
00:30:03,918 --> 00:30:04,168
Jacob: You

600
00:30:04,181 --> 00:30:06,761
Neil: think it's the developer
workflows where speed doesn't matter.

601
00:30:06,761 --> 00:30:11,611
I'm used to crafting a prompt, letting my
device cook for 10 minutes, 15 minutes.

602
00:30:11,801 --> 00:30:14,101
My benchmark test took an
hour and a half this morning,

603
00:30:14,101 --> 00:30:14,351
Hmm Hmm

604
00:30:14,454 --> 00:30:15,134
acceptable.

605
00:30:15,174 --> 00:30:15,454
So,

606
00:30:15,842 --> 00:30:16,142
Jacob: Yeah.

607
00:30:16,142 --> 00:30:16,582
It's like,

608
00:30:16,744 --> 00:30:17,802
Neil: may not matter for developer

609
00:30:18,102 --> 00:30:21,582
Jacob: Frank, if you're tossing, one
of those tasks like, you know, build an

610
00:30:21,602 --> 00:30:24,702
app, build whatever it is that you're
building these days, like your four tokens

611
00:30:24,712 --> 00:30:28,812
per second example of that, 180 billion
parameter model that you're running.

612
00:30:28,812 --> 00:30:31,572
Like, could theoretically just say,
"Hey, I'm gonna be gone for the

613
00:30:31,572 --> 00:30:33,047
weekend," like build me something

614
00:30:33,347 --> 00:30:33,667
Frankcx: Right

615
00:30:34,072 --> 00:30:37,842
Jacob: come back and you've got something
that's smarter and better and free so

616
00:30:38,061 --> 00:30:40,631
Frankcx: Especially these days when
you're saying, like, you're kicking off

617
00:30:40,631 --> 00:30:44,931
seven different Claude Costa sessions
and you're like, "Well, just go to

618
00:30:45,231 --> 00:30:45,525
Robert: Fleet.

619
00:30:45,605 --> 00:30:46,915
You've got a fleet, yeah

620
00:30:47,011 --> 00:30:48,831
Frankcx: You're like,
"Can you keep working on

621
00:30:49,059 --> 00:30:49,309
Robert: Does,

622
00:30:49,491 --> 00:30:50,021
Frankcx: please?"

623
00:30:50,129 --> 00:30:57,099
Robert: Does this equation change based on
what we saw from NVIDIA and Hugging Face?

624
00:30:57,109 --> 00:31:02,249
Now, maybe not in that exact
scenario, but thinking about

625
00:31:02,559 --> 00:31:09,449
consolidation in the industry does
that change your on any of this?

626
00:31:10,173 --> 00:31:12,403
Chauncey: Wait, so just kind
of rehashing that thought.

627
00:31:13,143 --> 00:31:18,733
Does NVIDIA potentially acquiring Hugging
Face affect this speed versus like

628
00:31:18,773 --> 00:31:19,403
They're only gonna

629
00:31:19,413 --> 00:31:20,293
be building for speed

630
00:31:20,593 --> 00:31:23,131
Robert: Yeah, but not that
specifically, but just in

631
00:31:23,131 --> 00:31:24,691
general industry consolidation.

632
00:31:24,711 --> 00:31:29,061
I mean, look at all the model
providers we have, the hardware

633
00:31:29,061 --> 00:31:31,071
manufacturers, the data centers.

634
00:31:31,121 --> 00:31:36,961
Like for example, Microsoft makes models,
build data centers, they provide services.

635
00:31:36,961 --> 00:31:39,181
I mean, are we gonna see more of that?

636
00:31:39,221 --> 00:31:42,311
And is that gonna affect your answer?

637
00:31:42,613 --> 00:31:44,943
Frankcx: there's some tools out
there that are like these routers

638
00:31:44,963 --> 00:31:48,403
now that can essentially look at…
They're like intelligence routers.

639
00:31:48,403 --> 00:31:51,173
They look at the question that's being
asked, and then they determine what

640
00:31:51,173 --> 00:31:53,873
model or what tool, this can run on.

641
00:31:53,883 --> 00:31:58,203
And I feel like within an industry, if
you had your own ability to route to

642
00:31:58,203 --> 00:32:02,043
local because you knew that job that
you were requesting based on the prompt

643
00:32:02,063 --> 00:32:07,713
could be handled by a local AI or a
cheaper model, then you would, right?

644
00:32:08,033 --> 00:32:12,823
So I feel like router, types of,
intelligence is gonna be involving

645
00:32:12,823 --> 00:32:14,623
local AI in the future as well

646
00:32:15,269 --> 00:32:15,749
Robert: And does

647
00:32:15,959 --> 00:32:15,969
Chauncey: Hmm

648
00:32:15,999 --> 00:32:18,179
Robert: equation change, I'm
gonna throw another curve ball

649
00:32:18,179 --> 00:32:20,179
here, with a three-tier model?

650
00:32:20,949 --> 00:32:26,029
If you have your device running a
local model, you have a server a

651
00:32:26,029 --> 00:32:31,449
server or blade in your house that
can run close to a frontier model,

652
00:32:32,259 --> 00:32:34,459
that change things even further?

653
00:32:36,883 --> 00:32:38,743
Chauncey: from a revenue perspective, like

654
00:32:38,863 --> 00:32:41,333
Robert: a, from a, what
are they gonna charge for?

655
00:32:41,491 --> 00:32:41,851
Chauncey: Yeah.

656
00:32:42,071 --> 00:32:43,431
Yeah, well that's, that's, yeah.

657
00:32:44,193 --> 00:32:44,411
Yeah

658
00:32:44,493 --> 00:32:47,043
Robert: I could probably get some
pretty serious speed off of it

659
00:32:47,946 --> 00:32:48,156
Jacob: You're,

660
00:32:48,405 --> 00:32:50,685
Chauncey: I mean it's, wonder
even, like, how we ended up making

661
00:32:50,685 --> 00:32:55,825
a decision, if you think about
even analogous to Office, right?

662
00:32:55,855 --> 00:33:00,565
At one point, we had Office running
locally, and, everyone questioned why the

663
00:33:00,565 --> 00:33:05,675
heck we moved, we being Microsoft, moved
to going to a cloud subscription-based

664
00:33:05,675 --> 00:33:09,355
model for, and now that that subscription
model is king for everyone, right?

665
00:33:09,355 --> 00:33:10,835
That is not just, a Microsoft thing.

666
00:33:10,885 --> 00:33:12,665
Subscription is king across the board.

667
00:33:12,915 --> 00:33:15,605
And so how do we start looking
at models in that same fashion?

668
00:33:15,605 --> 00:33:16,605
Are we reverting?

669
00:33:16,745 --> 00:33:17,365
I don't think so.

670
00:33:17,365 --> 00:33:21,495
I think there's still gonna be a play of
individuals that are gonna want access to

671
00:33:21,495 --> 00:33:24,805
the newest, the latest, the greatest, the
fastest of what the cloud has to offer

672
00:33:24,805 --> 00:33:27,635
and to that provides that scale, right?

673
00:33:27,685 --> 00:33:28,085
Frankcx: Yeah.

674
00:33:28,385 --> 00:33:29,328
Chauncey: where can the cloud grow?

675
00:33:29,328 --> 00:33:30,038
It can grow

676
00:33:31,675 --> 00:33:31,965
Frankcx: Yep.

677
00:33:32,635 --> 00:33:35,285
No, I think that's another show,
because I feel like there's this

678
00:33:35,285 --> 00:33:39,075
whole idea of, like, butts in seats
equals subscriptions times the number

679
00:33:39,075 --> 00:33:42,345
of people you have, and that doesn't
equal an AI model when it comes to

680
00:33:42,524 --> 00:33:42,944
Chauncey: Right

681
00:33:43,045 --> 00:33:46,285
Frankcx: But, hey, let's, let's
do a little vote around the horn.

682
00:33:46,345 --> 00:33:47,465
I'll jump in first.

683
00:33:47,785 --> 00:33:52,445
So I'm gonna support, Robert's,
theory that the $25 token is, a

684
00:33:52,445 --> 00:33:56,075
thing of the past, especially since
Grok and others are already leading.

685
00:33:56,135 --> 00:33:59,185
I don't know what that means towards
the economy and all the investment

686
00:33:59,195 --> 00:34:03,235
banking that takes place around AI.

687
00:34:03,235 --> 00:34:05,345
Hopefully they can pay their bills,
but, but that's the way I feel about it.

688
00:34:05,465 --> 00:34:05,835
Jacob?

689
00:34:07,131 --> 00:34:11,671
Jacob: I think that it's gonna get
expensive again before it gets cheaper.

690
00:34:12,461 --> 00:34:17,221
Local is here, we all think it, but I
think the market's not gonna be fully

691
00:34:17,221 --> 00:34:23,161
there and the costs are gonna go up
again, especially as licenses, like open

692
00:34:23,161 --> 00:34:27,841
source licenses are in the mix of like
what local models are able to do legally.

693
00:34:28,981 --> 00:34:32,921
I think that some of the Chinese AI firms
are going to… Well, we're seeing them

694
00:34:32,941 --> 00:34:36,271
add more, and not as open of licensing.

695
00:34:36,591 --> 00:34:41,241
So you're gonna see more commercial
customers want to-- they'll be paying

696
00:34:41,241 --> 00:34:44,471
more for it because it's in way
higher demand before it goes down.

697
00:34:44,811 --> 00:34:47,181
So I don't agree with
you yet, Robert, yet.

698
00:34:48,167 --> 00:34:48,277
Frankcx: All

699
00:34:48,441 --> 00:34:48,941
Robert: On the side

700
00:34:48,997 --> 00:34:49,737
Frankcx: Chauncey?

701
00:34:50,062 --> 00:34:51,912
Chauncey: Yeah, I'm on
the same boat as Jacob.

702
00:34:52,232 --> 00:34:55,332
Demand's gonna be high enough to continue
to maintain a higher price for now.

703
00:34:55,452 --> 00:34:56,362
Robert: Must be in marketing

704
00:34:57,710 --> 00:34:58,620
Frankcx: Yeah, they want the

705
00:34:58,920 --> 00:34:59,862
Chauncey: Or, yeah.

706
00:35:00,300 --> 00:35:01,150
Frankcx: And then, Neil

707
00:35:03,781 --> 00:35:06,481
Neil: Speed matters when
we're talking security.

708
00:35:06,981 --> 00:35:10,851
If we have bad actors using these
frontier models, we need access to

709
00:35:10,851 --> 00:35:12,671
these frontier models to combat.

710
00:35:12,871 --> 00:35:17,081
I'm going to agree with the stance
that frontier models deliver speed.

711
00:35:18,371 --> 00:35:20,301
The use cases for speed vary.

712
00:35:20,321 --> 00:35:22,931
I think there will continue to
be a premium for frontier models

713
00:35:22,941 --> 00:35:27,151
specifically to combat rising threat
actors leveraging the same models, us

714
00:35:27,987 --> 00:35:28,967
Jacob: And people will
pay for that, all right?

715
00:35:29,318 --> 00:35:31,558
Frankcx: All right, Robert,
you get the last word.

716
00:35:31,598 --> 00:35:31,968
Robert: I'm

717
00:35:31,972 --> 00:35:32,452
Frankcx: know where your

718
00:35:32,558 --> 00:35:33,378
Robert: I'm holding firm.

719
00:35:33,378 --> 00:35:37,688
Now, what we might see is different
tiers of speed you pay for.

720
00:35:37,888 --> 00:35:41,658
I think regardless of where it
goes, and I do think speed is gonna

721
00:35:41,658 --> 00:35:44,088
be, the first, shoe to drop here.

722
00:35:44,458 --> 00:35:49,208
Over the next six to 12 months, I think
we're gonna see some significant changes

723
00:35:49,218 --> 00:35:53,268
in how this looks, is good for us
'cause it gives us stuff to talk about

724
00:35:54,470 --> 00:35:54,710
Frankcx: Yep.

725
00:35:55,440 --> 00:35:57,350
right, let's, wrap up with on my device.

726
00:35:57,350 --> 00:35:59,540
Anybody have anything cool
that they're working on?

727
00:35:59,632 --> 00:36:00,482
Robert: Here's the short story.

728
00:36:00,782 --> 00:36:01,272
Frankcx: problem

729
00:36:01,572 --> 00:36:06,962
Robert: week I talked about my, struggle
to try and find out how we can get,

730
00:36:07,012 --> 00:36:09,762
on-device coding, a code agent running.

731
00:36:10,122 --> 00:36:14,392
And I of tripped into it or
backed into it by mistake.

732
00:36:14,732 --> 00:36:17,932
The short version of the story is
I loaded up a twenty-seven billion

733
00:36:17,932 --> 00:36:22,802
parameter model, connected it in,
thank you Jacob, into Copilot CLI.

734
00:36:23,502 --> 00:36:27,162
I typed in, "Let's make something,"
and then I got distracted.

735
00:36:27,652 --> 00:36:31,772
I turned on autopilot, which essentially
gave it full, permission to do

736
00:36:31,772 --> 00:36:34,132
whatever it wanted, and I walked away.

737
00:36:34,132 --> 00:36:37,272
And I came back, and it said, " Well,
you didn't respond in ten minutes, so

738
00:36:37,272 --> 00:36:42,902
I built you something." And it built an
online coding agent, and it said, " You

739
00:36:42,902 --> 00:36:47,592
should use a-- this MoE thirty billion
parameter three point eight Qwen

740
00:36:47,592 --> 00:36:50,882
model." and, and the rest is history.

741
00:36:51,282 --> 00:36:52,536
There's a lot more to it, but it

742
00:36:52,836 --> 00:36:53,776
Chauncey: Oh, that's awesome.

743
00:36:53,876 --> 00:36:57,546
Frankcx: All right, gentlemen, I
hope you all have a terrific weekend.

744
00:36:57,556 --> 00:37:05,036
So is no place like 127.0.0.1.

745
00:37:05,100 --> 00:37:05,820
Robert: hang with you fellas

746
00:37:05,886 --> 00:37:06,106
Frankcx: now

747
00:37:06,343 --> 00:37:06,703
Chauncey: See you.

748
00:37:07,120 --> 00:37:07,620
Jacob: Thanks, team

749
00:37:07,683 --> 00:37:08,093
Chauncey: Thanks guys