1
0:0:0,06 --> 0:0:2,06
Nik: Hello, hello, this is Postgre.FM

2
0:0:2,08 --> 0:0:5,8199997
My name is Nik, PostgresAI, and
as usual with me, Michael,

3
0:0:6,66 --> 0:0:7,16
pgMustard.

4
0:0:7,44 --> 0:0:8,18
Hi, Michael.

5
0:0:8,74 --> 0:0:9,5199995
Michael C.: Hi, Nik.

6
0:0:10,08 --> 0:0:13,5
Nik: And we have a very, very interesting
guest today, Michael Malis,

7
0:0:13,5 --> 0:0:18,7
who created pgrust, which has already,
I think, 5,000 stars

8
0:0:18,7 --> 0:0:22,38
on GitHub and a lot of noise around,
like a lot of buzz.

9
0:0:22,7 --> 0:0:24,439999
Hi, Michael, thank you for coming.

10
0:0:24,96 --> 0:0:25,76
Michael M.: Yeah, of course.

11
0:0:25,76 --> 0:0:27,04
Thank you for having me.

12
0:0:28,08 --> 0:0:32,56
Nik: So, of course, I think the
1st question should be how it

13
0:0:32,56 --> 0:0:33,78
all started, why?

14
0:0:33,84 --> 0:0:35,22
Tell us the story, please.

15
0:0:36,98 --> 0:0:40,52
Michael M.: So Jason and I were
looking for projects to work

16
0:0:40,52 --> 0:0:45,56
on, and we wanted to do something
that we knew super well.

17
0:0:45,86 --> 0:0:50,38
And initially, what we were focused
on was reliability and how

18
0:0:50,38 --> 0:0:53,14
can we help people make their websites
more reliable.

19
0:0:54,14 --> 0:0:56,7
And we were working with a bunch
of people.

20
0:0:57,1 --> 0:1:1,3
And the common pattern that emerged
was that A lot of the problems

21
0:1:1,36 --> 0:1:5,16
that caused reliability issues,
a lot of them stemmed from their

22
0:1:5,16 --> 0:1:8,22
database and how they were using
their database, whether it was

23
0:1:8,22 --> 0:1:10,54
like Postgres or Redis or some
other system.

24
0:1:11,12 --> 0:1:14,54
And so we started looking at what
are ways that we can solve

25
0:1:14,54 --> 0:1:15,26
this problem.

26
0:1:15,94 --> 0:1:19,64
And we were throwing around a couple
of things, and we ended

27
0:1:19,64 --> 0:1:23,66
up figuring out that AI is actually
pretty good at rewriting

28
0:1:23,66 --> 0:1:24,78
software at this point.

29
0:1:24,86 --> 0:1:27,1
This is back in April.

30
0:1:27,88 --> 0:1:31,84
And so we're like, hey, why don't
we actually try fixing these

31
0:1:31,84 --> 0:1:35,28
problems at the source and actually
trying to modify Postgres

32
0:1:35,98 --> 0:1:39,56
and fix a lot of the things that
cause people to have different

33
0:1:39,56 --> 0:1:43,2
reliability issues, whether it's
connections or it's like certain

34
0:1:43,2 --> 0:1:46,5
queries just take a really long
time or like 1 long-running query

35
0:1:46,5 --> 0:1:49,62
can take down your database or
vacuums as we're all familiar

36
0:1:49,64 --> 0:1:50,14
with.

37
0:1:50,28 --> 0:1:55,76
And so the idea behind pgrust,
the name's a bit misleading because

38
0:1:55,76 --> 0:1:58,14
it's not actually really about
Rust.

39
0:1:58,14 --> 0:2:0,52
I think that's actually the least
interesting part about it.

40
0:2:0,52 --> 0:2:5,7
It's actually more about how do
we leverage AI to re-architect

41
0:2:6,16 --> 0:2:10,14
Postgres and just build a much
better database.

42
0:2:11,38 --> 0:2:11,88
Nik: Okay.

43
0:2:12,12 --> 0:2:14,08
And when you say better, what do
you mean?

44
0:2:14,9 --> 0:2:15,4
Michael M.: Yeah.

45
0:2:15,66 --> 0:2:21,44
So there's a combination of like
different challenges that like

46
0:2:21,44 --> 0:2:23,94
I've just seen people repeatedly
have with Postgres and I actually

47
0:2:23,94 --> 0:2:27,78
wrote a blog post called it like
the 4 horsemen about a lot of

48
0:2:27,78 --> 0:2:29,24
the challenges that people have.

49
0:2:29,32 --> 0:2:32,64
And some of the most common ones
are like, you misconfigure your

50
0:2:32,64 --> 0:2:35,54
connection limit and like all of
a sudden like nothing can connect

51
0:2:35,54 --> 0:2:36,32
to your database.

52
0:2:36,74 --> 0:2:38,64
You have like JSON support.

53
0:2:38,64 --> 0:2:42,08
Tons and tons of people use JSON
but Postgres doesn't have statistics

54
0:2:42,54 --> 0:2:44,02
for how to actually query JSON.

55
0:2:44,02 --> 0:2:46,32
So when you try to do it, you're
going to get really bad query

56
0:2:46,32 --> 0:2:48,46
plans and everything's just going
to be really slow.

57
0:2:48,68 --> 0:2:52,66
There's like the wraparound vacuum
and you have more than several

58
0:2:52,66 --> 0:2:55,84
billion transactions and like all
of a sudden if your vacuum

59
0:2:55,84 --> 0:2:58,74
can't keep up your database falls
over and there's just like

60
0:2:58,74 --> 0:3:3,92
all these problems that have been
around for such a long time.

61
0:3:3,92 --> 0:3:7,94
And like Postgres is a great product
and like it is a really

62
0:3:7,94 --> 0:3:12,98
great system, but the way the Postgres
core team approaches things

63
0:3:12,98 --> 0:3:16,44
is they approach it in terms of
stability And how do we keep

64
0:3:16,44 --> 0:3:17,56
what we have today?

65
0:3:17,56 --> 0:3:19,14
Like, how do we keep it working?

66
0:3:19,2 --> 0:3:21,26
And how do we make sure we don't
break it?

67
0:3:21,42 --> 0:3:24,44
Because like, there's like billions
of Postgres instances out

68
0:3:24,44 --> 0:3:24,84
there.

69
0:3:24,84 --> 0:3:28,64
And so number 1 priority is just
how do we not break the existing

70
0:3:28,64 --> 0:3:29,14
stuff?

71
0:3:29,54 --> 0:3:32,98
Versus how do we actually fix the
things that are not working.

72
0:3:33,34 --> 0:3:36,84
And so pgrust, because it's a
new project, we can take a little

73
0:3:36,84 --> 0:3:40,64
bit of a different approach of
let's try to actually fix the

74
0:3:40,64 --> 0:3:42,08
things that aren't working.

75
0:3:42,44 --> 0:3:45,54
And maybe some of the existing
stuff, it's not going to be quite

76
0:3:45,54 --> 0:3:48,16
as reliable, at least upfront,
as Postgres is.

77
0:3:48,16 --> 0:3:52,86
But at least we'll be able to fix
some of the long-standing architectural

78
0:3:52,96 --> 0:3:54,82
issues that Postgres has.

79
0:3:56,48 --> 0:3:57,6
Nik: Yeah, it makes sense.

80
0:3:57,84 --> 0:3:58,94
How did you do it?

81
0:3:58,94 --> 0:3:59,76
It's a lot of...

82
0:3:59,76 --> 0:4:1,5
Obviously, it's with Claude, right?

83
0:4:1,72 --> 0:4:4,74
Michael M.: Yeah, it took a couple
of attempts to actually figure

84
0:4:4,74 --> 0:4:6,52
out what is the best way to do
this.

85
0:4:6,82 --> 0:4:12,72
And what ended up working really
well was with Opus, we were

86
0:4:12,72 --> 0:4:19,06
able to take each file of Postgres
and basically transpile it

87
0:4:19,06 --> 0:4:19,74
to Rust.

88
0:4:19,82 --> 0:4:22,12
And so if you look at our code
and look at the Postgres code

89
0:4:22,12 --> 0:4:24,06
side by side, it actually looks
very similar.

90
0:4:24,16 --> 0:4:27,24
Where like the functions are the
same, some of the details may

91
0:4:27,24 --> 0:4:30,24
be a little bit different, but
overall it's actually really close

92
0:4:30,24 --> 0:4:32,5
to just a straight rewrite of Postgres.

93
0:4:32,8 --> 0:4:35,42
And so we did this across all the
files.

94
0:4:35,9 --> 0:4:40,28
There were some things that had
to be changed to actually work

95
0:4:40,28 --> 0:4:40,92
in Rust.

96
0:4:40,92 --> 0:4:44,28
1 of the big ones was how memory
management is done, because

97
0:4:44,28 --> 0:4:47,38
Rust is very particular about how
you allocate memory.

98
0:4:47,86 --> 0:4:51,04
But this approach of going through
all the files, rewriting them,

99
0:4:51,04 --> 0:4:54,44
we had a bunch of Claude agents
like going at this in parallel.

100
0:4:54,72 --> 0:4:57,24
There were like some conflicts
between different files, but we

101
0:4:57,24 --> 0:5:0,62
were able to resolve those and
then get pgrust to actually work.

102
0:5:0,62 --> 0:5:4,08
And then from there, we focused
on the Postgres test suite, getting

103
0:5:4,08 --> 0:5:5,14
all that to pass.

104
0:5:6,04 --> 0:5:10,78
And we have something that actually
looks a lot like Postgres,

105
0:5:11,04 --> 0:5:12,08
but is in Rust.

106
0:5:12,8 --> 0:5:15,6
Michael C.: I have potentially
a minor question on the source

107
0:5:15,6 --> 0:5:16,1
code.

108
0:5:16,3 --> 0:5:18,7
Really interesting that you went
file by file and that worked

109
0:5:18,7 --> 0:5:19,2
well.

110
0:5:19,36 --> 0:5:22,74
1 of my favorite things in the
source code is all the is the

111
0:5:22,74 --> 0:5:27,34
comments and they describe in quite
a lot of detail how things

112
0:5:27,34 --> 0:5:32,62
work and I wonder if you've noticed
how it rewrote those.

113
0:5:32,78 --> 0:5:37,92
I can imagine the code being somewhat
reliable in terms of it

114
0:5:37,92 --> 0:5:41,8
being translated, but I can imagine
the comments actually being

115
0:5:41,8 --> 0:5:43,76
harder in some way to be trustworthy.

116
0:5:44,18 --> 0:5:45,74
What have you found on that front?

117
0:5:47,08 --> 0:5:49,9
Michael M.: Yeah, for the comments,
like Claude has been like,

118
0:5:49,9 --> 0:5:52,36
Claude is very verbose when it
comes to comments.

119
0:5:53,26 --> 0:5:56,48
And like, every file have a whole
like big comment about like,

120
0:5:56,48 --> 0:5:59,84
what's in the file that I've actually
had to give it feedback,

121
0:5:59,84 --> 0:6:2,7
like I had to give it feedback
to comment less.

122
0:6:2,72 --> 0:6:5,84
And I think there's a linter that says
files should not have more

123
0:6:5,84 --> 0:6:7,86
than 5% of the line shouldn't be
comments.

124
0:6:9,96 --> 0:6:13,7
I haven't paid quite close attention
to the actual content of

125
0:6:13,7 --> 0:6:18,28
the comments, but I do know Claude
is making a lot of comments

126
0:6:18,28 --> 0:6:19,2
in the new code.

127
0:6:19,64 --> 0:6:21,3
Nik: It does, yeah.

128
0:6:21,58 --> 0:6:24,78
So I remember, now it's version
0.2, right?

129
0:6:24,96 --> 0:6:25,46
Michael M.: Yep.

130
0:6:25,58 --> 0:6:29,72
Nik: Yeah, and I remember version
0.1, I think there was a claim

131
0:6:29,72 --> 0:6:36,0
that all tests pass, but when I
wrote some silly stuff, like

132
0:6:36,14 --> 0:6:39,86
obviously not SQL syntax, it accepted
it without any errors.

133
0:6:40,02 --> 0:6:40,94
Now it's not so.

134
0:6:40,94 --> 0:6:46,94
In 0.2 it works fine, so it looks
like Postgres syntax is implemented

135
0:6:47,12 --> 0:6:47,88
very well.

136
0:6:48,06 --> 0:6:52,86
Was it about some lack of tests
in Postgres tests for checking

137
0:6:52,86 --> 0:6:53,46
negative stuff?

138
0:6:53,46 --> 0:6:54,3
I don't know.

139
0:6:55,08 --> 0:6:58,42
Or it just wasn't complete implementation
before?

140
0:6:58,44 --> 0:7:1,56
What changed between 0.1 and 0.2?

141
0:7:2,58 --> 0:7:7,86
Michael M.: Yeah, so pgrust 0.1 was
largely rewritten by Opus.

142
0:7:8,44 --> 0:7:13,34
And Opus I found to be a fine model,
but it's a bit difficult

143
0:7:13,34 --> 0:7:13,94
to work with.

144
0:7:13,94 --> 0:7:16,86
They'll mislead you a bit about
how much work is actually completed,

145
0:7:16,86 --> 0:7:19,28
we would ask it to rewrite this
file and it would rewrite 0.5

146
0:7:19,28 --> 0:7:20,8
of the file and not all of it.

147
0:7:20,8 --> 0:7:24,96
And so we were working with this
underlying unreliable model

148
0:7:24,96 --> 0:7:28,84
and all models are unreliable to
various degrees, but with enough

149
0:7:28,84 --> 0:7:33,66
quality controls and layering on
top of it, we're able to get

150
0:7:34,3 --> 0:7:38,52
all of Postgres ported, and we're
able to get the Postgres regression

151
0:7:38,52 --> 0:7:41,6
suite to pass, but that's actually
a really low bar.

152
0:7:42,1 --> 0:7:46,44
The Postgres regression suite on
Postgres only has about 0.667

153
0:7:46,44 --> 0:7:47,22
code coverage.

154
0:7:47,68 --> 0:7:52,02
And so you can pass the regression
suite, but there's still 0.333

155
0:7:52,02 --> 0:7:54,02
of Postgres that isn't even being
tested.

156
0:7:54,52 --> 0:7:58,46
And on top of that, the regression
suites are largely more like

157
0:7:58,46 --> 0:7:59,62
functionality tests.

158
0:7:59,68 --> 0:8:2,72
That every feature in Postgres,
there is a regression suite for

159
0:8:2,72 --> 0:8:5,86
it that's like, there's 1 for hash
joins, there's 1 for lateral

160
0:8:5,86 --> 0:8:10,88
joins, and there's 50,000 of these
that basically just go through

161
0:8:10,88 --> 0:8:13,48
every single Postgres feature and
make sure it works to some

162
0:8:13,48 --> 0:8:13,98
degree.

163
0:8:14,14 --> 0:8:18,28
But they actually just do a couple
tests for each feature and

164
0:8:18,28 --> 0:8:19,34
then move on.

165
0:8:19,34 --> 0:8:22,28
And so the tests are more about
just making sure this feature

166
0:8:22,28 --> 0:8:27,34
is there and it exists and works
well, but doesn't actually really

167
0:8:27,34 --> 0:8:31,5
like you have a perfect implementation
of this feature.

168
0:8:32,46 --> 0:8:36,82
Nik: So you, there is an idea to
improve regression test suite,

169
0:8:36,82 --> 0:8:37,32
right?

170
0:8:37,82 --> 0:8:41,82
And this is already, this should
go to upstream, maybe, no?

171
0:8:42,72 --> 0:8:43,22
Michael M.: Yeah.

172
0:8:43,44 --> 0:8:47,3
So what we've been doing, the difference
between 0.1 and 0.2

173
0:8:47,3 --> 0:8:50,28
is 0.2 was run by Fable, which
we have a more reliable model

174
0:8:50,28 --> 0:8:54,6
now that's, is more consistent
about, okay, you ask it to port

175
0:8:54,6 --> 0:8:56,78
a file, it'll actually rewrite
the file completely.

176
0:8:56,92 --> 0:9:0,06
And so we do end up with something
that is just baseline, more

177
0:9:0,06 --> 0:9:1,02
similar to Postgres.

178
0:9:1,5 --> 0:9:5,5
And then before we released 0.2,
we started doing a little, a

179
0:9:5,5 --> 0:9:8,1
tiny bit of differential testing
and fuzz testing, and I can

180
0:9:8,1 --> 0:9:9,5
talk a lot more about that.

181
0:9:9,52 --> 0:9:14,18
But for 0.3, we're actually like
going to have a lot more of

182
0:9:14,18 --> 0:9:14,68
that.

183
0:9:15,42 --> 0:9:19,58
What we've thought of doing is
we could add more regression tests,

184
0:9:20,08 --> 0:9:24,24
but that, again, every single regression
test is only testing

185
0:9:24,24 --> 0:9:26,6
a small part of the functionality.

186
0:9:27,1 --> 0:9:30,48
And what we really want is not
just, hey, I can choose 100 tests

187
0:9:30,48 --> 0:9:31,8
and they'll all pass.

188
0:9:31,96 --> 0:9:35,24
I want to know is this thing actually
identical to Postgres?

189
0:9:35,84 --> 0:9:39,06
And some of the things we've been
doing, and I was very surprised

190
0:9:39,06 --> 0:9:40,52
by this, that it's even worked.

191
0:9:40,52 --> 0:9:41,74
I would have never thought this.

192
0:9:41,84 --> 0:9:46,22
We found this library called Kani,
which can take Rust code and

193
0:9:46,22 --> 0:9:50,2
C code and do formal verification
over the code.

194
0:9:50,66 --> 0:9:51,76
Nik: Without any execution?

195
0:9:52,2 --> 0:9:53,86
Michael M.: So it does symbolic
execution.

196
0:9:54,4 --> 0:9:57,74
So it'll run your code, but it'll
run it in a special way that

197
0:9:57,74 --> 0:10:0,3
it is more just picking out that,
hey, this is an if statement

198
0:10:0,3 --> 0:10:2,98
that is checking this condition,
as opposed to actually having

199
0:10:2,98 --> 0:10:4,34
variables go through it.

200
0:10:4,34 --> 0:10:6,98
And then I'll be able to convert
that into a format that you

201
0:10:6,98 --> 0:10:10,08
can then do formal verification
over that has the Rust, I'll

202
0:10:10,08 --> 0:10:12,92
have an expression that represents
all the Rust, and I'll have

203
0:10:12,92 --> 0:10:14,34
the same for the C code.

204
0:10:14,44 --> 0:10:17,54
And then we can do across all inputs,
are these 2 things gonna

205
0:10:17,54 --> 0:10:19,52
produce the exact same output?

206
0:10:19,9 --> 0:10:25,3
And so Postgres has 3000 different
user facing functions, everything

207
0:10:25,32 --> 0:10:28,7
from substring to regular expression
matching to there's like

208
0:10:28,7 --> 0:10:32,9
the, like you can calculate the
gamma function, there's a function

209
0:10:32,9 --> 0:10:33,58
for that.

210
0:10:33,68 --> 0:10:36,94
And so for about 1000 of these
functions, we were actually

211
0:10:36,94 --> 0:10:40,72
able to do actual formal verification
that the Rust code and

212
0:10:40,72 --> 0:10:41,86
C code are the same.

213
0:10:42,28 --> 0:10:45,98
But then beyond that, what we've
started, like formal verification

214
0:10:46,06 --> 0:10:47,12
doesn't work in all cases.

215
0:10:47,12 --> 0:10:48,3
It's actually narrow.

216
0:10:48,52 --> 0:10:51,94
And so what we've been doing for
the rest of the code base has

217
0:10:51,94 --> 0:10:55,2
been, we'll take the Rust code
and the C code and put them side

218
0:10:55,2 --> 0:10:59,74
by side, and then have a fuzzer,
like, code coverage guided fuzzer,

219
0:11:0,18 --> 0:11:4,04
generate millions of inputs across
the 2 and make sure that they

220
0:11:4,2 --> 0:11:5,74
behave exactly the same.

221
0:11:6,4 --> 0:11:10,52
And in the process of doing this,
not only did we like find over

222
0:11:10,52 --> 0:11:14,44
100 bugs in like pgrust, but we
actually found at this point

223
0:11:14,44 --> 0:11:17,34
a little bit over 20 bugs in Postgres
itself.

224
0:11:17,52 --> 0:11:17,78
Nik: I see.

225
0:11:17,78 --> 0:11:18,58
That's interesting.

226
0:11:18,74 --> 0:11:19,98
Have you reported them?

227
0:11:20,38 --> 0:11:20,74
Michael M.: Yeah.

228
0:11:20,74 --> 0:11:24,52
If you look, I think like end of
July, like 1st week of August

229
0:11:24,52 --> 0:11:29,28
on the like pgsql or like pgsql
bugs mailing list, like you'll

230
0:11:29,28 --> 0:11:31,98
see a lot of the form submissions
are actually from me.

231
0:11:32,24 --> 0:11:32,8
Nik: That's cool.

232
0:11:32,8 --> 0:11:33,44
I missed it.

233
0:11:33,44 --> 0:11:34,18
That's cool.

234
0:11:34,44 --> 0:11:35,52
Congrats, actually.

235
0:11:35,54 --> 0:11:36,1
That's good.

236
0:11:36,1 --> 0:11:36,6
Yeah.

237
0:11:37,12 --> 0:11:38,2
Michael M.: None of the bugs.

238
0:11:38,52 --> 0:11:43,2
I found 3 serious bugs, but all
of them had been found already

239
0:11:43,2 --> 0:11:43,82
by other people.

240
0:11:43,82 --> 0:11:48,66
I'd been testing on 18.3, 18.4,
and in 19 people had found them.

241
0:11:48,84 --> 0:11:53,16
There were a lot of not serious
bugs that people had not found.

242
0:11:53,76 --> 0:11:57,68
1 of my favorites is Postgres has
a quadtree implementation,

243
0:11:59,34 --> 0:12:2,32
and it will use floating points
to represent the positions in

244
0:12:2,32 --> 0:12:2,94
the quadtree.

245
0:12:3,48 --> 0:12:6,94
And there was a bug where, because
of floating point rounding

246
0:12:7,06 --> 0:12:10,38
and the arithmetic that Postgres
was doing, it was possible for

247
0:12:10,38 --> 0:12:13,94
a point to not be in any of the
4 quadrants.

248
0:12:14,44 --> 0:12:17,68
That the comparison would say,
this isn't to the left, this isn't

249
0:12:17,68 --> 0:12:19,5
to the right, and this isn't in
the center.

250
0:12:20,22 --> 0:12:22,22
And we're definitely doing like
very aggressive.

251
0:12:23,56 --> 0:12:25,86
Nik: This is in the SP-GiST, right?

252
0:12:26,18 --> 0:12:29,08
Michael M.: I think it was in,
I don't know exactly.

253
0:12:30,06 --> 0:12:37,12
I want to say it was in GiST, yeah,
but basically the amount

254
0:12:37,12 --> 0:12:39,72
of testing we're doing is so thorough
that we're actually finding

255
0:12:39,72 --> 0:12:43,66
these extremely hard to find bugs
and we're finding quite a few

256
0:12:43,66 --> 0:12:45,0
bugs in Postgres itself.

257
0:12:45,24 --> 0:12:48,18
Nik: This is very important because
it means that with AI you

258
0:12:48,18 --> 0:12:52,42
can, you basically have some mechanism
which is finding bugs

259
0:12:52,42 --> 0:12:52,8
Michael M.: all the

260
0:12:52,8 --> 0:12:53,58
Nik: time, right?

261
0:12:53,8 --> 0:12:57,54
And it can be maybe like, what
could it mean?

262
0:12:57,74 --> 0:13:2,6
Probably there's Postgres Buildfarm
for testing, which is used for

263
0:13:2,6 --> 0:13:3,3
all releases.

264
0:13:4,3 --> 0:13:8,68
With this work, it maybe should
be extended to involve more AI

265
0:13:8,68 --> 0:13:11,14
and find more bugs at scale, right?

266
0:13:11,2 --> 0:13:13,08
Michael M.: Yeah, there's a couple
things I think are really

267
0:13:13,08 --> 0:13:13,58
interesting.

268
0:13:14,06 --> 0:13:18,02
1 is that with Mythos, this whole
thing about Mythos can pose

269
0:13:18,2 --> 0:13:19,5
a security risk.

270
0:13:19,82 --> 0:13:22,36
The interesting thing about the
models is like, the security,

271
0:13:22,36 --> 0:13:26,0
the way that they pose a security
risk isn't that like, the model

272
0:13:26,0 --> 0:13:29,34
looks at your code and can find
a bug in the code.

273
0:13:29,44 --> 0:13:33,34
It's actually that the model can
build tools that can then find

274
0:13:33,34 --> 0:13:34,34
bugs in your code.

275
0:13:34,4 --> 0:13:36,98
Or like, I've been using Fable
and a lot of times it'll fall

276
0:13:36,98 --> 0:13:41,76
back to Opus 5 because it's, oh,
you found a memory bug, that

277
0:13:42,1 --> 0:13:43,5
seems like a security issue.

278
0:13:43,7 --> 0:13:46,8
Nik: I think, just From my experience,
I also found a couple

279
0:13:46,8 --> 0:13:47,94
of bugs with AI.

280
0:13:49,12 --> 0:13:50,78
A few were security-related.

281
0:13:51,34 --> 0:13:55,28
1 was officially registered and
was already patched.

282
0:13:56,04 --> 0:13:58,4
That 1 was just looking at the
code.

283
0:13:59,28 --> 0:14:1,8
I didn't use Mythos, even not Fable.

284
0:14:1,8 --> 0:14:5,78
It was like Opus 4.5 back then.

285
0:14:6,04 --> 0:14:8,76
And I just was looking at the code,
it was very old code, and

286
0:14:8,76 --> 0:14:12,34
just looking at it, it found SQL
injection, which is there already

287
0:14:12,34 --> 0:14:13,3
18 years.

288
0:14:13,9 --> 0:14:19,5
So in some contrib module, which
is not used directly, so nobody

289
0:14:19,5 --> 0:14:22,78
cared, but it was actually serious
because it's used as an example.

290
0:14:23,56 --> 0:14:28,02
So anyway, it was interesting that
just looking at code, it also

291
0:14:28,02 --> 0:14:28,66
can find bugs.

292
0:14:28,66 --> 0:14:31,7
But you're right, building tools
is even more powerful because

293
0:14:32,18 --> 0:14:36,26
just looking at the code, you can
miss complex relationship between

294
0:14:36,26 --> 0:14:37,78
various code pieces, right?

295
0:14:38,0 --> 0:14:38,4
Michael M.: Yeah.

296
0:14:38,4 --> 0:14:41,26
So with the models, like what you
can do is you can basically

297
0:14:41,26 --> 0:14:44,48
point them to be like, hey, I have
100 functions in this code

298
0:14:44,48 --> 0:14:48,0
base, write really thorough tests
for all of these different

299
0:14:48,28 --> 0:14:49,4
pieces of code.

300
0:14:49,7 --> 0:14:52,7
And so the way I like to think
about it with these models is

301
0:14:52,7 --> 0:14:56,88
you can, if you can get them to
do 1 task well, you now have

302
0:14:56,88 --> 0:15:0,56
a repeatable way to do that 1 task
100 times.

303
0:15:0,72 --> 0:15:3,24
And so if you can get a model to
be able to test 1 function,

304
0:15:3,24 --> 0:15:5,84
there isn't a reason you can't
get it to test 100 or 1000 or

305
0:15:5,84 --> 0:15:6,72
10000 functions.

306
0:15:7,68 --> 0:15:8,18
Nik: Yeah.

307
0:15:8,3 --> 0:15:10,9
How much of AI capacity have you
already used?

308
0:15:10,9 --> 0:15:14,28
Is it like just 1 $200 account
or no?

309
0:15:14,6 --> 0:15:16,66
I'm just curious, very curious.

310
0:15:17,5 --> 0:15:19,5
Asking for a friend, so to speak.

311
0:15:20,32 --> 0:15:23,94
Michael M.: The 1st version of
pgrust, the 0.1, which was done

312
0:15:23,94 --> 0:15:29,44
with Opus, which that spanned 5
different attempts, that cost

313
0:15:29,44 --> 0:15:30,52
about 100 grand.

314
0:15:31,04 --> 0:15:36,96
And then so far, with the new version,
like 0.2 and will become

315
0:15:36,96 --> 0:15:40,24
0.3 in total, we've put in probably
like another like 300 or

316
0:15:40,24 --> 0:15:41,14
400 grand.

317
0:15:41,76 --> 0:15:42,6
Where the...

318
0:15:44,44 --> 0:15:47,56
We're burning tokens so fast that
you can burn through a 5 hour

319
0:15:47,56 --> 0:15:51,18
quota on a $200 a month subscription
on the, in the order of

320
0:15:51,18 --> 0:15:52,08
like 15 minutes.

321
0:15:54,38 --> 0:15:57,26
It's because for a while, like
what we'll do is like, we'll have

322
0:15:57,26 --> 0:15:59,96
a 40 different Fable instances
like running in parallel, each

323
0:15:59,96 --> 0:16:1,82
like working on individual parts
of Postgres.

324
0:16:2,64 --> 0:16:6,02
And that just burns through credits
super fast.

325
0:16:7,74 --> 0:16:8,82
Nik: Yeah, yeah, that's impressive.

326
0:16:9,44 --> 0:16:11,28
So, yeah.

327
0:16:11,28 --> 0:16:13,1
And it takes also a lot of time,
right?

328
0:16:13,1 --> 0:16:16,64
Because like, it's just they are
like, you just need to wait

329
0:16:16,64 --> 0:16:20,68
sometimes a lot right what was
your the longest run or loop

330
0:16:20,68 --> 0:16:24,12
fully autonomous like days weeks
or something

331
0:16:25,16 --> 0:16:30,96
Michael M.: I usually well I've
had stuff like run overnight

332
0:16:31,16 --> 0:16:36,14
that worked but I think I try to
keep things on the order of,

333
0:16:36,14 --> 0:16:39,22
I try to break things down to like
tasks on the order of like,

334
0:16:39,6 --> 0:16:42,56
ideally like an hour, cause any
more than that, I find that like

335
0:16:42,56 --> 0:16:45,2
the chance of the model going off
and doing its own thing and

336
0:16:45,2 --> 0:16:45,8
like getting.

337
0:16:45,8 --> 0:16:46,3
Nik: Nonsensical.

338
0:16:46,5 --> 0:16:47,0
Yeah.

339
0:16:47,2 --> 0:16:47,7
Michael M.: Yeah.

340
0:16:47,9 --> 0:16:50,7
I found that like models, they
work really well.

341
0:16:50,98 --> 0:16:52,2
They enter 2 states.

342
0:16:52,2 --> 0:16:54,72
Like 1 is they like are just like
working really well and making

343
0:16:54,72 --> 0:16:55,54
forward progress.

344
0:16:56,0 --> 0:16:58,68
And then other times they'll just
start doing their own thing.

345
0:16:58,68 --> 0:17:1,16
They'll just start making like
a bunch of random edits to like

346
0:17:1,16 --> 0:17:3,24
code everywhere and not really
go anywhere.

347
0:17:3,24 --> 0:17:6,3
And like the more you can keep
models on like the productive

348
0:17:6,3 --> 0:17:9,44
side and not on the going in circle
side is better.

349
0:17:9,44 --> 0:17:11,96
Michael C.: I'm curious on the,
because that's a lot of money

350
0:17:11,96 --> 0:17:13,52
to spend on v0.2.

351
0:17:14,44 --> 0:17:18,04
Are you being public with how you're
funding it or like if you

352
0:17:18,04 --> 0:17:18,9
do have investors.

353
0:17:19,86 --> 0:17:20,28
Michael M.: Yeah.

354
0:17:20,28 --> 0:17:25,6
So far it's actually been entirely
funded by me where I had a

355
0:17:25,6 --> 0:17:28,86
previous startup, Freshpaint,
which did like pretty well.

356
0:17:28,86 --> 0:17:31,88
And I had the chance to sell some
of my equity and so far it's

357
0:17:31,88 --> 0:17:33,06
mostly been funded by me.

358
0:17:33,06 --> 0:17:35,52
We're hitting the point where like
it's starting to become like

359
0:17:35,74 --> 0:17:38,76
a bit unreasonable and we're starting
to look at other ways to

360
0:17:38,76 --> 0:17:42,16
like fund it and can we like
be more token efficient and things

361
0:17:42,16 --> 0:17:42,84
like that.

362
0:17:43,32 --> 0:17:44,66
Michael C.: That makes a ton of
sense.

363
0:17:44,8 --> 0:17:47,36
So yeah going back to the technical
stuff I was going to

364
0:17:47,36 --> 0:17:51,14
the test suite is really interesting
to me, and almost as a what

365
0:17:51,14 --> 0:17:54,0
can we learn from a Postgres perspective,
as well as how could

366
0:17:54,0 --> 0:17:55,3
pgrust become trustworthy.

367
0:17:55,58 --> 0:18:1,24
So, on testing the functions, I
can imagine a huge amount of

368
0:18:1,24 --> 0:18:5,92
the kind of user surface area is
going to be covered nicely,

369
0:18:5,92 --> 0:18:9,74
like especially for single user,
single query correctness.

370
0:18:10,64 --> 0:18:15,42
But I can't help but feel nervous
about like concurrency stuff

371
0:18:15,42 --> 0:18:19,2
and even Postgres and some other
databases that have had some

372
0:18:19,2 --> 0:18:22,36
interesting bugs come up when Jepsen
have got involved and run

373
0:18:22,36 --> 0:18:26,92
some like interesting tests along
the like isolation level side

374
0:18:26,92 --> 0:18:27,54
of things.

375
0:18:28,28 --> 0:18:31,5
Do you, does your testing approach
cover that yet?

376
0:18:31,56 --> 0:18:34,34
Or do you, are you going to have
needed like a slightly different

377
0:18:34,34 --> 0:18:35,54
approach for some of those things?

378
0:18:35,54 --> 0:18:36,98
How are you thinking about that?

379
0:18:37,58 --> 0:18:40,6
Michael M.: Yeah, we're basically
taking what is the modern and

380
0:18:40,6 --> 0:18:43,8
latest approach to testing and
trying to incorporate all of that

381
0:18:43,92 --> 0:18:45,36
into how we do things.

382
0:18:45,48 --> 0:18:47,68
And the 1st thing we've been doing
is we've started with the

383
0:18:47,68 --> 0:18:49,44
easy part of, okay, here's all
the pure functions.

384
0:18:49,44 --> 0:18:50,58
Here's the ones that we can verify.

385
0:18:50,58 --> 0:18:52,16
Here's the ones that we can just
fuzz.

386
0:18:52,44 --> 0:18:56,32
And then most recently I've been
working on how can we actually

387
0:18:56,32 --> 0:19:0,52
get full coverage of the database
where there's a lot of very

388
0:19:0,52 --> 0:19:5,58
stateful things that will only
trigger bugs in very rare circumstances.

389
0:19:5,68 --> 0:19:7,9
And if you actually look at the
issues people are reporting for

390
0:19:7,9 --> 0:19:11,04
pgrust, it's usually the combination
of 2 features of, oh, if

391
0:19:11,04 --> 0:19:15,04
I use a non-default collation
with certain string functions,

392
0:19:15,06 --> 0:19:16,52
it'll cause a bug.

393
0:19:16,56 --> 0:19:18,96
And so what we're trying to do
now is actually fuzz until we

394
0:19:18,96 --> 0:19:20,7
get to 100% code coverage.

395
0:19:21,14 --> 0:19:24,4
And then for the concurrency stuff,
once we have a fuzzer that

396
0:19:24,4 --> 0:19:29,16
can actually properly explore the
entire code base, we already

397
0:19:29,16 --> 0:19:34,7
purchased this tool called Antithesis,
which does Jepsen-style

398
0:19:35,32 --> 0:19:36,1
fault testing.

399
0:19:36,2 --> 0:19:37,82
It's actually a super cool product.

400
0:19:38,14 --> 0:19:44,04
They run your software inside of
a VM and then are able to inject

401
0:19:44,24 --> 0:19:46,68
faults into your running code.

402
0:19:46,68 --> 0:19:49,3
And so that can be everything from
the network is unreliable

403
0:19:49,3 --> 0:19:51,38
and your database can't talk to
its replica.

404
0:19:51,58 --> 0:19:55,08
Or it can be like, hey, your server
crashed and now is your data

405
0:19:55,08 --> 0:19:56,46
corrupted on disk or not?

406
0:19:56,46 --> 0:19:59,44
And then they actually do stuff
of they can, they have full control

407
0:19:59,44 --> 0:20:0,14
of the threading.

408
0:20:0,14 --> 0:20:1,16
It's fully deterministic.

409
0:20:1,62 --> 0:20:5,42
And so they can create all these
weird out of order sequences

410
0:20:5,66 --> 0:20:9,4
for your threads to try to just
throw the craziest stuff at your

411
0:20:9,4 --> 0:20:9,9
software.

412
0:20:10,2 --> 0:20:12,68
And right now we're working on
the fuzzer to actually explore

413
0:20:12,9 --> 0:20:13,58
the search space.

414
0:20:13,58 --> 0:20:15,98
And then once we have that, we're
going to give that to Antithesis

415
0:20:16,0 --> 0:20:18,9
to then try to come up with all
these crazy ways of running the

416
0:20:18,9 --> 0:20:21,6
code together to try to break
pgrust.

417
0:20:21,6 --> 0:20:24,14
And so I think that is, we've actually
been talking with a lot

418
0:20:24,14 --> 0:20:26,64
of database people through this
and basically they're all using

419
0:20:26,64 --> 0:20:29,48
Antithesis for their testing at
this point.

420
0:20:29,96 --> 0:20:30,3
Nik: Yeah.

421
0:20:30,3 --> 0:20:34,0
Speaking of testing, I remember
I noticed on Hacker News comments

422
0:20:34,0 --> 0:20:39,18
that on the thread that people
asked questions about was fsync

423
0:20:39,4 --> 0:20:43,44
on and did you do writes in those
tests and so on or is just

424
0:20:43,44 --> 0:20:47,42
only from memory, data fits in
buffer pool or something and

425
0:20:47,42 --> 0:20:48,9
nothing touches the disk.

426
0:20:49,2 --> 0:20:49,4
Michael M.: Yeah.

427
0:20:49,4 --> 0:20:52,0
So we have, if you look at the
code base, we actually have a

428
0:20:52,0 --> 0:20:55,44
crash simulator in pgrust that
mocks out the file system and

429
0:20:55,44 --> 0:20:57,8
then we'll try to just crash the
database at different points.

430
0:20:57,8 --> 0:20:59,44
And we've done a good amount of
testing.

431
0:20:59,6 --> 0:21:3,06
The thing is that even Postgres
does get these things wrong sometimes.

432
0:21:3,08 --> 0:21:6,76
That there was like a fsyncgate
almost like a decade ago now,

433
0:21:6,76 --> 0:21:10,32
I think, of like, yeah, fsync can
actually fail and Postgres

434
0:21:10,32 --> 0:21:11,78
actually didn't handle that case.

435
0:21:11,82 --> 0:21:15,22
Which be fair to Postgres, that's
a case that never ever happens.

436
0:21:15,58 --> 0:21:20,2
Nik: Honestly, I also think to
me going to Jepsen tests and multi-node

437
0:21:20,28 --> 0:21:24,66
basically network testing and so
on is maybe not the biggest

438
0:21:24,66 --> 0:21:26,08
priority for this maybe.

439
0:21:26,26 --> 0:21:31,16
In my opinion, I recently learned
that Postgres own tests are not

440
0:21:31,16 --> 0:21:34,26
checking properly disk outages,
for example.

441
0:21:35,24 --> 0:21:39,72
So this is like, what happens if
disk just has some problems?

442
0:21:40,46 --> 0:21:44,94
And there are some tools to improve
that in Postgres on codebase.

443
0:21:45,9 --> 0:21:46,94
So I'm very curious.

444
0:21:47,08 --> 0:21:49,64
This is purely a vibe-coded thing.

445
0:21:49,9 --> 0:21:54,86
You obviously created a good harness
to mock and so on, but still

446
0:21:54,86 --> 0:21:56,42
people reported some segfaults.

447
0:21:56,54 --> 0:21:58,16
It's obvious it's early stage.

448
0:21:58,42 --> 0:22:1,36
Will it reach some point when it's
very reliable so it can be

449
0:22:1,36 --> 0:22:2,2
put to production.

450
0:22:2,2 --> 0:22:5,68
So obviously you say yes, but what's
your like, why do you think

451
0:22:5,68 --> 0:22:5,9
so?

452
0:22:5,9 --> 0:22:7,0
Why do you think yes?

453
0:22:7,1 --> 0:22:9,56
Because I'm also very AI positive,
so to speak.

454
0:22:9,56 --> 0:22:13,34
I think actually in a few months
we will be able to revisit with

455
0:22:13,34 --> 0:22:15,0701
new models and polish it and so
on.

456
0:22:15,0701 --> 0:22:20,16
And It gives me so joy to work
with AI together, like human plus

457
0:22:20,16 --> 0:22:22,12
AI, we can achieve a lot.

458
0:22:22,36 --> 0:22:25,22
But still, there are so many doubts
around.

459
0:22:26,0 --> 0:22:30,56
And people think, for example,
it's impossible to achieve a point

460
0:22:30,66 --> 0:22:32,38
where it will be reliable.

461
0:22:32,96 --> 0:22:35,28
You started with reliability, right?

462
0:22:35,74 --> 0:22:40,16
Then you said hackers are too conservative,
but there's also

463
0:22:40,16 --> 0:22:43,7
they are too conservative to protect
that reliability thing,

464
0:22:43,7 --> 0:22:44,2
right?

465
0:22:44,28 --> 0:22:48,08
Now with vibe code, the thing fully
rewritten, won't it take

466
0:22:48,08 --> 0:22:51,44
10 years to achieve a very reliable
state?

467
0:22:51,82 --> 0:22:54,62
What makes you think it's possible
to achieve it much faster

468
0:22:54,62 --> 0:22:55,5
than 10 years?

469
0:22:56,74 --> 0:22:58,86
Michael M.: So I think for the
issues that people are reporting,

470
0:22:58,86 --> 0:23:2,8
I think 1 thing to clarify is that
so far for version 0.2, we

471
0:23:2,8 --> 0:23:6,22
had, or actually I think for version
0.2, that was actually before

472
0:23:6,22 --> 0:23:10,76
we had done even like most
of the fuzz testing I was talking

473
0:23:10,76 --> 0:23:12,84
about, where like we actually sorted
out those issues.

474
0:23:13,04 --> 0:23:16,06
The correctness work has like largely
followed version 0.2 and

475
0:23:16,06 --> 0:23:18,8
so far we've only covered like
15 or 20 percent of the codebase

476
0:23:18,8 --> 0:23:20,42
and like there's still a lot more
to cover.

477
0:23:20,46 --> 0:23:23,04
And so I'm like I'm totally not
surprised that people found issues

478
0:23:23,04 --> 0:23:25,08
in the version we released just
because like we hadn't actually

479
0:23:25,08 --> 0:23:26,62
hardened it that much.

480
0:23:26,84 --> 0:23:29,34
And to answer your question of
how we can actually get this to

481
0:23:29,34 --> 0:23:32,92
a place where people can trust
this, Back to what I said earlier

482
0:23:32,92 --> 0:23:36,66
about, it's not like the models
themselves are like reviewing

483
0:23:36,66 --> 0:23:38,66
all the code and like making sure
it's correct.

484
0:23:38,74 --> 0:23:42,1
What we're doing is like we're
building tools to actually make

485
0:23:42,1 --> 0:23:43,4
it, make sure it's correct.

486
0:23:43,48 --> 0:23:46,84
And we're doing it such a degree
that like we're able to do millions

487
0:23:46,84 --> 0:23:50,24
of inputs for individual functions
to the point that Postgres,

488
0:23:50,66 --> 0:23:53,86
this 30-year-old battle-tested
codebase, we're finding bugs in

489
0:23:53,86 --> 0:23:56,52
it that no 1 had found before,
even though people have been using

490
0:23:56,52 --> 0:23:57,94
it all this time.

491
0:23:58,04 --> 0:24:1,28
Nik: You find bugs in the old code,
but Maybe you have bugs which

492
0:24:1,28 --> 0:24:3,46
you haven't explored in new code
yet.

493
0:24:4,06 --> 0:24:5,88
So this is a really tricky question.

494
0:24:6,38 --> 0:24:8,86
What's your plan to prove it in
production?

495
0:24:9,02 --> 0:24:13,0
Just wait for people who will try
it in production on replicas,

496
0:24:13,0 --> 0:24:16,56
for example, because I always think
This thing can work as a

497
0:24:16,56 --> 0:24:17,46
physical standby.

498
0:24:18,22 --> 0:24:22,2
Is this a way or use some mirroring,
for example, PgDog has or

499
0:24:22,2 --> 0:24:26,28
other, maybe some others have,
to mirror traffic and see that

500
0:24:26,28 --> 0:24:27,18
it works well?

501
0:24:27,18 --> 0:24:30,48
Or what's your plan to prove with
production that it works?

502
0:24:30,84 --> 0:24:31,12
Michael M.: Yeah.

503
0:24:31,12 --> 0:24:34,54
So for proving it like actually
in production, our plan is to

504
0:24:34,54 --> 0:24:37,2
do exactly what you said and replicate
off of people's existing

505
0:24:37,2 --> 0:24:40,16
databases and be like read-only
standby.

506
0:24:40,68 --> 0:24:43,58
That gives people a lot of the
benefits of like what we're building

507
0:24:43,58 --> 0:24:47,9
in terms of this like ultra fast
columnar workload such that

508
0:24:47,9 --> 0:24:50,66
like people who are trying to do
analytics in Postgres and are

509
0:24:50,66 --> 0:24:53,72
struggling to do it, they can stand
up pgrust and then move

510
0:24:53,72 --> 0:24:55,12
to analytics over to that.

511
0:24:55,12 --> 0:24:58,24
And the really nice thing about
analytics is they tend to be

512
0:24:58,38 --> 0:25:2,32
like semi-production workloads
where either like it's internal

513
0:25:2,32 --> 0:25:5,5
usage and okay your database goes
down, okay my internal dashboards

514
0:25:5,5 --> 0:25:8,6
are broken for a little bit, it's
not the end of the world.

515
0:25:8,94 --> 0:25:12,52
And then over time as like we start
to get like battle tested

516
0:25:12,52 --> 0:25:16,86
in these not tier 1 database use
cases, that's I think, will

517
0:25:16,86 --> 0:25:18,96
give people more confidence that,
hey, this is actually something

518
0:25:18,96 --> 0:25:20,64
that, like, works well.

519
0:25:21,02 --> 0:25:24,7
Nik: So what I like in this idea
is that if it's, if pgrust

520
0:25:24,86 --> 0:25:29,48
is powering a physical standby,
and then it crashed, for example,

521
0:25:30,14 --> 0:25:32,22
but primary won't be affected,
right?

522
0:25:32,22 --> 0:25:34,7
Because, okay, some slot is not
used.

523
0:25:35,66 --> 0:25:40,24
And if it's working and you run
very long running query on it,

524
0:25:40,24 --> 0:25:44,76
usually it's a bad idea on regular
standbys because of hot_standby_feedback

525
0:25:44,76 --> 0:25:45,86
dilemma.

526
0:25:46,38 --> 0:25:49,64
If it's on, it will affect vacuum
on the primary.

527
0:25:49,64 --> 0:25:52,8
If it's off, it makes your node
basically single user.

528
0:25:52,8 --> 0:25:53,9
It will start lagging.

529
0:25:54,64 --> 0:25:58,94
So both choices are not okay if
you want analytics.

530
0:25:59,54 --> 0:26:2,86
But It's only true if those queries
last hours.

531
0:26:3,94 --> 0:26:8,54
If it's already, as your blog post
suggests, 300 times faster,

532
0:26:9,4 --> 0:26:14,14
then xmin horizon is not blocked
by hot_standby_feedback for long.

533
0:26:14,14 --> 0:26:15,7
It's blocked for a short time.

534
0:26:16,1 --> 0:26:17,36
This is very interesting.

535
0:26:17,78 --> 0:26:19,52
This looks promising.

536
0:26:20,64 --> 0:26:24,56
We have very fast standby, which
behaves exactly like Postgres.

537
0:26:25,6 --> 0:26:28,16
Like, primary thinks it's like
Postgres, okay?

538
0:26:28,74 --> 0:26:34,24
And we have ability to run some
aggregates and so on very fast.

539
0:26:34,74 --> 0:26:35,58
Michael M.: That's interesting.

540
0:26:35,92 --> 0:26:39,08
Actually, the thing I find really
interesting about pgrust is

541
0:26:39,08 --> 0:26:43,18
there's actually a whole bunch
of unexplored design space, new

542
0:26:43,18 --> 0:26:45,52
ways that you can do things.

543
0:26:45,9 --> 0:26:49,56
Where you mentioned, for instance,
when you have a replica, due

544
0:26:49,56 --> 0:26:54,44
to the way the WAL replay works,
there's all these like complexities

545
0:26:54,52 --> 0:26:57,74
about how did the vacuums work
and how does like the replica

546
0:26:57,74 --> 0:26:58,66
like fall behind.

547
0:26:58,66 --> 0:27:0,4
But I think there are options out
there.

548
0:27:0,4 --> 0:27:2,64
I haven't explored them thoroughly,
but there's options out there.

549
0:27:2,64 --> 0:27:5,64
If we change the file format and
the way the WAL replay works,

550
0:27:5,74 --> 0:27:8,2
maybe there's ways to actually
get around these problems.

551
0:27:8,74 --> 0:27:11,16
I think this is like a lot of what's
really exciting about pgrust

552
0:27:11,16 --> 0:27:14,24
is like, there's a lot of
like unexplored design area.

553
0:27:14,64 --> 0:27:17,84
And because it's this like experimental
new thing, there's actually

554
0:27:17,84 --> 0:27:21,24
lots of opportunity to explore
the design space.

555
0:27:22,48 --> 0:27:24,18
Nik: Right, but that's for sure.

556
0:27:24,52 --> 0:27:27,38
So experimental, you can go very
far.

557
0:27:27,74 --> 0:27:32,06
But grounding is when you put it
to production.

558
0:27:33,66 --> 0:27:36,84
This is the most interesting practical
piece here for me.

559
0:27:37,12 --> 0:27:40,9
Sounds like it's relatively low
risk to try it.

560
0:27:41,0 --> 0:27:42,08
There's a way, right?

561
0:27:42,44 --> 0:27:43,22
That's interesting.

562
0:27:44,12 --> 0:27:45,92
So, Yeah, that's cool.

563
0:27:45,92 --> 0:27:46,7
What else?

564
0:27:46,92 --> 0:27:49,68
I saw your blog post, I think
today, right?

565
0:27:49,68 --> 0:27:51,9
About JIT just in time.

566
0:27:52,04 --> 0:27:53,68
What attracted your attention?

567
0:27:53,68 --> 0:27:57,38
It was just switched off by default
finally, recently.

568
0:27:58,26 --> 0:27:59,62
Was this like, yeah.

569
0:28:0,1 --> 0:28:0,6
Michael M.: Yeah.

570
0:28:1,12 --> 0:28:4,2
As part of the Columnar stuff,
we looked at a bunch of the recent

571
0:28:4,2 --> 0:28:7,86
database literature and incorporated
a ton of stuff into pgrust.

572
0:28:8,36 --> 0:28:12,56
A lot of it actually coming from
the research into Umbra, which

573
0:28:12,56 --> 0:28:16,08
is this very experimental database
being developed by this team

574
0:28:16,08 --> 0:28:16,64
in Munich.

575
0:28:16,64 --> 0:28:19,2
And they've been publishing database
research for the last decade

576
0:28:19,2 --> 0:28:21,26
or more, and a lot of it comes
from them.

577
0:28:21,74 --> 0:28:25,44
But for JIT compilation, and this
is a lot of what I talked about

578
0:28:25,44 --> 0:28:28,18
in the post that came out this
morning, if you look at all the

579
0:28:28,18 --> 0:28:32,52
databases that have JIT compilation
in it, they either use LLVM,

580
0:28:33,14 --> 0:28:37,36
or they'll write C and C++ code
and then compile that.

581
0:28:37,36 --> 0:28:40,28
And the big downside of these approaches
is that there's actually

582
0:28:40,28 --> 0:28:44,72
a huge amount of latency overhead
to compile this code that for

583
0:28:44,72 --> 0:28:49,26
C and C++ or for LLVM, I think it's
50 milliseconds compile time.

584
0:28:49,4 --> 0:28:52,8
And so you're very limited in the
cases that you can use it.

585
0:28:52,8 --> 0:28:56,26
And you have this risk of if you're
the planner estimates the

586
0:28:56,26 --> 0:28:59,4
query wrong, you now have this extra
50 milliseconds of latency.

587
0:29:0,06 --> 0:29:4,84
But there was this approach to
JIT compilation called copy and

588
0:29:4,84 --> 0:29:7,64
I think it's like copy and fix
or something along those lines

589
0:29:7,64 --> 0:29:11,88
that was released in someone published
a paper on it in 2021

590
0:29:11,92 --> 0:29:13,94
that makes copy and patch.

591
0:29:13,94 --> 0:29:14,84
Yeah, that's it.

592
0:29:14,86 --> 0:29:16,72
Then it's much easier to write.

593
0:29:16,72 --> 0:29:19,28
Then you just have templates of
assembly code and then you just

594
0:29:19,28 --> 0:29:20,8
fill in the bits as needed.

595
0:29:20,8 --> 0:29:24,38
You can combine these templates
together to get the code you

596
0:29:24,38 --> 0:29:24,88
want.

597
0:29:24,96 --> 0:29:28,62
And so the pgrust JIT compiler
actually doesn't use something

598
0:29:28,62 --> 0:29:29,26
like LLVM.

599
0:29:29,26 --> 0:29:32,06
It actually just directly generates
the assembly.

600
0:29:32,42 --> 0:29:37,26
And so because of this, the compile
times are like 5 microseconds

601
0:29:37,84 --> 0:29:42,18
as opposed to 50 milliseconds for
LLVM.

602
0:29:42,56 --> 0:29:47,36
And this gives us a lot more, Like
we can JIT compile a lot more

603
0:29:47,36 --> 0:29:51,34
of the query that we actually,
for our executor, we JIT compile

604
0:29:51,34 --> 0:29:55,26
large parts of the query, not just
expressions, which Postgres

605
0:29:55,26 --> 0:29:59,44
will only JIT compile expressions
and tuple deforming, but we

606
0:29:59,44 --> 0:30:1,78
actually do a lot more than that.

607
0:30:1,88 --> 0:30:5,28
And so we get measurable speed
ups in many different places.

608
0:30:6,68 --> 0:30:10,56
Michael C.: Do you see any reason
Postgres couldn't use a similar

609
0:30:10,56 --> 0:30:11,06
approach?

610
0:30:12,26 --> 0:30:16,56
Michael M.: So 1 of the big, 1
of the challenges I see with Postgres

611
0:30:16,56 --> 0:30:22,0
adopting this approach is I Specifically
am targeting only the Graviton

612
0:30:22,2 --> 0:30:27,56
instruction set Whereas pgrust
is being built in a very different

613
0:30:27,56 --> 0:30:30,34
world where today most people are
running their databases in

614
0:30:30,34 --> 0:30:34,04
the cloud And for instance, there's
no Windows support in pgrust

615
0:30:34,04 --> 0:30:34,96
yet.

616
0:30:35,34 --> 0:30:38,42
And so I'm able to, because I know,
I'm just assuming that people

617
0:30:38,42 --> 0:30:42,04
are going to run this primarily
on AWS and GCP, I'm able to target

618
0:30:42,04 --> 0:30:46,14
1 specific CPU design and 1 specific
instruction set and focus

619
0:30:46,26 --> 0:30:48,76
just on that and make that work
really well.

620
0:30:50,28 --> 0:30:52,4
And for Postgres, I think it's
still possible.

621
0:30:52,44 --> 0:30:54,92
You could have the JIT compiler
only work in certain instruction

622
0:30:54,92 --> 0:31:0,26
sets, or you can have different
backends for different instruction

623
0:31:0,28 --> 0:31:0,6
sets.

624
0:31:0,6 --> 0:31:4,2
But it's an order of magnitude
more work to get that to work

625
0:31:4,2 --> 0:31:5,96
well than what I'm doing.

626
0:31:6,06 --> 0:31:9,38
Nik: And if, even if it's not in
cloud, they will run it on Mac

627
0:31:9,38 --> 0:31:11,14
Minis they have now, right?

628
0:31:14,72 --> 0:31:16,16
Should work anyway there, Yeah.

629
0:31:16,16 --> 0:31:19,2
Michael M.: I've seen multiple
people with like rooms of like,

630
0:31:19,2 --> 0:31:19,92
they'll have

631
0:31:20,8 --> 0:31:21,54
Nik: a bunch of...

632
0:31:21,54 --> 0:31:22,04
Yeah.

633
0:31:22,2067 --> 0:31:23,04
In the corner.

634
0:31:23,48 --> 0:31:23,8
Could be a good Postgres cluster.

635
0:31:23,8 --> 0:31:23,82
Yeah.

636
0:31:23,82 --> 0:31:24,84
So that's interesting.

637
0:31:25,46 --> 0:31:28,44
And you don't use GPT like Codex,
no?

638
0:31:28,44 --> 0:31:30,58
Like distributing workloads within?

639
0:31:31,52 --> 0:31:34,84
Michael M.: So in my experience,
actually like the, a lot of

640
0:31:34,84 --> 0:31:39,48
the early work we were doing with
Codex, like Codex 5.4 and Codex

641
0:31:39,48 --> 0:31:44,44
5.5, the thing that made me switch
to Claude is they released

642
0:31:44,44 --> 0:31:47,96
this feature called dynamic workflows,
which this is what was

643
0:31:47,96 --> 0:31:49,86
used for the Bun Rust rewrite.

644
0:31:50,14 --> 0:31:56,1
And dynamic workflows are, you
can basically, it will write code

645
0:31:56,14 --> 0:31:58,62
that orchestrates a bunch of different
agents.

646
0:31:58,94 --> 0:32:1,48
And so you can be like, Hey, I
have these 10 files.

647
0:32:1,48 --> 0:32:3,06
Can you rewrite all of them?

648
0:32:3,18 --> 0:32:7,2
And it'll write some JavaScript,
they'll then spin up 10 agents,

649
0:32:7,2 --> 0:32:10,58
they'll then rewrite each file
and then review it and fix any

650
0:32:10,58 --> 0:32:12,34
issues and then merge it.

651
0:32:12,34 --> 0:32:16,02
And so that level of orchestration
made it way easier to use

652
0:32:16,02 --> 0:32:16,52
Claude.

653
0:32:16,98 --> 0:32:21,8
And then in my experience, like
Fable just works, it's like way

654
0:32:21,8 --> 0:32:23,74
more reliable and trustworthy.

655
0:32:24,38 --> 0:32:28,22
1 of the actually like biggest
challenges I had with the rewrite,

656
0:32:28,48 --> 0:32:31,34
or like 1 of the things that's
actually really hard, was the

657
0:32:31,34 --> 0:32:33,06
Postgres regular expression engine.

658
0:32:33,4 --> 0:32:35,42
I don't know what it is about it.

659
0:32:35,8 --> 0:32:38,04
I think it's 10,000 lines of code.

660
0:32:38,04 --> 0:32:40,82
It's a decent amount of code, but
not too big.

661
0:32:41,0 --> 0:32:45,1
And I was doing this as a benchmark,
and I threw Sol at it on

662
0:32:45,1 --> 0:32:48,02
medium-level thinking, not the
highest thinking, but on a decent...

663
0:32:48,14 --> 0:32:52,68
Sol's a pretty good model, And
Sol wasn't able to rewrite it

664
0:32:52,68 --> 0:32:53,22
in Rust.

665
0:32:53,22 --> 0:32:54,44
I was like very surprised by that.

666
0:32:54,44 --> 0:32:57,98
I was like fully expecting that
given how smart Sol is, like

667
0:32:57,98 --> 0:32:59,94
it should just be able to do that.

668
0:33:0,06 --> 0:33:2,64
Whereas Fable was like a whole
level above that, where it was

669
0:33:2,64 --> 0:33:5,58
able to rewrite in Rust and then
it was able to do like performance

670
0:33:5,58 --> 0:33:9,62
optimization on it to make it a
tiny bit faster than the Postgres

671
0:33:9,62 --> 0:33:10,12
version.

672
0:33:10,13 --> 0:33:15,84
And so like, I think like it probably
is possible to use like

673
0:33:15,84 --> 0:33:19,28
Sol, but like I've just found the
difference in autonomy between

674
0:33:19,28 --> 0:33:23,42
the 2 models is like not worth
the difference.

675
0:33:24,92 --> 0:33:27,76
Michael C.: I think we need to
talk about licensing.

676
0:33:28,44 --> 0:33:29,06
I noticed

677
0:33:29,06 --> 0:33:29,62
Nik: you picked

678
0:33:29,62 --> 0:33:30,74
Michael C.: the AGPL.

679
0:33:31,58 --> 0:33:35,28
That's always controversial, but
at least open source, right?

680
0:33:35,28 --> 0:33:36,38
What, why pick that 1?

681
0:33:36,38 --> 0:33:37,98
I think I know why, but why?

682
0:33:38,8 --> 0:33:41,64
Michael M.: Yeah, we were looking
at the different options we

683
0:33:41,64 --> 0:33:42,6
had for licensing.

684
0:33:42,98 --> 0:33:44,6
And there's like 4 different options.

685
0:33:44,68 --> 0:33:46,92
There's you have MIT or like Postgres.

686
0:33:47,26 --> 0:33:51,66
You have like AGPL or like GPL,
where it's like the copyleft.

687
0:33:51,94 --> 0:33:54,96
You have what they call like source
available, like you have

688
0:33:54,96 --> 0:33:58,22
SSPL and like BSL, and you have
like closed source.

689
0:33:58,68 --> 0:34:2,16
And We knew we wanted to do something
open source, and it's now

690
0:34:2,16 --> 0:34:6,6
limited to the permissive or the
AGPL.

691
0:34:8,04 --> 0:34:13,54
And the thing was, we wanted something
that a third party couldn't

692
0:34:13,54 --> 0:34:17,64
just take and sell themselves to
the detriment of the work we're

693
0:34:17,64 --> 0:34:18,14
doing.

694
0:34:19,62 --> 0:34:23,26
Where there's lots of closed source
forks of Postgres.

695
0:34:23,26 --> 0:34:28,12
You have AlloyDB, you have Aurora,
now Neon is closed source.

696
0:34:28,14 --> 0:34:31,92
And none of that money that those
platforms make actually goes

697
0:34:31,92 --> 0:34:35,46
back, or like very, very tiny amounts
of it goes back to the

698
0:34:35,46 --> 0:34:37,06
actual Postgres core projects.

699
0:34:37,22 --> 0:34:40,12
And that was just something that
like I wanted to avoid with

700
0:34:40,12 --> 0:34:40,84
pgrust.

701
0:34:40,84 --> 0:34:44,06
Nik: But AGPL
doesn't protect you from selling.

702
0:34:44,06 --> 0:34:46,9
Anyone can still sell if they don't
modify code.

703
0:34:47,24 --> 0:34:50,28
And even if they modify code, they
just need to publish it.

704
0:34:50,28 --> 0:34:51,0
That's it.

705
0:34:51,74 --> 0:34:52,24
Michael M.: Yeah.

706
0:34:53,1 --> 0:34:53,32
Yeah.

707
0:34:53,32 --> 0:34:55,64
So at least the history for AGPL
is it is like extremely, basically

708
0:34:55,64 --> 0:34:56,88
no 1 has really done that.

709
0:34:56,88 --> 0:35:0,06
But like just the cost of needing
to open source your modifications

710
0:35:0,06 --> 0:35:5,78
to it has deterred like Amazon
and Google from offering any AGPL

711
0:35:5,78 --> 0:35:8,04
software on their platforms, at
least so far.

712
0:35:8,14 --> 0:35:8,76
Nik: Makes sense.

713
0:35:8,76 --> 0:35:8,92
So

714
0:35:8,92 --> 0:35:10,62
Michael M.: that seemed like the
right trade-off.

715
0:35:11,54 --> 0:35:12,04
Nik: Good.

716
0:35:12,44 --> 0:35:15,04
I know we're like almost out of
time.

717
0:35:15,1 --> 0:35:16,5
It was an interesting discussion.

718
0:35:16,5 --> 0:35:17,58
Thank you for coming, Michael.

719
0:35:17,58 --> 0:35:18,82
I enjoyed it very well.

720
0:35:18,82 --> 0:35:19,84
Good luck with your project.

721
0:35:19,84 --> 0:35:21,86
We will be keeping an eye on it.

722
0:35:22,06 --> 0:35:22,94
Michael M.: Yeah, of course.

723
0:35:23,0 --> 0:35:23,86
Thanks so much.

724
0:35:24,48 --> 0:35:24,84
Nik: Thank you.