1
0:0:0,06 --> 0:0:2,72
Michael: Hello and welcome to Postgres.FM,
a weekly show about

2
0:0:2,72 --> 0:0:3,6599998
all things PostgreSQL.

3
0:0:3,82 --> 0:0:6,22
I am Michael, founder of pgMustard
and I'm joined as always

4
0:0:6,22 --> 0:0:7,68
by Nik, founder of PostgresAI.

5
0:0:7,68 --> 0:0:7,8399997
Hey

6
0:0:7,8399997 --> 0:0:8,34
Nik.

7
0:0:8,48 --> 0:0:9,179999
Nikolay: Hi Michael.

8
0:0:10,2 --> 0:0:14,28
Michael: And today we have with
us David Steele, who is a significant

9
0:0:14,34 --> 0:0:17,26
contributor to PostgreSQL and the
creator and maintainer of both

10
0:0:17,26 --> 0:0:20,880001
pgBackRest, which we'll be talking
about today, and pgaudit.

11
0:0:21,1 --> 0:0:22,16
Welcome, David.

12
0:0:23,08 --> 0:0:23,939999
David: Thank you very much.

13
0:0:23,939999 --> 0:0:25,34
Good to be here, Nik and Michael.

14
0:0:25,68 --> 0:0:27,380001
Michael: It's a pleasure to have
you.

15
0:0:27,54 --> 0:0:30,22
Right, to get started, I wondered
if you could give us a little

16
0:0:30,22 --> 0:0:32,92
bit of history, a little bit of
the origin story, perhaps of

17
0:0:32,92 --> 0:0:33,74
pgBackRest.

18
0:0:34,32 --> 0:0:34,82
David: Sure.

19
0:0:34,84 --> 0:0:35,739998
That's easy enough.

20
0:0:35,739998 --> 0:0:39,44
Actually, pgBackRest was born
at the Dublin conference in 2013

21
0:0:40,68 --> 0:0:44,96
of conversations that Stephen Frost
and Cynthia Shang and I and

22
0:0:44,96 --> 0:0:48,48
Magnus Hagander and some others
were having At that time, we

23
0:0:48,48 --> 0:0:51,56
were working on a fairly large
database for the time, around

24
0:0:51,56 --> 0:0:52,06
50T.

25
0:0:52,64 --> 0:0:56,66
That doesn't sound so big these
days, but in 2013, that was a

26
0:0:56,66 --> 0:0:58,6
pretty significant database.

27
0:0:59,18 --> 0:1:1,44
You know, basically we needed to
make backups, of course, as

28
0:1:1,44 --> 0:1:2,14
you do.

29
0:1:2,32 --> 0:1:5,64
And just the available tools were
not up to the task.

30
0:1:6,02 --> 0:1:7,62
It was just simply too large.

31
0:1:8,0 --> 0:1:9,84
And we really need to be able to
do incrementals.

32
0:1:9,84 --> 0:1:12,24
And there was nothing that would
do incrementals plus compression

33
0:1:12,24 --> 0:1:13,94
at the same time, et cetera, et
cetera.

34
0:1:13,94 --> 0:1:18,48
So As you do in open source, we
decided that we would build our

35
0:1:18,48 --> 0:1:19,84
own thing.

36
0:1:20,38 --> 0:1:23,54
I remember originally it was going
to be a pretty simple project.

37
0:1:23,94 --> 0:1:27,62
Write it in Perl, keep it simple,
just be 1 file, so it would

38
0:1:27,62 --> 0:1:30,42
be easy to copy around and distribute
and etc.

39
0:1:30,86 --> 0:1:32,54
Yeah, that didn't last very long.

40
0:1:32,66 --> 0:1:34,26
Obviously it grew pretty quickly.

41
0:1:35,22 --> 0:1:36,4
So I built it.

42
0:1:36,46 --> 0:1:40,1
I built the initial software that
was usable in about 40 hours

43
0:1:41,3 --> 0:1:43,9
and to basically solve our initial
problem and then just kept

44
0:1:43,9 --> 0:1:46,92
building on it as we had problems
and bugs and other things I

45
0:1:46,92 --> 0:1:48,48
would build on it and build on
it.

46
0:1:48,48 --> 0:1:51,02
Convinced the company I was working
for, Resonate, which is an

47
0:1:51,02 --> 0:1:53,3
ad tech company, to open source
it.

48
0:1:53,76 --> 0:1:56,06
And then when I left there, I kept
noodling around at it.

49
0:1:56,06 --> 0:1:58,38
I took some time off and just kept
working on it.

50
0:1:58,38 --> 0:2:0,1
Got the restore functionality working
well.

51
0:2:0,1 --> 0:2:3,5
And then that's when I got hired
into Crunchy.

52
0:2:4,28 --> 0:2:7,36
At 1st they weren't super interested
in it but Stephen was.

53
0:2:7,66 --> 0:2:10,76
So I got a little bit of time to
work on it and then eventually

54
0:2:10,76 --> 0:2:13,68
got more time to work on it and
then eventually we hired Cynthia

55
0:2:14,18 --> 0:2:17,68
to come and work on it with me
and that went on for a while and

56
0:2:17,68 --> 0:2:21,34
then finally we decided to migrate
it to C and we did the whole

57
0:2:21,34 --> 0:2:26,14
C migration which was 2 years so
that was painful because we

58
0:2:26,14 --> 0:2:28,84
didn't we basically didn't write
any new features for 2 years

59
0:2:28,84 --> 0:2:32,24
we fixed bugs only And we could
only write a new feature if it

60
0:2:32,24 --> 0:2:33,98
lived entirely in the C code.

61
0:2:34,54 --> 0:2:37,44
And over time, it was more and
more possible for that to happen,

62
0:2:37,44 --> 0:2:38,8
but it was a tough migration.

63
0:2:39,22 --> 0:2:42,56
2 test suites, 2 of them, it was
just, I didn't even want to

64
0:2:42,56 --> 0:2:43,18
think about it.

65
0:2:43,18 --> 0:2:44,62
It was not a lot of fun.

66
0:2:45,04 --> 0:2:46,72
And that's basically the project
as it is today.

67
0:2:46,72 --> 0:2:48,4
Now it's, of course, we're rolling
along.

68
0:2:48,4 --> 0:2:49,34
It's a C project.

69
0:2:49,74 --> 0:2:54,18
We have features based on user
demand, based on chasing performance.

70
0:2:55,08 --> 0:2:56,46
Performance is always the big thing.

71
0:2:56,46 --> 0:2:58,58
Like how do we move a bunch of
data quickly?

72
0:2:59,04 --> 0:3:0,3
How do we do it efficiently?

73
0:3:0,78 --> 0:3:4,12
How do we keep the repo as small
as possible, block incremental

74
0:3:4,12 --> 0:3:6,38
backups, the list goes on and on.

75
0:3:7,66 --> 0:3:9,84
Nikolay: I have a question here about
history.

76
0:3:10,58 --> 0:3:16,32
Maybe you remember those discussions
should Postgres have Remember

77
0:3:16,32 --> 0:3:21,0
times of Slony and then Londiste
and there was discussion

78
0:3:21,04 --> 0:3:23,6
should Postgres have replication
inside.

79
0:3:25,48 --> 0:3:30,24
And there are many such discussions,
same for auto failover and

80
0:3:30,24 --> 0:3:30,92
so on.

81
0:3:31,3 --> 0:3:35,64
And For the case of replication,
the idea to have it in core

82
0:3:35,64 --> 0:3:36,14
won.

83
0:3:36,76 --> 0:3:39,52
For the case of auto failover, it
lost.

84
0:3:40,08 --> 0:3:41,92
Still auto failover is outside.

85
0:3:42,26 --> 0:3:46,18
I'm very curious what you think
about the idea to have full-fledged

86
0:3:46,4 --> 0:3:49,95999
backup solution inside core and
why it's not happening.

87
0:3:51,2 --> 0:3:53,94
David: 1 interesting historical
fact is, and 1 of the reasons

88
0:3:53,94 --> 0:3:56,32
why some of the other committers
were involved in really early

89
0:3:56,32 --> 0:3:59,62
planning, and part of the reason
for the migration to C was the

90
0:3:59,62 --> 0:4:4,08
idea was that we could actually
maybe make pgBackRest the core

91
0:4:4,08 --> 0:4:5,26
solution for backup.

92
0:4:5,82 --> 0:4:9,4
Now, this is not something that
was endorsed by core broadly,

93
0:4:9,66 --> 0:4:13,04
or it was just discussions that
I had with some committers, because

94
0:4:13,04 --> 0:4:15,28
they were interested in having
a more comprehensive solution

95
0:4:15,28 --> 0:4:16,02002
in core.

96
0:4:16,36 --> 0:4:18,3
Obviously, they wanted it to be
written in C.

97
0:4:18,94 --> 0:4:23,0
And so that was 1 of the reasons
that drove the adoption of C.

98
0:4:23,94 --> 0:4:27,54
As for actually having something
like pgBackRest in core, it's

99
0:4:27,54 --> 0:4:32,44
a little tricky because the pgBackRest
project moves a lot faster

100
0:4:32,44 --> 0:4:37,5
than core does So to have it on
the yearly cycle would be But

101
0:4:37,5 --> 0:4:41,24
let's say we had put pgBackRest
into core like from the beginning

102
0:4:41,24 --> 0:4:43,32
Maybe as soon as it was migrated
to C.

103
0:4:43,32 --> 0:4:46,88
It would not be nearly as far along
as it is now At the time

104
0:4:46,88 --> 0:4:50,8
we were doing 12 releases a year,
we went down to 6, and now

105
0:4:50,8 --> 0:4:53,44
we're currently at 4, which I think
is just about the right tempo

106
0:4:53,44 --> 0:4:54,96
for a project of this type.

107
0:4:54,96 --> 0:4:58,62
But being able to release features
4 times a year, get new stuff

108
0:4:58,62 --> 0:5:1,08
out there, get people trying it,
get people testing it, that

109
0:5:1,08 --> 0:5:1,56
kind of stuff.

110
0:5:1,56 --> 0:5:2,82
I think it's pretty important.

111
0:5:3,72 --> 0:5:8,86
And as pgBackRest gets more stable,
maybe it would be more appropriate.

112
0:5:9,4 --> 0:5:12,38
But at the same time, the project
has diverged significantly

113
0:5:12,56 --> 0:5:16,48
from, you know, we use a lot of
the same concepts as Postgres,

114
0:5:16,68 --> 0:5:20,66
MemContext, and error handling
looks a lot similar and et cetera,

115
0:5:20,66 --> 0:5:21,2
et cetera.

116
0:5:21,2 --> 0:5:22,62
But it's all pretty different.

117
0:5:23,04 --> 0:5:25,22
2, pgBackRest grew its own way.

118
0:5:25,52 --> 0:5:27,24
So it'd be very difficult to get
in there.

119
0:5:27,24 --> 0:5:30,78
I think the idea is pg_basebackup
was going to be that tool.

120
0:5:31,3 --> 0:5:35,04
Getting incremental backups into
pg_basebackup was a big step,

121
0:5:36,2 --> 0:5:38,06
but it's still not a complete tool.

122
0:5:38,6 --> 0:5:41,16
You have to have tooling on top
of pg_basebackup for it to

123
0:5:41,16 --> 0:5:44,88
be at all usable, especially for
incremental, because reconstructing

124
0:5:45,04 --> 0:5:48,1
all those incrementals, going and
fetching them and uncompressing

125
0:5:48,28 --> 0:5:51,26
them and getting them all ready
for the pg_combinebackup tool

126
0:5:51,26 --> 0:5:51,92
to run.

127
0:5:52,04 --> 0:5:54,84
That's a significant amount of
work and obviously it doesn't

128
0:5:54,84 --> 0:5:58,62
do anything with WAL archiving,
expiration, there's just the

129
0:5:58,62 --> 0:6:2,6
list of all the things that you
need to do and it's a pretty

130
0:6:2,6 --> 0:6:5,58
big list and it's intimidating,
honestly.

131
0:6:6,02 --> 0:6:7,36
You know how core goes, right?

132
0:6:7,36 --> 0:6:9,62
It's intimidating to even contemplate
getting something like

133
0:6:9,62 --> 0:6:10,58
that into core.

134
0:6:11,0 --> 0:6:15,72
So I think it would be a good idea,
but also Think about storage

135
0:6:15,72 --> 0:6:16,22
drivers.

136
0:6:16,78 --> 0:6:21,6
Support for pg_basebackup supports
POSIX, right?

137
0:6:21,6 --> 0:6:27,76
But most backup tools support S3,
GCS, Azure, SFTP, etc.

138
0:6:27,8 --> 0:6:30,04
Because those are the tools that
people are actually using.

139
0:6:30,04 --> 0:6:32,3
So all that would have to go into
core as well.

140
0:6:32,32 --> 0:6:34,5
All the storage drivers, people
would have to...

141
0:6:35,98 --> 0:6:37,5
Nikolay: Maybe not everything could
go.

142
0:6:37,5 --> 0:6:39,56
It's possible, Postgres is very
extensible.

143
0:6:40,2 --> 0:6:44,98
If just the core thing would go
to the core, but expose some

144
0:6:44,98 --> 0:6:47,94
interfaces, particular drivers
could stay outside.

145
0:6:48,88 --> 0:6:51,44
David: And in theory, that's
what we've done, a bit with pg_basebackup,

146
0:6:51,68 --> 0:6:53,86
but the amount of stuff that needs
to be done by the outside

147
0:6:53,86 --> 0:6:54,98
tool is huge.

148
0:6:55,28 --> 0:6:58,6
Now there are quite a few tools
that are based on pg_basebackup

149
0:6:58,94 --> 0:7:1,98
to do their page level slash block
level incrementals.

150
0:7:2,22 --> 0:7:4,46
Barman, obviously, is the most
well-known 1.

151
0:7:4,64 --> 0:7:6,68
I think pgmoneta is 1 of them.

152
0:7:6,96 --> 0:7:9,72
There's a newer 1 that I saw recently
that's also using it.

153
0:7:9,72 --> 0:7:12,48
So you can certainly build a tool
around pg_basebackup.

154
0:7:12,72 --> 0:7:15,9
And I think the idea would be that
you'd extend and extend pg_basebackup,

155
0:7:17,44 --> 0:7:20,42
and then people would take that
feature out of their tool.

156
0:7:20,48 --> 0:7:23,68
And eventually maybe pg_basebackup
would be a thing that does

157
0:7:23,68 --> 0:7:24,18
everything.

158
0:7:24,86 --> 0:7:28,7
But at the pace that it's actually
evolving, I would expect that

159
0:7:28,7 --> 0:7:32,24
to be able to reach feature parity
with something like pgBackRest

160
0:7:32,24 --> 0:7:35,26
or WAL-G or Barman in approximately
30 years.

161
0:7:36,34 --> 0:7:39,72
I'm exaggerating a little bit,
but if I look at the actual progress

162
0:7:39,72 --> 0:7:43,1
that pg_basebackup has made over
the years, that's where we are.

163
0:7:43,38 --> 0:7:46,12
Since that incremental, nothing's
been done with that, even though

164
0:7:46,12 --> 0:7:48,84
there are pretty huge performance
implications for restores.

165
0:7:49,24 --> 0:7:53,1
Any large, very large database
to restore with that page incremental

166
0:7:53,14 --> 0:7:56,58
format is you basically need, let's
say you've got a database

167
0:7:56,58 --> 0:7:59,98
that's a terabyte, you need at
least 2 terabytes to do the restore

168
0:8:0,64 --> 0:8:3,78
minimum, and it depends on how
many incrementals you have, so

169
0:8:3,78 --> 0:8:6,88
you could need 2, 3 terabytes to
do the restore.

170
0:8:7,44 --> 0:8:10,16
All the files need to be copied
down regardless of whether those

171
0:8:10,16 --> 0:8:11,46
blocks are used or not.

172
0:8:11,46 --> 0:8:14,44
Everything needs to be uncompressed
and then fed into pg_combinebackup

173
0:8:14,44 --> 0:8:18,3
which then rewrites everything
to a different location.

174
0:8:19,2 --> 0:8:23,94
So from a scalability standpoint,
it has a pretty serious problem.

175
0:8:24,52 --> 0:8:27,6
And we're 2 versions on from it
being introduced and no 1 has

176
0:8:27,6 --> 0:8:30,76
actually even thought about actually
addressing any of those

177
0:8:30,76 --> 0:8:31,26
issues.

178
0:8:31,72 --> 0:8:34,9
I don't want to harp on this too
much, but the main point is

179
0:8:35,24 --> 0:8:37,52
getting stuff into core is hard.

180
0:8:38,2 --> 0:8:39,16
It takes me years.

181
0:8:39,16 --> 0:8:44,02
I've been working on a very small
change for backup just to mark

182
0:8:44,02 --> 0:8:46,22
pg_control when a backup label
is required.

183
0:8:46,68 --> 0:8:49,76
So you do a backup, a backup label
is required to do the recovery.

184
0:8:49,76 --> 0:8:53,82
If the user deletes the backup
label, they can end up with corruption,

185
0:8:54,0 --> 0:8:54,8
silent corruption.

186
0:8:54,96 --> 0:8:56,26
It's quite annoying actually.

187
0:8:56,74 --> 0:9:0,48
And I've seen people who know Postgres
really well, hackers do

188
0:9:0,48 --> 0:9:2,56
this and not understand what happened.

189
0:9:2,56 --> 0:9:5,38
Nikolay: I saw it many times already
in various teams.

190
0:9:5,38 --> 0:9:6,1
It's annoying.

191
0:9:7,12 --> 0:9:10,58
David: So 2 years ago I introduced
a small patch for Postgres

192
0:9:10,58 --> 0:9:14,9
to just put a flag in pg_control
to mark it when we actually

193
0:9:14,9 --> 0:9:17,42
need a backup label.

194
0:9:17,42 --> 0:9:20,16
So if you start Postgres and backup
label's not there, it will

195
0:9:20,16 --> 0:9:22,7
just stop and say, no, you must
have backup label.

196
0:9:23,4 --> 0:9:24,44
Please provide.

197
0:9:25,12 --> 0:9:26,46
That was 2 years ago.

198
0:9:26,82 --> 0:9:30,98
It goes through review occasionally,
but really nothing there.

199
0:9:31,1 --> 0:9:34,18
In pgBackRest, we actually implemented
that about 3 years ago.

200
0:9:34,54 --> 0:9:37,84
What we do, this is a bit of a
hack but it works pretty well,

201
0:9:37,84 --> 0:9:42,02
what we do is we overwrite the
last checkpoint in pg_control,

202
0:9:43,1 --> 0:9:44,88
we write in the hex value DEAD.

203
0:9:46,18 --> 0:9:49,28
So if we get a report from a user,
because it will come up a

204
0:9:49,28 --> 0:9:52,0
Postgres that will say unable to
find checkpoint DEAD.

205
0:9:52,96 --> 0:9:55,18
And that's actually not a valid
checkpoint at all because it's

206
0:9:55,18 --> 0:9:57,44
under the 1st WAL segment limit.

207
0:9:58,08 --> 0:10:0,96
So, Postgres will mark it as invalid,
it will throw an error,

208
0:10:0,96 --> 0:10:4,34
and then when we get a report from
the user we immediately know,

209
0:10:4,74 --> 0:10:9,52
hey you tried to start this you
delete a backup label and they're

210
0:10:9,52 --> 0:10:12,26
like yeah I deleted backup label
and boom there you go.

211
0:10:12,34 --> 0:10:15,44
But I've been trying to get that
that I wrote that patch 2 years

212
0:10:15,44 --> 0:10:18,48
ago and I'm still hoping to get
that into Postgres.

213
0:10:18,48 --> 0:10:21,54
So that's like the speed at which
Postgres can operate sometimes,

214
0:10:21,54 --> 0:10:24,14
especially for people like me who
aren't really committers.

215
0:10:25,2 --> 0:10:28,04
So can you imagine trying to get
something the size of pgBackRest

216
0:10:28,2 --> 0:10:30,92
or complexity of pgBackRest into
Postgres?

217
0:10:31,46 --> 0:10:34,78
Nikolay: Yeah, I wish we could highlight
this, put a link to this

218
0:10:34,78 --> 0:10:39,02
patch and maybe drive some attention
to it because some people

219
0:10:39,02 --> 0:10:42,6
who participate in Postgres hacking
listen to us and maybe...

220
0:10:43,18 --> 0:10:45,88
I agree it's super annoying and
right now All the other tools

221
0:10:45,88 --> 0:10:46,82
they need to...

222
0:10:47,22 --> 0:10:51,52
Basically, if you create full copy,
full backup on a replica,

223
0:10:51,54 --> 0:10:55,22
you are responsible for placing
this backup label yourself.

224
0:10:55,52 --> 0:10:58,54
And it's also how restore works
in Postgres.

225
0:10:58,66 --> 0:11:2,72
I fixed recently, and This fix
was very quick because it was

226
0:11:2,72 --> 0:11:3,22
obvious.

227
0:11:3,34 --> 0:11:6,9
It was just a problem in code and
Postgres in the store path.

228
0:11:7,3 --> 0:11:11,24
How Postgres is working with the
backup label and also signal

229
0:11:11,24 --> 0:11:11,74
files.

230
0:11:12,12 --> 0:11:15,54
If you look how people use it,
they see some error in logs and

231
0:11:15,54 --> 0:11:17,72
then, okay, I will just delete
some file, right?

232
0:11:17,72 --> 0:11:21,04
To get it going and this is what's
happening all the time right

233
0:11:21,1 --> 0:11:25,68
but also I noticed pgBackRest also
requires maybe because of what

234
0:11:25,68 --> 0:11:29,9
this mechanism you explained well
when making backup on replica

235
0:11:29,9 --> 0:11:33,64
it requires connection to primary
And this surprised me a lot

236
0:11:33,64 --> 0:11:36,18
recently when I started to code,
right?

237
0:11:36,86 --> 0:11:39,4
David: Yeah, that actually doesn't
have anything to do with

238
0:11:39,4 --> 0:11:40,16
that specifically.

239
0:11:40,32 --> 0:11:43,58
The reason why we did that is you
get a better backup that way.

240
0:11:44,2 --> 0:11:45,04
Nikolay: Better backup.

241
0:11:45,06 --> 0:11:46,72
Not all backups are equal.

242
0:11:47,98 --> 0:11:49,16
David: You get a better backup.

243
0:11:49,16 --> 0:11:51,6
1st of all, there are some things
about backup from standby.

244
0:11:51,88 --> 0:11:54,48
I don't necessarily want to get
into that because it's pretty

245
0:11:54,48 --> 0:11:56,78
esoteric, but there's some things
about backups from standby

246
0:11:56,78 --> 0:11:58,86
that I don't 100% trust.

247
0:11:59,54 --> 0:12:1,32
But I do believe it works.

248
0:12:2,24 --> 0:12:6,2
The reason why a primary backup
is better, though, is you get

249
0:12:6,2 --> 0:12:6,88
a couple things.

250
0:12:6,88 --> 0:12:10,22
You get stats from the primary,
which is actually pretty nice.

251
0:12:10,32 --> 0:12:14,3
If you do a backup from standby
and you restore, you've got stats

252
0:12:14,3 --> 0:12:15,96
from, say, a standby.

253
0:12:16,62 --> 0:12:20,06
And if you're actually doing analysis
of statistics, it's pretty

254
0:12:20,06 --> 0:12:23,32
shocking to suddenly have like
all of your index patterns, scan

255
0:12:23,32 --> 0:12:25,54
pattern, read patterns, full change,
everything.

256
0:12:25,68 --> 0:12:29,02
So I'm not talking about, sorry,
table statistics like basically

257
0:12:29,18 --> 0:12:30,32
like scan statistics.

258
0:12:30,36 --> 0:12:31,78
Maybe I'm using the wrong word.

259
0:12:31,78 --> 0:12:32,08
Michael: Shoot.

260
0:12:32,08 --> 0:12:34,92
No, I think you, I think I understand
number of seq scans on a

261
0:12:34,92 --> 0:12:35,42
specific.

262
0:12:35,58 --> 0:12:37,8
David: Yeah, I think they're
both, I think they're both stats,

263
0:12:37,8 --> 0:12:39,9
but, but I'm not talking about
the planner statistics.

264
0:12:39,92 --> 0:12:42,76
I'm talking about like the actual
usage stats.

265
0:12:43,14 --> 0:12:46,4
Another thing is if you're doing
a backup on the standby, you

266
0:12:46,4 --> 0:12:49,86
can't actually finish the backup
and verify that all the WAL

267
0:12:49,86 --> 0:12:51,56
has reached the archive.

268
0:12:51,66 --> 0:12:55,8
You can, but you might have to
wait a day for that to happen

269
0:12:55,8 --> 0:12:56,46
or whatever.

270
0:12:56,58 --> 0:12:59,64
We feel it's really important to
make sure that this backup,

271
0:13:0,04 --> 0:13:2,66
When we mark the backup as done,
we want it to be done.

272
0:13:3,4 --> 0:13:6,76
Not hoping that someday this WAL
archive is going to arrive.

273
0:13:6,76 --> 0:13:9,18
We want it to be done at that moment.

274
0:13:10,08 --> 0:13:11,26
So that's pretty important.

275
0:13:11,68 --> 0:13:13,76
Nikolay: Do you mean the backup is
self-sufficient?

276
0:13:15,06 --> 0:13:18,0
It positively affects RTO, right?

277
0:13:18,04 --> 0:13:20,46
Like recovery time objective.

278
0:13:20,54 --> 0:13:23,12
It reduces the time needed to recover.

279
0:13:23,68 --> 0:13:27,18
David: Well, at the very least,
we want to know that the backup

280
0:13:27,18 --> 0:13:29,04
can be recovered to consistency.

281
0:13:30,06 --> 0:13:30,52
Right.

282
0:13:30,52 --> 0:13:33,06
Now, if you want to do point-in-time
recovery, that's going to

283
0:13:33,06 --> 0:13:35,8
be after the backup finishes, you'll
continue WAL archiving,

284
0:13:35,8 --> 0:13:37,36
you need to monitor that as well.

285
0:13:37,36 --> 0:13:40,24
But at the point where the backup
finishes, we want that to definitely

286
0:13:40,24 --> 0:13:40,86
be true.

287
0:13:40,92 --> 0:13:43,94
Other things like getting actual
logs from the primary are more

288
0:13:43,94 --> 0:13:45,86
interesting than getting logs from
a standby.

289
0:13:46,5 --> 0:13:49,12
Although I don't really recommend
that people put logs in their

290
0:13:49,12 --> 0:13:50,54
PGDATA directory at all.

291
0:13:50,54 --> 0:13:51,72
It's actually quite common.

292
0:13:52,28 --> 0:13:53,42
Nikolay: Hold on 1 2nd.

293
0:13:53,72 --> 0:13:54,84
Too many things here.

294
0:13:54,84 --> 0:13:57,66
1st of all, when you say logs,
it's WAL files, right?

295
0:13:58,28 --> 0:14:0,06
David: No, sorry, in this case
I do mean logs.

296
0:14:0,06 --> 0:14:3,34
So like basically textual logs
that are generated by Postgres,

297
0:14:3,34 --> 0:14:5,24
and a lot of people will put those
in PGDATA.

298
0:14:6,82 --> 0:14:7,96
Nikolay: Okay, I didn't get it.

299
0:14:7,96 --> 0:14:10,14
Why do we care about logs when
making backup?

300
0:14:10,24 --> 0:14:13,26
Ah, because we don't want them
to be backed up, right?

301
0:14:13,26 --> 0:14:15,24
David: We do if you have them
in...

302
0:14:15,24 --> 0:14:18,54
So If people are putting their
logs in PGDATA and they expect

303
0:14:18,54 --> 0:14:21,4
them to be from the primary, it
can be surprising when their

304
0:14:21,96 --> 0:14:24,14
logs from a standby or something
like that.

305
0:14:24,14 --> 0:14:24,84
I see.

306
0:14:25,08 --> 0:14:27,36
They're not seeing the information
that they expect to get.

307
0:14:27,36 --> 0:14:30,06
Maybe they have auditing turned
on in the primary and not on

308
0:14:30,06 --> 0:14:30,74
the standby.

309
0:14:31,5 --> 0:14:34,7
You recover the backup and now
all your audit records are gone,

310
0:14:35,28 --> 0:14:35,78
etc.

311
0:14:36,46 --> 0:14:40,28
So now I'm not a proponent of putting
logs in PGDATA on the

312
0:14:40,28 --> 0:14:43,46
primary, but there are people who
do this and they want to preserve

313
0:14:43,46 --> 0:14:43,94
those logs.

314
0:14:43,94 --> 0:14:44,88
That's another reason.

315
0:14:45,04 --> 0:14:48,4
So there's a whole raft of reasons,
but basically, so what we

316
0:14:48,4 --> 0:14:51,3
do is we copy everything that's
replicated from the standby.

317
0:14:51,38 --> 0:14:54,0
Oh, the other reason why we really
wanted to do this, although

318
0:14:54,0 --> 0:14:57,28
honestly we've never implemented
it, is once you've done this,

319
0:14:57,28 --> 0:15:0,98
you can actually parallelize a
backup across all available standbys.

320
0:15:3,06 --> 0:15:6,78
Because you're coordinating everything
on the primary, you wait

321
0:15:6,78 --> 0:15:10,58
for the standbys to reach the checkpoint
where the backup started,

322
0:15:10,58 --> 0:15:11,76
which we already do, of course.

323
0:15:11,76 --> 0:15:13,66
We do a bunch of consistency checks.

324
0:15:13,78 --> 0:15:17,14
And then you would be able to parallelize
your backup across

325
0:15:17,52 --> 0:15:21,1
all the available standbys and
really supercharge it.

326
0:15:21,38 --> 0:15:23,3
And it's not even that hard to
do.

327
0:15:23,4 --> 0:15:25,76
We just haven't really, there hasn't
been demand for it.

328
0:15:25,76 --> 0:15:27,78
And there's always 1000000 things
to write.

329
0:15:27,98 --> 0:15:28,68
You know how it is.

330
0:15:28,68 --> 0:15:30,04
So we just haven't gotten to it
yet.

331
0:15:30,04 --> 0:15:32,08
But that was another big reason
to do it that way, because it

332
0:15:32,08 --> 0:15:35,64
allows us to parallelize backup
in a way that otherwise wouldn't

333
0:15:35,64 --> 0:15:39,06
be possible if you're just backing
up from a single standby.

334
0:15:39,32 --> 0:15:43,1
Nikolay: There is a dilemma here
where, 1st of all, making full

335
0:15:43,1 --> 0:15:46,28
backup on primary, especially if
it's full, not incremental,

336
0:15:46,42 --> 0:15:48,42
full, It's huge stress for disks.

337
0:15:48,42 --> 0:15:49,66
David: Yeah, you don't want that.

338
0:15:49,86 --> 0:15:50,36
Nikolay: Yeah.

339
0:15:50,54 --> 0:15:54,58
So usually people have loaded to
some standby, and then the problem

340
0:15:54,58 --> 0:15:57,54
is also which standby, because
it's also stress.

341
0:15:57,54 --> 0:16:0,82
If it's a read replica, it's stress
for those reads as well.

342
0:16:0,94 --> 0:16:4,0
So distributing totally makes sense,
but there is also, like,

343
0:16:4,0 --> 0:16:7,84
since you talk and focus a lot
on corruption and consistency,

344
0:16:8,54 --> 0:16:12,46
I was always curious if we, for
example, have allocated standby

345
0:16:12,66 --> 0:16:16,98
only for backups, which I think
Crunchy Bridge had an issue, we

346
0:16:16,98 --> 0:16:19,84
had a customer there, Crunchy Bridge
had an issue when they made

347
0:16:19,84 --> 0:16:20,98
full backups on primary.

348
0:16:20,98 --> 0:16:24,52
I was super shocked because like
how come it was fixed then?

349
0:16:24,52 --> 0:16:29,0
And I think full backups went to
HA replica, which is for HA.

350
0:16:29,2 --> 0:16:31,22
It's allocated, it doesn't receive
reads.

351
0:16:31,4 --> 0:16:34,48
But then I'm curious, if some corruption
happens, it might happen

352
0:16:34,48 --> 0:16:36,84
only, there are many kinds of corruption,
right?

353
0:16:37,12 --> 0:16:41,08
I can think about, at least theoretically,
some corruption might

354
0:16:41,08 --> 0:16:43,98
be only on 1 node, but not on others.

355
0:16:44,44 --> 0:16:47,68
And if you always make backups
from only 1 node and you don't

356
0:16:47,68 --> 0:16:50,14
notice this corruption propagates
to all backups.

357
0:16:50,66 --> 0:16:54,44
So some rotation at least is good,
but combining and parallelization,

358
0:16:54,86 --> 0:16:55,94
it's also interesting.

359
0:16:56,14 --> 0:16:59,36
I'm just thinking when SQL Server,
for example, there is a mechanism

360
0:16:59,54 --> 0:17:3,48
to self-heal if, for example, primary
notices some pages are

361
0:17:3,48 --> 0:17:7,06
corrupted, it can grab them from
replicas, which was implemented

362
0:17:7,16 --> 0:17:7,98
long ago.

363
0:17:8,44 --> 0:17:9,64
This is cool technology.

364
0:17:10,02 --> 0:17:12,84
I'm curious about this dilemma.

365
0:17:12,88 --> 0:17:17,24
We want also distribute load not
to hurt our user traffic, but

366
0:17:17,24 --> 0:17:19,02
also what's happening with corruption.

367
0:17:20,5 --> 0:17:22,86
This is Pandora box of topics.

368
0:17:23,72 --> 0:17:25,52
David: That'd be an interesting
idea actually, too.

369
0:17:25,52 --> 0:17:29,54
If you were doing backups from
multiple standbys, you could compare

370
0:17:29,54 --> 0:17:30,06
and contrast.

371
0:17:30,06 --> 0:17:33,66
If there was corruption on 1 standby,
you could say, hey, why

372
0:17:33,66 --> 0:17:36,52
don't I try grabbing this file
from the other standby instead,

373
0:17:36,82 --> 0:17:40,02
fail it, send it back to the main
process, which would reissue

374
0:17:40,04 --> 0:17:42,38
that job on a different standby.

375
0:17:43,04 --> 0:17:44,44
That's a pretty cool idea.

376
0:17:44,76 --> 0:17:48,06
Nikolay: You talk about corruption,
like data checksums level

377
0:17:48,06 --> 0:17:50,64
of corruption but there may be
a higher level of corruption like

378
0:17:50,64 --> 0:17:54,2
index level, like logical level,
foreign keys, like a lot of

379
0:17:54,2 --> 0:17:54,62
stuff.

380
0:17:54,62 --> 0:17:57,44
David: Oh sure and your corruption
scenarios on primaries and

381
0:17:57,44 --> 0:18:0,8
standbys are quite different too
because the way things are being

382
0:18:0,8 --> 0:18:3,2
written out on a primary is actually
fairly different to the

383
0:18:3,2 --> 0:18:5,04
way it's being written out on a
standby.

384
0:18:5,38 --> 0:18:8,9
Also the standby is essentially
a single threaded operation.

385
0:18:9,64 --> 0:18:10,2
Let's not

386
0:18:10,2 --> 0:18:10,68
Nikolay: go there,

387
0:18:10,68 --> 0:18:11,46
David: it's terrible.

388
0:18:11,58 --> 0:18:15,08
Yeah, obviously it's terrible for
performance from the standpoint

389
0:18:15,18 --> 0:18:19,24
of any issues there might be with
locking and parallelism and

390
0:18:19,24 --> 0:18:22,7
other things on a primary, the
standby is not going to see those

391
0:18:22,7 --> 0:18:24,5
bugs, probably.

392
0:18:25,2 --> 0:18:28,76
And so there are whole classes
of bugs that can happen on a primary

393
0:18:28,78 --> 0:18:31,24
that could cause corruption that
aren't going to happen on a

394
0:18:31,24 --> 0:18:31,74
standby.

395
0:18:32,4 --> 0:18:36,3
So from that perspective, it might
actually be safer to do backups

396
0:18:36,3 --> 0:18:38,86
on a standby as well, just because
it's a simpler path.

397
0:18:38,86 --> 0:18:41,28
The data goes through a much simpler
path to get to disk on the

398
0:18:41,28 --> 0:18:45,04
standby than it does on the primary,
as a rule, unless you really

399
0:18:45,04 --> 0:18:48,96
have 1 writer on the primary, which
is not the way things tend

400
0:18:48,96 --> 0:18:51,9
to run these days, not for interesting
databases at least.

401
0:18:52,54 --> 0:18:54,98
So yeah.

402
0:18:55,08 --> 0:18:58,36
And as you say, the corruption
we're detecting is only checksum

403
0:18:59,38 --> 0:18:59,88
corruption.

404
0:19:0,04 --> 0:19:2,28
There are some projects out there
where people have combined

405
0:19:2,28 --> 0:19:5,04
pgBackRest with amcheck.

406
0:19:5,92 --> 0:19:8,76
So basically when it does recovery,
like basically you call this

407
0:19:8,76 --> 0:19:12,7
command, it automatically integrates
pgBackRest with some sanity

408
0:19:12,7 --> 0:19:16,4
checks afterwards and does amcheck
and some other things to check

409
0:19:16,4 --> 0:19:19,46
a higher level of consistency than
just the page checksums.

410
0:19:19,64 --> 0:19:22,08
Nikolay: Yeah, when you do recovery
it's a good point usually.

411
0:19:22,08 --> 0:19:23,98
This is what we usually do.

412
0:19:24,24 --> 0:19:28,04
Not all recovery attempts, but
some of them, percentage of them,

413
0:19:28,04 --> 0:19:31,04
should also check index health.

414
0:19:31,64 --> 0:19:33,52
David: This is really useful
in test recovery.

415
0:19:33,52 --> 0:19:36,32
So hopefully you're testing your
recovery, right?

416
0:19:36,5 --> 0:19:39,72
And that's a great time to do amcheck
because then you can have

417
0:19:39,72 --> 0:19:44,48
confidence in your emergency hair-on-
fire recovery that it actually

418
0:19:44,48 --> 0:19:48,62
does work because you practiced
it, you ran amcheck on all the

419
0:19:48,74 --> 0:19:49,7
practice runs.

420
0:19:50,02 --> 0:19:52,58
So at that point you should feel
pretty confident to run the

421
0:19:52,58 --> 0:19:56,32
production restore without having
to take the penalty of running

422
0:19:56,32 --> 0:20:0,26
amcheck at that time, if you've
been properly testing.

423
0:20:0,8 --> 0:20:4,82
Nikolay: I agree with everything,
but You cannot do it on all runs

424
0:20:4,82 --> 0:20:7,9
because full-fledged amcheck takes
a lot of time.

425
0:20:8,6 --> 0:20:10,92
David: Yeah, that's why I say
it's really for test restore,

426
0:20:10,92 --> 0:20:11,24
right?

427
0:20:11,24 --> 0:20:13,94
So you're testing things and it's
tricky, right?

428
0:20:13,94 --> 0:20:18,16
Okay, I guess you can say amcheck
takes 2 days, so we're going

429
0:20:18,16 --> 0:20:19,42
to do this once a week.

430
0:20:19,82 --> 0:20:22,72
And then we'll do more recovery
testing, but only run amcheck

431
0:20:22,74 --> 0:20:23,42
once a week.

432
0:20:23,42 --> 0:20:26,04
I think that's a perfectly valid
way to go about it, to think

433
0:20:26,04 --> 0:20:26,78
about it.

434
0:20:27,34 --> 0:20:27,84
Absolutely.

435
0:20:28,14 --> 0:20:33,1
Nikolay: I'm also very curious, did
you think about some more observability

436
0:20:34,3 --> 0:20:36,98
capabilities in pgBackRest to understand
that.

437
0:20:36,98 --> 0:20:40,22
My final goal is, for example,
I have 100 clusters.

438
0:20:41,58 --> 0:20:47,14
I just want to know actual RPO,
RTO all the time, measured somehow,

439
0:20:47,16 --> 0:20:50,78
during those restore tests, during
the backup process itself.

440
0:20:51,14 --> 0:20:55,18
And just to have some high-level
view, this is what's happening.

441
0:20:55,18 --> 0:20:55,94
That's it.

442
0:20:56,58 --> 0:21:0,12
Now you need to build a lot to
have that.

443
0:21:1,38 --> 0:21:5,32
David: RPO you can measure using
the information that pgBackRest

444
0:21:5,54 --> 0:21:7,26
provides in the Intel command.

445
0:21:7,66 --> 0:21:10,52
Because it'll tell you what the
most recently, how up to date

446
0:21:10,52 --> 0:21:12,6
you are on your WAL segments,
right?

447
0:21:12,6 --> 0:21:14,24
So that's going to give you your
RPO.

448
0:21:15,06 --> 0:21:19,3
So I'm 1 segment behind, 10 segments
behind, 5 segments behind.

449
0:21:19,78 --> 0:21:23,3
So at least you know what your
recovery point is based on the

450
0:21:23,3 --> 0:21:24,52
information that you get.

451
0:21:24,52 --> 0:21:26,98
RTO is a much trickier scenario
though.

452
0:21:26,98 --> 0:21:30,9
Nikolay: RTO is also tricky because
you think like technical guy

453
0:21:30,9 --> 0:21:32,48
and I'm also technical guy.

454
0:21:32,54 --> 0:21:33,98
We think about bytes.

455
0:21:34,54 --> 0:21:39,52
Business guys think about seconds,
hours of data loss.

456
0:21:39,52 --> 0:21:41,66
How many hours of data were lost?

457
0:21:41,78 --> 0:21:44,24
And this you need to translate
somehow.

458
0:21:45,04 --> 0:21:46,64
David: It's absolutely fair and
you're not going to be able

459
0:21:46,64 --> 0:21:51,36
to get that timing exactly from
just you could use for instance

460
0:21:51,36 --> 0:21:56,1
our repo-ls command to go and grab
information about the WAL

461
0:21:56,18 --> 0:21:58,92
and find out how much WAL you're
generating per 2nd.

462
0:21:59,44 --> 0:22:3,34
It's not going to give you exactly
the time, but it's going to

463
0:22:3,34 --> 0:22:4,38
be awfully close.

464
0:22:5,14 --> 0:22:8,3
So you could estimate, you're like,
OK, we're generating 5 WAL

465
0:22:8,3 --> 0:22:9,02
per minute.

466
0:22:9,56 --> 0:22:10,82
We're 2 WAL behind.

467
0:22:11,72 --> 0:22:15,2
And therefore, our recovery point
is this many bytes behind and

468
0:22:15,2 --> 0:22:16,6
this many seconds behind.

469
0:22:17,2 --> 0:22:19,6
So you would actually be able to
estimate that reasonably well

470
0:22:19,6 --> 0:22:21,8
using tools just that are available
on pgBackRest.

471
0:22:22,36 --> 0:22:26,2
Although actually estimating the
time of the WAL, you would

472
0:22:26,2 --> 0:22:29,44
have to roll your own for that
because that would have to be

473
0:22:29,44 --> 0:22:30,66
something you do for repo.

474
0:22:30,66 --> 0:22:33,08
So the nice thing about repo-ls
is it's a command we give you

475
0:22:33,08 --> 0:22:37,4
and it works on any repository
whether you're on S3, GCS, SFTP,

476
0:22:38,24 --> 0:22:38,68
POSIX.

477
0:22:38,68 --> 0:22:43,08
So you can use that tool to get
information about the repository

478
0:22:43,08 --> 0:22:46,92
that's not included in our JSON
info and it will work in any

479
0:22:46,92 --> 0:22:49,82
environment, on any repository,
on any storage.

480
0:22:49,82 --> 0:22:54,48
So you don't have to muck about
with AWS CLI over here and the

481
0:22:54,48 --> 0:22:58,32
Azure CLI over here and whatever
the GCS equivalent is and et

482
0:22:58,32 --> 0:22:59,28
cetera, et cetera.

483
0:23:0,06 --> 0:23:2,84
And then, okay, so you could sort
that out, I think, with the

484
0:23:2,84 --> 0:23:5,98
tools that are available with pgBackRest,
but RTO is still a

485
0:23:5,98 --> 0:23:7,28
thing that needs to be tested.

486
0:23:7,66 --> 0:23:10,16
We can't tell you RTO at backup
time.

487
0:23:10,24 --> 0:23:15,92
And since Postgres is famously
slow for recovery, Maybe not slow,

488
0:23:15,92 --> 0:23:18,52
but certainly things don't get
written as fast as they get written

489
0:23:18,52 --> 0:23:19,26
on the primary.

490
0:23:19,44 --> 0:23:20,72
That much we can agree on.

491
0:23:20,72 --> 0:23:23,16
Recovery actually, it goes pretty
well.

492
0:23:23,68 --> 0:23:25,44
And Postgres has gotten pretty
good at things.

493
0:23:25,44 --> 0:23:29,54
pgBackRest prefetches, WAL segments,
all this combined actually

494
0:23:29,54 --> 0:23:32,06
gives you a pretty good, pretty
good performance overall.

495
0:23:32,98 --> 0:23:35,6
But again, you're going to have
to test it to know how long it

496
0:23:35,6 --> 0:23:37,7
takes you to do recovery.

497
0:23:38,56 --> 0:23:41,38
Fun little example, there have
been, and this hasn't happened

498
0:23:41,38 --> 0:23:44,18
1 time, but the 1st time it happened
was in early days of

499
0:23:44,18 --> 0:23:46,94
pgBackRest, because our thing has
always been high volume performance,

500
0:23:47,56 --> 0:23:47,9
et cetera.

501
0:23:47,9 --> 0:23:50,06
So there are a number of people
over the years who have used

502
0:23:50,06 --> 0:23:52,74
pgBackRest because they can't use
replication.

503
0:23:53,56 --> 0:23:56,74
Because replication simply will
not keep up with their load ever.

504
0:23:57,54 --> 0:24:1,06
So what they do is they basically
continuously take and restore

505
0:24:1,06 --> 0:24:1,56
backups.

506
0:24:2,64 --> 0:24:5,86
And that way, they can measure
the RTO pretty well, actually,

507
0:24:5,86 --> 0:24:7,8
because they're constantly doing
these recoveries.

508
0:24:8,48 --> 0:24:9,96
And they get to a point where they...

509
0:24:9,96 --> 0:24:12,94
And then we'll see how long it
takes to get to the point where

510
0:24:12,94 --> 0:24:15,24
they were when they did the restore,
right?

511
0:24:15,24 --> 0:24:17,9
Because they can never actually
catch up, but they know how long

512
0:24:17,9 --> 0:24:18,68
it's going to take.

513
0:24:18,68 --> 0:24:22,4
If the primary fails, how long
will it be until the standby is

514
0:24:22,58 --> 0:24:24,41
up to date using WAL from the archive?

515
0:24:24,41 --> 0:24:25,58
A downtime will be this.

516
0:24:27,98 --> 0:24:29,56
A downtime will be this.

517
0:24:29,68 --> 0:24:32,04
And they know it exactly because
they do it every single day.

518
0:24:32,04 --> 0:24:35,38
And this is a use case that I've
seen multiple times now.

519
0:24:35,38 --> 0:24:36,42
People have brought it up with
me.

520
0:24:36,42 --> 0:24:38,14
It's a really interesting use case.

521
0:24:38,3 --> 0:24:41,9
Nikolay: So they are bumping into
this 100% single CPU situation

522
0:24:42,12 --> 0:24:44,36
problem of startup process on replicas.

523
0:24:45,14 --> 0:24:47,82
So that means also they live without
replicas.

524
0:24:48,3 --> 0:24:51,02
With these replicas, they cannot
use them because they are lagging

525
0:24:51,02 --> 0:24:51,92
basically, right?

526
0:24:52,06 --> 0:24:53,68
David: Yeah, they're pretty far
behind.

527
0:24:53,68 --> 0:24:55,66
So they really are just there to...

528
0:24:57,04 --> 0:25:0,2
I mean, obviously you can make
up scenarios where a replica like

529
0:25:0,2 --> 0:25:0,88
this could be used.

530
0:25:0,88 --> 0:25:4,12
Let's say you're doing reporting
and you know your replica lag

531
0:25:4,12 --> 0:25:8,42
is 6 hours, then you know by 8
AM you can start running reports,

532
0:25:9,14 --> 0:25:9,84
that kind of thing.

533
0:25:9,84 --> 0:25:11,18
So you can game it.

534
0:25:11,36 --> 0:25:15,48
But you're right, it severely reduces
the usability of the replica

535
0:25:15,92 --> 0:25:16,68
in that scenario.

536
0:25:16,68 --> 0:25:19,14
But they just need something because
they can't start from...

537
0:25:19,74 --> 0:25:22,18
They're just trying to minimize
the time it takes for them to

538
0:25:22,18 --> 0:25:23,9
get to running again.

539
0:25:23,92 --> 0:25:28,38
Because if they start from a restore
of their 50T backup, that's

540
0:25:28,38 --> 0:25:30,66
going to be however long that takes.

541
0:25:30,66 --> 0:25:33,9
pgBackRest is pretty fast, but
50 terabytes is 50 terabytes.

542
0:25:34,44 --> 0:25:35,8
And then they start recovery.

543
0:25:35,8 --> 0:25:36,44
Nikolay: 1 hour.

544
0:25:37,7 --> 0:25:41,76
I showed an article 37 terabytes
per hour.

545
0:25:41,82 --> 0:25:45,04
We can do it 1 hour, but it's local
NVMe to local NVMe.

546
0:25:45,04 --> 0:25:45,98
It's very different.

547
0:25:46,24 --> 0:25:46,74
David: Exactly.

548
0:25:47,22 --> 0:25:50,5
As usual, the backup is, you know,
I've gotten a lot of flack

549
0:25:51,36 --> 0:25:54,92
over time about using, we do a
lot of checksumming and we use

550
0:25:54,92 --> 0:25:58,34
pretty heavy duty checksums and
people are like oh those checksums

551
0:25:58,34 --> 0:26:1,24
are so expensive how can you stand
how expensive they are?

552
0:26:1,24 --> 0:26:4,0
They're not expensive once you
actually start pushing stuff.

553
0:26:4,0 --> 0:26:7,84
If you're writing to your local
SSD, then yeah, the checksum

554
0:26:7,84 --> 0:26:9,96
overhead is 5% in pgBackRest.

555
0:26:9,96 --> 0:26:10,76
I've measured it.

556
0:26:10,76 --> 0:26:12,02
We have tests for this.

557
0:26:12,26 --> 0:26:14,08
So yeah, it looks like a lot.

558
0:26:14,1 --> 0:26:18,54
But then you actually start pushing
stuff out to S3 and poof!

559
0:26:18,94 --> 0:26:19,9
It just disappears.

560
0:26:20,82 --> 0:26:22,32
You don't even see it anymore.

561
0:26:24,84 --> 0:26:28,68
The trick is getting this kind
of volume in a realistic environment.

562
0:26:28,94 --> 0:26:30,8
People don't store their backups
locally.

563
0:26:31,82 --> 0:26:33,28
At least they shouldn't be.

564
0:26:34,16 --> 0:26:35,66
That's the message we try to get.

565
0:26:35,66 --> 0:26:37,36
Nikolay: I should repeat the tests
with S3.

566
0:26:37,36 --> 0:26:39,86
I'm very curious about the level
of parallelization we can get.

567
0:26:39,86 --> 0:26:43,6
Of course, we need the machine
itself should be where we recover.

568
0:26:43,78 --> 0:26:45,28
It should have local NVMes.

569
0:26:45,54 --> 0:26:46,3
I'm very curious.

570
0:26:46,3 --> 0:26:47,08
David: We have a...

571
0:26:47,08 --> 0:26:47,78
Yeah, sorry.

572
0:26:47,78 --> 0:26:49,2
Let me just address that really
quickly.

573
0:26:49,2 --> 0:26:52,04
We do have a good, interesting
optimization coming, hopefully

574
0:26:52,04 --> 0:26:55,44
in the next release, if I can get
it reviewed, by July 10th,

575
0:26:55,44 --> 0:26:56,66
that's our feature freeze.

576
0:26:57,04 --> 0:27:1,82
It basically does, basically it's
adding prefetch for object

577
0:27:1,82 --> 0:27:5,74
stores and something I call I'm
sure other people do this but

578
0:27:5,74 --> 0:27:8,52
I haven't really figured out what
it should be called so I've

579
0:27:8,52 --> 0:27:11,04
been calling it readover hoping
someone will come up with a better

580
0:27:11,04 --> 0:27:11,52
name.

581
0:27:11,52 --> 0:27:14,8
But a lot of times we're reading
through say a bundle and we

582
0:27:14,8 --> 0:27:19,9
need n number of files or n-m number
of files and so we're basically

583
0:27:19,96 --> 0:27:22,86
starting a bunch of new reads so
we'll grab these 3 files and

584
0:27:22,86 --> 0:27:26,98
then there's a file we're skipping
so we'll start a new read

585
0:27:27,04 --> 0:27:29,9
grab 4 files and then we have to
skip 2 files so we'll start

586
0:27:29,9 --> 0:27:33,34
a new read and then do So what
we're doing is actually we have

587
0:27:33,34 --> 0:27:37,64
this thing now called read over
which is also under review and

588
0:27:37,64 --> 0:27:42,18
it will skip over a configurable
block of bytes.

589
0:27:42,34 --> 0:27:45,6
So if we're reading through the
file and suddenly And it's configured

590
0:27:45,6 --> 0:27:46,62
by default to 64K.

591
0:27:47,86 --> 0:27:51,02
So if we have this 1 small file
that's breaking up the read that

592
0:27:51,02 --> 0:27:55,34
we're doing, we'll just read it
and throw it away and then continue.

593
0:27:56,2 --> 0:28:0,04
And I picked 64 bytes because based
on my testing, that was,

594
0:28:0,36 --> 0:28:3,64
It's cheaper to read 64 bytes pretty
much in all scenarios, no

595
0:28:3,64 --> 0:28:6,68
matter what your storage type is,
whether it's reduced redundancy

596
0:28:6,82 --> 0:28:9,56
or call it the longer term storages
and stuff, I can't think

597
0:28:9,56 --> 0:28:10,12
right now.

598
0:28:10,12 --> 0:28:12,74
And so 64k is basically always
a win.

599
0:28:13,5 --> 0:28:18,22
Even if you're paying for egress,
which you will be in some cases,

600
0:28:18,34 --> 0:28:22,06
you're still going to pay less
if you just read the 64 bytes.

601
0:28:22,44 --> 0:28:25,3
And from a latency standpoint,
it's always a win.

602
0:28:25,52 --> 0:28:28,94
Because starting a new read is
just expensive in terms of time.

603
0:28:29,12 --> 0:28:31,56
And there's actually a cost to
starting a read as well.

604
0:28:31,56 --> 0:28:34,2
Once you've actually got a file
open you're reading and normal

605
0:28:34,2 --> 0:28:37,94
S3 storage You're not paying for
that egress But if you actually

606
0:28:37,94 --> 0:28:41,36
start a new request you do pay
for that new request Anyway, some

607
0:28:41,36 --> 0:28:44,02
interesting features that would
make the tests that you were

608
0:28:44,02 --> 0:28:48,26
talking just talking about Very
interesting to run Would be a

609
0:28:48,26 --> 0:28:50,64
very different scenario than pgBackRest
today.

610
0:28:51,1 --> 0:28:53,9
Nikolay: Longer storm to storage,
glacier, right?

611
0:28:55,3 --> 0:28:56,82
David: Yeah, no, I'm not thinking
of glacier.

612
0:28:56,82 --> 0:28:58,04
I'm thinking of the 1 in between.

613
0:28:58,04 --> 0:29:0,82
I actually set all my buckets to
transition to it automatically

614
0:29:0,88 --> 0:29:1,72
after 30 days.

615
0:29:1,72 --> 0:29:2,96
So just transition to that.

616
0:29:2,96 --> 0:29:6,4
You can still read it instantly,
but it costs.

617
0:29:6,5 --> 0:29:8,0
If you go and read it, it costs.

618
0:29:8,0 --> 0:29:11,2
And there was a, people have been
wanting to add this as a feature

619
0:29:11,2 --> 0:29:14,82
where you can go into the storage
class immediately for all backups

620
0:29:14,92 --> 0:29:15,6
and WAL.

621
0:29:16,16 --> 0:29:16,66
Nikolay: Yeah.

622
0:29:16,98 --> 0:29:18,04
Standard IA.

623
0:29:18,34 --> 0:29:19,2
David: IA, thank you.

624
0:29:19,2 --> 0:29:20,96
Infrequent access, that's the 1.

625
0:29:21,38 --> 0:29:24,96
And it's actually a bad idea to
write IA initially, because in

626
0:29:24,96 --> 0:29:27,32
a healthy repo, you're actually
reading the repo.

627
0:29:27,88 --> 0:29:30,96
And as soon as you read 1 backup
out of that repo, you've destroyed

628
0:29:30,96 --> 0:29:32,54
all the advantages of IA.

629
0:29:32,92 --> 0:29:33,76
It's all gone.

630
0:29:33,76 --> 0:29:37,16
The idea of IA is I've got an older
backup from 6 months ago

631
0:29:37,16 --> 0:29:41,1
that I'm holding for compliance
and now suddenly I need to recover

632
0:29:41,1 --> 0:29:41,6
it.

633
0:29:41,84 --> 0:29:42,34
Great.

634
0:29:42,54 --> 0:29:44,32
You do that once every 2 years.

635
0:29:44,64 --> 0:29:47,2
Now you're using IA correctly,
but if you're writing all your

636
0:29:47,2 --> 0:29:51,14
backups into IA initially, you'll
just end up spending more money.

637
0:29:51,14 --> 0:29:54,24
The only scenario where it works
out is where you write and never

638
0:29:54,24 --> 0:29:59,9
ever do any recovery or any WAL
recovery ever, which is just

639
0:29:59,9 --> 0:30:2,06
not a use case to me.

640
0:30:2,32 --> 0:30:3,9
Back to these

641
0:30:3,9 --> 0:30:7,5
Nikolay: cases you are familiar with,
I'm very curious how much

642
0:30:7,5 --> 0:30:13,58
of all data was written per 2nd
on those cases where it was like

643
0:30:13,58 --> 0:30:17,78
the problem like they cannot afford
replicas so if near 0 lag,

644
0:30:18,28 --> 0:30:18,78
right?

645
0:30:19,06 --> 0:30:19,56
David: Yeah.

646
0:30:20,22 --> 0:30:22,76
Nikolay: Some hundreds of megabytes
per 2nd, I suspect.

647
0:30:23,6 --> 0:30:24,64
David: It can be.

648
0:30:25,24 --> 0:30:26,82
It really depends on what you're
doing.

649
0:30:26,82 --> 0:30:30,66
Let's say you are, the simplest
possible case is you're replaying

650
0:30:31,72 --> 0:30:34,54
sequential writes on an unindexed
table.

651
0:30:35,66 --> 0:30:35,9
Right?

652
0:30:35,9 --> 0:30:37,62
That's your simplest possible case.

653
0:30:37,9 --> 0:30:40,68
That'll run like gangbusters and
that'll definitely go into hundreds

654
0:30:40,68 --> 0:30:42,08
of megabytes or gigabytes.

655
0:30:42,78 --> 0:30:45,76
Maybe not per 2nd, but hundreds
of megabytes per 2nd I think

656
0:30:45,76 --> 0:30:47,2
becomes pretty reasonable at that
point.

657
0:30:47,2 --> 0:30:50,08
There's still some, still quite
a bit of exchange to get the

658
0:30:50,08 --> 0:30:53,56
WAL segments, move them, rename
them, do this kind of stuff,

659
0:30:53,56 --> 0:30:54,78
but it's quite fast.

660
0:30:55,6 --> 0:30:59,56
But as soon as you start doing
interesting things like say updating

661
0:30:59,56 --> 0:31:3,76
indexes, then your write volume
is going to go down significantly.

662
0:31:4,3 --> 0:31:6,3
Then it's all going to be I/O latency.

663
0:31:7,28 --> 0:31:12,98
So the CPU will drop below 100,
you'll go into I/O wait, and you're

664
0:31:12,98 --> 0:31:14,24
going to be looking at latency.

665
0:31:14,24 --> 0:31:18,28
So if your CPU actually is 100,
you're probably doing really

666
0:31:18,28 --> 0:31:20,46
simple writes, generally speaking.

667
0:31:20,46 --> 0:31:22,84
And you can maintain a really high
throughput if you're doing

668
0:31:22,84 --> 0:31:23,13
that.

669
0:31:23,13 --> 0:31:26,64
If your CPU drops, if you've got
complicated index writes and

670
0:31:26,64 --> 0:31:29,48
stuff like that, your CPU will
generally not be 100 anymore.

671
0:31:30,2 --> 0:31:32,06
And you'll be into I/O wait instead.

672
0:31:32,5 --> 0:31:35,94
So I played around with this in
a bunch of different scenarios,

673
0:31:36,76 --> 0:31:40,78
trying to figure out ways to help
Postgres and maybe gain Postgres.

674
0:31:41,68 --> 0:31:43,52
Certain scenarios are just slow.

675
0:31:44,18 --> 0:31:47,04
Nikolay: It makes total sense because
if it's like simple, it's

676
0:31:47,04 --> 0:31:51,42
I/O bound and if on primary, we
have hundreds of cores doing this

677
0:31:51,42 --> 0:31:57,18
work, but on the back we have only
1 trying to catch up to do

678
0:31:57,18 --> 0:31:58,72
the same work, basically, logically.

679
0:31:59,5 --> 0:32:0,26
It's terrible.

680
0:32:2,02 --> 0:32:3,26
Yeah, that makes sense.

681
0:32:3,26 --> 0:32:3,84
Simple work.

682
0:32:3,84 --> 0:32:4,34
Yeah.

683
0:32:5,08 --> 0:32:9,0
David: And unfortunately, all
the cool work that Andres et

684
0:32:9,0 --> 0:32:13,34
al have been doing on async I/O,
Tomas, and et cetera, it's

685
0:32:13,34 --> 0:32:17,18
not going to help that scenario
very much in on the standby.

686
0:32:17,44 --> 0:32:21,1
It's great on the primary, but
it just makes a situation where

687
0:32:21,1 --> 0:32:24,86
the primary can write even more
data than the standby can ever

688
0:32:24,86 --> 0:32:25,92
possibly keep up with.

689
0:32:25,92 --> 0:32:28,78
Again, because single threaded,
so those async operations are

690
0:32:28,78 --> 0:32:30,72
only going to buy you so much.

691
0:32:30,72 --> 0:32:32,88
Direct I/O and stuff obviously
will help.

692
0:32:33,04 --> 0:32:37,44
So there are improvements being
made in the single threaded recovery

693
0:32:38,1 --> 0:32:40,44
So don't make no mistake I'm not
saying there aren't because

694
0:32:40,44 --> 0:32:43,26
there definitely are but they're
not keeping pace with the improvements

695
0:32:43,26 --> 0:32:47,5
that are being made on the primary
side Which are can be an order

696
0:32:47,5 --> 0:32:49,84
of magnitude greater than what's
happening on the standby.

697
0:32:50,74 --> 0:32:51,48
Nikolay: Yeah, exactly.

698
0:32:51,48 --> 0:32:56,54
I used, I think, Intel machines
in this experiment where I achieved

699
0:32:56,64 --> 0:33:0,8
30-something terabytes per hour
with pgBackRest, and I compared

700
0:33:0,8 --> 0:33:5,66
it to pg_basebackup, and I expected
200-300 megabytes per 2nd

701
0:33:5,66 --> 0:33:6,34
as usual.

702
0:33:8,1 --> 0:33:11,64
But since it was Postgres 18, I
saw 1 gigabyte per 2nd.

703
0:33:11,64 --> 0:33:12,84
I was super surprised.

704
0:33:13,52 --> 0:33:14,7
David: It's actually gotten better.

705
0:33:14,76 --> 0:33:15,72
I remember the article.

706
0:33:15,72 --> 0:33:15,9099
Yeah.

707
0:33:15,9099 --> 0:33:19,08
And that's jibes with what I've
been seeing in terms of overall

708
0:33:19,08 --> 0:33:21,6
I/O improvement, I/O improvements
in Postgres.

709
0:33:21,78 --> 0:33:26,2
So they're impacting everything,
make no mistake, but it's obviously

710
0:33:26,2 --> 0:33:28,22
that recovery bottleneck is still
there.

711
0:33:28,48 --> 0:33:31,08
And there's not much that pgBackRest
can do about that, except

712
0:33:31,08 --> 0:33:33,86
just make sure that Postgres always
has the WAL.

713
0:33:34,34 --> 0:33:37,04
As soon as Postgres requests a
WAL segment, we've got it right

714
0:33:37,04 --> 0:33:38,18
there waiting for it.

715
0:33:38,36 --> 0:33:41,16
And we got recommendations on how
to set up your storage so that

716
0:33:41,16 --> 0:33:46,56
we can do a move into the pg_wal
directory so it's as fast as

717
0:33:46,56 --> 0:33:47,44
it can possibly be.

718
0:33:47,44 --> 0:33:48,66
We're not doing a copy.

719
0:33:48,94 --> 0:33:51,46
We don't fsync things, we don't
et cetera, et cetera.

720
0:33:51,46 --> 0:33:54,86
So we're trying to run everything
as fast as possible, but there

721
0:33:54,86 --> 0:33:55,5
are limits.

722
0:33:57,44 --> 0:33:59,94
Nikolay: This makes total sense and
it's great insight.

723
0:33:59,96 --> 0:34:4,54
Like improving things around, you
might highlight some difficult

724
0:34:4,54 --> 0:34:7,74
problems and they might become
even more acute, right?

725
0:34:8,4 --> 0:34:11,18
So Everything improved so you have
better performance.

726
0:34:11,8 --> 0:34:15,42
You much faster you can achieve
the problem with lagging replication

727
0:34:16,4 --> 0:34:16,9
Yeah,

728
0:34:17,02 --> 0:34:20,74
David: it was definitely a time
before pgBackRest and

729
0:34:20,74 --> 0:34:25,68
other parallel WAL implementations
that were getting the WAL

730
0:34:25,68 --> 0:34:28,78
was the bottleneck and Postgres
kind of merrily sailed along

731
0:34:28,78 --> 0:34:31,36
and did its thing and it was just
WAL that was the issue.

732
0:34:31,56 --> 0:34:33,14
Now it's usually Postgres.

733
0:34:34,06 --> 0:34:38,3
Another place this expresses is,
if you're doing, like people

734
0:34:38,3 --> 0:34:40,6
have noticed and started pointing
out to me, although I already

735
0:34:40,6 --> 0:34:44,48
knew it, is that if you're doing
log shipping, you can do recovery

736
0:34:44,48 --> 0:34:46,94
much faster than if you are doing
replication.

737
0:34:48,46 --> 0:34:51,8
Because log shipping compressed
segments, which are prefetched

738
0:34:52,72 --> 0:34:56,08
and decompressed asynchronously
and then moved into the pg_wal

739
0:34:56,08 --> 0:34:59,76
directory is quite a bit faster
than streaming uncompressed WAL

740
0:34:59,76 --> 0:35:3,92
segments over the network from
the primary basically every day

741
0:35:3,92 --> 0:35:5,62
of the week if you have high volume.

742
0:35:6,66 --> 0:35:9,78
So that really brings into question
the idea of I brought up

743
0:35:9,78 --> 0:35:11,78
a standby and I want to do recovery.

744
0:35:12,28 --> 0:35:17,36
How do I do the best things where
I'm going to go and log ship

745
0:35:17,36 --> 0:35:21,02
as long as I can because that's
faster and also puts less load

746
0:35:21,02 --> 0:35:26,04
on the primary and then only switch
over to getting data from

747
0:35:26,04 --> 0:35:27,54
the primary when I have to.

748
0:35:28,08 --> 0:35:30,46
At 1 point you switch back and
stuff like that.

749
0:35:30,46 --> 0:35:34,24
Postgres doesn't currently manage
that very well and the rules

750
0:35:34,24 --> 0:35:35,54
are a little bit arcane.

751
0:35:36,54 --> 0:35:39,5
So some archivers, we haven't really
done this in pgBackRest

752
0:35:39,62 --> 0:35:42,18
yet, but there's certain things
you can do to basically, you

753
0:35:42,18 --> 0:35:46,5
can artificially give an error
to Postgres to make it switch

754
0:35:46,5 --> 0:35:48,7
back over to replication.

755
0:35:48,7 --> 0:35:49,6
Do you get what I'm saying?

756
0:35:49,6 --> 0:35:51,84
So let's say we get to pretty much
the end.

757
0:35:52,28 --> 0:35:53,6
Nikolay: You're moving very fast.

758
0:35:53,64 --> 0:35:57,5
So you're talking about replica
where both restore_command and

759
0:35:57,5 --> 0:36:0,04
primary_conninfo are configured,
right?

760
0:36:0,1 --> 0:36:0,6
Yes.

761
0:36:1,02 --> 0:36:5,04
And then you say Postgres has a
very strict precedence rule.

762
0:36:5,22 --> 0:36:8,68
I think it searches local pg_wal
directory 1st, then it uses

763
0:36:8,68 --> 0:36:11,68
streaming replication and only
then it uses restore_command,

764
0:36:11,68 --> 0:36:12,68
if I'm not mistaken.

765
0:36:12,88 --> 0:36:14,44
David: I believe that's the way
it works, yeah.

766
0:36:14,44 --> 0:36:17,4
Nikolay: And then you say you have
a way to fool Postgres to achieve

767
0:36:17,4 --> 0:36:22,58
what you want, and what you want
is to get WALs from archive

768
0:36:22,58 --> 0:36:26,48
because they are compressed and
you can parallelize it at restore

769
0:36:26,48 --> 0:36:30,28
command level and only switch to
primary_conninfo when to streaming

770
0:36:30,28 --> 0:36:31,92
replication in some cases.

771
0:36:32,3 --> 0:36:32,78
Right?

772
0:36:32,78 --> 0:36:36,54
And then you have to, with some
error, this I don't get already.

773
0:36:36,98 --> 0:36:39,76
I can imagine throwing an error
from a restore_command, I cannot

774
0:36:39,76 --> 0:36:42,34
imagine anything at primary_conninfo,
that's a problem.

775
0:36:42,36 --> 0:36:45,44
David: The situation is, let's
say you are, you've been doing

776
0:36:45,44 --> 0:36:47,96
log shipping because that's the
fastest thing to do, so you're

777
0:36:47,96 --> 0:36:49,06
getting WAL segments.

778
0:36:49,08 --> 0:36:52,28
Now, Postgres, if it's basically
decided on log shipping, isn't

779
0:36:52,28 --> 0:36:54,78
going to switch away from that
until something happens.

780
0:36:55,58 --> 0:37:0,48
The idea is, like the actual, the
WAL get routine, when it knows

781
0:37:0,48 --> 0:37:2,22
it's at the end of what it has.

782
0:37:2,48 --> 0:37:5,28
So let's say Postgres, so you get
to the end and now the next

783
0:37:5,28 --> 0:37:6,98
WAL segment arrives 2 minutes
later.

784
0:37:6,98 --> 0:37:10,2
Now you've essentially built in
2 minutes of lag because Postgres

785
0:37:10,2 --> 0:37:13,26
will patiently wait 2 minutes for
that next WAL segment to arrive

786
0:37:13,26 --> 0:37:14,7
as long as it doesn't get an error.

787
0:37:15,84 --> 0:37:19,74
So what you, in this situation,
what you can do is actually inject

788
0:37:19,74 --> 0:37:24,38
an error into the archive-get to
say, I know I'm at the end of

789
0:37:24,38 --> 0:37:26,9
the archive stream in the repo
because I can see it.

790
0:37:26,96 --> 0:37:30,56
I'll throw an error to force Postgres
back to replication.

791
0:37:31,12 --> 0:37:32,62
Nikolay: Yeah, streaming replication.

792
0:37:32,98 --> 0:37:34,44
I think I got it wrong.

793
0:37:34,44 --> 0:37:37,7
I think restore_command has precedence
over streaming replication.

794
0:37:38,18 --> 0:37:39,14
David: I think it has.

795
0:37:39,14 --> 0:37:40,32
Nikolay: I always forget that.

796
0:37:40,32 --> 0:37:44,68
But if the recipe you described
implies that.

797
0:37:45,24 --> 0:37:46,8
But it's an interesting recipe.

798
0:37:47,06 --> 0:37:49,24
David: And a lot of times, even
if the streaming replication

799
0:37:49,3 --> 0:37:53,48
were prioritized, let's say you
recover, restored a backup and

800
0:37:53,48 --> 0:37:56,7
you need to get WAL from 3 days
ago to start doing recovery.

801
0:37:57,04 --> 0:37:59,54
There's a very good chance that's
not going to be on the primary.

802
0:38:0,2 --> 0:38:2,96
So you're going to, you're going
to go over to log shipping at

803
0:38:2,96 --> 0:38:3,92
that point anyway.

804
0:38:4,64 --> 0:38:7,08
But the other thing you would want
to do is, let's say you start

805
0:38:7,08 --> 0:38:10,68
lagging too badly on the primary,
ideally Postgres would go back

806
0:38:10,68 --> 0:38:13,44
to log shipping because it's quite
a lot faster.

807
0:38:14,18 --> 0:38:16,88
Historically that wasn't true,
but with modern archivers, like

808
0:38:16,88 --> 0:38:21,46
something you'd find in pgBackRest
or WAL-G, we can do that

809
0:38:21,46 --> 0:38:22,7
a lot faster than replication.

810
0:38:22,72 --> 0:38:25,24
So in theory, if you're lagging
behind enough on replication,

811
0:38:25,24 --> 0:38:29,66
you switch back to log shipping,
catch up, then go back to replication

812
0:38:30,06 --> 0:38:32,52
in order to have as little lag
as possible.

813
0:38:32,98 --> 0:38:34,62
So the question is, how do you
coordinate?

814
0:38:34,82 --> 0:38:38,2
On the pgBackRest side, we can
game it to force pgBackRest

815
0:38:38,2 --> 0:38:42,12
to switch to streaming rep, but
we can't really get it to go

816
0:38:42,12 --> 0:38:44,94
back to log shipping.

817
0:38:45,06 --> 0:38:47,22
Now, these are discussions that
have been had.

818
0:38:47,22 --> 0:38:48,3
I don't know if we've...

819
0:38:48,42 --> 0:38:50,54
I can't remember if we've had them
on hackers, but it's something

820
0:38:50,54 --> 0:38:52,06
we've certainly talked about at
conferences.

821
0:38:52,06 --> 0:38:53,5
Nikolay: Yeah, that's an interesting
idea.

822
0:38:53,8 --> 0:38:57,16
It would be great to have more
control here and more flexible

823
0:38:57,66 --> 0:38:58,16
configuration.

824
0:38:58,5 --> 0:38:59,36
David: Some kind of...

825
0:39:0,02 --> 0:39:0,52
Nikolay: Sorry.

826
0:39:0,54 --> 0:39:1,1423
Go on.

827
0:39:1,1423 --> 0:39:2,68
We have huge delay because we are
on very different parts of

828
0:39:2,68 --> 0:39:3,58
the world, right?

829
0:39:3,84 --> 0:39:5,74
3 very different parts of the world.

830
0:39:5,74 --> 0:39:8,36
David: We are very geographically
distributed, that's for sure.

831
0:39:8,36 --> 0:39:12,16
So yeah, so on the 1 side, the
archiver would have to tell Postgres

832
0:39:12,16 --> 0:39:14,16
that it's caught up.

833
0:39:14,34 --> 0:39:18,28
So rather than throwing a nasty
error, we could say, we could

834
0:39:18,28 --> 0:39:22,3
have a code to return value 2 or
something to say, hey, I'm at

835
0:39:22,3 --> 0:39:25,38
the end of my, of the WAL in the
repository.

836
0:39:26,26 --> 0:39:28,2
Nikolay: Or just signal the lag.

837
0:39:28,26 --> 0:39:29,06
Signal the lag.

838
0:39:29,06 --> 0:39:30,36
This is my lag, That's it.

839
0:39:30,36 --> 0:39:34,54
And then the mechanism will decide
which precedence, adjust precedence

840
0:39:34,54 --> 0:39:36,3
based on knowledge about lags.

841
0:39:37,06 --> 0:39:37,7
David: Or that.

842
0:39:38,4 --> 0:39:40,52
And honestly, if you're going to
do that, the way to do that

843
0:39:40,52 --> 0:39:43,28
would be inside the archive_library
interface.

844
0:39:43,94 --> 0:39:44,14
Nikolay: And

845
0:39:44,14 --> 0:39:46,06
David: you're familiar with the
archive_library, right?

846
0:39:46,24 --> 0:39:46,72
Yes.

847
0:39:46,72 --> 0:39:50,22
We don't have 1 for pgBackRest
yet because it's not needed for

848
0:39:50,22 --> 0:39:53,76
performance in pgBackRest and
it doesn't really provide any

849
0:39:53,76 --> 0:39:54,44
other benefits.

850
0:39:55,34 --> 0:39:58,74
So we haven't really done it yet,
although now that several versions

851
0:39:58,74 --> 0:40:2,32
of Postgres support it, I'm looking
at probably adding that this

852
0:40:2,32 --> 0:40:3,94
year or next year at the latest.

853
0:40:4,44 --> 0:40:7,2
Because then maybe we could push
some of these ideas into Postgres

854
0:40:7,2 --> 0:40:10,74
to saying, hey, let's flip back
and forth between archive-get

855
0:40:11,0 --> 0:40:13,74
and streaming replication based
on what's...

856
0:40:13,74 --> 0:40:16,72
And passing stats back and forth
inside the archive_library would

857
0:40:16,72 --> 0:40:17,46
be easy.

858
0:40:18,26 --> 0:40:21,9
You don't need any fancy return
values or JSON or something.

859
0:40:21,9 --> 0:40:25,16
It would just be an API call that
you would use within Postgres.

860
0:40:26,46 --> 0:40:28,58
Nikolay: Postgres 15, when it was
introduced.

861
0:40:29,54 --> 0:40:33,28
I wanted to clarify very clearly
that we talk about fetching

862
0:40:33,28 --> 0:40:37,08
WALs and streaming, but after
that Postgres needs to replay

863
0:40:37,08 --> 0:40:40,36
and then we have this startup process,
100% CPU.

864
0:40:40,9 --> 0:40:43,68
This cannot be solved anyhow, right?

865
0:40:43,68 --> 0:40:45,56
So if you manage to...

866
0:40:45,78 --> 0:40:47,14
Usually problem is there.

867
0:40:47,18 --> 0:40:48,64
This is what I'm trying to say.

868
0:40:48,64 --> 0:40:53,36
It depends on the system, I think,
but even with a simple, regular

869
0:40:53,8 --> 0:40:57,98
streaming replication, which is
no compression, nothing, I usually

870
0:40:57,98 --> 0:41:2,22
don't see any problems with receiving
WAL and just putting it

871
0:41:2,22 --> 0:41:4,7
to file with WAL receiver, right?

872
0:41:4,86 --> 0:41:7,4
The problem is usually on the startup
process.

873
0:41:7,42 --> 0:41:9,66
This is my observations on multiple
systems.

874
0:41:9,8 --> 0:41:10,38
David: And that's true.

875
0:41:10,38 --> 0:41:11,76
A lot of it depends on your network.

876
0:41:11,76 --> 0:41:14,8
If you're running things in multiple
availability zones and stuff

877
0:41:14,8 --> 0:41:18,42
that, or across cloud providers
or whatever, that situation can

878
0:41:18,42 --> 0:41:19,26
definitely change.

879
0:41:19,54 --> 0:41:20,02
Nikolay: Yeah.

880
0:41:20,02 --> 0:41:21,42
But certainly if you've got,

881
0:41:22,06 --> 0:41:24,72
David: if you've got primary
and standby locally, but the network

882
0:41:24,72 --> 0:41:26,64
speeds these days, you're right.

883
0:41:26,88 --> 0:41:29,64
Streaming replication should not
be a problem in that environment.

884
0:41:30,18 --> 0:41:31,42
In this day and age.

885
0:41:31,78 --> 0:41:34,44
Nikolay: My observations are with
good network and single region.

886
0:41:34,44 --> 0:41:38,0
I agree with you if you introduce
network complexity, yes.

887
0:41:39,62 --> 0:41:43,56
Michael: That sounds all very clever
and like obviously needed

888
0:41:43,58 --> 0:41:44,88
at super high scale.

889
0:41:45,06 --> 0:41:48,96
We've talked to a lot of people
and it seems like there's quite

890
0:41:48,96 --> 0:41:52,26
a few projects in the space of
sharding these days and it feels

891
0:41:52,26 --> 0:41:56,08
like people that go the sharded
route sidestep this problem because

892
0:41:56,6 --> 0:42:0,04
each of the shards just has lower
volume, like by definition.

893
0:42:0,72 --> 0:42:3,94
So I'm wondering if it's just like
an orthogonal thing to

894
0:42:3,94 --> 0:42:7,8
pgBackRest, like sharding can happen
and each 1 can be backed up.

895
0:42:7,8 --> 0:42:10,4
Does it affect you in any way or
are you thinking along those

896
0:42:10,4 --> 0:42:11,34
lines at all?

897
0:42:11,58 --> 0:42:14,34
David: From a pgBackRest standpoint,
no, because it's really

898
0:42:14,38 --> 0:42:15,12
from our...

899
0:42:16,56 --> 0:42:21,18
So the things in this arena that
I know support pgBackRest are

900
0:42:21,18 --> 0:42:24,22
like Greenplum and Multigres.

901
0:42:25,24 --> 0:42:28,2
And in both those situations, it's
a very straightforward

902
0:42:28,2 --> 0:42:29,04
pgBackRest scenario.

903
0:42:29,76 --> 0:42:33,34
We're just doing recovery, they
give us a point in time to recover

904
0:42:33,34 --> 0:42:36,1
to, and that's up for them to figure
out how to get the shards

905
0:42:36,1 --> 0:42:39,0
back in sync, and then we just
give them what they ask for.

906
0:42:39,28 --> 0:42:42,84
Multigres looks enough like Postgres
that pgBackRest just is

907
0:42:42,84 --> 0:42:43,76
cool with it.

908
0:42:44,28 --> 0:42:48,1
On the Greenplum side, there's
a company that has created a fork

909
0:42:48,9 --> 0:42:51,44
that specifically supports Greenplum
that'll read its control

910
0:42:51,44 --> 0:42:56,04
files and do that and we've never
brought that into core mostly

911
0:42:56,04 --> 0:42:59,68
because I don't want to get into
Greenplum and testing Greenplum

912
0:43:0,4 --> 0:43:3,42
and also annoyingly the oldest
version of Greenplum is still

913
0:43:3,42 --> 0:43:7,26
based on Postgres 9.4, which we
expired a couple of years ago.

914
0:43:7,66 --> 0:43:9,66
So it's a little bit painful to
have.

915
0:43:11,04 --> 0:43:14,84
So that company is actually basically
supporting now, not only

916
0:43:14,84 --> 0:43:17,2
do they have their fork, but they're
also supporting versions

917
0:43:17,2 --> 0:43:20,28
of Postgres that we don't support
anymore so that they can still

918
0:43:20,28 --> 0:43:23,3
support Greenplum 6, which is based
on 9.4.

919
0:43:24,22 --> 0:43:26,6
And maybe that's something we could
bring into Core Postgres,

920
0:43:26,6 --> 0:43:28,92
but it's a resource issue.

921
0:43:29,54 --> 0:43:32,08
We just simply don't have the resources
for it and it's just

922
0:43:32,08 --> 0:43:35,14
easier to focus on open source
Postgres.

923
0:43:36,1 --> 0:43:38,36
When you drive people towards the
idea that if you make your

924
0:43:38,36 --> 0:43:41,68
stuff look like Postgres it'll
just work, if you muck about with

925
0:43:41,68 --> 0:43:43,2
pg_control, like the parts of...

926
0:43:43,2 --> 0:43:46,56
So we even have a feature where
you can add stuff to pg_control

927
0:43:46,56 --> 0:43:48,3
as long as you add it to the end.

928
0:43:48,86 --> 0:43:52,16
And we'll automatically figure
out, so you tell us your Postgres

929
0:43:52,16 --> 0:43:55,46
16, we'll automatically figure
out where the checksum lives,

930
0:43:55,84 --> 0:43:58,44
based on the size of your pg_control
and all that kind of stuff,

931
0:43:58,44 --> 0:43:59,62
and work all that out.

932
0:43:59,72 --> 0:44:2,72
But The contract is that the part
of pg_control that was part

933
0:44:2,72 --> 0:44:5,22
of that version of Postgres must
be the same.

934
0:44:5,66 --> 0:44:7,52
Because we read all over pg_control.

935
0:44:7,9 --> 0:44:10,9
It's extremely important for our
operations, so we can't just,

936
0:44:10,9 --> 0:44:13,7
you can't move things around, you're
just allowed to add to the

937
0:44:13,7 --> 0:44:14,2
end.

938
0:44:14,44 --> 0:44:16,56
So we do make concessions to forks.

939
0:44:16,56 --> 0:44:19,94
Things like it works with EDB,
their enterprise server, and other

940
0:44:19,94 --> 0:44:23,1
things that have new control versions
and stuff like that.

941
0:44:23,56 --> 0:44:25,38
As long as you kind of know what
you're doing.

942
0:44:25,38 --> 0:44:26,88
We call it a maintainer feature.

943
0:44:26,88 --> 0:44:29,9
So the idea is if you shouldn't
be mucking around with this feature,

944
0:44:29,9 --> 0:44:32,94
if you are not a maintainer of
a fork.

945
0:44:33,64 --> 0:44:38,8
And deploying pgBackRest for
that particular fork or documenting

946
0:44:38,86 --> 0:44:41,66
how pgBackRest should be used
with that fork or something.

947
0:44:42,08 --> 0:44:45,06
If you have to use these options
on a regular cluster, then you've

948
0:44:45,06 --> 0:44:46,58
probably done something wrong.

949
0:44:47,64 --> 0:44:48,84
You shouldn't be doing that.

950
0:44:48,84 --> 0:44:51,84
But obviously I can't control what
people do.

951
0:44:52,9 --> 0:44:54,44
That's part of the fun of it, right?

952
0:44:54,86 --> 0:44:55,36
Michael: Yeah.

953
0:44:56,14 --> 0:44:59,84
Talking of maintenance, is it a
good time to transition into

954
0:45:0,3 --> 0:45:2,86
how you have been maintaining it
over the years, kind of what

955
0:45:2,86 --> 0:45:5,52
the situation was, and it's changed
recently.

956
0:45:5,58 --> 0:45:7,86
So I'm wondering if you wanted
to share a bit of that.

957
0:45:8,24 --> 0:45:12,9
David: Yeah, maintenance has
always been a big part of pgBackRest.

958
0:45:13,04 --> 0:45:16,5
1 thing is we have a very comprehensive
test suite for a project

959
0:45:16,5 --> 0:45:17,04
of this size.

960
0:45:17,04 --> 0:45:18,58
We test on 5 different architectures.

961
0:45:18,8 --> 0:45:20,68
We test on a variety of distros.

962
0:45:20,68 --> 0:45:23,0
We have 100% unit test coverage.

963
0:45:23,0 --> 0:45:24,28
We have integration tests.

964
0:45:24,28 --> 0:45:28,6
We have, we test our doc code,
our test code, or we test everything.

965
0:45:28,84 --> 0:45:32,08
And just keeping that test suite
up to date is a bit of a challenge.

966
0:45:32,08 --> 0:45:33,08
It's not the core code.

967
0:45:33,08 --> 0:45:36,18
The core code doesn't really, they
don't, the C code doesn't

968
0:45:36,18 --> 0:45:38,44
break because rel 10 comes out.

969
0:45:39,16 --> 0:45:42,12
What breaks is our tests and the
documentation and all these

970
0:45:42,12 --> 0:45:42,72
sorts of things.

971
0:45:42,72 --> 0:45:46,5
So it's a non-trivial task just
to keep all that going.

972
0:45:46,56 --> 0:45:49,62
And then you have things like say,
Cirrus CI just going away

973
0:45:49,64 --> 0:45:50,78
on a month's notice.

974
0:45:51,5 --> 0:45:53,76
So we need to migrate off of Cirrus
CI.

975
0:45:54,06 --> 0:45:56,4
Luckily, we're already on GitHub
Actions, so we're able to do

976
0:45:56,4 --> 0:45:57,48
that fairly easily.

977
0:45:57,88 --> 0:46:2,54
When I say maintenance, I also
include bug fixes that come in,

978
0:46:2,54 --> 0:46:5,96
So we get a pretty regular stream
of bug reports.

979
0:46:5,98 --> 0:46:7,7
Some are bugs or some are not.

980
0:46:8,2 --> 0:46:11,76
Most of them these days tend to
be pretty weird edge cases.

981
0:46:12,52 --> 0:46:15,76
Enough people are using pgBackRest
that the Really obvious

982
0:46:15,76 --> 0:46:17,22
bugs get caught pretty quickly.

983
0:46:17,22 --> 0:46:20,1
If we release something and there's
bugs in it, I get reports

984
0:46:20,9 --> 0:46:22,18
the next day, almost.

985
0:46:22,64 --> 0:46:26,02
Bug fixes or maintenance, the documentation
is maintenance, interacting

986
0:46:26,06 --> 0:46:30,2
with the community and answering
questions and feature requests

987
0:46:30,24 --> 0:46:32,12
and other things like that's all
maintenance.

988
0:46:32,98 --> 0:46:35,1
So that's actually a pretty big
part of the job.

989
0:46:35,74 --> 0:46:39,48
And after I left, after I did not
transition to Snowflake, so

990
0:46:39,48 --> 0:46:42,26
Snowflake bought Crunchy Data, I did
not transition.

991
0:46:42,72 --> 0:46:46,48
I wasn't happy with their terms,
so I decided just to take some

992
0:46:46,48 --> 0:46:47,16
time and travel.

993
0:46:47,16 --> 0:46:49,84
I kept maintaining pgBackRest,
but at that point I was really

994
0:46:49,84 --> 0:46:51,06
just maintaining it.

995
0:46:51,78 --> 0:46:54,62
Which is fine, but it really got
to the point of thinking about

996
0:46:55,08 --> 0:46:57,54
really restarting development on
pgBackRest again.

997
0:46:58,38 --> 0:47:0,46
I was like, I think I might need
a job.

998
0:47:3,06 --> 0:47:5,84
And so I started trying to get,
I actually put sponsorship links

999
0:47:5,84 --> 0:47:9,2
and I've been including sponsorship
with every release since

1000
0:47:9,2 --> 0:47:9,9
last summer.

1001
0:47:10,56 --> 0:47:13,68
It was around for basically a whole
year, not getting a lot of

1002
0:47:13,68 --> 0:47:14,54
traction of course.

1003
0:47:14,54 --> 0:47:16,4
Then I started spending more time
on it.

1004
0:47:17,04 --> 0:47:20,72
Then I got to the point where I
realized maybe this whole sponsorship

1005
0:47:20,74 --> 0:47:21,82
thing isn't gonna work out.

1006
0:47:21,82 --> 0:47:26,36
And that got us to the recent crisis
where I announced that I

1007
0:47:26,36 --> 0:47:28,08
wasn't gonna work on pgBackRest
anymore.

1008
0:47:28,08 --> 0:47:30,26
And suddenly I got sponsorship.

1009
0:47:30,72 --> 0:47:32,8
And when I say suddenly, it really
was suddenly.

1010
0:47:32,8 --> 0:47:35,8
Within 3 weeks, pretty much everything
was worked out, and I

1011
0:47:35,8 --> 0:47:38,76
was able to make an announcement
that the project was going to

1012
0:47:38,76 --> 0:47:40,52
continue to be maintained, and
etc.

1013
0:47:40,96 --> 0:47:42,18
But not just maintained.

1014
0:47:42,34 --> 0:47:44,32
The key is adding new features.

1015
0:47:44,88 --> 0:47:47,64
Right now, I've been doing, on
pgaudit and pgBackRest, I've been

1016
0:47:47,64 --> 0:47:50,5
doing a lot of cleanup work because
they've been sitting for

1017
0:47:50,5 --> 0:47:53,48
a while, especially pgaudit wasn't
really getting a whole lot

1018
0:47:53,48 --> 0:47:54,1
of love.

1019
0:47:54,44 --> 0:47:57,88
So I've been working on that, but
I'll be done with that at the

1020
0:47:57,88 --> 0:47:59,78
end of this week, basically.

1021
0:48:0,26 --> 0:48:2,9
And then it's time just to dive
into big new features again,

1022
0:48:2,9 --> 0:48:4,38
which I'm really excited about.

1023
0:48:4,62 --> 0:48:5,86
I've got a lot of ideas.

1024
0:48:5,94 --> 0:48:8,34
The sponsors have a lot of ideas,
as you can imagine.

1025
0:48:8,6 --> 0:48:8,8
Nikolay: The

1026
0:48:8,8 --> 0:48:12,34
David: great thing is the ideas
that the sponsors have are

1027
0:48:12,98 --> 0:48:17,5
directly aligned with big ticket
items that have been on our

1028
0:48:17,5 --> 0:48:18,46
list for a while.

1029
0:48:19,16 --> 0:48:20,9
They're not asking for outlandish
stuff.

1030
0:48:20,9 --> 0:48:24,74
They're asking for repo-to-repo
backup, for instance, is a big

1031
0:48:24,74 --> 0:48:25,24
ask.

1032
0:48:25,64 --> 0:48:28,14
And that's been on the list for
a long time, because that's a

1033
0:48:28,14 --> 0:48:30,26
pretty important thing to be able
to do.

1034
0:48:31,34 --> 0:48:33,48
Streaming, doing streaming while
replication.

1035
0:48:34,84 --> 0:48:39,62
So you can have RPO 0 if you want,
without having a standby,

1036
0:48:39,62 --> 0:48:42,88
so the idea is you have RPO 0 without
a standby.

1037
0:48:43,26 --> 0:48:44,72
Nikolay: pg_receivewal, right?

1038
0:48:45,36 --> 0:48:47,8
David: Yeah, although we would
work directly at the protocol

1039
0:48:47,8 --> 0:48:48,3
level.

1040
0:48:48,46 --> 0:48:50,16
And in fact, we might...

1041
0:48:50,16 --> 0:48:52,68
I'm actually thinking, given that
we have this archive_library

1042
0:48:52,68 --> 0:48:56,34
thing, and given that pgBackRest
can just live in Postgres,

1043
0:48:56,58 --> 0:48:59,68
I'm not even sure if we need to
deal with the protocol level.

1044
0:49:0,18 --> 0:49:3,8
We could just sit there as a companion
to Postgres and stream

1045
0:49:3,8 --> 0:49:4,56
WAL out.

1046
0:49:4,9 --> 0:49:7,48
Nikolay: Isn't it expensive to sit
as Postgres?

1047
0:49:7,48 --> 0:49:9,64
Like, I don't get it.

1048
0:49:10,92 --> 0:49:12,34
What exactly is the idea?

1049
0:49:13,08 --> 0:49:16,82
David: So basically, instead
of a replication slot, we would

1050
0:49:16,82 --> 0:49:20,42
just be keeping up with the current
write WAL location.

1051
0:49:21,0 --> 0:49:22,66
Where are we currently writing
in WAL?

1052
0:49:22,66 --> 0:49:24,78
That's the point that we should
currently be tarping.

1053
0:49:26,0 --> 0:49:28,68
And although we would still need
a replica, if we really wanted

1054
0:49:28,68 --> 0:49:32,9
to say a synchronous replication,
RPO 0 stuff, we'd probably

1055
0:49:32,9 --> 0:49:34,9
still need a replication slot for
that.

1056
0:49:35,28 --> 0:49:38,44
The other thought we had is that
we could actually game, we could

1057
0:49:38,44 --> 0:49:44,02
create a replication slot but not
actually use it and just update

1058
0:49:44,02 --> 0:49:44,7
the row.

1059
0:49:45,6 --> 0:49:49,9
So basically do our own background
archiving and then update

1060
0:49:49,9 --> 0:49:55,86
the row in Postgres to our current
position, what we've actually

1061
0:49:55,86 --> 0:49:57,08
pushed out to storage.

1062
0:49:58,86 --> 0:50:0,62
And if I do this, I want to do
it right.

1063
0:50:0,62 --> 0:50:3,4
In a lot of cases, the things that
are doing the WAL receiver,

1064
0:50:3,4 --> 0:50:5,86
they're not actually, let's say
your eventual destination is

1065
0:50:5,86 --> 0:50:8,34
S3, they're not actually writing
packets off to S3.

1066
0:50:8,94 --> 0:50:11,88
They're storing the stuff locally
on whatever host the WAL receiver

1067
0:50:11,88 --> 0:50:14,18
is running on and then when they
get a full WAL segment, they're

1068
0:50:14,18 --> 0:50:15,8
pushing that to S3.

1069
0:50:16,56 --> 0:50:19,34
But what I would want to do here
is actually have a couple of

1070
0:50:19,34 --> 0:50:22,62
parameters that decide what your
acceptable lag is.

1071
0:50:23,0 --> 0:50:27,52
Let's say you're willing to have
3 seconds of lag, or so many

1072
0:50:27,52 --> 0:50:28,02
bytes.

1073
0:50:28,08 --> 0:50:29,24
And we can measure that.

1074
0:50:29,24 --> 0:50:31,56
In this situation, we'd be able
to measure either 1 of those

1075
0:50:31,56 --> 0:50:32,3
very easily.

1076
0:50:33,34 --> 0:50:37,9
So we would say, OK, if we get
to that point where we're going

1077
0:50:37,9 --> 0:50:41,14
to hit that lag, we'll actually
write a chunk of WAL out to

1078
0:50:41,14 --> 0:50:43,48
S3 and store it.

1079
0:50:43,48 --> 0:50:46,08
And the WAL archive-get routine
would actually be written so

1080
0:50:46,08 --> 0:50:47,42
that it would be able to...

1081
0:50:47,8 --> 0:50:50,24
Later these chunks would be assembled,
of course, reassembled

1082
0:50:50,4 --> 0:50:54,14
into a WAL segment to keep things
kind of tidy.

1083
0:50:54,28 --> 0:50:56,6
But the WAL get routine would
actually be able to understand

1084
0:50:56,6 --> 0:50:59,82
if you get to the end of the WAL
stream and all you've got is

1085
0:50:59,82 --> 0:51:3,72
these chunks, we'll be able to
read them out and reconstruct

1086
0:51:3,72 --> 0:51:6,1
them and send them Postgres.

1087
0:51:6,1 --> 0:51:8,8
So you actually have, because I
feel like the implementation

1088
0:51:8,8 --> 0:51:11,04
is out there right now, yeah, you're
getting the WAL off of

1089
0:51:11,04 --> 0:51:12,46
the primary, which is good.

1090
0:51:12,52 --> 0:51:15,16
Don't get me wrong, that's a pretty
important thing, but I want

1091
0:51:15,16 --> 0:51:18,84
to get the WAL all the way to
the repository, all the way to

1092
0:51:18,84 --> 0:51:19,78
the final storage.

1093
0:51:19,84 --> 0:51:25,08
So if everything goes away, and
all you've got is the S3 bucket,

1094
0:51:25,08 --> 0:51:30,74
then you still have the whatever
you set that to, sorry, then

1095
0:51:31,06 --> 0:51:34,24
you would have that, unless we
fall behind or other things happen.

1096
0:51:34,24 --> 0:51:35,42
Obviously, there's scenarios.

1097
0:51:36,02 --> 0:51:38,52
But I think it could be a lot more
efficient because then we

1098
0:51:38,52 --> 0:51:42,9
could figure out, hey, they're
generating a WAL at a ridiculous

1099
0:51:42,94 --> 0:51:43,44
rate.

1100
0:51:43,58 --> 0:51:46,96
We actually need to switch over
to just compressing and pushing

1101
0:51:46,96 --> 0:51:51,74
whole WAL segments and then come
back to chunking up portions

1102
0:51:51,74 --> 0:51:52,76
of this WAL segment.

1103
0:51:52,94 --> 0:51:55,64
And with the replication protocol
you don't really have that

1104
0:51:55,64 --> 0:51:59,76
option because you've just got
this fire hose that's sending

1105
0:51:59,76 --> 0:52:4,04
you data and what you really want
to do is say, no, this is too

1106
0:52:4,04 --> 0:52:4,34
much.

1107
0:52:4,34 --> 0:52:7,66
Also, it's a lot slower than just
compressing and shipping.

1108
0:52:8,94 --> 0:52:11,2
So, no, we're going to go back
to doing whole segments, and when

1109
0:52:11,2 --> 0:52:13,28
we get caught up on the whole segments,
we'll start doing the

1110
0:52:13,28 --> 0:52:16,02
most recent segment in chunks,
back and forth.

1111
0:52:16,16 --> 0:52:19,2
I think the best way to do that,
that I'm thinking, I'm always

1112
0:52:19,2 --> 0:52:23,44
an outside the box kind of guy,
is actually in an archive_library

1113
0:52:23,44 --> 0:52:27,66
sitting next to Postgres and just
asking Postgres, where are

1114
0:52:27,66 --> 0:52:31,12
you, where are you, you know, kind
of thing.

1115
0:52:31,12 --> 0:52:33,74
And we can actually monitor, you
know, check stuff on disk, so

1116
0:52:33,74 --> 0:52:36,02
we'd be able to see, oh, we've
just gotten a new WAL segment,

1117
0:52:36,02 --> 0:52:38,8
so everything in the old WAL segment
is fair game, etc., etc.

1118
0:52:39,06 --> 0:52:43,18
So there's plenty of heuristics
we can use to improve this.

1119
0:52:43,38 --> 0:52:45,6
So that's the way I'm leaning right
now, and that gives us the

1120
0:52:45,6 --> 0:52:50,76
archive_library and this WAL receiver
RPO feature at the same

1121
0:52:50,76 --> 0:52:50,86
time.

1122
0:52:50,86 --> 0:52:51,6
Nikolay: RPO control.

1123
0:52:51,74 --> 0:52:54,06
This is better RPO control, basically.

1124
0:52:54,34 --> 0:52:54,86
David: Yeah, better.

1125
0:52:54,86 --> 0:52:56,28
Sorry, I meant to say RPO 0.

1126
0:52:56,28 --> 0:52:59,94
So if we could actually gain this
so that we can update the node doing streaming

1127
0:53:0,26 --> 0:53:7,52
replication, but update the row
in Postgres for replication,

1128
0:53:7,64 --> 0:53:10,32
then we could actually do synchronous
replication this way.

1129
0:53:11,18 --> 0:53:14,24
And we could basically do synchronous
replication to S3.

1130
0:53:14,5 --> 0:53:17,78
If someone wanted such a thing,
yikes, It would be slow.

1131
0:53:18,16 --> 0:53:19,04
Nikolay: I'm very curious.

1132
0:53:19,04 --> 0:53:23,72
What is your opinion about, Barman
was the 1st tool which used

1133
0:53:23,72 --> 0:53:26,0
only streaming replication, right?

1134
0:53:26,0 --> 0:53:27,1
I remember this.

1135
0:53:27,5 --> 0:53:30,44
So what's, what is your opinion
on using streaming replication

1136
0:53:30,44 --> 0:53:31,16
for backups?

1137
0:53:31,8 --> 0:53:35,84
And this new CYBERTEC tool released
last week, pg_hardstorage,

1138
0:53:37,12 --> 0:53:40,68
I quickly checked they don't use
pg_receivewal as well, they

1139
0:53:40,68 --> 0:53:45,2
just implement protocol, they work
with protocol and they pretend

1140
0:53:45,2 --> 0:53:45,98
to be replicable.

1141
0:53:46,42 --> 0:53:47,08
Yeah, That's

1142
0:53:47,08 --> 0:53:48,0
David: what most people have
done.

1143
0:53:48,0 --> 0:53:51,42
I think Barman might still be using
pg_receivewal, but most other

1144
0:53:51,42 --> 0:53:53,3
implementations have gone straight
to the protocol.

1145
0:53:53,3 --> 0:53:56,54
The protocol's not that complicated,
to be honest.

1146
0:53:56,54 --> 0:53:57,84
Nikolay: Like all uppercase.

1147
0:53:58,34 --> 0:54:2,36
David: If you've got libpq, Once
you've got that set up, especially

1148
0:54:2,36 --> 0:54:5,1
if you're using libpq directly,
it's actually fairly trivial.

1149
0:54:5,34 --> 0:54:7,54
The replication protocol is very
simple, which it should be.

1150
0:54:7,54 --> 0:54:8,82
There's nothing wrong with simplicity.

1151
0:54:9,34 --> 0:54:10,46
I'm all about it.

1152
0:54:10,64 --> 0:54:17,22
I think ultimately it's a big bottleneck
making backups through

1153
0:54:17,22 --> 0:54:18,42
the streaming replication.

1154
0:54:19,82 --> 0:54:23,1
So for very large databases, it's
just going to be a pretty big

1155
0:54:23,1 --> 0:54:23,6
bottleneck.

1156
0:54:24,62 --> 0:54:27,28
Obviously someday we can add parallelism,
do all these things

1157
0:54:27,28 --> 0:54:28,2
and etc.

1158
0:54:28,62 --> 0:54:29,7
There's a lot of...

1159
0:54:30,28 --> 0:54:31,52
Well, they got compression.

1160
0:54:32,06 --> 0:54:32,96
So that's good.

1161
0:54:33,2 --> 0:54:34,06
Although I think...

1162
0:54:35,36 --> 0:54:37,42
Is the compression server side
or is it client side?

1163
0:54:37,42 --> 0:54:39,1
I think it might still be client
side.

1164
0:54:39,14 --> 0:54:39,88
No, it's server

1165
0:54:39,88 --> 0:54:41,78
Nikolay: side Okay,

1166
0:54:41,78 --> 0:54:42,68
David: I can't remember.

1167
0:54:42,72 --> 0:54:45,16
Anyway, that's actually a pretty
solvable problem But You've

1168
0:54:45,16 --> 0:54:48,08
got basically a scale of now you
need to back up something really

1169
0:54:48,08 --> 0:54:51,32
large and you're pushing everything
through this little pipe.

1170
0:54:51,7 --> 0:54:55,24
Nikolay: pg_basebackup has compression
since Postgres 15 and you

1171
0:54:55,24 --> 0:54:55,74
can...

1172
0:54:55,92 --> 0:54:56,14
Yeah.

1173
0:54:56,14 --> 0:54:56,6
Right?

1174
0:54:56,6 --> 0:54:57,1
Or

1175
0:54:57,8 --> 0:54:58,15
David: no.

1176
0:54:58,15 --> 0:54:58,5466
It does.

1177
0:54:58,5466 --> 0:55:0,04
I just can't remember which side
it happens on, whether it happens

1178
0:55:0,04 --> 0:55:1,56
on the server side or the client

1179
0:55:1,56 --> 0:55:2,06
Nikolay: side.

1180
0:55:2,08 --> 0:55:2,64
David: I think it is.

1181
0:55:2,64 --> 0:55:4,7
Nikolay: Compress server-gzip
option.

1182
0:55:4,74 --> 0:55:6,1
David: I believe it is, yeah.

1183
0:55:6,42 --> 0:55:7,86
Nikolay: But it's still single threaded.

1184
0:55:8,64 --> 0:55:10,64
David: That's 1 bottleneck taken
care of because you're not

1185
0:55:10,64 --> 0:55:12,96
pushing uncompressed data over
the network, which you definitely

1186
0:55:12,96 --> 0:55:13,94
don't want to do.

1187
0:55:13,98 --> 0:55:15,36
But it's still single threaded.

1188
0:55:15,66 --> 0:55:18,42
And that's a huge limitation on
the backup side, obviously.

1189
0:55:19,02 --> 0:55:22,0
And then on the recovery side,
we still have this problem of

1190
0:55:22,08 --> 0:55:26,2
all the data needs to be recovered
before anything can be done

1191
0:55:26,2 --> 0:55:27,78
if you're doing block incremental.

1192
0:55:27,78 --> 0:55:30,04
If you're not doing page incremental,
then you can make that

1193
0:55:30,04 --> 0:55:30,96
a bit more efficient.

1194
0:55:31,2 --> 0:55:34,72
You can stream the tar file from
S3 and decompress it as you

1195
0:55:34,72 --> 0:55:35,82
go and do various things.

1196
0:55:35,82 --> 0:55:39,24
But if you're doing page incremental,
which I think pretty much

1197
0:55:39,24 --> 0:55:43,14
everyone wants to do, that becomes
a huge bottleneck on the restore

1198
0:55:43,14 --> 0:55:43,38
side.

1199
0:55:43,38 --> 0:55:44,1
Nikolay: I'm sorry.

1200
0:55:44,64 --> 0:55:46,16
We never had so compressed.

1201
0:55:47,86 --> 0:55:51,96
I know we already are way over
time, but it's so compressed.

1202
0:55:51,96 --> 0:55:54,84
You compress so much knowledge
and I like it so much, but can

1203
0:55:54,84 --> 0:55:58,22
you explain, like elaborate page
level versus block level and

1204
0:55:58,22 --> 0:56:0,04
why you think it's important?

1205
0:56:0,94 --> 0:56:3,98
David: That's my confusion because
what was introduced in Postgres 17,

1206
0:56:4,18 --> 0:56:7,54
the page level incremental with
the WAL summarized and everything,

1207
0:56:7,8 --> 0:56:10,58
in pgBackRest we call that block
level incremental.

1208
0:56:11,3 --> 0:56:13,94
And the reason why we call it block
retal is we don't always

1209
0:56:13,94 --> 0:56:16,4
operate at the page, at page size.

1210
0:56:16,92 --> 0:56:18,58
We operate at block size.

1211
0:56:18,9 --> 0:56:22,54
And then we actually combine blocks
into what we call a super

1212
0:56:22,54 --> 0:56:23,58
block for compression.

1213
0:56:24,16 --> 0:56:29,54
So let's say I want to get 1 block
out, I might have to retrieve

1214
0:56:29,54 --> 0:56:32,22
3 compressed blocks or 4 compressed
blocks, depending on the

1215
0:56:32,22 --> 0:56:34,1
super block size, to actually get
that.

1216
0:56:34,12 --> 0:56:37,1
But the important thing is we can,
let's say we've got a 1 gigabyte

1217
0:56:37,12 --> 0:56:42,44
file and we need to recover 1 block,
we can go recover 3 blocks

1218
0:56:42,44 --> 0:56:47,56
to do that instead of 1000 blocks,
however many blocks are in

1219
0:56:47,56 --> 0:56:48,66
a 1 gig file.

1220
0:56:49,0 --> 0:56:53,2
I think it's usually 1, 200 at
our largest block size.

1221
0:56:53,26 --> 0:56:57,5
So we can go recover 3 blocks instead
of 1, 200 blocks, even

1222
0:56:57,5 --> 0:57:1,16
though we have to recover 3 blocks
just to write 1 block.

1223
0:57:1,16 --> 0:57:4,78
So the main thing is we're able
to go, let's say you've got a

1224
0:57:4,78 --> 0:57:8,32
standby or a primary that's failed
for some reason, you don't

1225
0:57:8,32 --> 0:57:8,98
know why.

1226
0:57:9,14 --> 0:57:10,74
So you're gonna do a delta restore.

1227
0:57:11,04 --> 0:57:15,42
So in a delta restore, we go and
we look at every file And for

1228
0:57:15,42 --> 0:57:18,64
the block level stuff, we chop
it up into pieces, and then we

1229
0:57:18,64 --> 0:57:21,5
do a hash on each piece, and then
we do a hash for the entire

1230
0:57:21,5 --> 0:57:22,0
file.

1231
0:57:22,3 --> 0:57:25,52
If the whole file hash matches
what's in the backup manifest,

1232
0:57:25,52 --> 0:57:26,42
then we're done.

1233
0:57:27,04 --> 0:57:27,78
Move on.

1234
0:57:27,9 --> 0:57:28,68
Next file.

1235
0:57:29,1 --> 0:57:33,4
If the whole file manifest doesn't
match, then we can go recover

1236
0:57:33,42 --> 0:57:37,4
what we call the block map, which
is a map that tells us where

1237
0:57:37,4 --> 0:57:39,68
all the blocks live in all the
backups.

1238
0:57:39,68 --> 0:57:43,86
You might have 10 incremental backups
since your last full, and

1239
0:57:43,86 --> 0:57:46,22
now we need to go figure out where
are the blocks, where are

1240
0:57:46,22 --> 0:57:50,14
the blocks that we need located
and go grab those blocks.

1241
0:57:51,0 --> 0:57:54,44
And in the other implementations
that have been done, you can't

1242
0:57:54,44 --> 0:57:54,84
do that.

1243
0:57:54,84 --> 0:57:58,48
You basically have to go recover
everything and some of them

1244
0:57:58,48 --> 0:58:2,06
will have good things like WAL-G
will, as it's reading it,

1245
0:58:2,16 --> 0:58:4,12
it'll know that it doesn't need
the block and it'll just throw

1246
0:58:4,12 --> 0:58:4,78
it away.

1247
0:58:4,82 --> 0:58:6,82
That kind of read over thing I
was talking about before.

1248
0:58:6,82 --> 0:58:9,16
So it'll just throw away the blocks
it doesn't need, but it's

1249
0:58:9,16 --> 0:58:10,46
still reading them out sequentially.

1250
0:58:10,48 --> 0:58:13,2
We're actually able to go random
access and pull out just the

1251
0:58:13,2 --> 0:58:14,16
blocks you need.

1252
0:58:14,64 --> 0:58:17,28
And It's extremely powerful because
we have that delta restore

1253
0:58:17,28 --> 0:58:20,26
concept where we can actually just
recover part of a cluster.

1254
0:58:20,74 --> 0:58:24,96
And that's where the power of those
checksums that we spent in

1255
0:58:24,96 --> 0:58:28,08
theory, 5% of our time, but actually
it's really much less than

1256
0:58:28,08 --> 0:58:32,44
1%, that's where that all comes
in on the pg_basebackup side

1257
0:58:32,44 --> 0:58:32,84
right now.

1258
0:58:32,84 --> 0:58:36,92
You basically have to recover all
the full and all the incrementals,

1259
0:58:37,54 --> 0:58:42,22
decompress everything all at the
same time, present that to pg_basebackup

1260
0:58:42,44 --> 0:58:45,98
that will then rewrite that into
another directory, sorry, pg_combinebackup

1261
0:58:46,26 --> 0:58:49,2
which will then rewrite
that into another directory.

1262
0:58:49,64 --> 0:58:53,94
So at a minimum to do recovery,
you're looking at double your

1263
0:58:53,94 --> 0:58:57,1
database size and actually more.

1264
0:58:57,66 --> 0:59:2,24
And you don't really know what
that more is, unless you've stored

1265
0:59:2,24 --> 0:59:2,42
it.

1266
0:59:2,42 --> 0:59:4,78
It's up to you to actually store
that information to find out

1267
0:59:4,78 --> 0:59:7,48
what is the total number of bytes
that I'm going to need to even

1268
0:59:7,48 --> 0:59:8,82
pull down the data.

1269
0:59:9,52 --> 0:59:13,48
On the pgBackRest side, when
we're doing a recovery, nothing

1270
0:59:13,48 --> 0:59:15,56
ever hits disk except in PGDATA.

1271
0:59:15,94 --> 0:59:18,58
So we're just reconstructing those
files inside PGDATA.

1272
0:59:19,06 --> 0:59:21,56
There's no spooling, there's no
spooling during backup, there's

1273
0:59:21,56 --> 0:59:24,84
no spooling during recovery, there's
no spooling during archive

1274
0:59:24,84 --> 0:59:27,32
push, there is spooling during
archive-get because we actually

1275
0:59:27,32 --> 0:59:30,84
prefetch WAL files and we store
them so that we can hand them

1276
0:59:30,84 --> 0:59:31,56
over to Postgres.

1277
0:59:31,56 --> 0:59:34,6
So we, we spool on that side, but
for the most part, we're 0

1278
0:59:34,6 --> 0:59:35,34
disk operation.

1279
0:59:36,74 --> 0:59:39,68
And to be really efficient, that's
the way you need to be looking.

1280
0:59:39,68 --> 0:59:43,52
And right now the design in pg_basebackup
is exactly not that.

1281
0:59:44,06 --> 0:59:45,4
A lot of disk I/O.

1282
0:59:45,74 --> 0:59:48,58
They've done some tricks with,
if you're on ZFS or other file

1283
0:59:48,58 --> 0:59:53,04
systems that can do copy-on-write
tricks and do other fun things

1284
0:59:53,1 --> 0:59:55,84
to try to minimize the number of
writes that you're going to

1285
0:59:55,84 --> 0:59:56,34
do.

1286
0:59:56,58 --> 0:59:58,94
I don't know if it actually minimizes
the space that you need.

1287
0:59:58,94 --> 1:0:0,86
On some file systems that might
be true.

1288
1:0:1,48 --> 1:0:3,74
Nikolay: Thank you for explanation,
it was great.

1289
1:0:4,12 --> 1:0:7,22
David: The whole idea was the
feasibility of pg_basebackup

1290
1:0:8,0 --> 1:0:11,3
using the streaming replication
as a backup tool.

1291
1:0:11,6 --> 1:0:14,64
And I think, yeah, you can do it,
but when you're working at

1292
1:0:14,64 --> 1:0:17,34
scale, it starts to introduce a
lot of limitations.

1293
1:0:18,16 --> 1:0:22,38
And from day 1, Gabriele Bartolini
wrote a pretty good thing

1294
1:0:22,38 --> 1:0:25,76
after I announced that I wasn't
working on pgBackRest about

1295
1:0:25,76 --> 1:0:29,82
his and my fundamental disagreement
about how this tool should

1296
1:0:29,82 --> 1:0:30,32
work.

1297
1:0:31,44 --> 1:0:36,36
And his idea is that the copy thing
should be owned by Postgres.

1298
1:0:36,46 --> 1:0:38,44
So Postgres should do all the file
copying.

1299
1:0:39,52 --> 1:0:43,54
And my idea was that no, the copy
should be built into the tool.

1300
1:0:44,32 --> 1:0:46,42
Because then we can do, the sky's
the limit.

1301
1:0:46,56 --> 1:0:47,94
We can do anything we want.

1302
1:0:48,42 --> 1:0:50,92
And I think the performance that
you can get from pgBackRest

1303
1:0:50,92 --> 1:0:55,44
speaks to the value of having that
complexity.

1304
1:0:55,68 --> 1:0:59,48
And it is, it's a lot of complexity,
don't get me wrong, to have

1305
1:0:59,48 --> 1:1:3,38
all that built into pgBackRest,
but the payoff is performance.

1306
1:1:4,82 --> 1:1:6,0
And also reliability.

1307
1:1:6,82 --> 1:1:9,1
We have checksums that we can check
on everything.

1308
1:1:9,14 --> 1:1:11,54
We know when the data that's coming
back is good.

1309
1:1:12,04 --> 1:1:15,78
That's 1 thing that kind of peeves
me is that if you, like let's

1310
1:1:15,78 --> 1:1:19,24
say you were doing pg_combinebackup
it doesn't actually verify

1311
1:1:19,44 --> 1:1:20,64
that the data is correct.

1312
1:1:21,62 --> 1:1:24,72
So if you want to verify, you actually
have to run pg_verifybackup

1313
1:1:24,72 --> 1:1:29,44
as a separate step before
you run pg_combinebackup.

1314
1:1:30,06 --> 1:1:33,04
And in my mind, those should always
be combined into a single

1315
1:1:33,04 --> 1:1:33,54
operation.

1316
1:1:34,2 --> 1:1:37,36
You're looking at the data, you're
verifying it, you're writing

1317
1:1:37,36 --> 1:1:37,92
it out.

1318
1:1:37,92 --> 1:1:40,68
And the cost for that, again, oh,
checksums are expensive.

1319
1:1:40,68 --> 1:1:44,56
Yeah, but if you're streaming the
data directly from S3 and checksumming

1320
1:1:44,6 --> 1:1:48,46
it and writing it out as it comes
in, the checksumming cost disappears.

1321
1:1:49,54 --> 1:1:50,82
You don't see it anymore.

1322
1:1:51,18 --> 1:1:55,16
If you copy all the data from S3
locally, decompress it, and

1323
1:1:55,16 --> 1:1:59,12
then start running checksums on
it, it looks really painful because

1324
1:1:59,64 --> 1:2:4,24
you weren't able to, the idea for
me is let's gain all the latency

1325
1:2:4,24 --> 1:2:5,98
that we can, the storage latency
that we can.

1326
1:2:5,98 --> 1:2:8,26
So while we're waiting for the
storage to give us something,

1327
1:2:8,42 --> 1:2:12,02
we'll send it the next request
to S3 and while it's thinking,

1328
1:2:12,26 --> 1:2:14,54
sending this data, we check some.

1329
1:2:15,18 --> 1:2:18,38
We're done with that block, the
next block is ready for us.

1330
1:2:18,9 --> 1:2:22,48
We pull that in, we start checksumming,
we've asynchronously

1331
1:2:22,56 --> 1:2:24,06
set the next request to S3.

1332
1:2:24,28 --> 1:2:25,94
So we've already got more data
coming.

1333
1:2:26,2 --> 1:2:29,54
In the next version of Postgres,
if everything goes well, we'll

1334
1:2:29,54 --> 1:2:30,36
also have prefetch.

1335
1:2:30,36 --> 1:2:34,06
So we won't be asynchronously fetching
1 block, we'll be asynchronously

1336
1:2:34,32 --> 1:2:36,44
fetching by default 4.

1337
1:2:37,36 --> 1:2:40,32
And just pulling that data in as
fast as we can, but even as

1338
1:2:40,32 --> 1:2:43,48
fast as you can get stuff over
the network, that cost of checksumming

1339
1:2:43,5 --> 1:2:46,02
disappears, But it gives you power.

1340
1:2:46,4 --> 1:2:49,26
And pgBackRest is the only thing
that has something equivalent

1341
1:2:49,28 --> 1:2:50,86
to delta restore.

1342
1:2:51,58 --> 1:2:53,5
And all of that is powered by the
checksums.

1343
1:2:54,16 --> 1:2:55,08
That's where all that comes from.

1344
1:2:55,08 --> 1:2:57,94
And then we have the block level
delta restore as well, which

1345
1:2:57,94 --> 1:3:0,86
is powered by the block level hashes,
which are actually XXHash,

1346
1:3:1,32 --> 1:3:2,52
not SHA-1.

1347
1:3:3,94 --> 1:3:6,36
Because the blocks are a maximum
of 88k.

1348
1:3:7,28 --> 1:3:9,72
Although you configure that, you
can also configure how much

1349
1:3:9,72 --> 1:3:12,98
of the XXHash you're going to use
for a particular block size.

1350
1:3:13,52 --> 1:3:16,38
But we did a lot of analysis on
this and looking at file systems

1351
1:3:16,64 --> 1:3:20,24
and how many bits they were using
to protect blah blah blah etc.

1352
1:3:21,1 --> 1:3:24,0
And basically came up with checksum
sizes that were appropriate

1353
1:3:24,0 --> 1:3:26,5
for each block size that we support.

1354
1:3:26,96 --> 1:3:29,82
The whole the way we do block incremental
is actually a pretty

1355
1:3:30,24 --> 1:3:31,1
complicated topic.

1356
1:3:31,1 --> 1:3:35,04
I did a talk on it once and it
was a complete disaster because

1357
1:3:36,82 --> 1:3:38,06
it was just too much.

1358
1:3:38,8 --> 1:3:40,02
Nikolay: You go too deep.

1359
1:3:40,08 --> 1:3:45,02
Truncating 32-bits you said, right?

1360
1:3:45,02 --> 1:3:49,46
Or truncating hashes you said,
right?

1361
1:3:49,46 --> 1:3:51,14
David: Oh yeah, so the 32-bits.

1362
1:3:51,54 --> 1:3:56,98
So what we do is we generate a
128-bit hash and then we use however

1363
1:3:56,98 --> 1:3:59,74
many bits out of that, well, bytes
actually, that we want.

1364
1:4:0,06 --> 1:4:4,58
The minimum number of bytes we'd
use for an 8k block is 6 bytes,

1365
1:4:5,22 --> 1:4:6,34
which is actually a lot.

1366
1:4:6,34 --> 1:4:13,84
That's 50% more than would be used
for 4k pages on ZFS, for instance,

1367
1:4:14,06 --> 1:4:14,76
or Btrfs.

1368
1:4:15,52 --> 1:4:18,58
So we go a little bit crazy on
the checksums because we just

1369
1:4:18,58 --> 1:4:20,6
we want everything to be right
and then when you're doing block

1370
1:4:20,6 --> 1:4:23,44
increment and we actually have
2 levels of checksum so we use

1371
1:4:23,44 --> 1:4:27,26
the block checksums to reconstruct
the file and Then we also

1372
1:4:27,26 --> 1:4:30,36
checksum the entire file and the
SHA-1 hash of the entire file

1373
1:4:30,36 --> 1:4:32,74
plus all the block checksums have
to match.

1374
1:4:33,54 --> 1:4:37,28
And the chances of that getting
a collision, you know, like a

1375
1:4:37,28 --> 1:4:39,8
piece of data that actually satisfies
the conditions, all the

1376
1:4:39,8 --> 1:4:42,54
block checksums matched and the
SHA-1 checksum of the whole file

1377
1:4:42,54 --> 1:4:43,04
matches.

1378
1:4:43,52 --> 1:4:46,16
We're talking heat death of the
universe probability here.

1379
1:4:46,16 --> 1:4:46,66
Right?

1380
1:4:46,8 --> 1:4:48,62
It's just simply not going to happen.

1381
1:4:48,74 --> 1:4:52,62
You're going to have disk corruption
1000000000 times before

1382
1:4:52,74 --> 1:4:56,78
anything like this ever fails,
as far as my math works out.

1383
1:4:57,26 --> 1:4:58,58
Michael: Well, it sounds reasonable.

1384
1:4:58,66 --> 1:5:0,98
Nikolay: It's a lot of interesting
topics, I feel.

1385
1:5:1,24 --> 1:5:8,44
And I like the direction you are
thinking, especially RPO control.

1386
1:5:8,44 --> 1:5:13,66
This is super cool thing because
I see the demand because Postgres

1387
1:5:13,66 --> 1:5:17,76
has been used in like critical
systems and people just don't

1388
1:5:17,76 --> 1:5:19,54
know what's happening there.

1389
1:5:19,54 --> 1:5:22,68
Like how, what's happening in the
case of disaster, right?

1390
1:5:22,68 --> 1:5:26,76
Like it's hard to answer simple
questions to business leaders.

1391
1:5:26,76 --> 1:5:26,96
David: Yeah.

1392
1:5:26,96 --> 1:5:28,14
It's not 1 size fits all.

1393
1:5:28,14 --> 1:5:28,32
Right.

1394
1:5:28,32 --> 1:5:28,78
I did.

1395
1:5:28,78 --> 1:5:30,06
You guys are familiar with ARIN,
right?

1396
1:5:30,06 --> 1:5:31,9
The internet number registry.

1397
1:5:32,38 --> 1:5:35,64
I know the people who run the database
over there, and they use

1398
1:5:35,64 --> 1:5:36,4
pgBackRest.

1399
1:5:36,58 --> 1:5:40,46
I worked with them for years in
another company, and they're

1400
1:5:40,46 --> 1:5:43,22
all RPO 0 stuff, right?

1401
1:5:43,26 --> 1:5:45,78
But their write volume is stupidly
low.

1402
1:5:46,1 --> 1:5:51,4
So they've got 1 primary, which
receives 100 writes an hour on

1403
1:5:51,4 --> 1:5:55,12
a really busy day, and then they
replicate that, and then you've

1404
1:5:55,12 --> 1:5:58,46
got the whole internet reading
to figure out what blocks are

1405
1:5:58,46 --> 1:6:2,28
allocated to who at any given point,
and where they should be

1406
1:6:2,28 --> 1:6:3,74
routed, that kind of thing.

1407
1:6:3,9 --> 1:6:8,1
So the read volume is pretty big,
the write volume is extremely

1408
1:6:8,14 --> 1:6:14,04
tiny, so they do synchronous replication
because they can't lose

1409
1:6:14,04 --> 1:6:17,28
anything and they have low enough
write volume they can get away

1410
1:6:17,28 --> 1:6:20,52
with it, so they would be a great
candidate for a pgBackRest

1411
1:6:20,68 --> 1:6:25,02
RPO 0 solution that allows them to
synchronously write to S3.

1412
1:6:25,42 --> 1:6:26,76
I bet they'd love it.

1413
1:6:27,32 --> 1:6:30,22
But for most people, that kind
of solution just isn't going to

1414
1:6:30,22 --> 1:6:30,72
work.

1415
1:6:31,62 --> 1:6:34,98
People ask me how I can work on
this software year after year.

1416
1:6:35,46 --> 1:6:38,18
And the reason is because the problems
are actually really interesting

1417
1:6:38,72 --> 1:6:42,34
and challenging and extraordinarily
complex in terms of the solutions

1418
1:6:42,48 --> 1:6:45,54
that we have to come up with to
solve these problems.

1419
1:6:45,54 --> 1:6:49,2
It's, it's, pgBackRest is a relatively
small project, but it's

1420
1:6:49,2 --> 1:6:49,6
dense.

1421
1:6:49,6 --> 1:6:50,74
The stuff that we do is...

1422
1:6:50,74 --> 1:6:52,2
Nikolay: Let me uncompress a little
bit.

1423
1:6:52,2 --> 1:6:55,38
You mentioned ARIN, it's American
Registry for Internet Numbers,

1424
1:6:55,38 --> 1:6:55,88
right?

1425
1:6:56,04 --> 1:6:56,4
David: Right.

1426
1:6:56,4 --> 1:6:57,26
That's the 1.

1427
1:6:57,64 --> 1:7:1,04
Nikolay: So it's basically the important
piece of internet, right?

1428
1:7:1,74 --> 1:7:3,06
David: Very important piece.

1429
1:7:3,2 --> 1:7:3,7
Yeah.

1430
1:7:3,72 --> 1:7:5,06
Very important piece.

1431
1:7:5,18 --> 1:7:8,3
Nikolay: It's interesting that you
Consider like volumes are not

1432
1:7:8,3 --> 1:7:11,08
huge and the streaming replication
and so on.

1433
1:7:11,76 --> 1:7:14,2
It's interesting and challenging.

1434
1:7:14,38 --> 1:7:19,9
But I have cases where it's like
just not answered.

1435
1:7:19,94 --> 1:7:22,92
RPOs are not answered and extremely
important cases.

1436
1:7:23,6 --> 1:7:26,88
Important companies also, they
simply don't know.

1437
1:7:27,44 --> 1:7:30,36
So if you build this, I think it
will be good.

1438
1:7:30,74 --> 1:7:33,3
It will be useful, helpful, and
so on.

1439
1:7:33,58 --> 1:7:35,82
David: It's a tricky thing though,
to some extent, when you

1440
1:7:35,82 --> 1:7:38,44
put those tools in people's hands,
because of course there's

1441
1:7:38,44 --> 1:7:42,94
some level in the hierarchy that's
going to say, always RPO 0.

1442
1:7:45,04 --> 1:7:46,72
Of course, we have to have synchronous
replication.

1443
1:7:46,72 --> 1:7:47,728
Nikolay: Why would we ever do that?

1444
1:7:47,728 --> 1:7:48,54
We cannot afford data loss.

1445
1:7:48,54 --> 1:7:50,88
We cannot afford, we need to use
the synchronous replication

1446
1:7:51,14 --> 1:7:52,06
and so on.

1447
1:7:52,7 --> 1:7:55,7
And also we cannot, we don't have
split brains.

1448
1:7:56,18 --> 1:7:58,6
You mentioned Gabriele Bartolini,
right?

1449
1:7:58,6 --> 1:8:3,2
I also have my opinion about CloudNativePG
and solutions he

1450
1:8:3,54 --> 1:8:4,86
chose in this tool.

1451
1:8:5,38 --> 1:8:9,64
Yeah, but like at high level we
are great, but if you look inside

1452
1:8:9,8 --> 1:8:13,6
you see data loss is possible,
split brains are possible, almost

1453
1:8:13,6 --> 1:8:14,1
everywhere.

1454
1:8:14,54 --> 1:8:17,64
And It's really extremely hard
to achieve and I like so much

1455
1:8:17,64 --> 1:8:23,44
you focus on these topics like
RPO, control, and especially corruption

1456
1:8:23,54 --> 1:8:25,22
and checksums everywhere.

1457
1:8:25,38 --> 1:8:26,1
It's great.

1458
1:8:26,74 --> 1:8:29,78
David: My whole job in life is
to protect the data, protect

1459
1:8:29,78 --> 1:8:30,6
everyone's data.

1460
1:8:31,1 --> 1:8:33,18
And that's what I think about all
the time.

1461
1:8:33,18 --> 1:8:36,68
Like how can pgBackRest be the
most reliable thing?

1462
1:8:36,98 --> 1:8:39,96
Even if it's not the fastest, or
I think it's generally the fastest,

1463
1:8:39,96 --> 1:8:42,28
but if it's not the most feature
rich, there are a couple of

1464
1:8:42,28 --> 1:8:45,4
features that other programs have
that we don't have, although

1465
1:8:45,4 --> 1:8:46,6
we're working on that.

1466
1:8:46,64 --> 1:8:50,52
But everything we introduce has
to be performant and ultimately

1467
1:8:50,74 --> 1:8:54,28
absolutely reliable people trust
the software to protect their

1468
1:8:54,28 --> 1:8:54,78
data.

1469
1:8:55,16 --> 1:8:58,28
Now they have an obligation to,
they need to, I get people at

1470
1:8:58,28 --> 1:9:2,06
conferences that will come to me
and say, you saved my job, pgBackRest

1471
1:9:2,06 --> 1:9:3,12
saved my job.

1472
1:9:3,12 --> 1:9:5,28
And I'm like, I really appreciate
you saying that, but actually

1473
1:9:5,28 --> 1:9:8,58
you saved your job because you
set up backup, you tested restores,

1474
1:9:9,06 --> 1:9:11,26
and when the big day happened,
you were ready.

1475
1:9:11,28 --> 1:9:12,98
Nikolay: Yes, I set up restores.

1476
1:9:12,98 --> 1:9:15,58
David: Because they'll tell me,
oh yeah, we set up weekly restore

1477
1:9:15,58 --> 1:9:16,74
tests like you recommended.

1478
1:9:17,32 --> 1:9:18,52
And I'm like, okay, great.

1479
1:9:18,52 --> 1:9:21,32
I can recommend these things, but
a lot of people don't do it

1480
1:9:21,34 --> 1:9:24,34
so if you actually go and do the
stuff that the backup people

1481
1:9:24,34 --> 1:9:28,38
recommend kudos to you you deserve
all the credit for saving

1482
1:9:28,38 --> 1:9:31,68
your job not me but don't get me
wrong I still like hearing the

1483
1:9:31,68 --> 1:9:35,18
stories It's great to know that
pgBackRest is useful and valued,

1484
1:9:35,46 --> 1:9:36,1
of course.

1485
1:9:36,34 --> 1:9:38,86
Nikolay: Untested backups are Schrödinger
backups.

1486
1:9:39,12 --> 1:9:41,3
They should not be considered proper
backups.

1487
1:9:41,38 --> 1:9:43,24
Yeah, backups must be tested.

1488
1:9:43,78 --> 1:9:45,76
David: I can't remember where
I found that, But I put that

1489
1:9:45,76 --> 1:9:49,98
on 1 of my very early talks, Schrödinger's
Backup.

1490
1:9:50,38 --> 1:9:53,48
It wasn't me, it wasn't original
to me, but I can't remember

1491
1:9:53,48 --> 1:9:54,14
where I saw it.

1492
1:9:54,14 --> 1:9:55,68
But as soon as I saw it, I was
like, yes.

1493
1:9:55,68 --> 1:9:57,42
Nikolay: It's an obvious idea, honestly.

1494
1:9:57,94 --> 1:9:59,88
Multiple people might invent it.

1495
1:10:0,48 --> 1:10:5,42
Anyway, way over time, I enjoyed
so much talking to you.

1496
1:10:5,74 --> 1:10:11,18
Final question, we discussed a
lot of technical detail, very

1497
1:10:11,18 --> 1:10:12,68
advanced, I must say.

1498
1:10:12,82 --> 1:10:16,4
I think it's okay if some people
don't get everything because

1499
1:10:16,4 --> 1:10:19,7
we definitely dived into multiple
areas very deep.

1500
1:10:19,92 --> 1:10:22,32
We also touched the situation about
sponsorship.

1501
1:10:22,64 --> 1:10:25,58
We haven't touched the situation
about maintainers.

1502
1:10:25,68 --> 1:10:29,6
What would help you to have a 2nd
big maintainer?

1503
1:10:30,04 --> 1:10:34,16
David: A 2nd big maintainer would
be good for a couple of reasons.

1504
1:10:34,16 --> 1:10:36,26
1, to help me write the big features.

1505
1:10:36,88 --> 1:10:40,24
These days, I'm the only 1 writing
big features and I'm just

1506
1:10:40,24 --> 1:10:40,94
1 person.

1507
1:10:41,4 --> 1:10:43,12
And the other thing is review.

1508
1:10:43,58 --> 1:10:45,3
So review is a big bottleneck for
me.

1509
1:10:45,3 --> 1:10:48,9
I can actually produce quite a
lot of code, but it's got to be

1510
1:10:48,9 --> 1:10:50,28
gone over carefully by somebody.

1511
1:10:50,28 --> 1:10:52,94
I also, I'm using LLMs now for
review as well.

1512
1:10:53,04 --> 1:10:56,14
So I'll ask cloud code review anything
I write before I send

1513
1:10:56,14 --> 1:10:58,94
it to any person just to catch
the really obvious stuff.

1514
1:10:58,94 --> 1:11:2,0
And these days the not so obvious
stuff is getting pretty good,

1515
1:11:2,64 --> 1:11:5,4
But I also need someone to bounce
ideas off of.

1516
1:11:6,26 --> 1:11:7,64
Another maintainer would be good
for that.

1517
1:11:7,64 --> 1:11:10,38
Someone who I could reliably chat
with and be like, so I was

1518
1:11:10,38 --> 1:11:13,02
thinking about X, Y, and Z and
I just, you know how it is when

1519
1:11:13,02 --> 1:11:16,84
you vocalize something, it becomes,
sometimes you don't even

1520
1:11:16,84 --> 1:11:18,54
need to get input from them.

1521
1:11:18,66 --> 1:11:21,96
Just by explaining the problem
to them, you, oh yeah, I know

1522
1:11:21,96 --> 1:11:23,22
what I need to do here.

1523
1:11:23,24 --> 1:11:24,88
It's so obvious to me now.

1524
1:11:25,44 --> 1:11:27,9
So that's why I really need another
maintainer.

1525
1:11:27,9 --> 1:11:29,4
We're working on that right now.

1526
1:11:30,02 --> 1:11:32,86
That should be, I don't wanna make
any announcements or say anything

1527
1:11:33,34 --> 1:11:35,74
regard to this, but this is something
that's actively being worked

1528
1:11:35,74 --> 1:11:41,58
on to get another maybe not full-time
maintainer, but very active,

1529
1:11:41,58 --> 1:11:42,78
very involved maintainer.

1530
1:11:43,92 --> 1:11:45,32
Nikolay: That's, that would be great.

1531
1:11:45,44 --> 1:11:45,94
Yeah.

1532
1:11:46,12 --> 1:11:48,74
David: And to some extent, it's
hard to know how big we could

1533
1:11:48,74 --> 1:11:49,6
scale this project.

1534
1:11:49,6 --> 1:11:51,56
It's an interesting, how many people
are interested.

1535
1:11:52,12 --> 1:11:54,02
Also how big are the problems we're
solving?

1536
1:11:54,02 --> 1:11:56,54
Can 2 people handle everything?

1537
1:11:57,44 --> 1:11:57,94
Probably.

1538
1:11:58,78 --> 1:12:1,06
Especially with tools these days,
helping.

1539
1:12:1,5 --> 1:12:3,58
We're getting more contributions
from the community.

1540
1:12:3,86 --> 1:12:6,42
Those are coming in a pretty regular
stream now so I don't have

1541
1:12:6,42 --> 1:12:7,26
to write everything.

1542
1:12:7,28 --> 1:12:9,44
I don't have to find all the bugs,
I don't have to fix all the

1543
1:12:9,44 --> 1:12:13,68
bugs, I don't have to do little
authentication tweaks for S3

1544
1:12:13,86 --> 1:12:14,94
or blah blah blah.

1545
1:12:15,04 --> 1:12:17,5
Most of that is coming from outside
contributors, and I just

1546
1:12:17,5 --> 1:12:21,0
review it, and if it's reasonable,
commit it.

1547
1:12:21,3 --> 1:12:24,34
So we're building community that
way as well, but if we can get

1548
1:12:24,34 --> 1:12:29,64
1 more person who's really looking
at it regularly, that would

1549
1:12:29,64 --> 1:12:30,06
be great.

1550
1:12:30,06 --> 1:12:33,52
And that would, of course, be in
addition to Stefan, with Data

1551
1:12:33,52 --> 1:12:36,5
Egret, who actually already spends
quite a bit of time on pgBackRest

1552
1:12:36,5 --> 1:12:39,18
as well, review and
testing, et cetera.

1553
1:12:40,96 --> 1:12:41,46
Nikolay: Cool.

1554
1:12:42,18 --> 1:12:42,68
Great.

1555
1:12:43,04 --> 1:12:43,44
Thank you.

1556
1:12:43,44 --> 1:12:45,04
I don't have any more questions.

1557
1:12:45,24 --> 1:12:49,18
I have, but I will post them, because
they are super technical

1558
1:12:49,2 --> 1:12:49,7
again.

1559
1:12:50,02 --> 1:12:51,22
Michael: David, thank you so much.

1560
1:12:51,22 --> 1:12:53,52
Thanks for joining us, but also
thanks for all the maintenance

1561
1:12:53,52 --> 1:12:54,86
you've done over the years.

1562
1:12:54,9 --> 1:12:55,7
David: Oh, absolutely.

1563
1:12:55,76 --> 1:12:57,22
It's definitely been my pleasure.

1564
1:12:58,14 --> 1:13:0,56
Michael: Congrats on getting all
the sponsors and Thanks to them

1565
1:13:0,56 --> 1:13:2,4
as well for keeping it going.

1566
1:13:2,4 --> 1:13:4,12
David: Yeah, like kudos to the
sponsors.

1567
1:13:4,12 --> 1:13:5,4
They have made this possible.

1568
1:13:5,58 --> 1:13:7,18
I'm just, I'm back to work.

1569
1:13:7,36 --> 1:13:10,02
I'm back to what I love doing and
everyone gets to benefit.

1570
1:13:10,12 --> 1:13:12,22
So I think this has worked out
really well.

1571
1:13:12,36 --> 1:13:14,42
So a really cool open source story.

1572
1:13:14,96 --> 1:13:16,88
Michael: Yeah, it is absolutely
good.

1573
1:13:16,88 --> 1:13:18,14
Great open source story.

1574
1:13:18,44 --> 1:13:18,82
Nikolay: Thank you.

1575
1:13:18,82 --> 1:13:19,9
Have a great week.

1576
1:13:20,2 --> 1:13:21,64
David: Yeah, thank you very much
for having me.