1
00:00:00,000 --> 00:00:06,580
Hi everyone. This week you might have noticed your favorite apps, Outlook, Strava, Steam, Slack, being out for some time.

2
00:00:06,719 --> 00:00:10,839
This was caused by an AWS outage. Today, we'll have a look at what happened,

3
00:00:11,019 --> 00:00:15,179
how did it affect us as Dataminded, and what can you do to prevent all of this.

4
00:00:15,320 --> 00:00:21,059
And with me to explain all that is Stijn De Haes, our technical lead at Conveyor. Welcome, Stijn.

5
00:00:22,279 --> 00:00:24,019
Hello. How are you doing?

6
00:00:25,280 --> 00:00:27,679
Better than Monday. Yeah, busy week?

7
00:00:28,519 --> 00:00:32,280
Well, it's okay. We had mostly under control.

8
00:00:33,240 --> 00:00:38,020
Today, we'll discuss three major questions like what happened at AWS, how did it hit us,

9
00:00:38,119 --> 00:00:41,859
and what can people do to prevent such an outage or to at least mitigate it?

10
00:00:42,779 --> 00:00:51,700
So what happened at AWS is they messed something up. As we engineers, we push things through

11
00:00:51,700 --> 00:00:56,979
production, right? And sometimes you make a mistake. So they made a small mistake with a DNS

12
00:00:56,979 --> 00:00:57,659
configuration. And then they made a mistake with a DNS configuration. And then they made a small mistake

13
00:00:59,679 --> 00:01:06,760
with a DynamoDB endpoint. And apparently, almost all of AWS has some connection with DynamoDB.

14
00:01:07,040 --> 00:01:13,340
You couldn't launch virtual machines in the North Virginia region. You couldn't do anything with IAM

15
00:01:13,340 --> 00:01:21,379
globally. You couldn't update your IAM roles, etc. So a lot of people experienced some hardship

16
00:01:21,379 --> 00:01:21,900
this Monday.

17
00:01:22,439 --> 00:01:24,540
Yeah. And how did it affect us?

18
00:01:26,519 --> 00:01:28,739
Mostly, we noticed it.

19
00:01:29,680 --> 00:01:33,219
Through our alerting system. We couldn't pull any new images on our Azure cluster.

20
00:01:33,760 --> 00:01:40,939
And then we also, some of our customers contacted us. We can't do Conveyor builds. So they couldn't

21
00:01:40,939 --> 00:01:48,700
deploy new packages on Conveyor. Everything on AWS was actually running smoothly for us. But it's

22
00:01:48,700 --> 00:01:52,120
just they couldn't deploy new projects at that point.

23
00:01:52,439 --> 00:01:59,000
Yeah. You mentioned already IAM failing in Virginia. What actually happened there?

24
00:01:59,680 --> 00:02:02,420
So IAM is Identity and Access Management.

25
00:02:02,900 --> 00:02:09,819
So it identifies virtual machines as having access to certain actions, right?

26
00:02:09,939 --> 00:02:14,500
So it gives you the power to interact with, I don't know,

27
00:02:14,560 --> 00:02:18,099
SQS and SQS messages to interact with S3.

28
00:02:18,719 --> 00:02:24,580
So your virtual machines or your ECS containers or your EKS containers,

29
00:02:24,960 --> 00:02:28,280
they need rights to do certain things in the AWS client.

30
00:02:29,680 --> 00:02:30,659
IAM is responsible for that.

31
00:02:30,780 --> 00:02:32,520
It has two parts.

32
00:02:32,639 --> 00:02:39,860
It has the configuration part where you configure which role can do which actions.

33
00:02:40,060 --> 00:02:43,560
So that's the configuring your IAM roles and policies.

34
00:02:43,840 --> 00:02:47,280
And it has a second part when your service is running, it fetches credentials

35
00:02:47,280 --> 00:02:51,280
so you can authenticate to do those actions with S3.

36
00:02:52,379 --> 00:02:56,439
So the global IAM configuration was completely down

37
00:02:56,439 --> 00:02:59,620
because that's hosted in the North Virginia region, in AWS.

38
00:02:59,680 --> 00:03:03,900
However, getting credentials and doing actions,

39
00:03:04,180 --> 00:03:06,740
there are regional endpoints for that.

40
00:03:06,860 --> 00:03:11,919
However, you have to either use an SDK that had a major version,

41
00:03:12,900 --> 00:03:15,919
I think around July 2022 or 2023.

42
00:03:16,060 --> 00:03:19,919
From that point on, every SDK who had a new major version

43
00:03:19,919 --> 00:03:22,000
then used those regional endpoints by default.

44
00:03:22,340 --> 00:03:29,000
But all other older SDKs still use the old North Virginia endpoint.

45
00:03:29,219 --> 00:03:29,659
Yeah.

46
00:03:29,680 --> 00:03:32,340
So you have to configure them to just use the regional endpoints.

47
00:03:32,419 --> 00:03:35,240
And that way you can't do updates to your IAM roles,

48
00:03:35,360 --> 00:03:40,300
but your services can still use other AWS services, right?

49
00:03:40,680 --> 00:03:43,659
So essentially IAM is hosted in Virginia,

50
00:03:43,800 --> 00:03:46,280
and this is for the configuration of your roles.

51
00:03:46,439 --> 00:03:49,659
So let's say I have a data job, it gets a specific role,

52
00:03:50,199 --> 00:03:53,199
which it can use to contact other services.

53
00:03:53,340 --> 00:03:55,460
And that configuration is part of IAM.

54
00:03:55,819 --> 00:03:58,919
My data job has the rights to this path on S3,

55
00:03:58,979 --> 00:03:59,560
or this bucket.

56
00:03:59,680 --> 00:04:00,099
Yes, correct.

57
00:04:01,340 --> 00:04:02,960
And the second thing you explained,

58
00:04:03,719 --> 00:04:08,039
whenever I do those actions and my role needs to get some tokens or credentials,

59
00:04:08,340 --> 00:04:10,479
this is fetched in a different way.

60
00:04:10,560 --> 00:04:12,699
This is not depending on this Virginia.

61
00:04:13,540 --> 00:04:15,819
So from North Virginia,

62
00:04:16,600 --> 00:04:21,899
this configuration is synced to all AWS data centers across the world.

63
00:04:22,040 --> 00:04:24,160
So this is globally synced.

64
00:04:24,180 --> 00:04:27,899
So updates that you were doing, you couldn't do updates,

65
00:04:27,959 --> 00:04:28,740
but even if you could,

66
00:04:28,860 --> 00:04:29,660
they couldn't be synced.

67
00:04:29,680 --> 00:04:31,319
So you're not syncing to the regional locations,

68
00:04:31,540 --> 00:04:35,279
but they still have the latest update and rights that you have.

69
00:04:36,860 --> 00:04:38,819
Okay. And this is why it was still functioning.

70
00:04:38,980 --> 00:04:41,819
But then if you would change your rights in IAM,

71
00:04:42,000 --> 00:04:45,279
that would not be propagated or not happen even.

72
00:04:45,540 --> 00:04:46,120
Correct. Correct.

73
00:04:46,399 --> 00:04:47,759
Yeah. Okay. I see.

74
00:04:47,819 --> 00:04:51,019
And the reason for this scale is, as I understand,

75
00:04:51,240 --> 00:04:54,959
many of AWS services, of course, rely on access management, all of them.

76
00:04:55,040 --> 00:04:58,339
And so you could not change any rights anymore.

77
00:04:59,680 --> 00:05:00,560
So it kept on working,

78
00:05:00,680 --> 00:05:03,519
but no changes anymore to your rights to use these services.

79
00:05:04,139 --> 00:05:10,800
So yeah, the annoying part about that one is that it is very hard to roll out updates

80
00:05:10,800 --> 00:05:16,959
or roll out new versions of your application to fix then the issues that were going on.

81
00:05:17,259 --> 00:05:19,399
So for example, we use Terraform.

82
00:05:19,540 --> 00:05:21,120
When you do a Terraform apply,

83
00:05:21,719 --> 00:05:24,600
well, it will try to see if the role needs to be updated.

84
00:05:24,759 --> 00:05:28,180
So we'll try to check the current configuration of the role.

85
00:05:28,300 --> 00:05:29,120
That wouldn't work.

86
00:05:29,800 --> 00:05:32,720
So depending on your configuration management tool,

87
00:05:32,879 --> 00:05:37,259
you weren't able to roll out a new release even of your application.

88
00:05:37,660 --> 00:05:39,120
Yeah, I see.

89
00:05:39,139 --> 00:05:41,060
But your application could keep on running.

90
00:05:41,680 --> 00:05:47,019
So at Dataminded, our product Conveyor allows you to run data workloads and schedule data workloads.

91
00:05:47,560 --> 00:05:49,500
How did this affect our product?

92
00:05:50,639 --> 00:05:52,180
So there was an outage.

93
00:05:52,319 --> 00:05:59,660
I don't remember when exactly, but years ago with regards to IAM in global.

94
00:05:59,680 --> 00:06:08,139
So at that point, we picked up the best practices of always setting the SDKs to talk to the regional endpoints.

95
00:06:08,680 --> 00:06:11,279
Luckily, this can be done with an environment variable.

96
00:06:11,879 --> 00:06:18,920
And as you know, Conveyor is a scheduler and runner for data applications using containers.

97
00:06:19,680 --> 00:06:21,819
Since it can be set with an environment variable,

98
00:06:22,100 --> 00:06:27,279
we configured an environment variable for all the jobs that our customers run.

99
00:06:27,759 --> 00:06:29,660
So all our customers automatically run the environment variable.

100
00:06:29,680 --> 00:06:32,420
So all our customers automatically have this configuration set correctly.

101
00:06:33,379 --> 00:06:41,180
So that means that most of almost all of the running jobs of our customers could properly run without issues.

102
00:06:42,180 --> 00:06:42,240
Yeah.

103
00:06:43,060 --> 00:06:47,980
Luckily, we also made a conscious decision a couple of years ago

104
00:06:47,980 --> 00:06:53,279
to copy our container images to our regional registry.

105
00:06:55,060 --> 00:06:57,199
So for our AWS customers,

106
00:06:57,579 --> 00:06:58,160
Yeah.

107
00:06:59,680 --> 00:07:03,800
we were also pulling containers from a local ECR registry and not from ECR public.

108
00:07:03,959 --> 00:07:07,379
So they didn't notice any issue there as well.

109
00:07:08,180 --> 00:07:11,259
Apparently during immigration, we forgot one minor service.

110
00:07:11,399 --> 00:07:15,040
So there was a bit of a hindrance there for our users.

111
00:07:16,560 --> 00:07:19,040
But mostly things were running fine.

112
00:07:19,779 --> 00:07:20,019
Yeah.

113
00:07:20,439 --> 00:07:25,959
So things that could have gone wrong were like we launch a data job,

114
00:07:26,180 --> 00:07:27,959
which is a Docker container essentially.

115
00:07:28,180 --> 00:07:28,600
Mm-hmm.

116
00:07:29,680 --> 00:07:31,019
And then you're pulled to your cluster.

117
00:07:31,279 --> 00:07:34,740
And if you would then have a dependency on, let's say, public ECR,

118
00:07:35,259 --> 00:07:38,860
which is hosted in Virginia, then you would not be able to fetch it.

119
00:07:39,699 --> 00:07:43,259
So if I understand correctly, the only fix that you needed to do

120
00:07:43,259 --> 00:07:47,379
was like an outdated container image that you were still referring to in public ECR.

121
00:07:47,620 --> 00:07:49,259
That was for our AWS customers.

122
00:07:49,620 --> 00:07:52,939
So for our Azure customers, we sadly didn't.

123
00:07:52,939 --> 00:07:55,420
There was one, we had a similar issue.

124
00:07:55,500 --> 00:07:59,579
A component, the networking component was pulling, not from public ECR,

125
00:07:59,680 --> 00:08:01,680
that wouldn't make sense, but from K.IO,

126
00:08:01,860 --> 00:08:04,519
which is the Red Hat Container Registry, I believe.

127
00:08:05,079 --> 00:08:05,560
Yeah.

128
00:08:05,779 --> 00:08:08,720
And they also had an outage because it's hosted on AWS.

129
00:08:09,639 --> 00:08:10,120
Yeah.

130
00:08:10,319 --> 00:08:15,319
So on Azure, we're going to do the exact same thing that we already did on AWS.

131
00:08:15,500 --> 00:08:21,439
We're going to ensure that we have every image replicated in the same region

132
00:08:21,439 --> 00:08:28,540
as our clusters are running so that we don't depend on an external service going down.

133
00:08:28,639 --> 00:08:28,779
Yeah.

134
00:08:30,799 --> 00:08:31,279
Okay.

135
00:08:31,899 --> 00:08:34,879
Let's maybe then talk a bit about what people can do to avoid this

136
00:08:34,879 --> 00:08:39,600
or to at least mitigate or mitigate the risk to be affected by an AWS outage.

137
00:08:39,679 --> 00:08:43,240
So what are the things you would recommend people do or take action on?

138
00:08:44,320 --> 00:08:47,899
So the first one is for this US East 1 outage,

139
00:08:47,960 --> 00:08:50,820
it's very clear that IAM is a global service,

140
00:08:52,439 --> 00:08:56,580
but you can mitigate the issue with using those regional endpoints.

141
00:08:57,340 --> 00:08:58,659
And that's rather easy.

142
00:08:59,059 --> 00:08:59,659
Second thing,

143
00:08:59,679 --> 00:09:01,720
you can easily remove that dependency on public ECR.

144
00:09:01,919 --> 00:09:08,639
If you're running on AWS, configure your EKS add-ons to use a local copy of public ECR.

145
00:09:08,799 --> 00:09:10,460
There are multiple ways to achieve this.

146
00:09:10,539 --> 00:09:14,379
You can manually copy them or you can ECR pull through caching.

147
00:09:16,320 --> 00:09:18,279
It's a service that AWS offers.

148
00:09:19,019 --> 00:09:22,960
Does that mean that when you pull in an image that it creates a copy?

149
00:09:23,320 --> 00:09:25,820
Yeah, it creates a copy on the fly.

150
00:09:26,179 --> 00:09:29,360
And the third thing is you also don't want to depend

151
00:09:29,679 --> 00:09:32,460
on external container registries too much.

152
00:09:32,740 --> 00:09:36,980
So again, you can think of this in two phases for the running or the building.

153
00:09:37,759 --> 00:09:42,000
For running, for sure, don't depend on external registries.

154
00:09:42,139 --> 00:09:46,620
We all know the thing that Docker Hub did a couple of years ago,

155
00:09:46,720 --> 00:09:47,720
introduced rate limiting.

156
00:09:48,000 --> 00:09:49,600
And I know why they did it right.

157
00:09:49,740 --> 00:09:51,559
It was costing them, it was a free service.

158
00:09:51,659 --> 00:09:53,379
It was costing them way too much.

159
00:09:53,840 --> 00:09:54,740
It makes sense.

160
00:09:54,960 --> 00:09:59,539
So you want to duplicate those as well to ensure that

161
00:09:59,679 --> 00:10:03,179
when they have an issue, your clusters don't have an issue.

162
00:10:04,480 --> 00:10:10,659
The biggest work is mostly in, if you're using the pull through cache,

163
00:10:10,899 --> 00:10:14,620
you still need to update where you're pulling your image from.

164
00:10:15,720 --> 00:10:17,460
And if you're manually copying them,

165
00:10:17,600 --> 00:10:20,159
you also need to update where you're pulling your image from.

166
00:10:20,299 --> 00:10:24,500
That is actually the most work in our experience.

167
00:10:24,679 --> 00:10:27,000
For every Helm chart you install in your Kubernetes cluster,

168
00:10:27,179 --> 00:10:28,240
you need to update those images.

169
00:10:28,399 --> 00:10:29,480
So that's the most work.

170
00:10:29,480 --> 00:10:29,659
That's the most work.

171
00:10:29,679 --> 00:10:30,320
That's the most work actually.

172
00:10:30,480 --> 00:10:31,000
Okay.

173
00:10:32,979 --> 00:10:35,759
And you had a fourth mitigation, I believe, right?

174
00:10:35,820 --> 00:10:38,159
Fourth one is prepare.

175
00:10:38,539 --> 00:10:41,639
It's basically prepare for outages.

176
00:10:41,700 --> 00:10:42,799
Outages will happen.

177
00:10:42,940 --> 00:10:45,559
And somebody, everybody can make a mistake.

178
00:10:46,300 --> 00:10:48,960
So you need to prepare and you need to ensure.

179
00:10:50,259 --> 00:10:55,000
What we like to do is we have a disaster recovery playbook.

180
00:10:56,059 --> 00:10:59,600
So basically it highlights the steps that we want to do.

181
00:10:59,679 --> 00:11:02,019
As a company, when a disaster strikes.

182
00:11:02,500 --> 00:11:05,220
For us, it's a three step approach.

183
00:11:05,559 --> 00:11:09,639
So basically the first one is the notification and activation phase.

184
00:11:11,020 --> 00:11:16,720
Basically, we notify the correct people that an incident is going on.

185
00:11:16,919 --> 00:11:17,480
Yeah.

186
00:11:17,559 --> 00:11:20,580
That means both internally and externally.

187
00:11:21,100 --> 00:11:23,740
And then we're investigating how bad is the issue?

188
00:11:23,940 --> 00:11:25,080
What is the impact?

189
00:11:26,000 --> 00:11:29,139
Do we need to go to our second phase, which is,

190
00:11:29,720 --> 00:11:32,360
recover to another region in our case?

191
00:11:33,519 --> 00:11:37,340
And the third phase is everything is open over now.

192
00:11:38,159 --> 00:11:40,019
It's called the reconstitution phase.

193
00:11:40,159 --> 00:11:43,539
But basically the simple thing is if you recover to another region,

194
00:11:43,659 --> 00:11:47,240
we need to go back to our primary region because that's where we want to host everything.

195
00:11:47,759 --> 00:11:50,399
And also we write a postmortem.

196
00:11:50,580 --> 00:11:52,480
So we prepare for the future.

197
00:11:53,320 --> 00:11:55,720
And how do you then test all of this?

198
00:11:55,980 --> 00:11:59,659
I mean, you can devise a playbook on disaster recovery.

199
00:11:59,679 --> 00:12:01,659
But how do you really prepare?

200
00:12:02,019 --> 00:12:07,759
That is another very important part of what we do is we test this every couple of months,

201
00:12:07,879 --> 00:12:11,480
minimal once a year using a tabletop exercise.

202
00:12:12,879 --> 00:12:14,159
Do you know D&D?

203
00:12:15,259 --> 00:12:16,240
Dungeons and Dragons?

204
00:12:16,399 --> 00:12:17,059
Dungeons and Dragons.

205
00:12:17,159 --> 00:12:22,159
Yeah, you basically all sit together on a table and you imagine that you're playing in this fantasy world.

206
00:12:22,920 --> 00:12:24,960
Well, this is a lot more boring.

207
00:12:25,019 --> 00:12:29,659
We fantasy that we're playing an issue that is going on.

208
00:12:29,679 --> 00:12:33,259
So we also have some kind of a dungeon master like role.

209
00:12:33,379 --> 00:12:40,620
We call it the tabletop facilitator, which prepared a session and said like,

210
00:12:40,740 --> 00:12:43,539
hey, this is the outage I prepared that is going to happen.

211
00:12:44,320 --> 00:12:45,940
And we're going to role play that.

212
00:12:46,600 --> 00:12:53,419
The idea is that they get the reflex of taking that runbook, following it, making a checklist,

213
00:12:53,559 --> 00:12:54,840
and it becomes a habit.

214
00:12:55,159 --> 00:12:56,460
It becomes easier.

215
00:12:56,700 --> 00:12:59,659
Secondly, it also helps with mental health.

216
00:12:59,679 --> 00:13:08,000
Mentally preparing for when an issue is happening, getting in that mindset of, well, we need to first ensure things are mitigated a bit.

217
00:13:08,100 --> 00:13:10,139
Do we need to go to that reconstitution phase?

218
00:13:11,220 --> 00:13:15,639
That's like a fire exercise that you do and you take some learnings.

219
00:13:15,639 --> 00:13:21,419
And in this case, you do it to prevent or to know how you would respond to an outage and follow the whole procedure.

220
00:13:21,480 --> 00:13:21,679
Right.

221
00:13:23,379 --> 00:13:29,240
So I'll try to summarize what you recommend is indeed a regional endpoints to have less dependencies.

222
00:13:29,279 --> 00:13:29,600
Mm-hmm.

223
00:13:29,679 --> 00:13:30,120
Global service.

224
00:13:30,460 --> 00:13:34,539
Then also the container registry that you want to pull locally.

225
00:13:34,679 --> 00:13:35,279
Mm-hmm.

226
00:13:35,799 --> 00:13:43,179
Thirdly, removing your dependencies on external dependencies like Docker Hub and these kinds of services that also rely indirectly on AWS.

227
00:13:43,620 --> 00:13:48,379
And then finally, the big thing, prepare for these things and do the tabletop exercise.

228
00:13:49,059 --> 00:13:49,340
Mm-hmm.

229
00:13:49,519 --> 00:13:50,120
Yes.

230
00:13:50,299 --> 00:13:50,360
Yeah.

231
00:13:50,779 --> 00:13:51,379
Okay.

232
00:13:51,440 --> 00:13:59,559
If we zoom out, we learned that there were many services or many products being out because more than 140 services of AWS

233
00:13:59,559 --> 00:13:59,659
Mm-hmm.

234
00:13:59,679 --> 00:14:02,980
were affected by the IAM outage.

235
00:14:03,840 --> 00:14:09,779
Then at Dataminded, we already took some mitigation steps before because we learned from the past that these things can go out.

236
00:14:09,940 --> 00:14:10,259
Mm-hmm.

237
00:14:10,259 --> 00:14:14,259
And then you mentioned four recommendations on how to mitigate that.

238
00:14:14,620 --> 00:14:19,919
And I believe you wrote all of them down in a blog post for people who are interested in more information.

239
00:14:20,019 --> 00:14:21,220
So I'll add them to the comments.

240
00:14:22,359 --> 00:14:29,139
And then the only thing remaining, I think, is to thank you, Stijn, for giving us a look into how you dealt with the AWS outage.

241
00:14:29,800 --> 00:14:31,240
And then, of course, to share your insights.

242
00:14:31,340 --> 00:14:32,100
So thanks a lot.

243
00:14:33,080 --> 00:14:35,600
Thank you too, Jonny, for being a great host.

244
00:14:36,300 --> 00:14:36,919
Thank you.

245
00:14:37,220 --> 00:14:37,779
All right.

246
00:14:37,820 --> 00:14:39,139
Thank you, everybody, for watching.

247
00:14:39,299 --> 00:14:41,960
Check out our other videos that we have published in our channel.

248
00:14:42,100 --> 00:14:43,559
And we'll see you next time.

249
00:14:43,779 --> 00:14:44,360
Bye-bye.

250
00:14:45,320 --> 00:14:45,720
Bye.

251
00:14:46,580 --> 00:14:46,980
Bye.