Dan Sadler (00:00): We have this mantra at Rootly that reliability is the product and comes from the technical elements plus the process elements and then hugely on the human and cultural side. Maybe the most impactful of the three is the culture side. How you really think about this thing, how it's talked about internally, again, how you praise and reward it. Ganesh Datta (00:18): What does that mean? What is a culture of reliability? Dan Sadler (00:21): Yeah, it's just people that actually care about this thing. I mean, the culture can be measured for sure through developer experience surveys and that sort of stuff. But really, I don't think you really want to measure culture. You know when your culture is, right? What the culture leads up to and what it produces is a final result. And so that's the thing that we have pretty intense metrics and analytics on. Ganesh Datta (00:47): You're listening to Braintrust by Cortex, where we explore how engineering leaders blend AI, platforms, and culture to build high performing software teams. I'm your host, Ganesh Datta, CTO and co-founder of Cortex, an engineering operations platform designed to help organizations continuously improve their operational maturity and reduce developer friction. In each episode, we go deep with CTOs, VPs of engineering, and technical leaders who've been in the trenches, navigating the tension between speed and quality, building reliability at scale, and figuring out how to lead through major platform shifts. Whether you're running a team of 10 or a thousand, this is your space to learn from people who've made the hard calls and live to talk about it. (01:38): Today, I have on with me, Dan, who leads engineering at Rootly. As you know, a lot of the episodes recently have been talking about reliability and building that kind of culture within your organization. And a lot of what organizations do when it comes to reliability is making sure that when things go wrong, they know what to do about it. But today we have on somebody who actually builds that part of the stack that lets people know, "Hey, something's going wrong, you should probably do something about it. " And I thought we could learn a lot about what it takes to build that kind of system from Dan. Great to have you on. Dan Sadler (02:06): Thank you. Really excited to be here. I can do a quick intro. So my name's Dan. I'm our VP of engineering at Rootly. Rootly is an on-call, an incident AI-driven on-call and incident management solution. So if you think about the two sides of that, the on-call portion is the classic on-call alerts coming in from an observability source, paging on-call engineers based on escalation policies, schedules, really the mechanism that lets people know when something is wrong. And then we start to treat our scope as the full lifecycle of an incident. So from the moment that you get alerted through to the resolution of the instance. So we manage everything in between. After the alert comes in, you get paged. Then we have this whole workflow automation product to help codify the incident response process, which is a massive portion of our product and is actually the thing we built first before we built on call. (02:56): I've been at Rootly for about two and a half years now. Come from a long career before that as a software engineer. And yeah, happy to dig into stuff around reliability, AI. Tons we can talk about. So stoked to be here. Thanks. Ganesh Datta (03:10): Sweet. Well, great to have you on. I thought maybe we could start with from the horse's mouth, what does reliability mean? How would you define reliability? Because obviously that's what we sell to people. Dan Sadler (03:21): This is what we do. This is what we do. Yeah. I mean, how would we define reliability? I guess there's the standard definitions around SLAs and uptime and all of that. And I think to me, reliability is almost more of a cultural thing. The way we treat reliability at Rootly is, yes, obviously we do measure SLAs and uptimes and SLOs and everything that goes into it from a reporting and technical perspective. But we have this mantra at Rootly that reliability is the product. Like you mentioned, I mean, we build a business critical product. Our job, you think about when our product is used, people use Rootly when shit is blowing up on their side. So this thing needs to work. We use our own product, so it needs to work for us too. So reliability to me is much more of a cultural thing. Comes from the technical elements plus the process elements and then hugely on the human and cultural side. (04:13): And that's something that we focus on a lot at Rootly. We really drive it from a culture perspective, from what do we praise and reward internally perspective. And yet that's how we've gone to a place where we can support such massive businesses on our on- call and incident response products. Ganesh Datta (04:29): Yeah, I love that. I love that you talk about the culture because it really is something you have to optimize for within the organization. I like to think about reliability as does your product do what you say it does and that bar may be different for different types of software for a social network versus an on- call product. Really doing what your product says it does could be different, but at the end of the day, that's what reliability means. It could mean quality, it could mean uptime, it could mean latency, all those things kind of fall under the bucket of reliability. But I think underpinning all of that is the culture of like, "Hey, this is a thing we care about. We care about holding ourselves to that bar of doing what we say we're going to do. " Dan Sadler (05:06): Of those three elements, the process side, the technical side and the culture side, and that's sort of the human side. I've always found that the first two are important, but maybe the most impactful of the three is the culture side. How you really think about this thing, how it's talked about internally, again, how you praise and reward it. You just won't get the good results on the process or technical sides if you don't have the right culture and people surrounding it. So anyways, yeah, just figured I'd add that because it's potentially more important on the culture side to predict the final outcome. Ganesh Datta (05:42): What does that mean? What is a culture of reliability, the nuts and bolts of it? Dan Sadler (05:47): Yeah, it's just people that actually care about this thing. I mean, I think it starts with the folks that you hire and it ends with how you praise and reward people. So there's a ton of our internal talk track and stuff that we chat about in all hands that is reliability focused. Again, we are a business critical product, so it's really, really important for us to continue to address those things over and over in all hands and one-on-one, some reviews and all that, that type of thing. I think the way you tell really is from the output. The culture is really a means to an end. The culture can be measured for sure through developer experience surveys and that sort of stuff. But really, I don't think you really want to measure culture. You know when your culture is, right? What the culture leads up to and what it produces is a final result. (06:39): And so that's the thing that we have pretty intense metrics and analytics on. Ganesh Datta (06:43): You guys are a business critical service. The reliability of a business critical service and the practices that go into it are probably, or I don't know, maybe they're not, are different from a organization that is building something that's not mission critical. So what do you guys do differently, if anything at all, that ensures that you hold yourselves to a much higher bar so that your customers can actually rely on you? What does that take? Dan Sadler (07:06): Yeah, for sure. So I think the way I've always thought about this is the nature of our product, being a business critical surface has fundamentally accelerated the timeline on a bunch of the platform reliability things and internal dev tooling reliability things that I don't think would be necessary in an organization that supports a different type of product. I think if you ... We're a Series A company. We've been a very successful and a very fast growing Series A company, but I think if you surveyed engineering leaders at other Series A companies that don't live in this reliability space where it is a business, like a truly, truly business critical product, they probably would not have invested nearly as early as we did in things like incredible automated testing and canary deploys and load testing with a little dose of chaos engineering in there to really stress the system and great observability with automated rollbacks. (08:04): I always knew we would have to invest in those things early at Rootly, but I think I even surprised myself at how early we did because they are incredibly important for us. And there's just a plethora of other products in the world where I think at series A time, you'd be like, "Why are you investing in this? " Go build product, go build features. You don't need any of that stuff. So all that stuff has come much, much earlier for us and has been a big part of our success. Ganesh Datta (08:32): Yeah. I mean, those kinds of practices, I mean, even in our space, we mostly hear customers who are much, much, much larger talking about canary deploys and automate rollbacks. They haven't even gotten that far yet. And do you think it's because reliability is a product feature? It's prioritized just like any other product feature? Is that part of- A hundred percent. Dan Sadler (08:53): I mean, like I said before, we have this mantra at Rootly that we'd repeat. Reliability is the product. It is a core feature. I think regardless of what features you have, you cannot be the best on-call products and reliability product in the market if you are not also the most reliable product. If it goes down when your clients need you, then it doesn't matter what features you have. It just like that goes all by the wayside. So reliability really, really is a feature. I guess Ganesh Datta (09:19): If I'm looking at a week, like a day in the life of an engineer at Rootly, what is happening in the day-to-day to make sure that reliability is top of mind? Do you guys do weekly operational reviews? Are people looking at things like SLOs every day? Are there things that people are doing on a daily basis or weekly basis that creates that for momentum? Dan Sadler (09:38): Couple different cadences for it. So we do have something called a weekly operational review where with our platform team, we go through just how the system performed over the past week, all of our core metrics and things that we're tracking, as well as looking for peaks in the system. When was our highest load across a variety of different areas over the course of the past week? How did our system respond to that? Could it have been better, extracting action items, all that sort of stuff? So that's on a weekly basis, we review with our platform team. And then on a daily basis, we do basically a 15-minute tour of the dashboards. And so we sit with the on-call engineer, and more and more they're doing this themselves these days. But when we started up, we were like, myself and our lead engineer on the platform team and whoever was on call at the time would go through the tour of the 10 dashboards that are most important to us and look at the key metrics. (10:31): And again, with a view to like, is there anything out of the ordinary today? Is there anything that we're not getting paged on and our monitors aren't going off on yet, but it's like an upward trend of something we need to be careful of with a view to extracting ash items and preventing bad things from happening before we get there. So those are the two main ones that allow us to have a steady hand on what exactly is going on and get ahead of things. I think the other thing is we have invested heavily in really good observability and really good monitoring. So these weekly operational reviews and daily tour of the metrics, they're reactive. They're like, we're going in and we're seeing what's happening. And I guess really, maybe that's more on the proactive side. The thing I was getting at is we've invested so heavily in our monitoring setup that I have a really high confidence that we're going to be notified of something that's really a big issue well in advance of it happening. (11:33): So the other piece of that was the more proactive side of like on a longer term time horizon, how can we get ahead of these things? But I also have confidence that if something bad is happening right now, we're going to know about it because we spent a ton of time tuning our monitors and reviewing those. We have a weekly review of alerts and monitors as well. Ganesh Datta (11:50): That's awesome. Well, actually those are two things that I would love to dig into a bit more. As you're saying this, it kind of sounds like you said reliability is the product and therefore these are things that a product manager might do. They might have their user usage dashboards where they're looking every day. It's like, here, there are anomalies and usage. Are there things that are spiking like retention metrics, things like that? And because reliability is a product, the engineers are the "product managers of reliability." And so you're kind of doing similar things where you're looking at, hey, what is the state of the product itself? And that just happens to be observability and reliability type stuff. Dan Sadler (12:26): I mean, this is honestly the coolest thing for me about working at Rootly is we're building this product for ourselves. I'm building this thing for myself. I've been on-call for 15 years. It fucking sucks. There's no way around it. Being on-call sucks. No one wants to be on-call. We get to try to make that better for people. I think we're doing a pretty good job of it. And when it comes to making product decisions and how we move forwards, we don't have massive dependencies on our product managers because we're truly building it for ourselves. We are our own ICP. So the best person to make this decision is the on-call engineer. So that's been a ton of fun for me, a ton of fun for our engineers, helps a ton with hiring too.That's such a strong selling point for engineers. (13:10): You get to build this thing for yourself, you get to use our own product, you build customer empathy because we are our own customer. That cycle, that life cycle of making improvements because you felt that the pain yourself is very strong. It's been a lot of fun. Ganesh Datta (13:26): Yeah. We get kind of the same thing. It's like we're building for developers, building for ourselves. And it gives you a different kind of motivation. And I think that kind of energy is really nice. I want to dive into the operational review that you mentioned, because I know a lot of organizations aspire to do those kind of operational reviews. What do you guys look at specifically?You said, okay, we look at these dashboards and metrics every single week, what exactly are you looking at? Dan Sadler (13:51): So we start with all of our infrastructure components and we look at the key metrics of each of those. So for instance, we'll look at our main database, we'll look at CPU, we'll look at throughput, we'll look at a bunch of other different in the weeds Postgres metrics that would be indicative of something bad happening or the lead up to something bad happening. And for each of our infrastructure components, we'll go through just what those metrics looked like over the week and where did they peak and how did the system respond. And then outside of that, we'll look at load. So we've looked at infrastructure, we'll move a level up to application layer and we'll see like, okay, what were the traffic patterns like over the past week? What kind of spikes did we have? And then within those spikes, how did the product respond? (14:35): How did our various infrastructure components and the actual experience that users had respond? And then we'll move over to alerts. So we'll look at all the monitors that went off over the course of the week and dig in to see what the patterns are. We have a cool LLM thing that we built that ingests all of the alerts over the past week from our API, runs an LLM on it and does some pattern matching and tries to figure out what themes and trends are. So that's been really cool. And something that we're thinking of like productizing HY products, so you don't need to pull it out of the API and build it yourself. So then we go through our monitors that went off and the idea there is to look for action items that A, increase the reliability of the product, B, get rid of bad signals before they become worse signals and C, reduce the pain of the next on- call engineer so that they get paged less. (15:33): Part of it is honestly to make on-call less noisy. And I don't want that to be misinterpreted. The goal is definitely not to like, "Hey, just tune your monitor threshold so it's not noisy anymore." Keep the threshold the way ... I mean, sometimes the thresholds need to be modified, but really it's about finding what is the correct threshold that gives us enough signal. And then if we continually go over that, what actually items do we need to do on the infrastructure side, code side, whatever, in order to keep ourselves safely under that threshold. Do Ganesh Datta (16:00): You guys use SLOs or is that part of this operational review for a separate kind of process? Dan Sadler (16:06): No, no. We go through all our SLOs in that meeting as well. Ganesh Datta (16:10): How do you make sure that the discussion doesn't, I don't know, veer off into weird or quote unquote unproductive things. If you're looking at the same dashboard three weeks in a row and it's like, okay, well, we had the same spike three weeks in a row. We should do something about it. Does it get prioritized? Do you have a specific agenda like, "Hey, this is exactly how we're going to talk about it every single week so that it doesn't just become a, I don't know, let's talk about the problem and we miss getting to the other three metrics type thing." How do you keep the meeting on track and productive? I Dan Sadler (16:41): Think it's nice to let the meeting run wild a little bit. Obviously we want to get out of there on time and have all of it be productive, but I find a lot of the productivity sometimes is in knowing when to let it run wild and then knowing when to reel it back in a little bit. In terms of the talking about the same thing three weeks in a row, if we've talked about the same thing three weeks in a row and we don't have either a prioritized action for it or an engine initiative for it or something, then something is very wrong there. I mean, we do talk about the same thing as week over week, but typically when you dig into them, as long as you've actually prioritized the actions that you pulled out to the last one, typically the nuances of what happened are a little bit different. (17:26): A classic one for us is like, oh, we got a big spike in alerts, we processed 8,000 alerts in a minute over a couple different customers that had spikes, or sometimes it's just a single customer that has a very large spike and we look through how did the system respond after that? And usually it's like it responds quite well because we put a lot of work into this, but there are many things that could be improved through our notification system, which is a series of SQS views and Lambdas and how all that works, things that we can tune there through how it interacts with the different databases that we have and our other data stores. There's just so much to ... Because the system is complex, there's so much to go through and how a single spike in alerts trends versus the system with the different permutations of configuration that different clients can have around how many people are on-call this time, how many alert reps these people have and just the different components of the system that are quite different client to client based on their configuration. (18:27): There's always, despite the fact that the signal looks the same at the beginning, "Oh, it's a spike in alerts." The nuances of it I find are always quite different. And I think one of the things we'd get out of that operational review meeting, even if there are no direct action items, which there always are, but say for instance there weren't, is a much deeper understanding of the system. There's no matter what, even if there's no action, you're like, "Huh, we had this spike for this client that had way more escalation policies than any other client. And on the first level of their escalation policies, these people have 20 people on the first level." That's very rare. That stresses our system in a very different way than someone that doesn't have that and has a massive spike alerts. So no matter what, I think we always learn something about the system, which increases our ability to respond to incidents when they happen and increases our ability to prevent things and debug things faster because we innately know how the system works. Ganesh Datta (19:19): That's super interesting. And sorry, I'm diving super into the weeds on this meeting because I think it's something a lot of companies want to do and then maybe they don't actually get around to it for a variety of reasons. Who is in this meeting? Is this senior leadership? Is it staff engineers? Is it everyone? What is the scope of the actual meeting? Dan Sadler (19:39): It's not anyone from senior year leadership other than myself. It's me. It's our core lead engineers on our platform team, and it's whoever is on-call for the week from our various application layer teams. So it's a collection of engineers. So it ends up being a decent size meeting, but we typically get really good conversation and I find it hasn't gotten big enough that it's distracting and we are definitely getting value out of it. Ganesh Datta (20:06): In that middle layer, you mentioned the bottom layer is infrastructure, and I think that's relatively self-explanatory, the kinds of things you might want to look at. The application layer I think is really interesting. Are you guys looking at aggregated stuff like latency across all APIs or you think about it more from a user journey perspective, like alerts and is it like specific slices of the product that you guys are looking at in terms of metrics and monitors from the application layer, or is it a combination of application metrics and ... Yeah. Dan Sadler (20:34): It's a combination, but I think our mindset, the view has always been more from the user experience side of things, from what is the result in the product. Taking a big step back, I think that a lot of engineers have forgotten that the job is not to write code. The job is to produce business value. And from that perspective, yeah, you got to be in the technical weeds for sure. But if you're not finding a way to equate the thing that you're building or the thing that you're doing back to the business value, the user experience, something product related and have customer empathy for the folks that are using it, even for folks on the platform too. I think that if you're not doing that, then you've lost something along the way. So we try to keep that tie really, really close in many different ways out really, but this is one of them where in that meeting it's ... After we talk about the infrastructure components, when you talk about application layer stuff, everything is equated back to like, okay, so what would a user have experienced in this scenario? (21:34): Was it slow for them? Did they experience errors? Were there retries? Did they get paged? The answer is always yes. Yeah. So equating it back to the product implications and what the user experience would have been like is one of the many ways that we try to keep engineers super close to the implications of their work from a product perspective, from a business perspective, from a revenue perspective, which has always been one of my core tents as an engineer. Ganesh Datta (21:59): I'm guessing your SLOs are probably a mix of the two as well, like user customer focused SLOs and then probably some platform level stuff as well. Dan Sadler (22:06): Yeah. Yeah, correct. Ganesh Datta (22:07): That makes a lot of sense. You mentioned earlier about this feeling of confidence. You guys have confidence in deploying things and making changes. And I think that's a really interesting ... It's an interesting word because confidence is the lack of anxiety. It's like, I'm not worried about making a change and those two things kind of go hand in hand. I'm curious, you guys initially launched your on-call product. How did you make sure that you didn't have anxiety about launching such a critical piece of software? What gave you confidence like, "Hey, this thing is ready to go? Dan Sadler (22:39): " I mean, again, from a product perspective, it's load testing. If we can test this thing under the type of load that we were expecting to receive on day one, which keep in mind was very small. We built on-call as an experiment. Is this thing going to breathe massive life into our business? We're not really sure. Turns out it has ... Obviously can't share revenue metrics, but it had a massive spike in terms of the revenue of our business. It's been really, really cool to watch. So I mean, I would say there wasn't any anxiety. It was definitely an anxious time rolling out a new product. We started slow, so there weren't a ton of clients on our on-call product right off the bat, but we had done some pretty rigorous load testing and again, using it ourselves, can we use this product ourselves? (23:30): Is it good enough for us? Then yeah, we should be able to extend this to other folks. And I think when we started building our on-call product, that is when we really started developing the next level for us of operational maturity and platform maturity. And that gave me a lot of confidence. I won't say that we didn't have it beforehand, we did, but just not nearly to the same extent as we did when we started building our on- call product. And so the things that we've talked about, that operational review and reviewing metrics and just doing things that a mature engineering org does, keep in mind, we were doing this in early series A days, which I don't think you talked to a ton of other engine leaders of early days series A products or series A companies, that's probably not in the wheelhouse at that point. Ganesh Datta (24:21): Definitely not. Dan Sadler (24:23): So doing all those things a lot earlier than we thought we would need to give us a lot of confidence for sure. And then it was like rolling it out, you gain more confidence as you gain more users on the platform. I won't say nothing broke in the first year. Obviously it did, but we learned a ton along the way. We had a great time doing it. We built an incredible team around it, and I couldn't be more happy with where we are now from a product perspective, from a impact on the business and revenue perspective, and from a reliability perspective. Ganesh Datta (24:55): I'm sure a lot of people have a question about, if your own product goes down, what do you guys do? What is your failsafe? People ask, if Slack goes down, do you guys send emails to each other at Slack? What do you guys do? Dan Sadler (25:06): We've put a lot of time and thought into this. So yes, we do have a fallback on-call product that we use and the way that it gets triggered is we have monitors. We put a lot of time into defining monitors that if they go off, would be indicative that Rootly on-call is potentially down. And those monitors go through a fallback paging product, which pages our engineers in a very similar way to our own on-call product. There's a lot of internal debate around, is this the right way to do things? Should we just use another on-call product? But we were really, really bullish on the advantages of dogfooding your own product, which we had always done building our incident response product, which is still a reliability product and reliability is paramount there, but there's a lot less volume and throughput on that product. (25:51): So it's less of real time, must be 4-9's availability where our on-call product is real time. It's always, always, always got to be up. I think it's really tough actually in other verticals where you're not building for engineers, you're not in internal dep tooling in some way to build up excitement with engineers and build up that empathy and build up that knowledge of the product because it's not for them. The best engineers are very, very capable of this, but it's just a lot easier for us because it's for us. Ganesh Datta (26:22): Yeah. You have to work a lot harder as an engineering leader, as a leadership team generally to translate that context of customers and users into developer empathy. Because one of the things that we talk a lot about in engineering leadership is like, how do we give people an internal motivation to work hard? And one of the things we talk about a lot is in a startup, urgency is really, really important on the intensity and the pace and stuff, but you don't want to burn people out. It's a marathon, not a sprint. And so how do you create that culture? And one way to do that is if you can really create a link between the user and then drive that empathy in the engineering team, then people will want to put time and energy into working on it. And obviously if it's a thing that you're building for yourself, it's a lot easier to do that. (27:00): So I think I'm sure that plays a huge role. But you did mention a thing at the very beginning around part of a culture of reliability is rewarding that kind of behavior. So how do you guys reward that kind of behavior? Is it performance reviews? Is it compensation? Is it something else entirely? What does rewarding that behavior look like in practice? Dan Sadler (27:22): It's a combination of a bunch of things. We talk about it a ton in one-on-ones. We talk about it in performance reviews. And I think a big portion of this comes in public facing praise as well. So we have a feedback and kudos channel in Slack that we call things out. I try to be extremely active in there to call out good behaviors when we see them. That's a big way that we, across the entire org, not just on the engineering side, we call it behaviors that we like and keep those rolling. And I think a lot of this gets set at onboarding time as well with really driving in the values that we have as a company and what's important here and how we want people to operate. So there's a lot of work that's been done to define those values and make sure that they're dispersed in a way at onboarding time and reminded in a way throughout anyone's tenure at the company that is thoughtful and keeps people motivated and driven in the same direction. Ganesh Datta (28:24): You're mentioning that you guys talk about this at all hands or kind of org wide stuff like reliability metrics as well. Does that play into people feel like their reliability work is getting shown on the big screen or whatever? Does that play a role? Dan Sadler (28:38): Yeah. Oh yeah, for sure. I mean, we talk about reliability and present a version of the weekly reliability report guard at all hands. And so yeah, I think there is a certain element of seeing your work on the big screen and being proud of that. I think the other thing is we talk a lot about revenue as any company should. And if you take that, if you look at that in a little bit more of a granular way, why do we accrue revenue? How does this happen? Well, we have an amazing sales team, we do, we have a very feature reached product, we do. But because we are such a business critical product, people will not buy our product if we don't have an incredible track record reliability wise. And so I think again, it's about finding ways to tie the impact of people. It's about finding ways to tie people's work to the bottom line impact that it has. (29:34): And that yes, you can do through talking about reliability and the metrics there at all hands and raising the importance of those metrics to be something that we really care about and we present to the full company and all of that. At the end of the day, we talk about revenue a lot too, and there's just no way you get to where we are revenue-wise without being an extremely reliable product. So that's something that we talk a lot about too in engineering is like the ARR is not just a sales number, Sure. ARR is engineering's metric today. Ganesh Datta (30:02): I love that. And it goes back to what you were saying earlier, people lose sight of the fact that engineering, software engineering is about generating business value and ARR is the easiest way to measure. Especially if you're a startup, it's the easiest thing you can measure of, are we actually building things that matter and that people actually want to buy at the end of the day? So I think that's a really good way of putting it. I know we're almost out of time here. I want to wrap up with a topic that is top of mind for everyone, and I talk about this in every episode, the impact of AI on software. And particularly, I'm curious about your take on this because one of the things that we saw in our own data and report released recently is that people are moving faster, they're shipping more code, but reliability is going down. (30:45): Qualities is decreasing. Incidents and regressions are happening relatively more frequently, almost like a one-to-one without how much faster we're moving. And a lot of it I'm sure is because of coding assistance and we're just doing more stuff and we understand less of it over time. What is your stance on AI coding assistance and stuff internally? Do you have different processes around it? Do you even use AI coding assistance? How do you think about the role of AI within your organization? Dan Sadler (31:13): So first off, we see the trend that you just described in all of our customer data, for sure. We see incidents are going up. There's more and more every day. And the only thing that we can reasonably attribute it to is the rise of LLMs and AI coding assistance. My take on it and our take at Rootly is you would be crazy to not use these tools, to not leverage the new technologies that we have at our disposal in order to make one portion of the software delivery pipeline faster. You'd be kind of crazy not to use those. Velocity is very important to us. We need to move fast. We have tools that allow us to do that faster and to write code faster. And so yes, we have had a big push on this. We're using all the different tools. We are continuing to experiment and push the envelope on what is possible there. (32:07): The caveat to all this is that the demands on reliability do not go away. And if anything, they just continue to increase for us. So if you think about what are the mechanisms that you would need to have in place in order to guarantee the same level of reliability in a world where code is being produced much faster in a world where engineers ... I'm not saying now, but maybe a year, two years from now, where engineers maybe don't even understand or have even looked at all of the code that goes out to production. What are the things that you would put in place now in order to prepare for that future? I think the answer is a lot of the stuff that we've already talked about actually. The answer is incredible coverage for automated tests, load testing and a healthy dose of chaos engineering. (32:58): Canary deploys so that we can deploy it to a very small subset of our traffic first with good observability and good automated rollbacks before we roll out to everyone else. So all these tooling things around the writing of code are what allow us to continue to shift reliably and deploy it to production safely in a world of writing code being much faster. The other interesting thing I think here is that if you think about all of the steps in the delivery pipeline, writing code is only a small portion of it. Everyone has a different process, but it's some version of talk to clients, gather requirements, write some specs, get some designs going, talk to engineering, write an architecture document, then you write code. Then you got to test, you got to validate, you got to deploy to maybe you have staging some non-local cloud, non-production environment, then you got to run through CI, then you got to deploy. (33:57): There's like a million things after that you got to observe after the fact, you got to iterate. So we're talking about, we have optimized a very small portion of the pipeline. I named maybe 10 things there. I would argue that coding isn't one tenth of that distribution from a time perspective. There's probably a bit more than that. But even if that time goes to zero, say code is now free, it's ubiquitous, you pick it up, the writing code time has gone to zero. Software delivery is still incredibly slow because we've optimized only a single portion of that pipeline. So that's what the stuff that we're thinking about and working on now is like, how do we optimize both ends of that pipe with incredible CI, lightning fast tests, a bunch of the other stuff that we've already talked about. And then the precursor to writing code is optimizing the whole product and design workflow. (34:43): So that's been really, really interesting for us to think about how do we leverage tools now? How do we maintain reliability in this world where code is being written much faster? And in a future world where I can't say for sure, but I'd say it feels more likely than not that the tools continue to get better and more and more code is produced and engineers get farther and far, but their abstraction layer between them and the code gets farther and farther away. Ganesh Datta (35:08): Yeah, that makes a ton of sense. And I fully agree. It goes back, like you said, everything we were talking about is the idea of confidence. It doesn't matter who's writing the code, are we confident that this quote is going to do what we set it out to do? And if it doesn't do what it does and it has customer impact, are we well set up to mitigate and/or figure out that something has gone wrong? And if we do, then it's okay if we ship something that doesn't work because we're going to catch that very, very quickly. It kind of goes back to the idea of thinking about the customer, like what is the expected impact on the customer? What is the value you want and can we capture that? And maybe over time, the lessons that you shared on this podcast might be relevant to every organization. (35:46): If the feature, quote unquote, feature development becomes commoditized in the coding loop, then maybe reliability is the feature for more organizations as the human feature. It's a thing that more humans are focused on because that is what you can control on the outside of the software development cycle. And so maybe more organizations will have to do this kind of like, "We're going to keep an eye on stuff because we don't really know what's going on in the code and this is the stuff that we can control." Dan Sadler (36:11): I would not be surprised if you look at the YC batches next year to see a bunch of companies trying to optimize the tooling on either sides of writing code. Ganesh Datta (36:23): Yeah. Dan Sadler (36:24): I would not be surprised at all. I don't think my thoughts here are particularly new. I think a lot of people have said the same thing. As the time it takes to write code gets smaller and smaller, the things that matter are on either end of writing code. So how do we optimize those things? Yeah, super interesting to think about. Ganesh Datta (36:43): And careers and companies will be made there, like people who are really good on either side of that equation is going to become more and more important the better you are at that kind of thing. Well, so many lessons here today. And I think the big takeaway was if you're building a product that's mission critical, think about reliability as one of your feature reliability is the product, and that kind of dictates how you build software and how you think about that reliability within the organization. Very interesting. I think a lot of folks have been curious about the internals of an organization that delivers their reliability in some way. And so Dan, it was great to have you on. Thanks so much for coming and sharing all these learnings with the audience. Thanks, Dan Sadler (37:24): Man. Appreciate it. This was fun. Ganesh Datta (37:32): Thanks so much for listening to this episode of Braintrust. If this resonated with you, do me a favor, share it with another engineering leader who's wrestling with these same challenges. And if you want to continue the conversation or learn more about how we're thinking about engineering operations platforms at Cortex, reach out to us at cortex.io. Thanks for listening, and we'll catch you on the next one.