Arjun Iyer (00:00): Some people just merge after running some basic unit tests. That's it, right? They just merge after that. Some engineering organizations do a lot more testing pre-merge. They run integration tests and other kinds of tests as well. Very few run end-to-end system tests. Very few because it's a technical problem that that's what our solution does. Ganesh Datta (00:27): You're listening to Braintrust by Cortex, where we explore how engineering leaders blend AI, platforms, and culture to build high performing software teams. I'm your host, Ganesh Datta, CTO and co-founder of Cortex, an engineering operations platform designed to help organizations continuously improve their operational maturity and reduce developer friction. In each episode, we go deep with CTOs, VPs of engineering, and technical leaders who've been in the trenches navigating the tension between speed and quality, building reliability at scale, and figuring out how to lead through major platform shifts. Whether you're running a team of 10 or a thousand, this is your space to learn from people who've made the hard calls and live to talk about it. Very excited to have you on. For the listeners who don't know, Arjun was part of my batch in Y Combinator back in 2020 when we were very, very naive young startup founders. (01:28): And here we are almost seven years later, lots of things we've learned and obviously building in dev tools, we've seen a lot of things changing and Arjun has been in the validation verification space for software for a very long time. So very excited for this conversation. Arjun, thanks for joining me today. Arjun Iyer (01:45): Oh, it's a pleasure, Ganesh. Yeah, it's been a while and it's great to reconnect and reminisce about our old days at going through YC and talk about what we learned along the way. So pretty excited to be here. Ganesh Datta (01:58): Very excited to hear about... I think you had a very unique perspective on the space having seen humans trying to figure out whether what they're building is working to now whether agents can figure out whether what they're building is working and excited to talk about that transition. So before we dive in, we'd love to hear a bit more about your background and maybe tell us a little bit about Signadot and your journey. Arjun Iyer (02:23): Sure. Yeah. Just for listeners, I'm Arjun Iyer. I'm the CEO and co-founder of Signadot. We are a validation framework for the software development lifecycle for distributed applications. So that's like anytime there's an application with complexity, these include cloud native and otherwise, where there's a lot of dependencies between components that makes it very hard to validate code changes, that's where we help. So that's the core focus of the company and the journey so far. And you're absolutely right. So when we started, there was no AI coding tools, so the focus was on making developers more productive. Shift left a lot of the validation that was happening very late in the process. And now with agents doing all of the coding, the things have changed quite dramatically. It's been a tailwind for us because there's 10X more code now being produced and it's not only engineers that are producing code, it's like marketers and designers and product managers. (03:32): Everybody's producing code. So all of that code needs to be validated before it hits production. So that's where we come in and we make it much easier for all this code to be validated in a much faster and more efficient, cost efficient way, as well as in a more high quality way. So it's kind of three dimensions that we always try to balance, which is cost effectiveness of the validation as well as the scale at which you can validate. Ganesh Datta (04:02): When you say validation, maybe for listeners, if you can expand on that a bit more, both for humans and agents, what does it mean to validate my changes using your platform? Arjun Iyer (04:12): Yeah. When you validate, the answer we are trying to address is, okay, my coding agent has written this piece of code. Can this run in production? That's kind of the main question that we are trying to address. The answer to that question is, okay, how do you validate that this code can run in production? It spans a whole gamut of validations or tests, including just basic unit tests, which is part of the code itself, but there's a lot of things that happen beyond just unit test. Does this work with other services? Are the contracts being adhered to when this code ships or based on the changes that have happened? Is it performant? Is it secure? Does it scale? All the stuff that needs to be answered before I can reliably push this to production. So that's what we mean by validation. And traditionally, all of this happens very late in the software development life cycle and our whole mission is like, okay, can we make it happen as early as possible? (05:17): Because the moment you shift left and the moment you make it happen earlier, the cycles are much faster. It's cheaper to validate, it's more effective, developers love it because you get that fast feedback loop. And so that's the real focus here is to shift left as much of the validation as possible. Ganesh Datta (05:38): So maybe walk me through the end-to-end scenario for, I'm a developer, I'm working in Claude Code, I open up a change. What does that look like if I'm validating that change using your platform? What actually are the series of steps that happens? Arjun Iyer (05:50): Yeah, so our focus is right now mostly on the pre-merge phase, which is basically split into traditionally it's local development and then the PR phase. And typically they're called inner and outer loops typically, but even that is sort of changing with agents playing a part in both. Those are the two phases we really focus on. In the inner loop, it's all about, so agent has written some code. Can I now validate it with the rest of the system? Not only in isolation, which is unit test and everything, that's table stakes that needs to be done anyway, but can I do much more than that? Can I actually test my code end-to-end? I may have changed a couple of backend services and maybe even a front-end service, but can I test the whole thing end-to-end even before I publish a PR? So that's what the local development looks like. (06:49): And then on the PR phase, it's all about really having that extensible regression suite. So I've done my local testing, I'm confident that this feature works, but now I need to make sure that I haven't broken something else, I haven't regressed anything, and all the suite of tests that catch that pass in my PR. And again, the key here is pre-merge validation. (07:13): So your post-merge is kind of, again, a slow path. So that's the traditional path is like, okay, you deploy a code to a staging environment and that's where you run the whole suite of tests. But we are saying that, "Hey, you run it pre-merge and that's how we allow you to do that using the solution that we have." Ganesh Datta (07:31): Got it. So you're saying mocks and unit testing are not enough? Arjun Iyer (07:35): Yeah, it's not enough because they give you a false sense of confidence. It's easy for sure. Mocks do have their place. I do want to make sure that I'm not saying don't do mocks. Mocks do have their place, but they're not enough. So there's a lot of things that go beyond the realm of what you can do with mocks, and you do need a realistic environment where you're actually testing the whole feature or the whole capability end-to-end, including the dependencies. You may have dependencies with other services, with databases, with message queues, with systems like durable workflow systems like Temporal. So it's a very complex system when you have these distributed applications. So it's about testing with everything together that makes it really, really powerful and gives developers and agents the confidence that this code is ready to merge. Ganesh Datta (08:26): That makes sense. When you talk about the validation, you're referencing the point of pre-merge a lot. Why is that point of merge so important? Arjun Iyer (08:35): Yeah, so that's a great question because that's a very pivotal moment in the software development life cycle where we believe that... I wrote about this recently as well, is merge is a contract, right? When you merge code to a trunk branch, what you're telling the engineering organization is, "Hey, this code is ready, this code is ready to go to production and you can actually... Other developers and other teams can actually pull it from trunk, build on top of it, build their own changes on top of it." And so that's why it feels like that is a very pivotal moment where before you merge, unless you validate it thoroughly, you're leaving a lot on the table and it's only you're delaying the inevitable, which is finding things later in the SDLC, which is again a tediously slow process. Developers hate it. When I have to rework my PRs, it's the worst thing ever. (09:33): Developers hate that sort of experience because once I merge my PRs, mentally I'm done. Onto the next thing. I want to go into the next thing. Whereas two days later, if I find a bug that PR led to, now I have to go and debug that and unwind it or maybe fix forward. It's a tedious process because I don't even know if it's my PR that's the faulty one. So troubleshooting itself is a problem because there's so many PRs that go and land into that staging environment and unwinding that whole drama of the root cause of which PR led to that issue on staging is not revealed. We have been there several times. Ganesh Datta (10:14): Yeah, that makes sense. When you say a merge is a contract, that implies some sort of set of things that are invariants at that point, like by merging, I am satisfying the contract that all these things must be true before it makes its way into production. So working backwards from the concept that merging into a trunk branch implies that a certain set of checks are done, implies that the contract is a thing that is checking all of those things. In your experience, do a lot of organizations have a clear contract today of this is exactly what a merge means or are they still in the process of designing that contract and making sure that people adhere by it? Arjun Iyer (10:54): Yeah, it's very ad hoc from what I've seen across the industry. Some people just merge after running some basic unit tests. That's it, right? They just merge after that. Some engineering organizations do a lot more testing pre-merge. They run integration tests and other kinds of tests as well. Very few run end-to-end system tests, very few, because it's a technical problem that that's what our solution does. So it's kind of a gamut of different degrees as to which how much they've done pre-merge validation, and that's what we want to change. Our hypothesis has always been like, "Hey, once you merge, you should be able to go to production pretty soon." That's kind of the contract. And contract, like we talked about, it's not only does this work, but does it perform? Does it scale? Does it work well with other services? Everything. It's like everything is covered in that contract. Ganesh Datta (11:49): You've talked about the four layers of merge confidence, one of them being testing against real systems. If an organization maybe doesn't have as mature of a thought process around a merge contract, where do you recommend people start to think about what the right contract is for them? Arjun Iyer (12:07): So that, again, depends a lot on the application that you're developing and how complex or how distributed the application itself is and where the failure modes are. That typically comes from your historical record of how many production incidents did you discover in the last 10 months? How many bugs did you discover on staging? That would be a great place to start. So you work backwards from there and figure out, okay, where are the bugs actually being discovered? Production issues are the worst, of course, but even bugs that you find in staging are pretty bad because again, it leads to very bad rework of the PRs that I talked about. So that would be a great place to start and get an understanding of like, "Hey, these are the places where we feel like the bulk of the bugs are accumulated and then work backwards and figure out, okay, how could I catch this even doing local development or doing the PR phase?" And that would be a great place to start and then you can build from there. That's what we usually advise our customers is look at what actually is happening, where are your failure modes happening today, and then work backwards. (13:19): And so that's kind of more of a quality angle, but there's also a developer productivity angle as well. Ganesh Datta (13:24): Has it changed with AI coding agents what that contract needs to be or has it made having such a contract more important in a way that wasn't before or is it just amplification of an existing problem? Arjun Iyer (13:36): Yeah, I would say it's an amplification of an existing problem because the amount of code is huge compared to what it was before. And it's very difficult to now not have a process around it because you're going to be like... So what we've seen is a lot of our customers and prospects, they're producing a lot of code, but it's just sitting at the PR phase. It's just waiting for somebody to review it. And manual code review is very hard, so people start using agents for code reviews as well, but that gets them a little bit further, but still it's not enough. Still, they will discover issues on staging or worst case in production. So that's kind of where we see the... Yeah, it's definitely amplification of what happened, but now it needs to be something that if agents can actually self-correct and self-test and self-validate, that we believe is the real ultimate goal here where you're giving agents themselves the capability to run a lot of a wide range of validations and self-verification routines, which allow them to self-correct and allow them to move autonomously through the SDLC. (14:54): So that's kind of how we see this evolving. Ganesh Datta (14:58): Walk me through that a bit more because a listener might hear like, "Oh, you're asking me to do another phase of validation before PRs. We already have enough checks and our CI builds are slow and you're here asking me to do one more thing before I merge something." But what you're describing here, it sounds like actually this may increase your velocity in the PR cycle if you can have the right validation step. How does having this validation phase or this verification loop make things faster? What does that look like in practice? Arjun Iyer (15:27): That's a great question because it basically comes to the gist of it, is that independent deployability. That's kind of the key idea here, which is not a new idea. This has been talked about even in the context of microservices and distributed applications. So essentially when you make a code change, can I push that code change all the way to production without waiting for anybody else? Because that gives me velocity. So I would rather push a small amount of code more frequently rather than do a big amount of code and batch it and do it like a batched testing and then release the waterfall model, which was traditionally how it was done. You would push all of these changes, (16:12): All of this gets deployed into a staging environment, and then you'll have a bug burn down chart that goes on and on and on, and that's a slow way of doing things. So now, and this was always the case, agents have just made it, I would say more imperative. Now without having this, it's going to be a log jam on staging because you're going to generate so much more PRs and it's just going to sit on staging and that burndown is going to take forever because even debugging, if staging fails and you merge a thousand PRs into staging, you have no clue why that staging is broken. So basically it's impossible to debug that and it's going to be a blame game. The front end developer will be like, "No, it's a backend issue. Backend will be like, no, it's some other issue." So there's a hot potato thing going on there. (17:02): Who wants to own this? Who broke staging? That becomes the question of the hour. So I think that's the wrong model. I think the key aspect is if an agent can produce code and can validate that code in a realistic environment and push all the way to production, I think that's where velocity is derived. (17:21): So I think that's going to be a key metric, even for you folks, in the engineering scorecards and other tools, that would be a key metric that would be great for somebody to look at. Ganesh Datta (17:36): That makes a lot of sense. I mean, it's this idea of how do you help an agent close a loop? What was a bottleneck before was, like you said, deploying a staging, validating things, unwinding things. But if you can pull that back earlier in the cycle and give agents the ability to self-validate that their changes are done, you can actually close that entire loop and make it possible to move to a more dark factory or software factory where agents can more autonomously get those changes into production because you have more trust in the validation gates that are in place before making it to production. Is that the right way to think about it? Arjun Iyer (18:09): Yeah, absolutely. I think because now the agents are forcing this function. This was always the ideal goal state. Continuous delivery was always a term that everybody wished for, but it was impossible to achieve that. It was good in principle, but it was impossible to achieve that. But now agents make it imperative. You have to go there because otherwise you're just not going to be able to scale. And if your competitor is able to scale shipping daily because these agents are producing more in software factory fashion, like their own autonomous loops going all the way to production, now there's no way game over. You lose competitive advantage if you can't do that. Ganesh Datta (18:49): Is what you're describing, I guess, starting to blur the lines between the inner loop and the outer loop? Because it used to be like you have the outer loop of validation and CI and things like that, and you have the inner loop of all the things actually writing and delivering that code to the outer loop in the first place. But we're saying that actually those two things are more connected than they were before. Is that the right way to understand it? Are those things coming closer together? Is there as clear of a delineation between inner and outer loop anymore? Arjun Iyer (19:17): Yeah, that line is getting very blurry now and it comes back to the same question. Like I said, previously the model was you ship code into a staging environment and run a regression suite that runs for five hours every night. Especially if you're a big company that has a lot of features and capabilities, the regression suite tends to grow and grow and grow. But now the right model would be don't run 10,000 tests, run the top 50 tests that pertain to the code change that the agent just did. So that's the model that scales. So the agent made a code change and obviously I'm totally advocating for smaller code changes. Don't make a huge code change and that's not the right model. Small code change more frequently shipped, but now this also reduces the scope of tests that you need to run because the scope of change is small and the agent knows what has changed. (20:11): So it can actually pull in, let's say you have an inventory of 10,000 tests, it can pull in the top 50 that makes sense to run because of this code change that's relevant to this code change. So you run that small unit of... And that can be gamut. I'm not saying only run unit test, that's not the model. You run integration tests, end-to-end tests, performance, even security, all the tests, but it relates to the code that has been changed. You're not running a black box thing where no matter what code you touch, you always run the 10,000 number of tests that you have, which is a very bad model. You just waste time. You're just wasting compute and everything just running the same test over where the change has nothing to do with the tests that you're running. So that's kind of the model that we see and you're absolutely right. These tests were usually run in some kind of CI loop or some kind of background loop in a scheduled manner. Now these can be run much earlier. And this goes very well with the agentic autonomous software factory model where agents can actually select even the test because now the intelligence layer allows the agents to select the right tests based on the code changes that it just did. So it can pull in the right subset of tests to run, run it and then self validate. If the tests fail, it goes and fixes the code and then keeps running the test till all the tests pass. And then it's a very clean way, fast way to production. Ganesh Datta (21:36): Do you still recommend that people run the full gamut of tests in their CI environment or basically is it you have your agent run the subset of tests and validations in the local inner loop that matter and then have CI run everything end to end? Or do you recommend that people try to optimize the entire end-to-end cycle itself? Arjun Iyer (21:59): Ideally I would choose the latter because even in the CI, I would say you don't need to run all the whole gamut of tests because that's very inefficient. So I think that's where the intelligence layer of the agents can come into play even in the CI thing. And that's why I feel like the inner and the outer loop sort of merge because it becomes just one thing. It's just a validation layer. It doesn't matter because agents increasingly run in the background. It's not running on my laptop anymore. At least not... It's standing. I can do whatever I want. I can run some agents on my laptop, some agents I just kick off in the background in the cloud. And so I don't even know what they're doing in the sense until they completely produce output. So they need to be independent. So if they're changing code in the background, they themselves should be able to run, choose the subset of tests that they need to validate that piece of code. (22:54): Of course, there's always going to be human oversight. I think we believe in that. It's not going to be completely like, oh, this agent ships something to production and then who's responsible for that? So I think the responsibility always falls on the human. (23:08): So there's always a human that signs off this landing in production, but the agents can go a long way before that human step needs to come in. Ganesh Datta (23:17): Yeah. One of the interesting points you just made is you have this... It's not just humans on your laptop writing code, but it's potentially background agents opening PRs who are ideally validating their changes against the contract that you have. But it sounds like, going back to the earlier point, the merge is really important. Taking a step back from that, the PR is really important. We're talking a lot about the PR as a unit here. You've talked about the PR being more important as a unit of demand than individual developers as the unit of demand. What does that mean? What do you mean by unit of demand and what does it mean for PRs to be more of a focus as you're describing? Arjun Iyer (23:57): Yeah, because now the whole... Previously, the number of developers put an upper bound on the amount of code that would get typically pushed to production on let's say a monthly basis or something like that. So if you were to measure how many PRs landed in production, it would have a direct correlation to the number of developers. Typically, we used to see maybe between five to 10 PRs a developer per month. (24:22): That was what we used to see. But now that math is not valid anymore. We are seeing tiny teams with just 10, 20 people, 10, 20 developers pushing a lot more PRs to production. So now the unit is not developers anymore. It's kind of like how much of the agentic SDLC have you adopted within your organization? The more agentic you are, the more you're doing things using agents, the number of PRs is limited by that. And there's really no limit to that. In the sense I can spin off 10 background agents for the 10 ideas that I had in the last 10 minutes. And so that skyrockets the number of PRs. And I think that's why the PR becomes the most focal point of... That becomes a focal point because that's the unit that you want to really govern. You want to measure, you want to go on, you want to be safe, and you don't want agents willy-nilly producing PRs and pushing them to production. So you need to have that validation really built in and you need to be able to scale. So if you have an engineering team of let's say 50 developers and these 50 developers are running background agents and they're producing, I don't know, thousands of PRs a month, this cannot scale unless you have a platform where you can actually build this validation sort of infrastructure that can scale cost efficiently because that comes back to the solution that we have. But essentially that's kind of the idea is that now if PR is a unit of governance, you need systems or you need infrastructure that can actually scale at the scale of PRs now. So that every unit of code needs to have validation infrastructure and the validation routines themselves that can run now concurrently at scale. (26:21): So that becomes the real, I would say, the guarding factor. Ganesh Datta (26:25): I love that distinction because recently, shameless plug, I released a framework called the Drive framework, which is really focused on how do we bring... Where does the human in the loop go now with more agentic software development? Because we're not reviewing every single PR anymore. It is unsustainable over time. The SDLC is only becoming more and more agentic, and so you kind of have to look at the entire system as an observable unit. And secondarily, a lot of the ways in which people have tried to measure the SDLC in the past have been developer oriented. There's a lot of frameworks out there which are focused on developer productivity frameworks. And when you have background agents capable of opening PRs on automations, for example, we have background agents that look at our flaky tests every week and just open PRs to fix them. Who wrote that PR? (27:12): There's clearly more output from the organization, but when you look at an individual developer, it's not attached to that set of background agents. And so the atomic unit, I totally agree, is no longer the developer itself. The biggest possible atomic unit is the organization. You start at organizational effectiveness, and the smallest atomic unit that we can control is the individual change, like the PR, like you're saying. And that has ramifications on both ends, both on what do you actually measure to understand if your system is healthy and it's not developer metrics. And the corollary there is like, okay, well, if PR throughput is the general then assessment or even for organizations that are slightly more mature, like deploy frequency is also kind of a similar atomic unit in some ways, then it's like, okay, what is the bottleneck that is holding that thing back? And anything that is not focused, it's like Goldratt's theory of constraints. (28:08): Anything that's not focused on the number one bottleneck is not going to have any impact. So you should focus on that key thing. And so if we're saying that, okay, the PR is the atomic unit, you have to make it easier to get that PR through. Trying to make an individual engineer 5% more effective with coding agents is not actually going to lead to any better outcomes. Making it possible for agents to ship 50 times the PRs autonomously is actually going to be the difference maker. Is that the right way of thinking about it? Arjun Iyer (28:32): Yeah, no, absolutely. I think that's kind of where the SDLC is also evolving is I think that it's going to be key how agents and developers or humans, could be even product managers and other personas, how agents work with them in unison. I think that's going to define the next, I would say, the next generation of software development. And the key part is how is that coming to fruition in harmony? Because on one side, you do need the agents to be autonomous and self-healing and self-correctable or being able to correct its own code. On the other side, you need humans to be accountable. So there needs to be that sort of balance there. And that's going to be the next systems and frameworks that attack that factor would be the real way that productivity will be unlocked from a software development point of view. Because what we are seeing right now from our customers and prospects is you're generating so much more code, but it's not all making it to production. (29:39): So the actual productivity is not really being realized. It's just like, okay, more code. But more code is not who wants more code, right? In fact, I want less code, but I want more meaningful changes that actually bring value to my customers, to my end users, and being able to have a system where I can scale this at the agentic level, with humans playing a very core part because I think that finally the accountability will come back to the human. Ganesh Datta (30:10): Yeah. How does that play out in practice? We were talking a lot about giving agents the ability to validate their own changes. And I think in the past it was expected that whether on staging or whether on more PR centric representations of the code changes, humans were in the loop trying to validate things and you would try to automate through end to end tests and things like that. But there was some element of humans trying to design the verification steps. Where do humans live now? If agents are closing the loop on their own, you're giving them the tools to close that loop, what are humans looking at? Why do we need humans in that loop? Arjun Iyer (30:45): So humans, obviously they serve a higher purpose in the sense, even when it comes to coding, I think even deciding what to build, that's the first step where the humans will be most, they will own that because what to build is a very multifaceted question. It's not only a technical question, it's a business question, it's a product question, it's like market question. So you need to really have a well overall understanding of what would drive value for your customers. And I don't think agents are there yet. I don't know. In the future they might be playing a bigger role there, but at least now I feel humans are going to be the ones dictating what to even build and let agents do the grunt work. That's kind of how I see it. Humans are still using their thought process and the evolved maturity and the whole information that they have about the market, the product, the customers, and they tell the agents what to build and then let the agents do the grunt work, build the code. (31:50): And then also in the design, I think humans will be involved, how to build it in a way that is scalable, that's performing. All the best practices of software engineering still needs to come from a human, even though agents will obviously get better at that. And then on the validation side, humans have an intuition about where the cracks could occur. So that can be encoded. And again, use the agents to do the grunt work. I want to write a playwright test or I want to write some kind of API test or a Kafka test that tests my Kafka message flows end to end. I know exactly what to write as a developer, as a human. Let the agent do the grunt work, let them write the test, but I will dictate what to do. And how to do it is of course the agent can, they've gotten pretty good at that. (32:43): So I think that's where the synthesis is and also the collaboration is. And when something fails, for example, when something fails, I think the agents need some kind of way to select, okay, what arsenal of tests do I have at my disposal? And that list of things can come from humans because I have some domain knowledge where I know that, okay, I've encoded, these are the failure points that are most common in my domain, in my technology stack. And so the agents can again benefit from that. So again, I look at this as human agent collaboration Ganesh Datta (33:20): Yeah Arjun Iyer (33:21): That's happening in a very seamless way. And let the agents do what they're good at, which is producing code because that's what they're expert at doing. And let the human focus on higher level thought process, really thinking about what to even build and how to design for scale, how to validate this piece of code that just could have different failure points and then have that final authority to push something to production. I think humans have to have that. Otherwise, I can't say, "Hey, my agent pushed it." And if I have a huge production issue, who's accountable for that? So that accountability I think is always going to come back to the human. So that's why I feel that some people say, "Hey, we don't need developers anymore." I'm like, "No, you do." Ganesh Datta (34:09): 100%. (34:11): And like you said, I think going back, I was talking about the Drive framework, one of the things that I think is really important now with AI is as the development of software becomes more of a black box, it becomes very important to be very explicit about the contract with the customer. The merge contract is basically saying that this code change does what it says it does and what it says it does is meaningful to the customer. We can't lose sight of there's some line between that change and what the customer cares about. And so from a validation standpoint, the human has that context about what the customer cares about. So even if we allow the agent to own the exploration within the constraints, the constraint setting is the human task as well. This is the thing the customer cares about. These are the boundaries of this particular problem. (34:54): And then the agent can go figure out how to validate within that guard. Okay, I'm going to go write 50 playwright tests that validate all the different possibilities within this constraint space. But who defines that constraint? It's the human at the end of the day. That's how I think about that. Arjun Iyer (35:08): Absolutely right. (35:09): I think the human has, like I said, the context, it's a multifaceted context that it's very difficult for agents to have that. Over time, I can see some of that context being fed back to the agents, but still there's going to be a human driving the whole loop. And it's all loops everywhere. Even if there's a production issue, again, it's a loop back. Why did that production issue happen? First diagnose it, of course, root cause it with logs and observability data. And then that opens a PR. But again, that PR has to go through the same validation loop. (35:45): And then you either hot patch production or you maybe have some kind of short-term fix until the long-term fix is figured out. But again, it's a loop from production. Similar to that, you will have a loop from staging as well. And so it's all loops coming all across. And you'll have loops from say a product manager might give you some feedback once the feature is before even merged saying, Hey, this is not what I asked for or this is not what I had in mind. So again, there's a feedback loop coming from the product manager or maybe designers or somebody else or some other stakeholder in the company. And so it's all like how you have this optimized closed loop feedback system that's going to dictate the next generation of SDLC. That's kind of what we believe in. Ganesh Datta (36:34): I totally agree with that. Maybe a last question for you. A lot of organizations are in various points of this journey of agentic software development. If folks are thinking about maybe for the first time trying to close this loop and shift some of this validation further left, what advice do you have for people that are starting out in this journey today? Arjun Iyer (36:55): I think definitely it's look at your existing metrics and look at your existing state of the system. I think that would be the best place to start. Don't design or adopt processes and tools in a vacuum. Look at what your top three problems are today, where your bottleneck is, why aren't you shipping to production fast enough or where are the bugs showing up? Why is the developer productivity so low even though your agents are producing or rather why is the shipping velocity so low even when your agents are producing so much code? And that sort of dictates like, okay, where do you want to invest your time in? But shift left should be a key concern because the agents now, the center of gravity is changing to agents, which is basically like the inner loop. So that's where the real sausage is made, if you will. (37:49): So that's where you want to focus on, really make that loop very highly efficient. Whether you're running locally, you're running agents locally or you're running in the background, make that loop very efficient and also make that PR loop very, very efficient. That would be my advice in terms of getting that loop and start thinking in terms of feedback loops. That would be another advice that I have, which we really strongly believe in as well because it's all feedback loops all over, all the way to production and beyond. So I think once you have a system where you can have these agents close the loop, whether it's a production issue or it's a local issue, as long as the loop is being closed very efficiently, I think you'll be in a very good position. Ganesh Datta (38:35): I love that. Focus on things that actually are impacting your SDLC, focus on closing the loop and it's all about producing leverage. Exactly. So Arjun, thanks so much for joining me again from Signadot if you're interested in this kind of validation and verification framework. So thanks so much for joining me on the podcast today. Arjun Iyer (38:52): Happy to, super excited and thanks so much for the opportunity. Ganesh Datta (39:01): Thanks so much for listening to this episode of Braintrust. If this resonated with you, do me a favor. Share it with another engineering leader who's wrestling with these same challenges. And if you want to continue the conversation or learn more about how we're thinking about engineering operations platforms at Cortex, reach out to us at cortex.io. Thanks for listening and we'll catch you on the next one.