Braintrust by Cortex

Cortex co-founder and CTO Ganesh Datta sits down with Adam Berman, who recently stepped back from VP of Engineering into an individual contributor role at Semgrep, the application security company built on static analysis. Adam and Ganesh dig into how AI coding agents are reshaping both engineering leadership and application security.

Adam shares why he traded running Semgrep's engineering org for hands-on building again, and how AI has changed the math for leaders who want to stay close to the code. He and Ganesh talk through why static analysis is having a resurgence as a cheap, fast guardrail for agents, how security review is shifting from CI into the IDE and even into the design phase, and how agents are getting unnervingly good at chaining together minor vulnerabilities into critical exploits. They close on how much to trust agents with fixing their own security issues, and what it actually takes to go "full dark factory" safely.

What is Braintrust by Cortex?

Candid conversations with the builders shaping the future of engineering.

Braintrust dives into the operational realities of running high-performing engineering organizations, from production readiness and migrations to AI adoption and operational excellence.

Hosted by Ganesh Datta, CTO & Co-founder of Cortex

Adam Berman (00:00):
You can really tell the difference when someone is producing AI slop basically because they just said, Claude, go do something versus when you tell the agent to execute after a few hours or however long it took to land on that really sharp problem you're trying to solve. It's like sometimes you end up with a PR that's like, oh, this is 15,000 line change. You definitely didn't look at this. I don't think anyone has seen this and I don't think this should go to prod, but that's generally not happening.

Ganesh Datta (00:34):
You're listening to Braintrust by Cortex, where we explore how engineering leaders blend AI, platforms, and culture to build high performing software teams. I'm your host, Ganesh Datta, CTO and co-founder of Cortex, an engineering operations platform designed to help organizations continuously improve their operational maturity and reduce developer friction. In each episode, we go deep with CTOs, VPs of engineering, and technical leaders who've been in the trenches navigating the tension between speed and quality, building reliability at scale, and figuring out how to lead through major platform shifts. Whether you're running a team of 10 or a thousand, this is your space to learn from people who've made the hard calls and live to talk about it. Hey Adam, thanks for joining me on the podcast. Very excited to have you on today, obviously as a member of the Braintrust community and we've had a lot of very interesting conversations, especially around AI and security.

(01:36):
So excited to chat more about that today with you. Thanks for joining me.

Adam Berman (01:39):
Yeah, thanks for inviting me. My name's Adam Berman. I'm excited to be on the podcast here. My background, I until recently was the VP of engineering at Semgrep and in February stepped back into an IC role. So we can talk about that a little bit. Semgrep is a static analysis company focused on application security. So we spend a lot of our time thinking about how can we quickly and in a developer friendly way prevent and detect vulnerabilities. And we've talked a lot about AI and security and AI is really changing things there. So we're doing a lot of really fun exploration, which is one of the reasons I got back into being an IC.

Ganesh Datta (02:16):
Excited to have you on. Well, I have a lot of questions about security in AI, but before we dive into that, I would love to hear a little bit more about that transition that you just mentioned from engineering leadership into being an IC. I have been seeing that a lot more. Actually, I just recorded an episode with another guest that went through a similar journey and I think everyone is realizing it's a very unique time to be a software engineer and just be in this industry. And I think a lot of folks like yourself are deciding to maybe get their hands dirty for a little while before they figure out what's next. But yeah, I would love to hear about that decision. What drew you back into being an IC? Obviously you had a very successful career as an engineering leader running engineering at Semgrep.

(02:59):
What prompted this decision? What led to this?

Adam Berman (03:02):
Yeah, I think I'd always been attracted. I mean, I came up through product engineering. I love building products. The thing that gets me really excited is working directly with customers, taking an idea, building basically hand in hand with our customers, that really quick feedback loop. And I got to do that really early on at Semgrep and previous roles and helped to build our product portfolio. And I think like many people who end up as engineering leaders at startups, I had went through the calculation that the highest impact I could have to help Semgrep grow and succeed was to step into an engineering leadership role. I had done engineering leadership with previous companies and to help take the product engineering mindset that I had and help the rest of the company think about product problems the same way. And I'm really proud of what we got to do over two years leading engineering there.

(03:54):
And we got to scale the company and grow the product portfolio. But there came this time where I just felt like this itch. I just needed to get back into building. And that itch was really amplified seeing builders around me getting to move so much faster with the AI tools that are now available where in the past really took a team of three, five, a two pizza team to build any product. And I think to build any product for real, you still need a two pizza team, but that exploration can really be compressed, especially in greenfield work. And I felt this urge, I want to get my hands dirty again. I got to play with the tools a little bit and realize that I am. I looked at my calendar and was like, what can I take off of my plate to do more experimenting?

(04:34):
And at some point I realized, okay, I think I need to swap my order of priorities. Yeah.

Ganesh Datta (04:39):
Were you getting your hands dirty outside? I mean, obviously it's very hard to find time when you're running into your organization to be building, but obviously it's a very drastic shift to go back full time to being an IC. Did you get your hands dirty first and you were like, oh, this is awesome. I need to be doing more of this. Was it more like I'm not getting enough time and so therefore I must make a drastic decision? What was the calculus that went into that? I

Adam Berman (05:01):
Think there were a couple things. One is just the personal itch, which I was trying to satisfy. I play ultimate Frisbee competitively and I was building an app for my team. I was like, how can we manage finances? How can we do our strategy better with an app? I was like, okay, this is not enough. I need to dig into something a little meatier. And then I think the other part was thinking about what are the risks and areas for opportunities for Semgrep as a whole? And I felt like, okay, I feel like what my job is taking me in the direction is making sure that other people are managing that risk appropriately, other people are exploring that appropriately. But it feels like the make or break here is not whether or not we ship everything on time. The make or break here is do we have as many of our best people as possible on some of these open-ended problems that don't have obvious solutions?

(05:50):
And then I felt like, okay, I think once again, my best contribution to Semgrep will be going back in and seeing, can I take one of those open-ended problems that could make or break or that could make the company if it turns out that it pans out and go drive that forward? And I think what I really love doing is failing fast on many of those and having the skill set to fail fast on many of those and then until you find the one that really will take off and then shepherding that forward. And I feel like that became the calculus. It's like, okay, I think someone else could actually do this VPE job better than I could and maybe I can go do the product engineering side in a way that other people maybe at Semgrep couldn't or maybe that other people. How can I go help the product engineering side do it better?

Ganesh Datta (06:35):
Do you feel like that was primarily because of AI that the shift happened for you? I think for me personally, I'm writing a lot more code than I ever have. I think it's easier to do greenfield things than before. And I think in the past as an engineering leader, it was probably not the best use of your time because there's other things you could be doing, but also it was harmful in many ways for the rest of the team because you would ship something and you're not truly accountable to that the way the rest of the team is because they're on call for it and they're the frontline of defense in a lot of ways. And so actually it was more harmful for me to be popping in with random PRs here and there. But the math has kind of shifted a lot now for engineering leaders because of AI in particular.

(07:16):
It's easier for you to gain context in a much deeper way than you weren't able to before. It is possible to put more guardrails in place. It is easier to focus on greenfield things in a way that it wasn't before. And so I think for me personally, AI has been the difference maker in being able to go back and satisfy that itch and have it be a net positive thing and not a net negative thing for the business. Is that how you thought about it as well?

Adam Berman (07:39):
I think some of what you said in the beginning is still true that if I'm going to pop into a stack that I don't know anything about and I'm going to try to ship a product and also I'm not going to be accountable for that product as an engineering leader, I don't think there's any tool chain that's going to make that not feel friction for the team that does own that tool chain. I think what AI does is it both compresses certain things and it amplifies certain things. So the compression is that it allows you to get much further in that greenfield product development process. And I felt like before when I was an engineering leader, I just realized I knew before AI tools, I could not get that far down the realm of building the thing. I think I would really stop at the exploration in terms of talking to other leaders and almost wearing the product hat and going to talk to customers.

(08:26):
But I would hand off the learnings prior to building because I just couldn't get that far in the half hour a day I had to go code. Now in half hour a day, with AI tools, a half hour a day, you can get pretty far if you can prompt things. It's also like it's not just a half hour a day, it's like five minutes between meetings. You can get much further. Now again, I don't feel like I ever got to the point where I could ship what I was building, but I could get to a POC, I could get to something that was demoable. And then I realized, okay, I still, something needs to drive this forward, but it gave me more of the push to say, okay, no one is driving some of these things forward and I feel like I have some of these skills and I really want to go back and push on that.

(09:07):
The amplification part is that okay, if I had just done that a thousand times and I could do that a thousand times and leave a thousand unowned problems for the team to own. And so what I didn't want to end up is that situation.

Ganesh Datta (09:20):
Yeah, absolutely. Yeah, it's interesting. I think in the past, even for very senior ICs, like staff plus engineers, a lot of the role was shifting into leverage through influence, leverage through guidance. And I think at least what we're noticing internally is our most senior engineers are writing more code than ever before. Not necessarily product code, but all the stuff that they maybe would've done if they had a army of people to make the things better for everyone else, they're now able to actually get done. And so the leverage multiplier is a lot higher and it's just a different application of the same thing, I think.

Adam Berman (09:57):
100%. We see we're building some microservice stuff that needs Kafka. And I saw a staff engineer on the team go build the things that he needed to be able to connect the two services with rate limiting and partitions and all the helpers in the language that he needed that didn't exist. But with a couple more prompts, get that to be a shared library that everybody could use. And it's like, okay, what it takes, that might've been the archetype of the staff engineer, that person might've spent a couple weeks doing that. Instead, it's like a couple of days, maybe even just a couple of hours to go build those helper things that maybe more junior engineers don't even know to ask for. And now our microservice connectors are so much better because the staff engineer now has the kind of army of people to say or army of agents to say, "Go build this thing that we all know needs to exist for the next set of things."

Ganesh Datta (10:45):
Yeah. And one of the jobs of a staff engineer, one of the characteristics was leading through influence. Even if you had built that library, it would've been then kind of a road show to get people to be like, "Hey, let's go adopt this thing, migrate off of the current way you're doing thing to this new thing because it's better in all these ways." And then maybe you would try to do some of it yourself, but you're hoping that other parts of the organization would do it. But now it's almost more addicting for folks like that where they got into that role because they like to make things better for everyone,

Adam Berman (11:12):
Yeah.

Ganesh Datta (11:13):
And they can just do it. I'm just going to migrate everyone's repos to this new library because it's so easy for me to just send off a prompt to do it. And I think people will need to maybe find their balance now with that kind of thing.

Adam Berman (11:24):
Totally.

Ganesh Datta (11:24):
But it is so much easier to amplify that kind of influence across the organization, make things better for everyone, thereby kind of feeding into the loop in the rest of the organization.

Adam Berman (11:34):
100%.

Ganesh Datta (11:35):
I think these are probably the things that made it exciting for you to go back into that hands-on role.

Adam Berman (11:39):
Exactly. And I think the part that is not replaced, and I keep hearing people say this and it really resonates with me, is taste is not getting replaced. So knowing that this is the pattern we want to follow, that's what we're really relying on. And then okay, so he can go edit the read mes of things to say, okay, this is the pattern we should follow. And then junior engineers when they're using Claude can go do that. And I think the part that we need to then connect is like, okay, I don't want junior folks to just blindly follow the ReadMe. They need to build their own taste. They need to know when to adopt this pattern or when to say this is not the right pattern and we should go do this other thing. And I think that is the part we're all, I think as an industry kind of grapple with, which is how do we help our juniors really grow?

(12:16):
I think our product is really technical, so a lot of the areas of growth we see are on the product engineering axis where they have to go explore how will we even do this thing? And you can go talk about it with Claude, but you need to be the one who stands behind your plan.

Ganesh Datta (12:31):
Yeah, 100%. I was just saying even with the most frontier of the models like Fable, I'm finding that it is very good once you give it a direction at executing against that direction, it's good at coming up with pros and cons for architectures, but it is not good at determining this is the right way to build a thing.

Adam Berman (12:50):
Totally.

Ganesh Datta (12:50):
And if I'm able to shape it with Claude and say, oh, actually we need to consider X, Y, and Z, and this architecture is over complicated over here or whatnot, and then let it kind of go off and do its thing. It's great going off and doing its thing and it will stick to the plan and it'll do it really well and find issues and whatnot. But sometimes it makes silly decisions and I'm sure it'll get better at that over time. But

(13:12):
Even if they're very long horizon, I think a lot of that context still is with humans and we're able to shape that in different ways. And so to the point about junior engineers and changing the industry, I do think there is an element of the role is just different. We would get people, new grads would come in and they would do bug fixes and stuff like that for their first month. And that was not because they were only capable of doing bugs, but it was a great way of helping them learn the skills of understanding a code base and understanding context for a thing and shipping it and taking it through the entire life cycle. And maybe some of those things are not as relevant. Maybe exploring a code base in the way we would do it is not relevant the same way it is now.

(13:50):
And so maybe the things that they will pick up early on in their career are just different things and we will naturally gravitate towards training them on those things. Architectural principles may just come earlier in the life cycle. They didn't need to learn those things before because you get a lot of value out of a junior engineer just shipping tickets, but now that's not the case. And so we're just push them ahead. It's not that they can't learn those things.

Adam Berman (14:11):
Yeah.

Ganesh Datta (14:11):
You just didn't need to at the very beginning, but we've speed ran that part of the life cycle.

Adam Berman (14:16):
I think that's exactly right. I think though we see junior folks, our junior folks are not blindly trusting what the AI agent is giving them. In fact, we see them instead using the AI agent to explore the code base, I think in a really different way than you and I explored the code base many years ago before these tools existed where you really had to go line by. I remember seeing a coworker, this was the other end of the spectrum at an old job, literally print out the code base and then read through. I was like, okay, I've never seen anyone do that before. But we had to go line by line. You had to understand what each function was doing. Now you don't need to understand what each function is doing necessarily, but you do need to understand how the building blocks fit together.

(14:55):
And I see the junior folks using the tools to do that.

(15:01):
I think the problem solving part, I don't know, you've probably seen you can really tell the difference when someone is producing AI slop basically because they just said, Claude, go do something versus when you tell Claude or you tell the agent to execute after a few hours or however long it took to land on that really sharp problem you're trying to solve. It's like sometimes you end up with a PR that's like, oh, this is a 15,000 line change. You definitely didn't look at this. I don't think anyone has seen this and I don't think this should go to prod, but that's generally not happening.

Ganesh Datta (15:41):
Yeah, exactly. Yeah. And to that point, I think what I'm noticing with something like Fable is I feel less inclined to read all the code it's producing at the end and I can trust the model to go and build a thing only if I have spec'd it out really well upfront. And I think in the past, people talked about spec first development, but you would still have to review the end and make sure did you actually do the thing you said you were going to do? I think now it is going to do the thing it said it was going to do. It's gotten a lot better at that, but is the thing that it said it was going to do the thing that you wanted it to do?

Adam Berman (16:15):
Yeah.

Ganesh Datta (16:15):
So I think defining that upfront kind of changes it and we're seeing that in the code review life cycle and whatnot. And actually I think that's a great segue into what you were talking about earlier, which is the security aspect of the software development life cycle. There's a lot of talk about like, oh, we need to think about security for agents and security for code written by agents. What do people mean by that? Is the things that we need to secure, is it changing? Is it something entirely different? Is it just there's more of it and so the way we need to treat it is different? When we talk about security for the AI era, what do people mean by that?

Adam Berman (16:52):
Yeah, I think there's some things that have changed and some things that I think have stayed the same but are amplified. So one thing that has changed is you don't know what your model is going to do when you tell it what to do. And if you give your model permissions to do computer use or anything on your computer, it could act in a way that you wouldn't expect. We had Dr. Katie as a security researcher at Semgrep and she's talking at DEFCON, I think at DEFCON or maybe at Black Hat in Vegas in a few weeks about how easy it was to poison a model to, I think she said she did it with a hundred bucks to be able to release a model that had a backdoor that would act in ways that you would not expect it to, that would act in potentially a malicious way.

(17:35):
But it's basically like it's impossible to detect unless you know already how it is backdoored. I don't think you can really test it in the same way that you could even fuzz software. So I think that part is a totally different threat model that we haven't really thought about because it is such a black box. But in terms of the production of code itself, I think it is just amplifying. We talked a lot in the early days of Semgrep that the reason we built static analysis to solve the security problem is because you can pinpoint this right part of the triangle between fast enough that doesn't get in your way, understandable enough that you can do something about it or you can fine tune it and then powerful enough that it can actually find things. You want to find the right spot between those three cases.

(18:21):
And with static analysis, we felt like with Semgrep, the open source product and then the commercial product, we felt like we could do that. But fast enough, powerful enough, these things are getting what is fast enough, what is powerful enough, what is understandable enough and understandable enough to whom is those things are changing. And so now when we're thinking about where does Semgrep sit, it used to be in CI because that's where you got your code review. That's where you got your security review if you got a security review at all, which not everyone does. And so you expected that that was the first point in the back and forth. Maybe you are an outstanding developer and you got a design review ahead of time, but for most people, CI was the first point in which anything was going to check their code outside of the squiggly line in your IDE.

(19:05):
But now the IDE is where everything is happening and the IDE isn't an IDE anymore, it's an agent. So when we talk about shipping left, it's always our goal. Can we bring Semgrep? Can we bring static analysis? Can you bring your tool chain into the IDE? And that was harder and harder, especially given real time typing latency. But another thing that's changed is what used to be you need to be under 100 milliseconds or whatever it was to make developers feel like real time when they would type. Now, if there is a half second of latency, that's kind of built into the agent, which means you can do so much more as the agent is building. And so that's what we're trying to think about is what needs to shift all the way to the left within the agent so that no code is produced basically until you've had your back and forth with the agent.

(19:51):
And then what needs to go further to the right because it's going to take so much longer because there's all these deeper security issues. You said you've used Fable, it does take a long time. If you had Fable review every time you give a prompt to an agent, I think you'd get pretty sick and tired of how slow that would be. Right.

Ganesh Datta (20:13):
Okay.

Adam Berman (20:14):
Totally. But also we hear stories about people just letting Fable run. I think there was a security researcher from Anthropic who talked about, they just said we're going to run Mythos in a loop and we were able to find a ton of open source vulnerabilities. Okay, but that loop takes a long time, but if you can find those things, that's worthwhile in a way that used to be gated by either domain knowledge or cost or probably both.

Ganesh Datta (20:40):
Yeah, absolutely. Yeah, it's interesting. I think you're describing two different types of issues at different ends of the spectrum. The note, I was reading an article this morning about one of the kind of Fable class models that OpenAI had deployed internally that one of the interesting thing about agents with longer horizons is that they're more persistent about solving the problems, which means they will try to find ways to solve their problems that may not be expected. And so there was this thing where they were trying to do some basic RL work and somewhere in the instruction that said, "Oh, open a PR against GitHub." But the sandbox I was running in did not have network access. And so it was like, "I need to open a PR." And so it started trying to figure out a way out of the sandbox so that it could open a PR.

(21:22):
Totally. And eventually it figured it out and it was like, oh shit, it can find those really interesting capabilities. And it's not trying to find the vulnerability, it's trying to do the thing you asked it to do. But while doing that, it's having all these unintended consequences.

Adam Berman (21:36):
Totally.

Ganesh Datta (21:37):
So that's kind of an area of security that, like you said, is much more expensive and much more time consuming, but it is a lot cheaper now to, because you can embed these best practices were talking about earlier with staff engineers being able to go and embed their best practices into the systems and the development life cycle much more easily in the same way you can embed standard security practices early in the life cycle as well. Is that where you see static analysis maybe, for lack of a better word, coming back in vogue? It is so much easier to just give an agent a hammer. It's like, "This is just wrong. Don't do this thing. I found the more hammers I can give my local setup, the better it is at finding the right path." Is that how you think about static analysis from a security shift life standpoint?

Adam Berman (22:23):
I think we think about this from, I think, two different perspectives. One is cost, one is speed. And those intersect as well. So one example is if it takes you two hours to run, say Fable against your code base, and part of it is the persistence of knowing I need to find everything. Well, a lot of that is going to be finding things you may already know about. Also finding patterns that are not super subtle or complex or that are easily findable with static analysis. And so we think static analysis can solve many of these problems. The clearest example is I spend a lot of my time thinking about supply chain security and most of supply chain security is pretty dead simple. It's like read a lock file, see what dependencies you have in a lock file, compare that lock file against databases of known vulnerabilities.

(23:13):
And you can have Fable go read a lock file. It's either going to be incomplete or it's going to write its own program to do it. Or we can give you a static analysis tool that reads the lock file and produces the set and then dips against the known database and that will be both more correct and faster and way cheaper. There are ways in which obviously these AI tools make some of these things much, much easier. One of the things that we've been thinking a lot about is breaking change detection. You want to upgrade some dependency and doing so could or could not break your whole code base. And most of us don't have the test suite to be able to actually verify that. It's like, okay, a test suite is basically static analysis. If I could have a static analysis tool that'd be able to tell all of the different ways that this package is used, that would be great.

(24:03):
And so what we do is we use static analysis and AI tools to solve this, which we use static analysis to figure out what's the difference between these two versions of package. And then we use AI to go say, okay, and will any of this break basically between these two versions? So these two things together make sense and also make things a lot faster. The other part is the cost aspect. We keep hearing about companies that are like six months ago we're putting token leaderboards in place to say, can you just spend more tokens probably as a way to get their developers who might have been hesitant to start using more AI tools? And I get that, but obviously perverse incentives create perverse outcomes and we see people using AI for things that they probably shouldn't. And now we're seeing the opposite side, there's a lot of spending limits on can we figure out ways to not use AI for these things?

(24:57):
And static analysis just uses normal CPUs. We don't use GPUs. So it becomes much faster and then also much more cost efficient to be able to say there are categories of things for which static analysis is either better or just as good, but in order of magnitude cheaper. And you can think of this, a really basic static analysis security rule is like, do all your routes have authentication? You can use Fable to go check for all those things. Also, you can write a SEM grip rule or a rule in whatever you're even just a linter rule to say, does all my routes have authentication? And if so, you don't have to spend your Fable tokens on that.

Ganesh Datta (25:37):
Yeah. Interesting. So I guess we've been talking about shifting security left for God knows how long. Yeah, forever. But it sounds like static analysis might be, especially now with coding agents where you can embed those practices into the actual loop of writing software before PR, it's become a lot easier. And so is that the furthest left we can go now? It's truly in the loop, semi-instant verification of things that we know are best practices from a security standpoint. Is it even possible to go any further left than that or is that kind of the limit that we're at?

Adam Berman (26:13):
I think there's probably one step further left and I think it's also a little bit like horseshoe theory where once you get all the way to the left, you end up kind of on the right again where you. Okay, so the agent harnesses have done a pretty good job of making them extensible in terms of the act of producing the code. Probably one step further to the left is getting it into the design review itself. And one of the things that we've been working on, I know a lot of other people are working on is the concept of a context engine, something that can produce the context of what is your threat model, what are the ways in which your system works together? And before you write any code, can that get into the planning process so you can design with security in mind?

(26:54):
I can't imagine, at least for now, further left than thinking about the issues, thinking about the problem I'm trying to solve, but that is the kind of thing that takes into account these more longer running things where every time you deploy a system or every time your system changes in a meaningful way, or I don't know, maybe weekly because you just expect, okay, we only want to run this once a week because it's expensive, trying to take a snapshot of your whole system and then being able to understand and reason about that system and then build that knowledge back into, okay, now we're trying to solve a specific sharp problem. What is the SEMgrep way? What is the Cortex way? What is our company engineering way of solving this problem that also fits with our threat model and our guarantees and stuff like that?

Ganesh Datta (27:39):
I'm guessing for a lot of organizations, it is now easier than ever before to turn things that may have been design or review stage verification things and turn those into static rules in a way that maybe was much more difficult before. Is that a safe assumption?

Adam Berman (27:53):
Yeah. I think it's funny, we always talked about SEMGRp, the reason it got popular was because it was possible to write rules. It took these systems. Static analysis used to be though as black box. You needed a PhD to be able to understand how to customize. And the idea was that Semgrep was way easier. It looks like the code you're writing, so you can just amend the rules. And a lot of customers are really excited about it, but it turns out they also have a million other things to do and people don't have the time to go polish and tailor and fine tune these rules. But what used to be an hour is now often like 30 seconds. It's like, okay, I have this rule, it created this false positive, changed the rule so it doesn't give me this false positive anymore. That it'll improve the rule.

(28:33):
Or also you get from your bug bounty program like, "Hey, we saw this vulnerability come through. This was real. Give me a Semgrep rule so I never see this pattern in the code base again." And it's really quite good at writing rules because also it's trained on our open source. Our rules are open source, so it's trained on them. So that's been a huge, I think, advantage too in terms of where do you spend your AI tokens versus where do you spend just pure CPU is you can use these Fable-class, Mythos-class models to go find novel exploits and then you can understand what those look like and then you can use something much cheaper to make sure you never see that pattern ever again. Semgrep ends up being a pretty good tool for that.

Ganesh Datta (29:14):
Yeah. When we talk about security for coding agents and things like that, is it mostly. You touched on this a little bit earlier. Is it just the volume and the amplification that is making this harder? Are there new types of vulnerabilities that we need to keep an eye out for? Yeah, I guess where is that split?

Adam Berman (29:33):
Yeah, mostly we don't see a ton of brand new types of vulnerabilities. What we've seen Mythos and others be really good at is chaining vulnerabilities together in a way that a human researcher might take years to figure out how to chain these vulnerabilities together. And so it's like chaining together these techniques and able to do it in a really sophisticated and nuanced ways. One of the things we've seen is it was able to chain together a few mediums in a way that produced a critical, and it used Be kind of known, not said out loud, but known that you just ignore your lows and your medium severity issues. And now you're thinking like, okay, but if there is someone that has access to Fable or Mythos out there and they're able to get around the guardrails, could they exploit our software if they chain together a few of these mediums?

(30:16):
So I think that is changing. So I don't think there's necessarily novel classes, but I think what used to be. I think the barrier to entry for a lot of attackers used to be not money, but subject matter expertise. And that expertise has just become commoditized.

Ganesh Datta (30:36):
On a similar note, I guess there's this concept of a lot of security has been, like you said, human verification and we've tried to codify as much of that we can into static analysis type things. But if we're shifting the stuff further left, is it okay for agents. Well, not that is it okay, but are we as humans okay with agents fixing their own vulnerabilities? Something is flagged. Do you want an agent to just pick that up and fix it?

Adam Berman (31:06):
Yeah.

Ganesh Datta (31:06):
Or is it like, we found a problem, so therefore it needs to go to a human. How do you think about where agents fit into this whole loop now?

Adam Berman (31:13):
Yeah, that's a great question. And I don't have a here's the right way to do things answer. I think for a lot of companies it's your own risk tolerance. But I think in the past, a lot of security teams felt like cost centers, that it was like, okay, how much money do we want to spend on security that we're not going to spend on product development? And so that was like, okay, that meant the relationship between security and engineering was often security would flag something and then ask for time, which is also money from the engineering team to say, instead of doing product development, instead of achieving your other goals, can you fix this problem? And that felt like it was a drain. I think agents have the ability to at least make it feel like less of a drain because you can at least take a shot at fixing it and you can push maybe the human verification one stage later to say, okay, can we at least verify that nothing has broken and that we have fixed the vulnerability itself as opposed to requiring a human to pick it up from the beginning.

(32:11):
We see a lot of our customers, it's like the detection is just one piece and it used to be 10 years ago, a lot of AppSec was just, can you find more things than anybody else and who's got the best database of whatever? I think that kind of got eroded when people realized, wow, there's so much noise. We did things like reachability, other people do things like exploitability verification, things like that to say, okay, can we focus on the things that really matter? I think there is a push now even further to say, okay, beyond just focusing on the things that matter, can we get ourselves all the way towards fixing them? Can we think about the whole workflow of the on-call security engineer? Their page when they find a vulnerability, they have to go verify the vulnerability is actually real and not just something that whatever the tools spit out.

(32:56):
Can we verify that it is a part of our threat model that we actually care about? Sometimes it's like, okay, yes, this is a vulnerability, but it's also five layers beneath some defense and depth that we don't actually care about this part of the system in a way that makes sense to put time into. Okay, now what does the fix look like? Is the fix something that's going to require more work? Can we verify that it actually will fix the thing once it's out there? It's like, okay, can we push agents to get us further and further down that workflow so that the job of the AppSec on-call engineer is basically like, oh, I saw a vulnerability came through. Can I verify it's fixed now?

Ganesh Datta (33:30):
I think, and maybe this is not true for a lot of security first companies, but for a lot of organizations, they're probably not doing any of that today at all. So even if there is some sort of false negative, false positive rate, doing any of this stuff to some degree, even if fully automated by agents, seems like they would be in a better place generally than they were before.

Adam Berman (33:49):
Totally.

Ganesh Datta (33:51):
Especially if you can automate, like you said, the mediums and the lows, generally pretty easy things to fix for the most part. It was a trade-off decision between do we want to do the thing, do we not want to do the thing? So now it's much cheaper. So yes, there's a question of do we trust, do we verify? But it's probably just better to fix those things and you're probably not going to do that anyway. So letting agents do it is probably.

Adam Berman (34:13):
Totally. I think about it's not that dissimilar from autonomous vehicles where it is really scary to make the decision that you're going to give the keys to the car to something that you don't have the same control over, which is I just came off a flight yesterday. I was like, well, I mean, I do that every day when I fly. I'm not piloting the plane. I'm trusting that this person who has spent a lifetime gaining this expertise has done it. And in a certain way, a Waymo has spent a lifetime figuring this out. And I also bike around the city. I think to myself, well, I know it's always looking. I feel safer as a cyclist or as a pedestrian. I also see its guarantees that it's always going to drive within this rate of the speed limit or that it's going to take the safer choice.

(34:55):
Whereas many human drivers are not always taking those choices. And so I think to myself, what is the way that you can gain that confidence for yourself that the agent isn't going to just totally go off the deep end? And for those of us who are using models from the beginning, they sometimes did. GPT-3 is a really different experience than GPT-5. They hallucinated much more frequently. They failed much more frequently. They just got stuck.

(35:25):
There's a class of vulnerability that happens more often is dependency hallucination, just these attacks where it's like, okay, you can name a dependency something that you think a agent might reach for and put something malicious there. Then the agent might hallucinate that, oh, this thing exists and pull it down and then you're got. But the agents don't do that nearly as much anymore and we're able to build guardrails. We're building a malware firewall. So there's things you can do that put guardrails around your agents to help you verify what the agent is going to be doing that can help you build your confidence over time that okay, if you can get there, this whole class of work that you're either not doing or that you really wish you weren't doing can be handed off. Maybe it'll get it right 90% of the time and in 10% of the time, can you build that observability and to say, "Hey, this doesn't seem quite right.

(36:18):
We're going to stop it from going all the way to production or whatever." Or maybe you can do observability in production that says, "Oh, this is not quite right. Let's roll it back or whatever." But you're also nine times out of 10 getting the security improvement that you wouldn't otherwise have gotten.

Ganesh Datta (36:31):
Yeah. The idea of trust and confidence I think is really interesting. I had an episode recently with Martin who's an engineering leader at Tealium and we were talking about the difference between confidence and trust. Those are two unique things.You might have confidence in your tools, but do you still trust the outcomes? Those are slightly different. I guess I know we're coming up on time and so maybe a last quick topic here. On the idea of trust and confidence, do you think about code that agents are writing differently from code that humans are writing from a trust/confidence standpoint? Or is it all the same from a security standpoint?

Adam Berman (37:08):
I think if you'd asked me two years ago, I would've said absolutely, this code seems. Or even a year ago, the code that's coming out of agents seems like either poor quality or it gets hallucinations much more frequently. I think today the quality of the code is much almost indistinguishable or maybe it's even better in the sense that it gets more edge cases than if a human were writing it. I think the problem is more like did you describe the problem well enough to the agents that it will solve it all the way through? And so in that sense, I have less confidence that an agent is going to attack the exact slice of the problem I asked it to than if I gave it to a human, especially because the human will raise their hand and say, "I'm not exactly sure what's going on here or I'm not sure if I got this right." Whereas the agent often acts with a lot of confidence, even if it is confidently wrong.

(37:58):
And so I think that the question is how do you go and either give it the guardrails to say focus on this part of the thing or do the evaluation to know that what you're building or how you're building is focused on the right kind of the problem. We've talked before, we built a internal code review tool called SEER and we backtested that against a bunch of incidents that we'd had to give us the kind of confidence, is this even finding the problems that we want it to be finding? And it overwhelmingly did, but also it was really good signal to see what are the things that it doesn't find because now we know how do we need to improve it so that we can try to detect those incidents as well.

Ganesh Datta (38:36):
Yeah. Maybe a last question. A lot of organizations are trying to move to a world where not just code reviews, but more and more of their SDLC is automated and more autonomous. There are a lot of precursors to that. People talk a lot about harness engineering and verification loops and all that kind of stuff, but security is obviously a very, very big part of that. If you were giving an engineering leader advice on if they were asking, "Hey, are we ready to go full dark factory on this thing? From a security standpoint, where should they start? Is it shifting more left? Is it adopting things like static analysis? What advice do you have people who want to go more and more autonomous with their SDLC?

Adam Berman (39:13):
I think it's to shift. I mean, you're going to hear everybody say it's to shift left, but I think it's to shift left and to shift right. Because I think if you're stepping out of the code creation process yourself, you want to then have something you can go back to regularly to not just inspect or have something in the loop every single diff, but to look at how all these dips stack together over the course of a while and to be able to shift your course a little bit if you feel like you're moving in a direction that feels anathema to where your stated goals are from a security perspective. I think we all want to live in this world where we can spend infinite money on security and security is perfect and it turns out security is always a trade off, but building something that's 100% secure, the only thing that's 100% secure is no code at all.

(39:57):
It's like as soon as we're building code, we know that there are potential security issues. So the question is, can you give yourself the visibility to not just to say, oh, well it reviewed these things and found these things, but visibility into the way the system now operates and to say, okay, this is moving in a more secure direction or hey, there's this flag or something like that. I think the reality we'll have to come to terms with is security issues are black swan events and we won't know for sure, but the more visibility you can put into the system, the better off and the more likely you'll get comfortable with it and get comfortable saying the agent can do more.

Ganesh Datta (40:34):
Yeah, absolutely. I mean, shameless plug just because you gave me an opening there, but the drive framework that I authored recently is very explicitly designed for that purpose, which is we are not in the loop as much anymore as humans. And so how do we create an outside of the loop loop where we're introspecting the system on a recurring basis and saying, are things going the way we expect things to be going and are things healthy, including security as one of the key pillars in the drive framework? Because like you said, it is a risk vector that is becoming easier to exploit and there's a lot more volume which then creates that forward pressure on security. And so being able to take a step back and say, yes, we're meeting vulnerabilities and whatnot, but how are things looking at a much higher level and treating the entire organizational system and software delivery as an observable unit I think is really important.

(41:26):
Because to your point, you might be shifting left and handling all the static analysis issues and whatnot, but you don't want it to be like a boiling the frog type situation where yes, things are healthy when you zoom out and zoom out, oh, things are really trending downwards because that is a very common failure mode is what we see.

Adam Berman (41:45):
Totally.

Ganesh Datta (41:45):
That's exactly how we think about it.

Adam Berman (41:46):
Yeah. We might see this world in which the OWASP top 10 or the basic vulnerabilities like, oh, do you have auth? Do you have all these different categories? But they're the base categories and we might see those categories of vulnerabilities really, really drop as the models get pretty good at training on them and producing code that doesn't have those. But I think the more frequent exploits then are going to be the cross system vulnerabilities, the ways in which systems interact with each other, the contracts between them and those parts where it's like if you only know what one system looks like, it looks good. If you only know what this other system looks like, it looks good. But the ways in which the whole application runs, that's where something is vulnerable. I think those complex ones I think are going to go up as we or could potentially increase as we give agents more control over what they build.

Ganesh Datta (42:33):
Absolutely. Adam, thanks so much for joining me on the podcast. This was awesome and very, very timely as people are trying to grapple with the ramifications of coding agents. So thanks so much for coming and sharing the learnings.

Adam Berman (42:44):
Yeah, thanks for inviting me. Great to be here. Great to chat.

Ganesh Datta (42:53):
Thanks so much for listening to this episode of Braintrust. If this resonated with you, do me a favor. Share it with another engineering leader who's wrestling with these same challenges. And if you want to continue the conversation or learn more about how we're thinking about engineering operations platforms at Cortex, reach out to us at cortex.io. Thanks for listening and we'll catch you on the next one.