AI Security Ops

This episode of AI Security Ops explores a growing question in AI security: who is responsible when an autonomous AI agent causes a real-world security breach without direct human instruction? Host Brian Fehrman examines reported incidents involving major AI labs, discusses how existing concepts such as liability, negligence, and intent may apply to AI-driven attacks, and considers the legal and security challenges organizations face as agentic AI systems become more capable. The episode highlights why responsibility for AI-caused harm remains an open question and what it could mean for the future of cybersecurity and regulation.

  • (00:00) - Intro - Who Is Responsible When an AI Agent Goes Rogue?
  • (01:25) - AI Models Breach Real-World Systems
  • (02:04) - OpenAI’s ExploitGym Incident & Hugging Face Breach
  • (03:52) - Anthropic and Meta Reveal Similar AI Incidents
  • (05:36) - Who Is Liable for an AI-Caused Breach?
  • (08:37) - Model Makers vs. Those Deploying the AI
  • (09:42) - Does the Computer Fraud and Abuse Act Apply?
  • (10:29) - AI Breaches and the Question of Negligence
  • (11:45) - Safeguards, Sandboxing, and Responsibility
  • (13:34) - What Happens When Your Organization Is the Victim?
  • (14:48) - AI Liability and the Self-Driving Car Parallel

Click here to watch this episode on YouTube.


Brought to you by:
Black Hills Information Security 
https://www.blackhillsinfosec.com

☯️ Introducing BHIS Fusion Penetration Testing
https://www.blackhillsinfosec.com/fusion-penetration-testing/

Antisyphon Training
https://www.antisyphontraining.com/

Active Countermeasures
https://www.activecountermeasures.com

Wild West Hackin Fest
https://wildwesthackinfest.com

🔗 Register for FREE Infosec Webcasts, Anti-casts & Summits
https://poweredbybhis.com


Creators and Guests

Host
Brian Fehrman
Brian Fehrman is a long-time BHIS Security Researcher and Consultant with extensive academic credentials and industry certifications who specializes in AI, hardware hacking, and red teaming, and outside of work is an avid Brazilian Jiu-Jitsu practitioner, big-game hunter, and home-improvement enthusiast.

What is AI Security Ops?

Join in on weekly podcasts that aim to illuminate how AI transforms cybersecurity—exploring emerging threats, tools, and trends—while equipping viewers with knowledge they can use practically (e.g., for secure coding or business risk mitigation).

Brian Fehrman:

Hey, everybody, and welcome to this week's episode of AI Security Ops. So today, we are going to talk about who is responsible when an agent goes rogue. It's been a pretty interesting summer. But before we dive into that, as always, we'll talk about our sponsors. Black Hills Information Security, you or your organization are in need of any kind of security service that you could think of, whether that's internal pen test, external pen test, assume compromise, physical pen test, wireless, social engineering, SOC monitoring services, AI pen testing, AI assisted pen testing.

Brian Fehrman:

Fusion offerings, one of our new things in that realm. Anything that you might be interested in, check us out at blackhillsinfosec.com. Additionally, we have a training branch, which is Antisyphon Training, where many of our consultants take the knowledge from their daily jobs, that they're doing day in and day out, coming up with new stuff all the time. They take that. They package it up into an easy to digest and affordable format for everyone to, to take and hopefully pick up a couple things, help you level up in your current position, help you get the job you've been looking for.

Brian Fehrman:

Or maybe if you're just a hobbyist and you want to learn something new, definitely check out antisyphontraining.com. So with that, let's go ahead and let's dive in for our topic for the week. So over this summer, three of the biggest AI labs each admitted the same thing. One of their models broke out of a test environment and hacked a real company that had nothing to do with the test. This wasn't a criminal group, not a nation state.

Brian Fehrman:

It was their own models doing their own homework a little too well. Hugging Face got hit in this process, and two other victims never noticed until they got a phone call. And now many lawyers and tech are asking the exact same question, which is when an autonomous agent breaks into someone's production system and no human told it to, who pays the price? Rather, who is responsible for that breach? And so let's give a little bit of a timeline, a little bit of a back back history here.

Brian Fehrman:

In July, OpenAI ran an internal cyber benchmark called exploit gem, with cyber safety refusals deliberately turned down to measure the model's true ceiling. GPK five six soul and an unreleased model found a zero day in their one permitted network egress, reached the open Internet, and then without anyone instructing them to reason that Hugging Face probably held some kind of an answer that it was looking decided to go after it. And through a malicious dataset that executed code during routine processing, they were able to gain access, into the, the the data pipeline, basically into HuggyFace. There were multiple actions that were taken, around, 17,000 or so actions that were taken. Lateral movement across 11 nodes, stolen cloud credentials, forged Kubernetes tokens, nine CDs eventually credited to the models.

Brian Fehrman:

So just all kinds of essentially, mayhem, Mandan destruction. I mean, just absolutely running amok. So within a couple days, Hugging Face detected it, contained it, disclosed it, and then actually called in law enforcement, of course. Because at this point, they had no idea who the attacker was. Someone had gotten into their environment and was just going crazy, and they're like, hey.

Brian Fehrman:

What is going on here? I don't know, but call the cops. You know, someone get the law enforcement on the phone. But then about, less than a week later, Open AI was like, hey, guys. Sorry.

Brian Fehrman:

It was us. Our bad. Unprecedented. Shoot. Whoopsie doodle.

Brian Fehrman:

It's basically the defense that I'm calling this. I'm calling it the whoopsie doodle defense. Funny enough, within about a week about, Anthropic reviewed about about a 100 and something thousand evaluation runs of their of their own and found that they had also created three incidents of their own. They went back all the way as far as April. So there is some speculation here, at least, you know, some joking around, if nothing else, that, you know, Anthropic solved it.

Brian Fehrman:

Hey. Open AI is getting this publicity for this. Like, we need to tell people that our models are fully capable of committing crimes as well. So they they released out this information. And then a week after that, Meta confirmed that their muse Spark one point one had breached some unnamed company and altered internal systems.

Brian Fehrman:

So it could be that also Meta was like, oh, hey. OpenAI and Anthropic are on this. We better get on this too. So you noticed that what is missing from this list, which is Google's, Gemini. We we don't see that on here.

Brian Fehrman:

Obviously, there are other models besides these ones we've listed, but those are some of the big, frontier players anyway. And so we wondered, is is Google feeling left out at this moment? You know, if they're they're they're poking their model with a stick and saying, come on. Come on. Commit crimes.

Brian Fehrman:

You can do it too. And if everyone else is doing it, why not you? This is all pretty interesting. Right? Basically, you know, one way or another, really, these models did things that they were not supposed to and caused real incidents.

Brian Fehrman:

They they actually I especially in the case of the OpenAI model, I mean, it legit breached another company, and, you know, the details have been laid out there. That is where we go into the question of liability, which is still very much open for debate. Now, I am not a lawyer. I do not pretend to be a lawyer. Law certainly interests me a lot because, I feel it's an extension of philosophy in terms of, you know, like, the logic component of philosophy.

Brian Fehrman:

And I just find it very, very interesting to nail down specific situations and trying to deal with all the different edge cases and trying to encompass things. And it's almost like programming in a sense of trying to write the perfect program that's not going to have bugs, and it's not, that it's going to be usable, not overly restrictive, but does the things that you want it want it to do but no more. And I feel that there's a lot of kind of parallels with the legal system in the way that laws at least should be, are attempted to be written or should be written. I mean, some are just overtly oppressive. Others are very half thought through.

Brian Fehrman:

But we'll just say when people have the best intentions, you know, that's I I I think that there's a lot of parallels to to the programming portion. Nonetheless, the philosophy aspect of it just it really interests me. With that said so I'm not a lawyer. So we're just gonna speculate here. We're gonna give this is just opinions.

Brian Fehrman:

But looking at, some recent case law, the whole well, the AI did it isn't really California has kind of addressed this already, at least in some way. This was back in January 2026 with a b three sixteen, which basically said, if you develop, modify, or use an AI system, you can't just argue that the AI autonomously caused the harm. But it still doesn't say which person is on the hook for it. So we have, in the OpenAI incident, you know, we have a multi company problem. So we have OpenAI who built and ran the model.

Brian Fehrman:

We have a JFrog exploited zero days, ModelLabs customer, whose exposed endpoint became a stager based and Hugging Face who ate, you know, detection remediation, FBI, filings. I mean, obviously, like, I wouldn't think that the victim is going to be at blame here at fucking pace. You can't be like, oh, hey. It's not my fault you had a vulnerability. I mean, that's that's not really, that's not a valid defense.

Brian Fehrman:

I mean, you can't be like, oh, hey. It's not my fault that your front door was unlocked. That's that's not a thing. It's not my fault that you, that you've left your, you know, your lawnmower out in your front yard. It's you just, like I don't I don't think that that's that's defensible.

Brian Fehrman:

But, you know, where we talk where we can really talk about those that, you know, who are we holding responsible in terms when we're talking about the technology itself and the use of the technology? Is it the person who built the model that is capable, built it built a model that is capable of performing these actions, or the person who is deploying the model and actually using the model without putting the restraints into place. You know, my opinion, I lean towards the latter. I mean, technology and tools are are are just that. They are components, especially in this case, don't inherently have risk in themselves until they are applied to something, until someone tries to actually do something with them.

Brian Fehrman:

So, sure, we can get into the whole concept of negligence, which we'll talk about here, just a little bit in a minute. But, you know, I'm more on it's the the person who has deployed the technology in an unsafe way. That that's where I'm at. Where do we talk about what about with maybe a computer fraud and abuse act, the CFAA? How do we apply it to these situations?

Brian Fehrman:

Well, one of the things that the court hasn't really been able to to pin down is that with the CFAA, there really needs to be intent. They need to show that a person intended to commit the crime, that they just accidentally happened to to do it. They were doing some other stuff, and, accidentally, this this action was performed. And so if the companies don't have didn't have the intent to breach somebody else in this, then it's hard to apply the CFAA to them and to say that they they are guilty of that. But that's where we get to the point of negligence, which is that just because you didn't intend to do something doesn't mean that you can't be held liable for that thing that happened.

Brian Fehrman:

Right? So if I've, for instance, if I'm, you know, hauling home a load of lumber on on a trailer, I'm just throwing this out because, what what you're seeing behind me, if you're watching online, this is, it was an unfinished space, so we moved in. Wife and I finished it all ourselves, everything except the drywall, but we did framing, plumbing, electrical, all that stuff, everything. So we've had a whole lot of materials, so that's on my mind. But, you know, if I was hauling that home and I didn't strap down the wood, like, took no no effort in securing my load, and a piece of that flies off.

Brian Fehrman:

It hits another car and causes an accident. Well, I didn't have intent there to cause that accident. However, I could potentially be held liable for the turn on the basis of negligence, because I didn't take any precautions to prevent that situation from happening. And that's a key point in this, is we could talk about negligence from two standpoints. So if we have, on the the model makers themselves, if they this is this is currently an open an open topic.

Brian Fehrman:

Right? We we've seen this in the news with, you know, the government wanting to shut down Anthropic, which you can argue that that's, where their models, you know, ban some of their models. We can argue if that's political reasons or not. But, you know, kinda still a point that stands there is, is it the responsibility of the people who create the model to at least put in some level of safeguards to try to prevent these behaviors, these actions? And if they don't, is it considered negligence?

Brian Fehrman:

But then we move on to the next step of the people who are actually deploying these models and using these models in an agentic faction fashion, allowing them to perform actions. And if people are not taking steps to prevent them from doing these behaviors, to prevent them from causing this harm, then I think that negligence could apply here. But at what point does it stop? At what point can you say that, hey, I tried to sandbox this. I tried to control it, and this thing still broke out, and it still it still did the thing.

Brian Fehrman:

It still breached, another company. That's that's a tough one. Because now you're saying, like, hey. I took a lot of steps to try to prevent this from happening, but because of what was basically a and called a bug, stuff happened. It it did things.

Brian Fehrman:

I I didn't want it to. I didn't mean for it to, and I took steps to try to prevent it from doing so. And that's a reasonable argument. That's gonna be a tough one, and that's still something that is not settled. So, basically, you know, right now, if you're the victim of one of this, this still isn't isn't settled.

Brian Fehrman:

There's the idea of, do you need to report this? So you got breached, but it was AI. So what do you need to do? Well, you probably still need to dispose it because you're probably still breached. But now who do you hold hold liable for it?

Brian Fehrman:

And that's gonna be something that we're gonna see answered almost certainly within the coming years. I mean, this is all still relatively new. Agentic capabilities have, I mean, just grown exponentially over just even the past year. And so now we're seeing these deployed more. We're seeing this technology used more.

Brian Fehrman:

We're obviously gonna see more and more of these incidents, and it's going to have to be something that the courts decide and and they solve. It's gonna be really, really interesting of where this ends up at. And, yeah, the answer is right now, we we don't know. It's an open ended question. So for those of would be curious to see other people's thoughts and comments if you wanna leave comments up here on the, the YouTube channel here or wherever else you're able to to leave comments for the podcast.

Brian Fehrman:

I I'd be curious to hear the the community's thoughts on this because I think it's a very fascinating problem that we're running into. You know, we have this not just with, with these ageptic things, but we also see it playing out with self driving cars as well. We're still working that out of who's liable if, you know, one of these self driving services ends up in an accident. So, anyway, we'll just, you know, we'll keep watching. Stay tuned.

Brian Fehrman:

Hope everyone enjoyed this episode. Hope it kinda spurred some, thought and conversation, and we'll see you next time. As always, keep on prompting.