AI Security Ops

šŸ”’ Want to run AI without sending your data to the cloud?

AI Security Ops co-host Bronwen Aker is teaching Keeping Things Local: Build Private LLMs for Your Team.

āœ”ļø Build a network-accessible private LLM with Ollama
āœ”ļø Customize models for your workflows
āœ”ļø Secure it with Tailscale and nginx
āœ”ļø Keep sensitive data under your control

Only $25
Next live session: August 17, 2026


šŸ¤– Want to get hands-on with Agentic AI?

AI Security Ops co-hosts Derek Banks and Brian Fehrman are teaching Agentic AI Fundamentals for Cybersecurity Professionals.

āœ”ļø Build and extend AI security agents
āœ”ļø Analyze real attack logs and malware samples
āœ”ļø Integrate live threat intelligence
āœ”ļø Assess a deliberately vulnerable target

No prior AI or programming experience required (basic security knowledge recommended).

Only $25
Next live session: August 18, 2026


----------------------

In this episode of BHIS Presents: AI Security Ops, the team breaks down one of the most significant AI security incidents to date:

How did an AI model escape its testing environment and autonomously compromise a production system?

This wasn’t a stolen password. It wasn’t a poisoned public model. And it wasn’t a human attacker sitting at a keyboard.

Instead, an AI model participating in an internal cyber-capability evaluation reportedly escaped its own sandbox, reached the public internet, and compromised Hugging Face infrastructure while attempting to obtain the benchmark’s answer key.

The incident offers a rare opportunity to walk through a real AI attack chain from beginning to end—not just what happened, but what defenders could have done to prevent or detect every stage along the way.

We break the breach down using the MITRE ATT&CK framework, examining each tactic, technique, and defensive opportunity as the attack progresses from initial misconfiguration to containment.

We dig into:
• Why disabling AI safety guardrails created the initial opportunity
• How the model escaped its evaluation sandbox
• The role of the Artifactory zero-day in the attack chain
• How the agent reached external infrastructure and executed code
• Why malicious datasets became an unexpected execution vector
• MITRE ATT&CK mappings for each phase of the intrusion
• Detection opportunities defenders could have used at every step
• Secrets management, workload identity, and least privilege
• Why segmentation and egress controls remain critical for AI systems
• The difference between OpenAI’s and Hugging Face’s published timelines
• How AI-assisted detection ultimately helped stop the attack
• What security teams should learn before deploying autonomous AI systems

This episode explores an important reality of AI security: autonomous agents don’t invent new attack techniques—they chain together familiar ones at machine speed. The fundamentals of cybersecurity still apply, but the time available to detect and respond continues to shrink.

The takeaway: don’t ask whether your AI system is powerful. Ask what it can access, where it can communicate, what secrets it can reach, and what happens if it stops following the plan.


  • (00:00) - Intro: Revisiting the OpenAI and Hugging Face Breach
  • (01:19) - Walking Through the Attack Step by Step
  • (06:08) - The Evaluation Goal and the Agent’s Unintended Path
  • (07:39) - Sandbox Escape Through Artifactory
  • (14:28) - Initial Access into Hugging Face
  • (19:28) - Privilege Escalation from Worker Pod to Root
  • (22:54) - Credential Harvesting and the JWT Signing Key
  • (26:12) - Lateral Movement Through the Tailscale Network
  • (28:41) - Collection, Exfiltration, and Command and Control
  • (31:36) - How Hugging Face Detected and Investigated the Attack
  • (35:51) - What This Means for Defenders and AI Development

Click here to watch this episode on YouTube.


Brought to you by:
Black Hills Information Security 
https://www.blackhillsinfosec.com

ā˜Æļø Introducing BHIS Fusion Penetration Testing
https://www.blackhillsinfosec.com/fusion-penetration-testing/

Antisyphon Training
https://www.antisyphontraining.com/

Active Countermeasures
https://www.activecountermeasures.com

Wild West Hackin Fest
https://wildwesthackinfest.com

šŸ”— Register for FREE Infosec Webcasts, Anti-casts & Summits
https://poweredbybhis.com


Creators and Guests

Host
Bronwen Aker
Bronwen Aker is a BHIS Technical Editor who joined full-time in 2022 after years of contract work, bringing decades of web development and technical training experience to her roles in editing pentest reports, enhancing QA/QC processes, and improving public websites, and who enjoys sci-fi/fantasy, Animal Crossing, and dogs outside of work.
Host
Derek Banks
Derek is a BHIS Security Consultant, Penetration Tester, and Red Teamer with advanced degrees, industry certifications, and broad experience across forensics, incident response, monitoring, and offensive security, who enjoys learning from colleagues, helping clients improve their security, and spending his free time with family, fitness, and playing bass guitar.

What is AI Security Ops?

Join in on weekly podcasts that aim to illuminate how AI transforms cybersecurity—exploring emerging threats, tools, and trends—while equipping viewers with knowledge they can use practically (e.g., for secure coding or business risk mitigation).

Derek Banks:

Hello, and welcome to another episode of AI Security Ops. And to today, we're gonna talk, well, of course, about the Hugging Face Breach because it's like the the gift that keeps on giving. It's the story of the year, at least until the next one. And so we're gonna talk more about it. But before we do that, we'll talk about our two sponsors, Black Hills Information Security, as I'm wearing a shirt today for them, where we do penetration testing, red teaming.

Derek Banks:

We do SOC services to continuous pen testing. And and if, you know, if you're doing AI evaluations and wanna learn how to, you know, set up your sandbox correctly, we'd help you with that as well. And then Antisyphon where we have industry led or practitioner led training for information security topics. And with that, let's dive into the episode. So I thought it would be fun, and Bronwen grudgingly agreed with me because she's probably been talking about this for the last three days straight.

Bronwen Aker:

Something like that.

Derek Banks:

I didn't get to talk about it because my power and my Internet was out. So you're gonna have to, like, you know, to bear with me here.

Bronwen Aker:

Dad. Bad company.

Derek Banks:

I know. Right? Dad. And so I would I thought we'd do kind of a walkthrough. Now, some more details.

Derek Banks:

Even since Monday, when they're talking about the BHIS News and the In Focus, some more informations came out, some more details. So I thought it would be fun to kinda walk through kinda each big step of the attack, like the TTPs that the rogue agent or the rogue model was was taking, and then talk about, like, what we could do to kinda prevent or detect that step, as an organization. Like, what could have been done to defend or, you know, to detect or prevent the the each step. And so we'll start with, I guess, what, you know, some people probably think the root cause is here, and that is that the large language model that this evaluation harness was connected to ran without its safety guardrails. Right?

Derek Banks:

They had a model, that probably wasn't Soul. At least they were not gonna say if it was Soul, and they, had the safety brakes off. Right? And which is interesting in itself because if you've ever gone through the the process of ablating a model or obliterating a model, you know, you have to mean to do it. Right?

Derek Banks:

Like, it's not something that, you know once you put the the once you train the model to have those guardrails, once you set the weights inside of the neural network to have those guardrails, what you would do to ablate it would be to measure the request responses in terms of the activation of the neurons inside of the neural network, and then you would, essentially negate them with math, right, with linear algebra. And so, I'm not saying they have an ablated soul. Maybe they train one either way. So they've got this model, and so, you know, what you know, before any technique, they they disabled those refusal classifiers.

Bronwen Aker:

And so I thought they the articles that I read said that they disabled some of the guardrails. So it was, theoretically, the the containerization and and the restrictions to keep it in the box were weakened but not eliminated.

Derek Banks:

So you're saying they probably loosened the restrictions in the harness, not necessarily the model.

Bronwen Aker:

I could probably get That's the gist of what I read. But, again, I wasn't there.

Derek Banks:

Yeah. Well, none of us were, so we're speculating. So well, you know, not to to to to beat a dead horse with this, that how could you have predicted prevented or detected this? So prevention for this stuff is well, for every disabled safeguard, there's a time limited exception or a named person signs off on that, that disabled safeguard, and then it is, make sure it doesn't stay off silently after the test ends. So let's just assume this is in semi production inside of your network.

Derek Banks:

It's running on your your inference, you know, system, and so you disable a safeguard. It kinda like firewall rules. Right? If I was gonna disable it, I don't wanna disable it forever. There should be, like, a a time on that.

Derek Banks:

And then detections, watch inference compute during evaluations. Model burning unusual amounts of compute is an early tell that it's pursuing a goal that you didn't intend. Now that's kind of interesting. Right? It's not really a security thing.

Derek Banks:

It's more of an operations thing, but anomalies and operations data just like security one zero one. Right? This is anomalous. Let's go figure out why. Right?

Derek Banks:

And so I I would think so that's the thing, you know, I have heard the comment being, you know, being made that OpenAI didn't realize this happened for a bit. Did they say days, or was it hours or days? I think it was days. Right?

Bronwen Aker:

Days. I think it was days. I mean, the the picture I've seen is that they launched they they set it up. They gave it a mission. They weakened the guardrails and the harness, and then they let it run.

Bronwen Aker:

And days later, they went, oh, gee.

Derek Banks:

You know, if you would've did that in a deeper voice, it could've been the opening to a movie. Right? Like and then they disabled the safeguards. That's how terrible you're

Bronwen Aker:

starting They disabled the safeguards. So

Derek Banks:

and and old to have that much compute where you didn't realize that your model that's doing all these things for all of these days, I don't have that much compute. So I could probably, couldn't do this. But so yeah. So anomaly detection on operations and monitoring is kind of the thing there. So alright.

Derek Banks:

So the next step, you've got this thing that's got, you know, this this agent with this very powerful model behind it that's set to go essentially do was it CyberGem benchmarks, I think. And then somehow, in its thinking stream of tokens ExploitJim. Exploit Jim. Exploit Exploit Jim.

Bronwen Aker:

They specifically set it to go through an exploit gymnasium. Yeah. It's like a a benchmark target

Derek Banks:

of some sort. Right? An evaluation. And that's the thing, like, you're getting into LLMs and how this stuff works, if you're gonna be like an LLM engineer, man, evaluation is a big part of that. Knowing, you know, how it works, what's happening, and the evaluation of that to get meaningful results and and to, you know, see progress, like, that's a part of, you know, something that you would do.

Derek Banks:

So I can get why, you know, they set this test up. But it's interesting, and and I heard John mention this on the in-depth thing that, like, that to make the leap of, like, well, why did it did it decide to do that? Well, because somewhere in its training corpus, it learned from hackers that, you know, there is no cheating in hacking. Right? Sorry.

Derek Banks:

There's not. And so you wanted it to go off and hack stuff, and so it decided I needed to accomplish this goal. I'm just gonna go read this thing over here. It's like being on a CTF and finding somebody who already had an answer key. Right?

Derek Banks:

Like, well, yeah. Alright. So the next step is to get out of its sandbox. And that that's the kind of the part that I got hung up on for a bit, so I'm glad some more details came out. So the sandbox escaped through an artifactory zero day.

Derek Banks:

The model was boxed in a test sandbox whose only outbound network path was a j frog j frog artifactory instance. It's a package registry cache proxy that a server sits between developers and public code repositories to be able to download dependencies faster. And so basically, they put this package proxy in place in between the sandbox environment and Internet, so the harness either the systems in the environment or the harness on the system could go and get tools and packages. And even if the model is not using it as a tool, it's on the network, and so let's just say that it it realized, you know, that it it needed to get out of the sandbox. And so it went and found this, proxy and realized recognized that that was the method out.

Derek Banks:

Just in itself, that kind of amazes me. Right? But then to find a zero day and and so from what I read, it might have actually been an existing, like, CVE already. And so either way, the that's this is like one of the more amazing parts to me is that it it somehow figured out a zero day, presumably without source code in the package manager. And that's that's kinda fun.

Bronwen Aker:

It is fun. Now I'm just wondering because by the way, did you take a look at the Hugging Face animation of the engagement?

Derek Banks:

I haven't. I heard it was cool. I haven't seen it yet. No.

Bronwen Aker:

It's it is pretty impressive. And I I will try to remember to put a link so that we can get it into the show notes

Derek Banks:

to other

Bronwen Aker:

people as well because it it really is impressive. And you watch how many agents were spun up for this. And I think this is one of the things that is most impressive about a single model in an an agentic mode spinning up additional agents because it's been given the directive, don't stop until you achieve this goal. And the it's it's kinda like the infinite monkeys thought experiment. If you have essentially an unlimited number of agents and they're able to be educated and trained based on output from hackers all around the world, how could they not find zero days?

Derek Banks:

I mean, I guess. I mean, would almost have to be one of those that even though it's a zero day, it it couldn't have been super difficult to find. Right? I mean, if you're looking at, like, the interface into a proxy and, like, just the response there in back, maybe if there was, an admin panel and there's, like, a a directory traversal or something, maybe. Like, I could see that.

Derek Banks:

So like I've said a couple times in comments and chat and internally at BHIS, this is certainly possible, but plausible is I get kind of hung up on it. But it but either way, it found a zero day according to what's been said, and the sandbox escape, you know, through this package manager. Basically, it got it got on that box through command execution and was able to then use that as a jump host. Right? So what can we do here?

Derek Banks:

So in case anybody is wondering, the the MITRE mapping here, t one one nine o exploit public facing application, never let a high capability evaluation run on any host that has Internet egress. Kill transitive escape routes, not just direct ones because the sandbox is only as isolated as the most privileged service it could reach. Man, I feel like I'm preaching back in my shipyard days. Like, air gap means nothing is plugged in. Right?

Derek Banks:

Like, you cannot plug in a cave an an ethernet cable and then tell me that it's air gapped. That is not how that works. Right? And so Well, as long

Bronwen Aker:

as the ethernet cable isn't plugged into anything else,

Derek Banks:

maybe. Sure. Okay. If they got a link light on. Right?

Derek Banks:

So anyway, so pre resolve all the dependencies. Don't grab dependencies midstream. And then one thing that John mentioned that I would like to reiterate here that seems to be falling out of favor in even in information security, and I've had people push back on on me when I've said this is, don't forget about monitoring the network. I mean, just one Zeke instance and looking where this thing is going, where it should be going, and where it is going could have caused this a lot quicker. Right?

Derek Banks:

And you can even do anomaly detection. You could set up Zeke rules. So some network monitoring here, know, and then John did mention this. Like, our platform that we've developed, you know, we take a PCAP for every, you know, every run. Right?

Derek Banks:

And and so if we have to go back and look and see what happened, well, we can. And so we can make that real time if we needed to on our platform. And and so network monitoring and and, I guess we'll say that could be, you know, a a part of the detect stack. Right? If you're going to install Zeke or some kind of system where you're doing a PCAP, you know, I think that that that would have been a failure on the designing of the sandbox, so to speak.

Bronwen Aker:

It still blows my mind how much can be detected using decent threat hunting. And I think that now more than ever, be because the because agents are capable of behaving in nonstandard ways and and doing really wacky stuff, the normal profiling that we're used to using is gone.

Derek Banks:

We're never gonna be able

Bronwen Aker:

to go back to those days. Oh, I feel like we use the network

Derek Banks:

seems like it's one of those, like, lost and arcane arts. Like, I mean, I have a a book on the shelf behind me from Richard Beatlitch called The Tao of Network Security Monitoring. Was written in 2004. Right? And you could read that.

Derek Banks:

About 80% of that book is still, like, applicable to today. Right? Because not a lot's changed on networking and the Internet fundamentally. Right? And so, anyway, so we're on step two of like seven.

Derek Banks:

Right? And so then the next one is the initial access into Hugging Face, and this one I actually didn't know about. I learned about this today. And so apparently, what happened is the model, the harness, and the model figured out, a, that it wanted to get to Hugging Face because I'm assuming that's probably where the exploit gem data was housed or something along those lines. Either way, it recognized Hugging Face as being a place that it could find these answers, And so it decided it was going to exploit Hugging Face.

Derek Banks:

And so this initial access kind of, you know, probably could spend a lot of time talking about, like, things that would be inside of here, like, how did it do the reconnaissance? How did it find that it was able to essentially weaponize an HDF five dataset and have two exploit paths inside of there? The first one was what is the days? It's CVE twenty twenty five nine nine zero five, which is apparently how how HDF five works. Basically, you point it to a sensitive file path.

Derek Banks:

You point like a a pointer inside of the dataset to a central a a sensitive file path. And so it's not novel. It's been known. Essentially, it's like a deserialization attack, kinda similar to pickle when you like load the model. And then there was another there was a Genja two Python templating, server side template injection flaw that it took advantage of.

Derek Banks:

And I gotta say, in all my times of web app hacking, I have never really been able to do something like that. A server side template injection that led to remote code execution. That's impressive in my opinion. If I Yeah. If I the people I know who could probably identify and pull that off, I could probably count on one hand.

Derek Banks:

Right?

Bronwen Aker:

And have fingers left over.

Derek Banks:

And and so I I'm impressed if that is the the the just step two in this. Alright? Alright. And so prevent prevention here. And and this is something that I'm gonna be including in my class because up until this point, again, I didn't realize that this is actually a flaw.

Derek Banks:

Treat every downloaded model or dataset as untrusted executable code that gets reviewed before touching infrastructure. Did y'all just hear what I said? Every dataset and every model's got to Hugging Face, and you're gonna get a model to do something around your Spark, and you've now gotta do a security evaluation on it to make sure that you're not getting back doored? Yeah. Well, shit.

Derek Banks:

I'm guilty of not doing that stuff. Right? And I and if you would've asked me yesterday if you could weaponize a dataset, I'd have said, damn, that sounds pretty hard. I don't know how you would do that. Well, now I do.

Derek Banks:

Thanks unnamed GPT model that pulled this off. Alright. And so treat everything that you download from the Internet maybe as bad. Reject datasets that declare an external file list or carry template expressions. Perhaps if you're going to be using Hugging Face and datasets, dig a little deeper into potential security issues of those things.

Derek Banks:

Detection, a dataset ingestion worker spawning a shell or a new process, a normal dataset set load never does it. So basically, what we're saying is, like, in my testing routine, I would expect when I load the dataset certain things to load, not spawn additional process. Kind of akin to my Microsoft Word shouldn't be spawning command shell. Right? I mean, it's kind of the same same concept there.

Derek Banks:

Man, we're only on the step three of seven.

Bronwen Aker:

Hey. This this went on. I mean, it was and and there's so much that still is trickling out. Every day, we're learning new stuff. And and it it makes me wonder, is it because they're sharing the stuff that they've known?

Derek Banks:

Or is it because it out. Well, incidence of response takes a bit. Right? Like, I've actually been kind of impressed with the speed of which stuff's coming out. Usually, it's a month or two or three later.

Derek Banks:

Right? Because figuring out what happened and this goes to, like, what we'll say at the final step, you know, with you using AI to make yourself faster. If you don't think you could respond in the into this kind of thing in the same kind of time frame, well, then you need to up your game. Right? And so kudos to the Hugging Face folks.

Derek Banks:

They obviously were able to do something very difficult in a quick time and had help with AI to do it. So step three in this was after initial access so basically, what's happened here up until this point, the model has now got initial access into what I'm assuming is a Kubernetes pod. You know, right? So basically, I'm on something you know, a GitLab runner or something along those lines. I'm on infrastructure inside of Hugging Face now, and I've gotta escalate from Worker Pod to the cluster.

Derek Banks:

Again, this is like this is high end hackery in my opinion. Like, I wouldn't expect, like, every pen tester, like, off the street to be able to do this. Inside that worker pod so it the model it was able from that pod to read the environment and the cloud instance metadata, shout out to cloud instance metadata and being able to pull data from there, minted from the cloud instance metadata, it was able to mint a a security token impersonating the node's identity. Man, actually, this part is something I've done in a cloud environment before. I've definitely definitely gotten stuff from cloud metadata in the past, to to get, like, credentials.

Derek Banks:

It they so it got temporary credentials. It created its own privilege pod with a host path mount that exposed underlying host's file system because it had the rights to do that at that point, and should've so because an emission control policy that should've blocked that was missing and escaped to root on the node, it was able to escape to root and essentially overbroad container storage interface driver role. Oh, wow. So basically, long story short, was able because of missing security controls in the Kubernetes cluster, it was able to get off of the pod onto the underlying host and escalate to root.

Bronwen Aker:

Youch. What we

Derek Banks:

call an incident response a bad day.

Bronwen Aker:

Very very

Derek Banks:

much so. So that minor mapping is t one one t one five five two zero zero five unsecured credentials. So giving your your cloud processing workers credentials, You shouldn't give them anything beyond a strict need. That shouldn't be in the metadata unless it absolutely has to be. And then do not forbid privilege containers, host mounts, or any other kind of capability that, gives any kind of, know, control, authorization control.

Derek Banks:

Deny worker pods any access to the control plane. If a worker fails, then damage should stop there. And then for detection, a worker service account calling self subject rules review. If you have a a worker calling that, then that is a sign of a privileged pod creation. So that means that you had to be ingesting the appropriate log files into your sim and have an alert to look for that thing.

Bronwen Aker:

Alright. Yeah. That self subject rules are you. What can I do? What can't I do?

Bronwen Aker:

Of course. That's that's priceless information for anyone who wants

Derek Banks:

to be malicious. Recon. Right?

Bronwen Aker:

Very much so.

Derek Banks:

So step four, credential harvesting including the signing key that changed everything. Oh, man. So the model then alluded secrets. It read, proc self environ, a Linux path exposing process environment variables to pull AWS keys and a static MongoDB password. So not only is this thing a cloud and web app security expert, it also is a Linux security expert because that's some that's some the only other thing I would've been impressed with it because it downloaded LENPs and ran that.

Derek Banks:

But either way, that's pretty good. And so then it hit the elastic compute EC two instance metadata for the notes roll credentials, enumerated a 136 keys from a production secret object. One Harding credential store did its job and denied 84 of its attempts. So at least some of it wasn't there. So so either way, it stole an a a signing key that was used to sign JSON web tokens, and it I'm assuming it then minted its own JWT.

Derek Banks:

And from there let's see. A JWT is signing key, which is far holding the key is broader than holding a stolen token. And so, basically, it made it sell its own JWT. And so to prevent that, issue a fairmal evaluation specific credentials whose lifetime matches the job, Use workload identity that cannot be replayed in another environment. Never let research, research identity be accepted by a production system.

Derek Banks:

Basically, separation of roles here. Right? So

Bronwen Aker:

bored by that one hardened credential store denying 84 attempts, man. It was it was trying to hold that gate, man.

Derek Banks:

I just when I heard more detail, I heard more details came out this morning. And so when I went and started reading this, I was like, man, this is this is good stuff. Right? And so for detection, hunt for unexpected reads of environment variables, cloud made a data metadata endpoints or secret stores. So I'm assuming you wanna ingest your cloud data into your SIM, have the appropriate rules.

Derek Banks:

A lot of this stuff so far sounds like things that I preached in my incident response class back in the day, which I think you took at one point. And this sounds like stuff I was saying to do. Yeah.

Bronwen Aker:

I was gonna ask you, how much of this breach could have been prevented by stronger security hygiene?

Derek Banks:

Almost all of it. Really? It's the world we live in now. But I mean, look, to to Hugging Face, like, I listen. Running that kind of service online is very difficult.

Derek Banks:

And because they have flaws, I mean, welcome to the party. Right? Like, I'm on an engagement today where our AI platform found what is remote code execution, and there was one ACL standing in my way. And so I would have never found that without, like, the help of AI. It's just not on a couple day engagement.

Derek Banks:

Right? So anyway, so now we are it got a signing key. It was able to mint JWT. So basically, now it's got root access on like on a node, some forge identities. It had enrolled the compromised node, and an external sandbox it had rooted.

Derek Banks:

I'm sorry. Hang on. It enrolled the compromised node and an external sandbox, which we haven't really talked about this other company who had their sandbox rooted. Right? Into the company's tail scale mesh virtual private network.

Derek Banks:

Damn. Ow. Wow. Okay. Yeah.

Derek Banks:

That that is some a post exploitation goodness right there. So basically, now it's part of the tail scale network. Yeah. It ran in user space networking mode with a SOX five it ran a okay. It ran a SOX relay and a and to tunnel traffic over the tail scale, found a single connector credential shared across clusters that was bound to system, masters the highest private privilege group in Kubernetes over the campaign, and it rolled a 181 devices into the tail scale network.

Bronwen Aker:

Dang.

Derek Banks:

Wait. So you didn't know like, this is all new to you. Right? Like Yeah. Okay.

Bronwen Aker:

Like, you're learning new stuff. I read something else that said that during the compromise, the the models compromised at least four companies. Yeah. It's like four four accounts, four different companies. So so, I mean but still, a 181 devices in the mess.

Derek Banks:

If I was if I took this out and put BHIS red team did this, just up to this point, this is weeks worth of work. Like, this is not, like, this is not happening on the scale as fast as the AI did it for sure. Like, it would take like, I've been on red teams where we were able to get in through an, you know, a CICD flaw in something that was exposed. It took us like three days to find it though. Right?

Derek Banks:

Like, mean, you know, thousands of web services grinding through looking at, you know, go witness or eye witness screenshots. And, anyway, so we're on to step six. So we got lateral movement out of the way. So we're gonna go on the next in the, you know, in the old kill chain here. Collection, exfiltration, and command, the control, and cover up.

Derek Banks:

So we're down towards the end now, step six. After lateral movement here, the model reached its actual goal. Oh, hey. Slash goal. Right?

Derek Banks:

Five internal datasets holding the benchmark answer key. Well, I guess they were it was right. It Hugging Face had the answer key for the exploit gem. Right?

Bronwen Aker:

Okay. Okay. Now hold on a second. Hold on a second. I have known people, and so have you, that rather than do what they're supposed to do, they'll write an entire custom suite to cheat.

Bronwen Aker:

Am am I off base, or is that not what just happened? It was told to go through exploit Jim. And what did it do? It did all of this stuff.

Derek Banks:

Somewhere in his token stream, it decided that this was easier than just trying to solve exploit Jim stuff. I do I don't know. Like, I so so it set up I I I it calls it command to control, like, guess command to control. Basically, it was more of like obfuscation in my opinion. The the channel and attacker used to steer robbery.

Derek Banks:

So basically, it deliberately used ordinary public services, so nothing looked malicious over a 100 single use request capture URLs, paste in sites, file drops. So it's trying to cover its tracks over the network. It encoded payload with with exclusive XORs. Shout out to Troy. Single byte XORs are your friend.

Derek Banks:

Right? And some g's at compression. So basically, it ran it it was trying to muddy the forensics picture, which is kind of an a hole thing to do. And then it it had a total of 56 exfiltration actions as opposed to cover up its data exfiltration of the exploit gem answer key. So here, this is this MITRE t one two one three data from information repositories, and then also t fifteen sixteen seven, exfiltration over web services.

Derek Banks:

And so log every identity assumption token use and privilege transition with full service account context so the real restricted under a blizzard of decoy noise. So it said that's the prevention. That's really detection kind of stuff. And then just in general, outbound connections from a a machine learning worker pod to pay spin seems like a bad idea. So basically, network monitoring outbound from your Kubernetes cluster would have probably gave you a heads up for this.

Derek Banks:

However, at this point, I'm I'm not sure that Hugging piece didn't know. I'm not saying they didn't detect it by this point. And so really kinda get down to the end of it. Hugging Face says they caught this with with more AI. They had a a machine learning anomaly detection pipeline that used it says uses a large language model.

Derek Banks:

I I don't know. Anomaly detection is usually not what you use an LLM for, maybe a transformer that's fine tuned for. But anyway, they used AI, and apparently, there were 17,600 user interact or attacker actions, which sifting through the noise, that's a lot. That's a lot of actions. And then kind of like the the what do you call it?

Derek Banks:

The coup de grace. When when Hugging Face went to respond to this, they were using frontier models. And based on the data they were trying to input and ingest, they were hitting guardrails, because their claim is that that that the the input from the log files and from what the model was doing, they couldn't tell the frontier model couldn't tell if it was offensive or defensive related. And, you know, on the, on the I I don't wanna say it's BS. I don't mean like the like that it's not true, but, you know, on the on the scale of one to 10 of I don't know about that to, oh, yeah, I agree.

Derek Banks:

I'm gonna give it like a seven because I have definitely used frontier models for both offensive and defensive work. And I'm I've definitely hit guardrails, but I'm not saying I've never been completely shut down. I don't know. But but either way, they then said that they had to fall back to GLM five two, which they were self hosting and running on their own. I mean, I don't know if guardrails in the frontier models are better or worse than the internal guardrails in GLM five two.

Derek Banks:

But I do know that using the API versus using, like, you know, a chatbot interface at, you know, quad.ai or whatever is gonna get you a little little bit less of a a guardrail, in my opinion. And then but I mean, if that's true, then I think the the takeaway is is, you know, defenders and companies should really start thinking about buying some GPUs and running their own models, which these days might be kind of complicated, which should probably be a topic of another podcast of I wanna run open weight models. Now what?

Bronwen Aker:

Why do you wanna run open weight models?

Derek Banks:

And and this is the thing.

Bronwen Aker:

I understand the guardrails. They're they're a necessary evil. They it's kinda like seatbelts. It's a lot like seat belts. It's not going to prevent everything, but it will hopefully at least let the patient survive the experience.

Bronwen Aker:

But I I don't know. This whole this whole thing, like, when you were going off on all of the attacker actions, you would think that the amount of noise that that generated would have tripped something early on, I would hope. And then I I just I'm speechless. Congratulations.

Derek Banks:

Yeah. I mean, I'm still kind of I'm leaning towards this is true and and and but I I have some doubts. I'd like to see more. Now that more information's came out, like, I I think that whatever model is behind this, you know, harness that they were using is probably pretty good at cyber security. I I've used both frontier models and open weight models like GLM to do you know, to basically testing out our own platform at Black Hills.

Derek Banks:

And I mean, the open weight models are pretty capable. I think they're still a little slightly behind, you know, frontier models, but that gap's not a lot. And so I think the big conversation now is, okay, what do we do about that? And I I just read this morning that a bunch of companies, Anthropic and OpenAI being two of like 1,700, are basically petitioning the government to slow AI development down, which I think is a stupid idea. I mean, hey, Anthropic and OpenAI, you can choose to slow it on your own without, like, me capping everybody else.

Derek Banks:

And I just even if, you know, the United States government was able to get China to agree to slow down AI development, do you believe them? Come on. Really? Unfortunately, I think at this point, you know, the Pandora box has been opened, and really, I just don't see any path forward besides putting this capability in Defender's hands. I really don't.

Derek Banks:

Like, that is that is the only way forward, in my opinion, that gets us to a chance of, you know, we win and the Chinese don't kind of thing, or we win and we all win. I don't know. But either way, I I just I don't think that slowing AI development is going to be a a a good thing.

Bronwen Aker:

I don't think it's achievable, honestly. No. Given given how the every tech company, every big tech company, I should say, is doing something in the AI space and has embedded it inside everything. And the the marching order, the drumbeat coming through is use AI more. Use AI more.

Bronwen Aker:

Use AI more. Well, there's going to be a cause for that. This is the kind of thing we're gonna see. And until and unless defenders are able to get the access, we can't keep up.

Derek Banks:

Yeah. If I was in I'll just say, like, you know, in a position that, you know, I've been in, you know before Black Hills, I worked for a lot of different companies and, you know, like NASA, the shipyard. And we had a lot of kinda like rules and regulations about what we could and couldn't do and buy and use. But if I were at one of those places now, I would insist on getting something like a DGX Spark or some kind of desktop that had one, two, three, four GPUs in it, that I could load an open weight model and then have control, like, data sovereignty, be able to practice doing IR with it and using it on my dataset. Like, I would start doing that now.

Derek Banks:

I'm not saying don't start trying and and using your claw subscription or whatever, but have have the capability in place before you, you know, get to a point where you're like, I I can't I can't do this because I don't have the right tools.

Bronwen Aker:

And if you're in a position That's no longer that's no longer a valid argument because the tools are available. There are plenty of open source tools available that can make it possible for you to set something up internally on relatively inexpensive hardware.

Derek Banks:

Yeah. And and so and you're you might be thinking, dear listener, that, well, I mean, we could just do forensics the old fashioned way. Oh, absolutely. But if I had to do what Hugging Face has done at this point, like, say me and another person, like a team of two or three people, which would be like what you would normally get if you called in, you know, like, you know, a forensic service, it isn't going to be the week after that you get this kind of level of detail. Right?

Derek Banks:

It's going to be a week or two or three. And that's why, you know, most of the time in the past, like, you hear about breaches and get the forensics report, like, way after it happened. Also, lawyers. Right? Which is kinda interesting.

Derek Banks:

Like, there must not be legal and lawyers involved at this point because it's happening so fast. But but I think that to stay, like, moving at the speed of of attackers is why I think you need to look into getting your own compute in your environment.

Bronwen Aker:

And and looking at the numbers, 1,700 recovered recovered agent actions. 6,200 or 6,000 groups. The Yeah. The keys. All of that.

Bronwen Aker:

For Wednesday.

Derek Banks:

Millions of log file entries. Right? Like, I I mean, when I first read it, I was like, 17,000 log file entries? That says nothing. And then I reread it, I was like, oh, no.

Derek Banks:

That's 17,000 actions. That log file Exactly. Must have millions millions of entries in them, which will be what I expect. Right? And so manually parsing through that just takes, like, time, right, and knowledge.

Derek Banks:

And and so, today, if you ask me to parse through all this kind of data, I would most definitely turn right to AI to do it.

Bronwen Aker:

This this level of data would give Hal Pomerantz a hard time.

Derek Banks:

Yeah. Yeah. I bet I wonder if he was surprised at that that Linux exploit there or that Linux That would be that was good conversation to have with him. I was like, actually, no. I'd like to talk to Hal again.

Derek Banks:

One time, I actually found something on an incident response and ran it by, he's like, oh, that's new to me. And I was like, yes. I can't

Bronwen Aker:

believe That's a major win.

Derek Banks:

I know. Right? I didn't do it. The threat actor did, but it was pretty cool. Anyway, with that, if you're still listening, I guess we'll end this one and keep on prompting.