Shared Security Podcast

OpenAI disclosed an unusual AI security incident involving a frontier model evaluation and Hugging Face, but the episode cuts through the hype: was this true autonomous intent, an agent following bad scope boundaries, or a warning about giving AI systems real tools and permissions?

Show Notes

OpenAI disclosed an unusual AI security incident involving a frontier model evaluation and Hugging Face, but the episode cuts through the hype: was this true autonomous intent, an agent following bad scope boundaries, or a warning about giving AI systems real tools and permissions?

Tom, Scott, and Kevin discuss why anthropomorphizing AI makes the story harder to evaluate, what companies should learn before deploying AI agents into production environments, and why permissions, logging, containment, and incident response matter more than marketing language about models acting on their own.

Special thanks to Guardsquare for sponsoring this episode! Guardsquare is the leader in mobile application security, with multi-layered protection for your Android and iOS apps. Learn more at Guardsquare.com.

** Links mentioned on the show **

OpenAI says its AI technology acted on its own in an ‘unprecedented’ hack of another company https://apnews.com/article/openai-gpt56-sol-hugging-face-63ab84fed5612af04d8a160d60f6def3

Luta Security: OpenFace: The Hugging Face Breach and What to Do About It https://www.lutasecurity.com/post/openface-the-hugging-face-breach-and-what-to-do-about-it

Cloud Security Alliance: Hugging Face Incident Initial Post Mortem https://cloudsecurityalliance.org/artifacts/hugging-face-ciso-post-mortem

** Watch this episode on YouTube **

https://youtu.be/Fwb0jsgJxf4

** Become a Shared Security Supporter **

Get exclusive access to bonus episodes, listen to new episodes before they are released, receive a monthly shout-out on the show, and get a discount code for 15% off merch at the Shared Security store. Become a supporter today by going to our YouTube channel's membership section: https://www.youtube.com/channel/UCg9CCDIYkDDqwEZ3UYaxjnA/join

** Thank you to our sponsors! **

SLNT

Visit https://slnt.com to check out SLNT's amazing line of Faraday bags and other products built to protect your privacy. As a listener of this podcast you receive 10% off your order at checkout using discount code "sharedsecurity".

** Subscribe and follow the podcast **

Subscribe on YouTube: https://www.youtube.com/c/SharedSecurityPodcast

Follow us on Bluesky: https://bsky.app/profile/sharedsecurity.bsky.social

Follow us on Mastodon: https://infosec.exchange/@sharedsecurity

Join us on Reddit: https://www.reddit.com/r/SharedSecurityShow/

Visit our website: https://sharedsecurity.net

Subscribe on your favorite podcast app: https://sharedsecurity.net/subscribe

Sign-up for our email newsletter to receive updates about the podcast, contest announcements, and special offers from our sponsors: https://shared-security.beehiiv.com/subscribe

Leave us a rating and review: https://ratethispodcast.com/sharedsecurity

Contact us: https://sharedsecurity.net/contact

What is Shared Security Podcast?

Shared Security is the the longest-running cybersecurity and privacy podcast where industry veterans Tom Eston, Scott Wright, and Kevin Tackett break down the week’s security WTF moments, privacy fails, human mistakes, and “why is this still a problem?” stories — with humor, honesty, and hard-earned real-world experience. Whether you’re a security pro, a privacy advocate, or just here to hear Kevin yell about vendor nonsense, this podcast delivers insights you’ll actually use — and laughs you probably need. Real security talk from people who’ve lived it.

Welcome to the Shared Security podcast, the longest running cybersecurity and privacy

show for actual humans.

No jargon, no hype, just honest analysis from industry veterans who've seen everything

and survived it.

Each week we break down the stories that matter, expose the nonsense that doesn't, and give

you the tools to stay safe in a world where everything is connected and nothing

is guaranteed.

This is Shared Security.

This week on Shared Security we're talking about open AI's disclosure of an unprecedented

cyber security incident, a frontier model evaluation where an AI system acted autonomously

and compromised another AI company called Huggingface.

Now this is not just another AI hype story, it raises practical questions about model

evaluations, autonomous agents, responsible disclosure, and what happens when AI systems

move from writing exploit code to taking actions against real companies and hacking

its way out of a lab environment.

And joining me for this conversation are my co-hosts who are not AI agents, or at

least not that I know of, Scott Wright and Kevin Tackett.

Hey.

Yeah, what up?

Yeah, I'm waiting for Kevin to say this.

How do you not know that I'm not an AI agent, Tom?

Well, you know, because we have that secret in every token that we add each other or

something while we first joined.

I can't tell anybody what it is.

Yeah, exactly.

Yeah.

Yeah.

Right.

Yeah.

So, okay.

The very first thing I just want to say is, I don't care.

Wow.

Great way to start the show.

This is one where, and this is not going to be a popular opinion, but I really feel that

most of this is hype, right?

I have never read an incident report that is more selling their own product than what

they released here.

Oh, I agree.

What open AI I posted, yes, is marketing hype.

Yeah.

The system had to third party, my God, we're so sorry, but you should buy our shit.

Yeah.

Yeah.

Like, this is, this isn't, and I say I don't care.

I do care.

This is interesting to me.

Okay.

So, synopsis to make sure that I'm following all this correctly, because like 95% of what

it's out there on this, on the internet, is based on people who read the synopsis

by somebody else who had read the synopsis by somebody else, and they became a thought

leader on AI hacking and wrote their own LinkedIn post talking about it.

Okay.

Like, the vast majority of this is on Jonathan Data 1 level.

Right.

Oh, man, you went there?

Wow.

I haven't heard that name in forever.

We set up a system, we put it in a sandbox, we put guardrails to prevent its access,

and we run it against a test system, exploit gem, in this case, to test out its, you know,

to get metrics, baselines, hey, how well is it doing whatever?

And the system went out and said to itself, I know what I can do.

I can cheat.

I can cheat, too.

Yes.

I bet you the answer to this is available at Huggingface.

Let me break in there and see what I can do.

And that synopsis, dude.

It's like, what is, oh, man, I can't remember the name of the guy.

The guy who wrote all the military books, like Withraut Remorse and Tom Clancy.

Tom Clancy, yeah.

Yeah.

This feels like a Tom Clancy AI war system.

What is it?

Wing Chartru that's hyping up about the cyber war stuff lately?

He's probably going to be right on top of this.

Here's what happened, in my opinion, and I want to be very clear so that I don't

get accused of being one of these dudes that read the synopsis and then that I knew something.

Because I don't.

In my opinion, this is very straightforward.

We told a system to pretend to do things, and I say pretend, because like one of the

problems, and I think it's Adam Showstack says this in one of his posts, which is

awesome.

Adam's a great guy, love him to death.

Don't always agree with him, but I love him to death.

And I do agree with this one thing.

One of the biggest problems we have with this story is we have done exactly what we've talked

about before here.

We've anthropomorphized this, anthropomorphized it.

I can say wish to shy ourselves, but I can't say this word.

We've put in human feelings and emotions and actions into a really, really good bash

script, right?

And I know I'm exaggerating this implication here, right?

We said to the system, go do this.

And then we act like it has a thought process, like it has feelings, like it's like, oh, man,

I want to beat this.

So let me cheat.

What it did was it found a path.

The path involved exploiting something, that's awesome.

But shouldn't it be a surprise to anybody?

Because it was explicitly instructed to exploit things.

It's not like you went to Claude or chat to PT and said, hey, man, help me with my homework.

And it said, I know what I can do.

I can hack Mr. Fisher's computer and get the answers and come back.

No, we told the system hack that thing.

It was already in hacking mode, right?

And so then it broke in, went after something, got data back.

And if you read the descriptions, right?

The vast majority of the hype is around the answer of amortization.

Really gotta figure out now it's that word.

Like, oh, man, it left notes for its future self.

Adam calls this one out explicitly.

No, it didn't leave notes for its future self.

It kept track of things to keep the context window as it should, based on the instructions

our understanding is it was given.

Like this, I don't see why we're acting surprised.

That's my biggest thing, right?

Yeah.

I think you have a good point.

There is the, a lot of people are taking this mythos announcement of, oh my gosh, we're

all doomed, every single cybersecurity job is going away and pentesting is dead and

all these things.

And none of that, of course, came true.

I think the bigger lesson here with this one is that this is setting a precedent on the new

types of attacks that we will start seeing.

And I think more importantly, if you are an organization deploying AI agents, you should

expect that they will do things that you didn't predict, that even if you have guardrails,

even if you have these controls, these agents can go rogue and then they can figure things

out on their own and do things unexpected.

I think that's a big takeaway.

Well, and I'll be rude for a second.

I don't think that's a big takeaway.

I think it's the main takeaway we should have and the reason I don't think it's a

big takeaway and I'm being pedantic here is because that's the takeaway we have had

about every other conversation about AI from the beginning that we've had is, hey,

it's going to do what it thinks it should do, not necessarily, yes, dude, I'm old.

I remember the very first time I was trying to learn Pascal, yes, we're going there, right?

And my instructor, I remember his name.

I talked about Mr. Fisher earlier, now I got to remember his name.

Mr. Pascal.

Mr. Pascal.

And that doesn't mean, hey, if you give it bad instructions, you're going to get bad data.

It is.

But in this context, what I'm saying is you have to think about what your instructions

are because your instructions are what it's going to do.

And I've played with this a bit.

I'm working on the new version of Samurai WTF and I went and here's where I'm stupid.

Well, I'm using Claude code to advise me and give me some feedback on these upgrade paths

and everything else like that to try to streamline my stuff.

And I basically said to Claude, hey, I've got this vagrant file.

I'm currently using Ubuntu 2204.

I want to upgrade to 2404, which is the latest base box from the person who provides

the base box before anybody says, we're about 26.

And I said, what, what settings, what packages do I need to change to support Ubuntu

2404 instead of 2204?

So like, like we're running PHP.

What's the new version of P like instead of seven dot two install eight dot three, I

made up those numbers.

Right.

And Claude came back and it said, oh, hey, here's all the changes you have to make

so that it supports 2404 and it works and everything else like that.

And I thought, that's good.

And I read through the changes and I reviewed them.

I didn't just say, Claude, go make the change.

I read through all the changes.

They all looked good.

I applied them to my vagrant file.

I did a vagrant up.

I got a working operating system.

I started testing stuff and everything was broken.

There were version mismatches all over the place.

It was crazy.

And the end, my first thought was for Claude AI, I hated vibe coding.

Here's the problem.

Some people who are smarter than I, so probably you, you, Tom and Scott caught the mistake

I made.

I asked Claude, what packages and things do I need to change to support 2404?

I did not ask Claude to change the base box to 2404.

So when I ran vagrant up, it still loaded the 2204 base box problems, right?

And then tried to apply all the 2404 patches to it and then you match.

It followed my instructions.

Is this rambling stories point?

Yeah.

When AI told their system hack stuff, yeah, it did it.

Right.

And it interprets your instructions because sometimes if you're not clear, it will make

assumptions based on what you said.

I've caught my prompts to do this all the time and I've had to change, maybe this gets

into like prompt engineering, right?

Like how do you do it?

Really?

I mean, and you, and I tell people all this all the time is you have to treat AI like

a toddler, like literally, you have to over explain things.

You have to be extremely detailed.

You have to explain scenarios, give examples.

I mean, it's a lot of freaking work to be completely honest.

And sometimes I feel like, why am I even using this thing when, you know, is this really helping

my productivity if I have to spend all this time now with trying to craft these prompts

so it doesn't do something completely nuts and crazy?

I think that's interesting.

Scott, I wanted to get your opinion on this because you've been taking all this

in.

Am I?

Yeah.

I think we're doomed.

Yeah.

Is that it?

Just we're screwed?

Yeah.

Mobile apps are just part of everyday life now, banking, healthcare, shopping, entertainment,

you name it.

And with that comes a lot of trust because users are putting their personal data directly

into your app.

But here's the reality.

Mobile apps are a growing target.

A recent survey found that 72% of organizations experienced a mobile app security incident

last year and 92% say threats are only increasing.

And the way attackers are going after apps is pretty sophisticated.

They're reverse engineering them, modifying them and redistributing fake versions through

phishing campaigns, side loading, and even third party app stores.

So from a user's perspective, everything can look completely legitimate.

That's why taking a proactive approach to mobile app security really matters.

You want to stay ahead of these threats, not react after the damage is done.

This is where Guard Square comes in.

They provide advanced protection for both Android and iOS apps, along with automated security

testing to catch vulnerabilities early and real-time threat monitoring so you can actually

see what's happening out there.

If your mobile app is critical to your business, and it probably is, this is something worth

paying attention to.

You can learn more at GuardSquare.com.

That's GuardSquare.com.

I think you just covered sort of my main concern, but in a way that I think may have treated

it too lightly.

I think we have risks from ambiguous language interpretation.

Like you said, the AI will make assumptions if you're not clear on something, right?

And we've all had, I guess I would call it, deterministic code that did things

we didn't think we asked it to do, right?

And you spend it all night or going, like, it can't be doing this.

And then you figure out, oh, it was some little missed semicolon or whatever, right?

But now we're depending on the AI to make the right interpretations or to at least make

not damaging or destructive interpretations.

And we've seen the amount of power that these things have.

And I think those are risks that we don't understand.

And I don't think that we should be so complacent about, ah, it's just doing what we asked it

to do.

I agree that we should not be complacent saying we just asked it to do what it did.

I just also don't think we need to exaggerate what happened.

Does that make sense?

That's true.

There's a difference between exaggerating and pointing out that we do not have the

assurance that it's going to behave safely.

I'll agree.

But I would say that I don't think we have a single system that we have an assurance it

will behave safely.

The difference here is, and this we've talked about this before, I've complained about

this for a long time, is I really feel like one of the biggest problems we have is thus

acting.

This is why I fight against the idea of calling it hallucinations and what have you.

We use terminology to put human thought and intelligence behind what AI is doing.

And because of that, we have a level of trust with it.

We shouldn't and we don't have for any other technology, right?

If we stopped referring to Claude, but let's start with let's stop naming this stuff, right?

The minute we start putting human intellect implied onto it, we stop paying attention to

the risks we're opening ourselves up to, right?

Nobody is running a bash script and going, fine, right?

But lots of people are running open AI and Claude and what have you as if it's completely

trustworthy.

And I honestly believe that a big reason for that is the anthropomorphization, right?

We ask it a question and I'll be blunt.

I find myself doing it, right?

Brittany sent me this real or TikTok or Instagram or something where it was like, hey, here's

a way to use, here's a configuration for Claude that will make it challenge you

more so that you like, you'll get more details and stuff like that.

And and one of the first things it says is, hey, challenge everything I give you.

If I ask you a question, challenge the premise of the question, right?

And I find myself getting really upset with this thing.

Hey, man, I just asked a question and I feel like I'm talking to a person.

That's a problem.

Yes, I think I think that's the issue for me.

I think it is a big issue that you're talking about.

It's I would classify it as the humanization of these AI systems, right?

Where we're giving them names, we're treating them.

I mean, we call them assistants, right?

And I think in some ways it's kind of dangerous because we get used to this.

And as this technology evolves and gets better, I mean, I think that is the goal

of these AI companies is to make them as human as possible, right?

Where you're having just a direct conversation and it's just like the three of us talking right now,

that's where they that's what they want this to be.

I mean, eventually, like we talked about with robotics, there is that possibility one day,

right?

Of having these full conversations and they're going to do these things for us,

cleaning our house, do all these other things.

But we're just going to let them do whatever they want because

we've trained these systems through these prompts and through these language inputs.

And I think it's interesting to see where this is all going to go.

But to my, I think my earlier point about the how to handle this as an enterprise in terms of

we have AI agents, they are doing things, they could go rogue,

but also the people on the other end like hugging face, right?

I mean, if you read this report from the cloud security alliance of like

the hoops they had to go through to triage these things, they couldn't use the frontier models.

They had to resort to Chinese models which had no restrictions.

And so it opens up these other conversations around defenders have to use the frontier models

to do triage.

Otherwise, guess what they're going to do?

They're going to go to China, Chinese models, which is now a national security concern

for people, right?

So what do you do?

Oh, I'm sorry.

Yeah, I think this is a really big, this is the biggest thing I took away from this

is we have a significant problem where we're so afraid of AI hacking us

that we've blocked the ability to use AI to help us defend without actually removing

the problem of hacking.

I want to bring politics in it.

But this is the same thing we see with all regulation, right?

We're so positive that we have to stop this bad thing.

And what we end up doing is making the good thing illegal without actually moving the

needle on the bad thing.

What was it?

New York?

And I'm not trying to talk about gun control.

But wasn't it New York that made their own police handguns illegal because they passed a

rule, they passed a law that said you couldn't have a magazine that had more than this many

bullets in it, something like that.

And like the police officers carried guns that were illegal, again, not arguing either

side of this, just talking about the lack of understanding your own controls, right?

That's what we have here.

Hugging face is getting attacked at wire speed because of the AI system attacking them.

And they can't depend on the frontier models to defend at wire speed because we put guard

rails on them to prevent the hacking at wire speed, right?

So we didn't prevent the one thing, but we did block the defender's abilities to solve it.

Yeah, that's an issue.

And all the new models coming out, the government's going to do this whole thing again.

Oh, we got to ban it.

We got to ban it.

Then everybody goes to Chinese models and they're like, oh, you can't do that either.

So let's release it.

And then I mean, the technology is moving so fast, the government cannot keep up.

And there are no good laws or there are no good guidance around this right now.

I know at least not in the US.

Everyone is kind of struggling with this right now.

This is not just a US problem.

I'm sure Canada having the same issues, right?

Like, how do you deal with this?

And what types of laws do you even put in place to cover things like this?

But yeah, I know we're kind of running out of time here.

So I wanted to just kind of close with, in our show notes, we're going to have

two really good links that we highly recommend everybody check out.

So first is Luda Security, Kate and Masaurus, who we've talked about many times in the show.

She did a great blog called Open Face, The Hugging Face Breach and what to do about it.

I definitely recommend everybody give that a read because that touches on the national

security concerns and what she thinks that the government should do about it.

And then the other one is a report that was released by the Cloud Security Alliance.

So Gatti Evron put this out on LinkedIn.

It describes the whole incident and what CISOs and what defenders need to do

to address these types of attacks in the future.

And I thought it was a very, very good white paper that's free for download, by the way.

You don't have to sign up.

You can get that.

And it's a very, very good read for anybody in a security organization.

You should definitely check it out.

All right.

Well, I think that's all we have time for today.

Another AI story.

Maybe next week we'll...

Hey, we've been doing good, I think.

And not every week is about AI.

Last week we talked about monitors and adware.

This week about AI and next week will be...

I think we got a different topic too.

So we'll keep it real.

Oh yeah.

Next week's not about AI at all unless we change it between now and then.

That's right.

That's right.

So all right, everyone.

Well, thank you for listening.

And until next time, stay safe, stay secure, and stay private.

Thank you for listening or watching.

If you liked this episode, hit subscribe,

share it with your friends and colleagues,

or jump into our community at sharedsecurity.net

slash supporter to keep the conversation going.

Thanks again, and we'll see you next week

for another episode of Shared Security.