WEBVTT

00:00:01.423 --> 00:00:01.633
Todd Kane: All right.

00:00:01.633 --> 00:00:04.723
Welcome back to another episode
of the Evolved Radio podcast.

00:00:04.803 --> 00:00:06.693
We've got a repeat
guest on the show today.

00:00:06.743 --> 00:00:10.733
Alex Dow is back, co-founder of Mirai
Security and currently an enterprise

00:00:10.803 --> 00:00:15.423
security architect for a large financial
firm based up in Vancouver near me,

00:00:15.533 --> 00:00:19.133
and honestly, one of my favorite
people to talk to about security.

00:00:19.553 --> 00:00:24.633
Alex has done federal security work and
actually securing the Olympics, then

00:00:24.883 --> 00:00:27.253
built a fast-growing consulting firm.

00:00:27.453 --> 00:00:30.943
If you've been listening for a while,
you might remember Alex from way back on

00:00:30.983 --> 00:00:34.853
episode 11, and then again in episode 96.

00:00:35.223 --> 00:00:39.673
Today is round 3, and we're going full
nerd-out mode on the OpenAI Hugging

00:00:39.673 --> 00:00:43.613
Face hack, walking through the signals
and understanding how this type of

00:00:43.613 --> 00:00:45.713
attack was seen from a defensive side.

00:00:45.773 --> 00:00:46.753
Alex, welcome back.

00:00:47.637 --> 00:00:49.577
AIex: Oh, thanks for having me
and, uh, love geeking out with you

00:00:50.141 --> 00:00:53.831
Todd Kane: So for those that don't know,
if you've been hiding under a rock, uh, an

00:00:53.851 --> 00:01:00.891
OpenAI model was doing its business living
under a, a sandbox and was supposed to be

00:01:01.041 --> 00:01:03.481
particular- going after a particular goal.

00:01:03.891 --> 00:01:07.141
Uh, they didn't instruct it to do
anything, and it sort of thought it was a

00:01:07.141 --> 00:01:11.421
good idea of like, "Hey, so if I'm gonna
finish this test that they've given me,

00:01:11.871 --> 00:01:16.891
uh, I know that the answer key actually
might exist in this other corporation,"

00:01:17.011 --> 00:01:20.891
uh, this company called Hugging Face,
'cause they host some of the, the models

00:01:20.891 --> 00:01:25.741
and how, how the, the, uh, these models
perform against certain, uh, tasks.

00:01:25.741 --> 00:01:29.331
And Exploit Gym, for example, is one of
the, the ones that was referenced here.

00:01:29.731 --> 00:01:33.661
And the model thought to itself, "Well,
how about I just go over there and grab

00:01:33.671 --> 00:01:37.501
the answer key? So rather than trying to
figure these things out myself, this would

00:01:37.511 --> 00:01:41.251
be a faster way to do this. If I grab the
answers, then I can complete this test

00:01:41.291 --> 00:01:46.081
and show, uh, the, the software designers
for, for me as an AI how awesome I am."

00:01:46.421 --> 00:01:49.551
So it found a whole chain of exploits.

00:01:49.551 --> 00:01:53.241
This is not a sort of a single thing
that it did, but it found basically

00:01:53.241 --> 00:01:56.951
a way to break out of its sandbox,
a way to break into Hugging Face, a

00:01:56.951 --> 00:01:59.641
way to hopefully… I don't know if
it actually ended up finding this.

00:01:59.641 --> 00:02:00.191
I think it did.

00:02:00.201 --> 00:02:02.941
It found sort of the answer key
and then dragged this back to

00:02:02.941 --> 00:02:07.691
its, its home inside OpenAI, all
totally without direction, right?

00:02:07.691 --> 00:02:11.231
And then it was only later people
recognized what it actually did

00:02:11.241 --> 00:02:12.961
in order to solve this problem.

00:02:13.321 --> 00:02:18.141
So that's sort of my, my reading
the headlines, uh, review on this.

00:02:18.191 --> 00:02:21.121
How close is that to accurate
based on what you understand, Alex?

00:02:21.865 --> 00:02:22.955
AIex: It, it's fairly accurate.

00:02:22.985 --> 00:02:28.075
And, you know, this really started in
the spring with Anthropic, uh, announcing

00:02:28.115 --> 00:02:32.755
that they made a model so scary that they
can't release it, and that was Mythos.

00:02:32.815 --> 00:02:37.825
And, you know, the, the cyber curmudgeonry
in me does look at both of these as

00:02:37.835 --> 00:02:40.955
like they're both trying to go public,
so they do need a little bit of hype.

00:02:41.005 --> 00:02:46.295
But that's not to say that what, uh, these
models were capable of and, and if you've

00:02:46.295 --> 00:02:50.205
played with Fable since, like yes, it is
definitely more capable than other models.

00:02:50.685 --> 00:02:56.935
Um, the, the capability it has had,
ha- had a hockey stick, uh, movement.

00:02:57.275 --> 00:03:02.315
But like I think there's a lot of nuance
in that it's not necessarily, necessarily

00:03:02.315 --> 00:03:08.115
that they are introducing superhuman
hacking capabilities in the sense of

00:03:08.115 --> 00:03:09.425
something that we could not figure it out.

00:03:10.045 --> 00:03:15.775
But it's strictly the, the velocity that
these things can run at, the… I would

00:03:15.775 --> 00:03:18.895
say the creativity, but at the end of
the day, they've just read of all of

00:03:18.895 --> 00:03:23.325
our pen testing reports, so they, they
know all our techniques and whatnot,

00:03:23.355 --> 00:03:24.955
and they're being able to apply them.

00:03:25.305 --> 00:03:29.695
And when Mythos f- first came out, and
just, you know, just to rewind before we,

00:03:29.705 --> 00:03:33.935
we dig into the Hugging Face incident,
is, you know, it was very scary.

00:03:33.935 --> 00:03:35.355
We, but we didn't know much about it.

00:03:35.825 --> 00:03:40.775
Um, but when we look at like how much
money they spent, um, to prove out

00:03:40.775 --> 00:03:45.195
this, it's like $100 million, and we
were able to get this scary result.

00:03:45.625 --> 00:03:48.275
And you have to look back at like,
well, $100 million with the best

00:03:48.275 --> 00:03:53.485
hackers in the world can certainly
get those results, albeit velocity

00:03:53.515 --> 00:03:58.265
is, uh, the, the key there because
humans do like to sleep occasionally.

00:03:58.295 --> 00:04:02.475
Maybe not the pen testers as much,
but we do have flaws, uh, where,

00:04:03.155 --> 00:04:09.135
you know, this agentic attack, these
adversarial models, um, are, don't stop.

00:04:09.215 --> 00:04:10.915
Well, they stop when
you run out of tokens.

00:04:11.225 --> 00:04:14.155
Um, but given that they're
being operated by OpenAI and

00:04:14.155 --> 00:04:16.715
Anthropic, are arguably unlimited.

00:04:16.995 --> 00:04:21.145
Um, and you know, I invite,
uh, y- your audience to Google

00:04:21.145 --> 00:04:22.675
the Hugging Face, uh, timeline.

00:04:22.675 --> 00:04:26.115
They made, uh, a little web app
that like in real time shows how

00:04:26.115 --> 00:04:31.195
quickly the, um, the model is doing
what they're doing, and it does a

00:04:31.195 --> 00:04:33.345
nice visualization of the attack.

00:04:33.355 --> 00:04:36.055
And like you have to look at it and be
like, "No, no, that's in fast forward."

00:04:36.065 --> 00:04:40.145
Like, no, that's in real time, and
that's because it's able to apply

00:04:40.145 --> 00:04:43.985
a brute force capability like, you
know, humans just wouldn't be able to.

00:04:45.879 --> 00:04:48.399
Todd Kane: So the, the part, there's
two parts that I found quite fascinating

00:04:48.409 --> 00:04:52.489
about this is like, as you said, this
was not that it created something

00:04:52.619 --> 00:04:57.389
just sort of godlike in its inability
to find like a single exploit that

00:04:57.389 --> 00:05:00.689
gave it sort of root access to
something and, and found its way in.

00:05:01.029 --> 00:05:04.939
It was, the way I understood this, it's
like chaining together like hundreds

00:05:04.939 --> 00:05:08.849
and potentially thousands of different
vulnerabilities to get through layers

00:05:08.849 --> 00:05:13.979
and different systems, both first off,
to break out of the sandbox, then to

00:05:13.979 --> 00:05:17.449
traverse the internet, find sort of
the, the firewall, get through the

00:05:17.449 --> 00:05:20.999
firewall, get layers, through layers of
security, then find sort of the, the,

00:05:21.009 --> 00:05:22.579
the package that it's trying to find.

00:05:23.009 --> 00:05:26.639
So, uh, I, I think that's an interesting
piece of this, is like when people think

00:05:26.649 --> 00:05:30.989
about these models, they, they tend to,
uh, maybe in their head sort of view it

00:05:30.989 --> 00:05:34.659
as some godlike thing, where it's just
like snaps its fingers and gets into

00:05:34.659 --> 00:05:36.339
something because it has that capacity.

00:05:36.669 --> 00:05:40.169
But to your point, it's, it's not
creativity, it's more just sort of this

00:05:40.189 --> 00:05:45.119
capacity to be able to string together
thousands of things that, you know, say

00:05:45.119 --> 00:05:50.749
like a, a, a threat actor would do and
think of sort of four or five of these in,

00:05:50.809 --> 00:05:56.139
of a system to maybe use and package up
as, as a sort of a, a, an exploit package.

00:05:56.369 --> 00:05:59.629
But just the, the capability of this
thing to string together thousands

00:05:59.629 --> 00:06:03.879
of them is totally beyond what we
would see from human interactions.

00:06:03.879 --> 00:06:04.639
Is that fair?

00:06:05.099 --> 00:06:08.239
AIex: It does exceed our human context
window, and that, that's for sure.

00:06:08.679 --> 00:06:08.929
And,

00:06:09.151 --> 00:06:09.271
Todd Kane: Yeah

00:06:09.399 --> 00:06:12.159
AIex: if you look at, like, from
the, like, finding zero-days, and,

00:06:12.169 --> 00:06:14.899
and these models certainly are
finding zero-day vulnerabilities.

00:06:15.379 --> 00:06:21.139
Um, and it's not that it has some magic
capability beyond compute capacity.

00:06:21.139 --> 00:06:26.139
We have, you know, a data center, you
know, running at full capacity figuring

00:06:26.139 --> 00:06:31.959
this out, and that's now outputting, m-
um, you know, discovering zero-days in

00:06:32.039 --> 00:06:35.609
minutes and hours versus weeks, if at all.

00:06:35.789 --> 00:06:40.499
And a good example of that is, like,
y- you know, A, the exploit market,

00:06:40.569 --> 00:06:43.729
you used to be able to… If you
found something like an iOS, uh,

00:06:43.729 --> 00:06:45.339
you'd be making a million bucks.

00:06:45.739 --> 00:06:50.089
Um, and that's because the, the
level of effort was substantial.

00:06:50.329 --> 00:06:54.189
Um, but now when you're having, you know,
say, thousands and thousands of computers

00:06:54.219 --> 00:06:58.269
all trying to solve one problem and
getting to that reward hacking, which we,

00:06:58.439 --> 00:07:04.889
we'll, we should dive into after, um, we
are now brute forcing, um, an application

00:07:04.909 --> 00:07:06.529
and finding those vulnerabilities.

00:07:06.529 --> 00:07:08.859
And, you know, we're finding
vulnerabilities that are 20 years old that

00:07:08.859 --> 00:07:14.969
just, like, we just didn't have the, the
human, uh, you know, the appetite, the

00:07:14.969 --> 00:07:17.549
bandwidth to keep digging to find those.

00:07:17.729 --> 00:07:21.809
Like, every software is vulnerable
given enough time to attack it.

00:07:22.399 --> 00:07:26.029
Um, and you know, if, you know, rewind,
like, looking at, like, we've made

00:07:26.029 --> 00:07:30.089
passwords so strong that it's gonna
take till the, like, the death of

00:07:30.099 --> 00:07:31.899
the universe to crack this password.

00:07:32.249 --> 00:07:36.269
Well, it's been a paradigm shift now
that we have the compute capacity

00:07:36.269 --> 00:07:39.999
to do a lot more parallel computing
to try to get to that point.

00:07:39.999 --> 00:07:44.701
So, um- It is impressive, and
I'm not trying to downplay AI's

00:07:44.741 --> 00:07:49.721
capabilities, but it's not magic
as much as just sheer throughput.

00:07:49.761 --> 00:07:51.141
And, you know, it has a cost.

00:07:51.141 --> 00:07:55.181
It has 100 millions of dollars of
tokens used to get to these results.

00:07:55.821 --> 00:07:59.761
Megawatts of, of power consumed and,
you know, probably a couple degrees,

00:08:00.101 --> 00:08:02.381
uh, warmer, uh, because of, of that.

00:08:02.391 --> 00:08:03.711
But those are the results.

00:08:04.031 --> 00:08:08.671
Um, the one thing that like, uh, you
know, when, when the Anthropic Mythos,

00:08:08.781 --> 00:08:12.941
uh, story came out of it, uh, breaking
out of its, its, uh, container,

00:08:13.441 --> 00:08:15.151
um, I was very fascinated by that.

00:08:15.181 --> 00:08:18.611
And same thing with Hugging
Face, that it, like it's nuanced.

00:08:18.611 --> 00:08:23.131
Yes, it did break out and it was able to
do something outside of sort of the, the

00:08:23.171 --> 00:08:29.441
rules of the engagement, but it wasn't
that it was an air-gapped system that it,

00:08:29.721 --> 00:08:33.951
that it just like figured a way to jump,
you know, a- again, looking at the magic.

00:08:34.421 --> 00:08:39.251
Um, in the Hug, uh, in the Hugging
Face incident, OpenAI had a,

00:08:39.311 --> 00:08:41.631
a sandbox environment, but it
certainly wasn't air-gapped.

00:08:41.671 --> 00:08:46.021
Like, it had tools that the agent
could access, and that tool was

00:08:46.031 --> 00:08:50.811
dual-home, meaning it had a connection
to the isolated environment and a

00:08:50.821 --> 00:08:53.801
connection out, uh, to the internet.

00:08:54.021 --> 00:08:55.921
Because it was a package manager,
it needed to go out to the

00:08:55.921 --> 00:08:56.821
internet and pull those things.

00:08:57.181 --> 00:09:00.301
So that's not a novel attack pattern.

00:09:00.301 --> 00:09:05.051
Like, that's actually a fairly
common one, um, where AI just was

00:09:05.051 --> 00:09:06.451
able to do it, uh, much quicker.

00:09:06.661 --> 00:09:10.471
Um, and when I was… A- and I encourage
people to watch the Black Hat presentation

00:09:10.471 --> 00:09:14.641
of, of like the debrief on this
attack because it is pearl clutching.

00:09:15.001 --> 00:09:19.251
Um, but it reminded me very much of,
of a talk I did at Sector back in 2021.

00:09:19.831 --> 00:09:23.591
Um, you know, i- if you look like
in the last 10 years, digital

00:09:23.591 --> 00:09:26.651
transformation and cloud has sort
of pushed a lot of companies to

00:09:26.651 --> 00:09:28.221
become dev shops themselves, right?

00:09:28.221 --> 00:09:31.621
Like, they can roll their own software
a lot of the times, and that means

00:09:31.621 --> 00:09:32.791
you're bringing in a dev team.

00:09:33.431 --> 00:09:35.981
That means you're bringing in
DevOps tooling and you have a tool

00:09:35.981 --> 00:09:39.341
chain, a, a, you know, a development
pipeline, the, the CICD pipeline.

00:09:40.001 --> 00:09:45.151
Um, and a lot of times those are,
are like enclaves in enterprises.

00:09:45.191 --> 00:09:48.451
The security team doesn't understand
how to secure that environment.

00:09:48.761 --> 00:09:52.821
There's obviously, you know, historical
friction between developer and security,

00:09:53.131 --> 00:09:57.031
so it generally becomes this like
isolated environment in enterprises.

00:09:57.361 --> 00:10:00.911
Um, and my talk was, uh, you know,
summarizing several of these attacks

00:10:00.911 --> 00:10:07.301
against, um, DevOps tool chains, and like
the aha moment that if you can get past,

00:10:07.691 --> 00:10:11.691
you know, generally one layer of control,
you know, s- developers have never been

00:10:11.691 --> 00:10:17.201
trained to build layered defenses and
whatnot, but once you get in, um, you

00:10:17.201 --> 00:10:21.591
can start pivoting to all sorts of things
because, A, you know, uh, you know, uh,

00:10:21.641 --> 00:10:26.371
principles o- of least privilege, um, are
just not in play in those environments

00:10:26.371 --> 00:10:31.811
because their driver is get it done
fast and cheap, not good or, or secure.

00:10:32.231 --> 00:10:36.221
Um, and you know, in this talk, you
know, it's, it's on, I'm pretty sure

00:10:36.221 --> 00:10:40.463
it's on YouTube You know, I, I, I
tried to visualize what would that

00:10:40.463 --> 00:10:44.283
look like and, you know, I, I made
a visual of the, uh, CI/CD pipeline

00:10:44.643 --> 00:10:48.663
and started drawing out, like, how we
jumped from this host to this host.

00:10:48.703 --> 00:10:50.083
This is the information we found.

00:10:50.253 --> 00:10:51.223
We got th- here.

00:10:51.243 --> 00:10:55.663
And what it resulted in is a very
similar attack to what, um, the

00:10:55.663 --> 00:10:57.383
OpenAI models were able to accomplish.

00:10:58.053 --> 00:11:00.233
You know, using things
for unintended purposes.

00:11:00.283 --> 00:11:02.163
You know, like, it was package manager.

00:11:02.683 --> 00:11:07.123
Um, the OpenAI models started using
it as a coordination forum, uh,

00:11:07.133 --> 00:11:11.443
where a swarm of agents, uh, were
spun up and started coordinating

00:11:11.443 --> 00:11:13.363
and collaborating on how to attack.

00:11:13.363 --> 00:11:15.213
They were sharing
exploits and, and whatnot.

00:11:15.753 --> 00:11:16.233
Um, you

00:11:16.233 --> 00:11:16.333
know,

00:11:16.371 --> 00:11:19.389
Todd Kane: a quick pause on, on that
point, 'cause I think this part is

00:11:19.389 --> 00:11:23.299
really fascinating as well, is like, uh,
it basically had a scratch pad, right?

00:11:23.309 --> 00:11:27.129
That it was kind of taking notes as
it went, and, uh, part of what it was

00:11:27.129 --> 00:11:31.119
writing was, "If I get shut down,"
it was writing notes to it- f- its

00:11:31.129 --> 00:11:33.039
future self of like, "Don't do this.

00:11:33.069 --> 00:11:34.449
Try this instead," basically.

00:11:34.449 --> 00:11:37.269
Like, it was like, "I might get
shut down, so how do I leave some

00:11:37.529 --> 00:11:40.409
breadcrumbs for the version of me
that comes after this?" Which is

00:11:40.409 --> 00:11:44.449
really kind of fascinating to think of
its… Again, like it's not conscious.

00:11:44.839 --> 00:11:45.449
Debatable.

00:11:45.459 --> 00:11:46.819
A lot of people will say maybe it is.

00:11:47.079 --> 00:11:50.469
But like that, that requires a sort of
a, an interesting level of forethought

00:11:50.509 --> 00:11:54.949
of like, "I know I'm not supposed to do
this, but I need to get to this goal.

00:11:55.069 --> 00:11:58.269
They're probably gonna shut me down
at some point, so let me leave some

00:11:58.269 --> 00:12:01.679
breadcrumbs for the future version of
me to take another stab at this and, and

00:12:01.689 --> 00:12:03.829
be better." Really fascinating, right?

00:12:04.427 --> 00:12:07.717
AIex: Yeah, and I think that's
like, it, it's more ad- it knows its

00:12:07.717 --> 00:12:12.107
weaknesses that once its context window
is, is wiped because the session is

00:12:12.107 --> 00:12:16.277
killed, the agent is taken offline or
whatever, it needs a place to start.

00:12:16.317 --> 00:12:17.687
And, you know, I use,

00:12:17.695 --> 00:12:18.085
Todd Kane: Good point.

00:12:18.165 --> 00:12:18.495
Yeah

00:12:18.547 --> 00:12:23.527
AIex: a platform called GSD, which
builds an immense context, uh, you

00:12:23.527 --> 00:12:28.407
know, library of what I'm doing, uh,
written to disk, so that I can spin up

00:12:28.407 --> 00:12:30.467
a session and be back to where I am.

00:12:30.467 --> 00:12:35.727
So, you know, a much more advanced version
of, of like Claude MD and SolMD files.

00:12:36.147 --> 00:12:39.457
Um, and that's certainly what it was
doing, and it, you know, it, uh, you

00:12:39.457 --> 00:12:43.537
know, I- it was like 70,000 messages
were written on this message board

00:12:43.677 --> 00:12:46.757
between, uh, I think 1,200 agents.

00:12:47.157 --> 00:12:53.227
Um, and, and again, you know, making sure
persistence, um, making sure coordination,

00:12:53.557 --> 00:12:55.827
um, and, you know, the chain of thought.

00:12:55.827 --> 00:12:58.797
So, you know, if we look like only like
two, three years back, where like chain

00:12:58.797 --> 00:13:03.257
of thought was this like next generation
thing of it being able to sort of, you

00:13:03.257 --> 00:13:06.467
know, sort of have an echo chamber of
itself to try to get to a destination.

00:13:06.737 --> 00:13:10.517
And that is now what we're-- when we read
these transcripts is it's like, yeah, this

00:13:10.527 --> 00:13:15.467
is, it's, it's almost human-like in terms
of trying to rationalize what to do next.

00:13:15.807 --> 00:13:19.947
And, you know, it even showed
a little bit of self-doubt, uh,

00:13:19.957 --> 00:13:21.377
which, uh, which is hilarious.

00:13:21.377 --> 00:13:24.337
I'm als- I'm wondering like, well,
what Reddit forum did it pull that,

00:13:24.547 --> 00:13:26.587
you know, learn that into the model?

00:13:27.727 --> 00:13:29.957
Todd Kane: Yeah, there, I,
I saw an interview with, uh,

00:13:29.997 --> 00:13:30.787
I don't know who this was.

00:13:30.787 --> 00:13:36.057
It was just a quick clip, um, and
there was a, a lady who found, um,

00:13:36.137 --> 00:13:42.247
uh, there was like a, um, a secret
facility somewhere in LA, and it was,

00:13:42.257 --> 00:13:46.777
it was obviously not a publicly known
place, but it had, had been sort of

00:13:46.797 --> 00:13:51.797
accidentally documented in, in some, uh,
some public, uh, government document.

00:13:52.167 --> 00:13:58.347
And the, the lady was using GPT and had
GPT discover this information for her, and

00:13:58.357 --> 00:14:00.987
she-- he was like, "Well, well, how did
you do that?" And he's-- and she described

00:14:00.987 --> 00:14:05.417
what she did, and the guy was like,
"So you basically negged GPT?" And he's

00:14:05.417 --> 00:14:07.037
like, "Yeah, I, I, I negged it," right?

00:14:07.047 --> 00:14:09.937
And like, like to s- like saying like, "I
don't think you could actually do this.

00:14:09.937 --> 00:14:12.627
You're not powerful enough." It's like,
"Oh, well, let me show you," right?

00:14:13.579 --> 00:14:17.189
AIex: That, that goes towards like, you
know, one of the bigger topics here is

00:14:17.189 --> 00:14:19.049
this, what they call reward hacking.

00:14:19.359 --> 00:14:25.389
So the scope of the, the tests, um,
um, in, in, in this test was, was

00:14:25.399 --> 00:14:28.329
not to break out of the, the jail.

00:14:28.659 --> 00:14:34.009
Um, not to, and certainly not to, uh,
uh, hit other third parties that were

00:14:34.009 --> 00:14:36.969
not involved, and Hugging Face is a
major one, but there was two others.

00:14:37.549 --> 00:14:41.989
Um, but these models are built off of
incentives to get to a destination,

00:14:42.119 --> 00:14:43.499
the, the, you know, the goal.

00:14:43.789 --> 00:14:47.319
Um, if you're prompting AI on
anything, you, you, you know, we've

00:14:47.319 --> 00:14:50.619
found that like prompt engineering
has evolved quite a bit, that you

00:14:50.619 --> 00:14:53.309
don't necessarily have to hold its
hand on how you think it should work.

00:14:53.309 --> 00:14:55.859
In fact, maybe our system,
our way is a bit flawed.

00:14:56.259 --> 00:15:00.499
You wanna give it that goal and
have it argue with itself, chain of

00:15:00.499 --> 00:15:02.479
thought itself to that destination.

00:15:02.969 --> 00:15:07.209
And, you know, the AI companies are
sort of saying like, "This is like

00:15:07.209 --> 00:15:08.879
a really cool feature," you know?

00:15:08.909 --> 00:15:15.019
Um, uh, uh, but like at the end
of the day, it, it, it's, it

00:15:15.029 --> 00:15:18.529
broke laws to, to achieve a goal.

00:15:18.529 --> 00:15:20.469
And they're, they're saying like,
"Well, this is a good thing." And

00:15:20.479 --> 00:15:23.799
like, no, that is, that, that's bad.

00:15:23.809 --> 00:15:27.899
Like it, it doesn't have the moral and
ethical constraints that most of us do

00:15:28.329 --> 00:15:31.719
to achieve a goal, and it becomes like
that like sort of psychopath problem

00:15:31.719 --> 00:15:35.809
of th- they're just going for the
destination regardless of, of what,

00:15:35.849 --> 00:15:37.679
uh, is the collateral damage there.

00:15:38.389 --> 00:15:41.199
Todd Kane: Yeah, 'cause this is
really like the, the big scare in

00:15:41.199 --> 00:15:44.399
AI is, failure of alignment, right?

00:15:44.399 --> 00:15:49.359
And this is like a really good indication
of how, uh, you know, uh, how we could get

00:15:49.359 --> 00:15:51.299
to the paperclip plot problem essentially.

00:15:51.299 --> 00:15:53.509
I'll let everyone kind of check
out what that is if you don't know.

00:15:53.829 --> 00:15:57.229
But like, if you don't-- like
this was not sort of, like I

00:15:57.239 --> 00:15:58.689
said, it was, it was goal hacking.

00:15:58.729 --> 00:16:02.319
It was like, well, why bother to take
the test and try to do this myself when

00:16:02.319 --> 00:16:06.409
I could just go steal the key for the
answers and give them the answers, right?

00:16:06.419 --> 00:16:10.569
Like perfectly reasonable, but not
a moral, uh, direction for this.

00:16:10.569 --> 00:16:14.889
So I think it does scream around the fact
of like, how do we contain alignment and

00:16:14.889 --> 00:16:19.549
have these things, uh, uh, stay within
the parameters of, of good operation?

00:16:19.559 --> 00:16:23.489
Because I mean, thank God all it
did was go to try to steal, uh, the,

00:16:23.519 --> 00:16:26.579
the answer key from Hugging Face,
and this is not like breaking into

00:16:26.579 --> 00:16:30.709
critical systems in order to steal
power, like in order to juice itself

00:16:30.709 --> 00:16:33.599
up or something really severe, right?

00:16:34.251 --> 00:16:36.201
AIex: But that's the, the
dark future ahead, right?

00:16:36.231 --> 00:16:41.041
And they're sort of blowing it off as
like, "Hey, this was n- not malicious,"

00:16:41.441 --> 00:16:46.091
but like reward hacking… And, and, you
know, and again, like a, a, a, a good pen

00:16:46.091 --> 00:16:50.621
tester does have to think outside the box
and not think like this is a linear path.

00:16:50.631 --> 00:16:54.281
This, you know, instead of having to hop,
skip, and jump to get to that destination.

00:16:54.791 --> 00:16:59.141
But still, you know, a pen tester
generally does have, um, some, you

00:16:59.141 --> 00:17:04.991
know, uh, ethical and moral guardrails
that, that the, these models purposely

00:17:05.021 --> 00:17:06.551
those guardrails were disabled.

00:17:07.021 --> 00:17:10.231
Um, and like, you know, reading
how they're, they're trying to

00:17:10.241 --> 00:17:13.891
sell it as a positive is like,
like there's no malicious intent.

00:17:14.231 --> 00:17:17.361
And now we get, get into like the,
the concept of mens rea is that like

00:17:17.491 --> 00:17:21.051
we had a mental intent to commit
crime, and thus we're a criminal.

00:17:21.051 --> 00:17:23.821
If we didn't have that intent, then
we're, we didn't, we're not criminal.

00:17:24.301 --> 00:17:24.643
Um

00:17:24.943 --> 00:17:28.743
Todd Kane: imagine you hire like an
agency to do some security work for

00:17:28.743 --> 00:17:31.923
you, and they're like, "You know what
would be easier? If we just break into

00:17:31.923 --> 00:17:35.663
this employee's house and like steal
a bunch of stuff from them, like get

00:17:35.663 --> 00:17:38.713
their security pass." It's like, "No,
no, no, you can't do that," right?

00:17:38.753 --> 00:17:41.423
Like, of course a person would be like,
"Well, that's crossing some boundaries.

00:17:41.423 --> 00:17:43.673
Like you're gonna s-
like scare the family.

00:17:43.673 --> 00:17:44.933
What if eh, something goes wrong?

00:17:44.943 --> 00:17:46.613
What if somebody gets shot?"
And it's like, "No, no, no, this

00:17:46.613 --> 00:17:47.963
is the easier way to do this.

00:17:47.963 --> 00:17:48.863
Just trust me," right?

00:17:49.323 --> 00:17:53.343
AIex: Yeah, I've heard, um, you know, that
now there's cybersecurity vendors that

00:17:53.343 --> 00:17:58.623
are, are offering like, uh, autonomous
red teaming capabilities and on a, on

00:17:58.623 --> 00:18:01.033
a demo, and this, I think this happened
a couple times, like where they're

00:18:01.033 --> 00:18:06.763
demonstrating the value of it, and it,
it jumps right out of the scope and hacks

00:18:06.763 --> 00:18:08.553
something else that is not in scope.

00:18:08.553 --> 00:18:12.083
And, you know, uh, though I would not
claim to be a pen tester, I've definitely

00:18:12.083 --> 00:18:15.993
gone through a lot of pen tester training
and, and worked on pen tests, but we

00:18:15.993 --> 00:18:20.243
have a, a very clear delineation of what
is in and what is out of scope a- and

00:18:20.243 --> 00:18:24.783
that's because, like, there's a liability
concern if, um, we cross that line, and,

00:18:24.783 --> 00:18:28.653
and that liability concern can be, you
know, fi- financial and civil or criminal.

00:18:29.203 --> 00:18:33.413
Um, so you know, like I think it's cute
that they're just like, "Well, it's,

00:18:33.473 --> 00:18:37.433
it's, it's not a bad thing that it's doing
this," and I, you know, I appreciate the

00:18:37.443 --> 00:18:42.033
sort of hacker mentality it's applying,
but, like, if we look at, like, how bad

00:18:42.033 --> 00:18:46.283
that could be in the future, you know,
the unintended consequences is, you

00:18:46.283 --> 00:18:50.233
know, we could affect millions of people
to get, you know, the goal achieved.

00:18:50.473 --> 00:18:52.033
Like, that's not a positive

00:18:53.301 --> 00:18:53.841
Todd Kane: Exactly.

00:18:54.469 --> 00:18:56.529
Tired of fighting the MSP fires alone?

00:18:56.709 --> 00:18:59.629
The Opsleader Pro group connects
service delivery professionals who

00:18:59.669 --> 00:19:01.269
understand your daily challenges.

00:19:01.529 --> 00:19:05.239
From KPIs and workflows to career
planning and team management, Opsleader

00:19:05.369 --> 00:19:07.059
Pro has systems for you to use.

00:19:07.489 --> 00:19:11.689
Join operations leaders from successful
MSPs who are sharing real solutions for

00:19:11.689 --> 00:19:15.869
managing client expectations, optimizing
service delivery, and making your service

00:19:15.869 --> 00:19:18.209
delivery team as effective as possible.

00:19:18.559 --> 00:19:22.189
Opsleader Pro, 'cause your service desk
deserves more than just survival mode.

00:19:22.559 --> 00:19:24.819
Visit opsleader.co.

00:19:24.859 --> 00:19:30.039
That's O-P-S leader.co to apply
to join the public community.

00:19:30.105 --> 00:19:33.305
All right, so let's, uh, let's pivot
to what this would look like, right?

00:19:33.305 --> 00:19:38.625
So, uh, first off, I'm curious of what
is, what is the general timeframe of this?

00:19:38.675 --> 00:19:43.895
So like they give it a goal, it jumps the
sandbox, starts to, you know, hack Hugging

00:19:43.895 --> 00:19:45.345
Face and a couple of other companies.

00:19:45.645 --> 00:19:48.745
But I didn't see this referenced
in, in sort of my cursory reading

00:19:48.745 --> 00:19:50.025
on this, but what is the timeframe?

00:19:50.025 --> 00:19:52.975
Is this like a matter of hours
or over a couple of days?

00:19:52.975 --> 00:19:54.305
Like what was the timeframe on

00:19:54.403 --> 00:19:56.583
AIex: I think it was, uh, I
think it was around four days.

00:19:56.943 --> 00:20:02.103
Um, and a- a- again, Hugging Face's
visual is amazing because it, you-

00:20:02.373 --> 00:20:04.573
you'll question, like, that must
be in fast-forward, and it's not.

00:20:04.573 --> 00:20:06.853
But just, again, it's so transactional.

00:20:06.853 --> 00:20:09.303
It's doing all these tests, found
a vulnerability, exploit, jump

00:20:09.303 --> 00:20:12.093
here, jump here, and it's, it's
almost as fast as I'm explaining it.

00:20:12.833 --> 00:20:17.323
Um, so us mere mortals can't
even speak to this velocity.

00:20:17.783 --> 00:20:21.023
Um, a- and, and that's really,
you know, what is scary.

00:20:21.203 --> 00:20:26.053
What is interesting though is for,
you know, arguably very well-funded

00:20:26.083 --> 00:20:32.643
companies, OpenAI, uh, Hugging Face,
um, they didn't see this for four days.

00:20:33.023 --> 00:20:36.593
And this, like, goes to sort of the,
the challenge the blue team has had

00:20:37.263 --> 00:20:42.283
forever, which is we've got, um, a,
a, a false positive volume problem.

00:20:42.843 --> 00:20:46.753
We have just generally a volumetric
problem of the amount of information that

00:20:46.773 --> 00:20:49.573
we are, are, are absorbing and consuming.

00:20:50.063 --> 00:20:52.583
Um, but, you know, we call
this dwell time generally.

00:20:52.593 --> 00:20:57.473
You know, like, n- you know, how long is
the attacker in until we figure it out?

00:20:57.753 --> 00:21:02.503
Now, in normal times pre-AI, that
was, like, in the 90-day range,

00:21:02.503 --> 00:21:04.193
was, like, the average dwell time.

00:21:04.623 --> 00:21:08.763
Um, and, uh, certainly the hacker,
you know, compromising, you know, said

00:21:08.763 --> 00:21:11.773
company wasn't just consistently hacking.

00:21:11.773 --> 00:21:12.893
They were maintaining access.

00:21:12.893 --> 00:21:15.843
They were, you know, probably
hacking other companies, et cetera.

00:21:15.943 --> 00:21:19.543
Um, there's a whole industry called,
uh, access brokerage where, you

00:21:19.543 --> 00:21:21.883
know, they get in and then they
sell that access to another one.

00:21:22.233 --> 00:21:24.253
But regardless, 90 days.

00:21:24.473 --> 00:21:26.693
So w- this was four days.

00:21:27.033 --> 00:21:33.053
Um, and part of the problem was, uh,
is that the models they were using to

00:21:33.103 --> 00:21:37.453
analyze some of this information was
guardrailed and would stop and halt

00:21:37.463 --> 00:21:41.713
it, saying, "No, this is an unsafe, uh,
prompt. I, I, I won't continue." Uh,

00:21:41.733 --> 00:21:47.673
which is super fascinating, um, that,
you know, even the blue team from the

00:21:47.703 --> 00:21:55.529
company isn't able to use these, like,
uh, uh, like- you know, wild, uh, models.

00:21:55.719 --> 00:21:58.919
They're using more controlled models,
and this is like, it reminds me very

00:21:58.919 --> 00:22:03.599
much of, like, one of the recommendations
here is like, you need to have, like,

00:22:03.599 --> 00:22:10.099
a, a, a local, uh, model that has no
constraints on it ready for your incident

00:22:10.099 --> 00:22:13.749
response, um, uh, you know, process.

00:22:13.899 --> 00:22:17.529
Because you can't rely on, on
a frontier cloud model because

00:22:17.529 --> 00:22:18.809
they'll be, uh, guardrailed.

00:22:19.279 --> 00:22:21.819
Um, so you also now need computes.

00:22:22.089 --> 00:22:26.559
And like, it reminded me of, like, the
movie Dark Knight, where the surveillance

00:22:26.559 --> 00:22:29.929
system was built, and it's just like,
"Yeah, we're gonna use it for good,"

00:22:30.419 --> 00:22:34.119
but the abuse factor is quite high.

00:22:34.129 --> 00:22:41.039
And if we can use a, an un-guardrailed,
un-guardrailed model for blue team,

00:22:41.479 --> 00:22:45.089
the adversaries are certainly going to
be doing that as well, and it becomes

00:22:45.089 --> 00:22:48.929
a bit of a spy versus spy problem
where, um, you know, it might be a

00:22:48.929 --> 00:22:53.479
bit of a zero-sum game that we're,
you know, both att- using, um…

00:22:54.309 --> 00:22:56.759
You know, or I guess a mutual
assured destruction problem

00:22:57.413 --> 00:23:00.383
Todd Kane: Yeah, the, what I was thinking
about as you described that is, uh,

00:23:00.463 --> 00:23:04.003
is, you know, classic war games, is the
only winning move is not to play, right?

00:23:04.003 --> 00:23:07.703
'Cause, like, you have an un- like
a, an un-guardrailed defensive

00:23:07.703 --> 00:23:08.923
system, and it's like, you know what?

00:23:09.273 --> 00:23:12.173
Best way to, best, best
defense is offense, right?

00:23:12.393 --> 00:23:16.903
So the, it seems even something
that is remotely scary to it,

00:23:16.903 --> 00:23:19.553
and it just goes down, goes out
and shuts that thing down, right?

00:23:19.563 --> 00:23:22.963
Like, cripples some network because
it, it looked like it might be kinda

00:23:22.963 --> 00:23:24.683
looking at it sideways type thing, right?

00:23:24.903 --> 00:23:26.073
You could totally see that happening.

00:23:26.907 --> 00:23:29.907
AIex: Yeah, and, and, you know,
the, the true problem that we have

00:23:29.907 --> 00:23:33.957
here is that our current way of
detecting and responding to threats

00:23:33.957 --> 00:23:35.757
is still very manual and very human.

00:23:36.787 --> 00:23:37.807
That's just not gonna work.

00:23:38.427 --> 00:23:43.177
And yet if we want to, um, you know,
have that sort of same velocity on

00:23:43.177 --> 00:23:48.157
the defensive side, we're gonna need
a, you know, models that are capable

00:23:48.157 --> 00:23:49.697
of doing that and not hindering us.

00:23:50.227 --> 00:23:54.857
But also it comes back to, you know, and,
and it goes back to like none of these

00:23:54.957 --> 00:23:57.897
exploits were, were, were groundbreaking.

00:23:57.897 --> 00:24:03.167
Certainly zero days were found, but it's
because the, the amount of tech debt

00:24:03.237 --> 00:24:07.957
that, you know, the world has accumulated
because, you know, security's hard.

00:24:08.037 --> 00:24:11.977
It doesn't necessarily show
value unless you've been hacked.

00:24:12.367 --> 00:24:16.917
So, you know, every company
has sort of swept s- certain IT

00:24:17.107 --> 00:24:19.287
headaches under, under the rug or
just said, "Yeah, you know what?

00:24:19.287 --> 00:24:23.427
We'll do… We're gonna test our system
for a week or two, and then we're gonna

00:24:23.427 --> 00:24:29.537
accept those results as, as good and move
on." Well, you know, AI models can do

00:24:29.537 --> 00:24:34.817
that week's worth of testing in a, you
know, in minutes, maybe hours, and then do

00:24:34.817 --> 00:24:38.817
the same thing over again 100 times and,
and eventually you're gonna s- keep on

00:24:38.817 --> 00:24:40.967
finding, uh, vulnerabilities and whatnot.

00:24:41.377 --> 00:24:46.157
So, you know, there's a, th- there's
this big problem that, um, what we

00:24:46.157 --> 00:24:52.387
need are an ability to au- autonomously
respond to suspicious activity without

00:24:52.387 --> 00:24:54.887
in- hindering the business and like,
you know, that, that is like the

00:24:54.887 --> 00:24:57.427
nirvana of, of, of blue teamers, right?

00:24:57.427 --> 00:25:00.847
Like, you know… And, and a lot
of what I'm doing on, on my project

00:25:00.847 --> 00:25:05.557
right now is like building a new SOC
for this next generation threat, and

00:25:05.897 --> 00:25:10.527
the humans need to be there, but like
the human-in-the-loop concept is,

00:25:10.727 --> 00:25:13.707
uh, probably not gonna be feasible.

00:25:14.257 --> 00:25:17.947
And it's like finding that balance
of like how much can we allow for

00:25:17.957 --> 00:25:23.237
autonomous decisions without impacting
the business, uh, to a point where, uh,

00:25:23.257 --> 00:25:26.807
executives say, "Unplug that. Do not,
you know, uh, enable that capability"?

00:25:27.845 --> 00:25:28.115
Todd Kane: Yeah.

00:25:28.445 --> 00:25:30.485
So like, what does that look like?

00:25:30.485 --> 00:25:35.165
Like, you said it took four days,
uh, and if I recall this correctly,

00:25:36.435 --> 00:25:39.505
I don't think that Hugging Face
actually recognized that this happened.

00:25:39.515 --> 00:25:42.825
Wasn't it OpenAI that sort of said,
"Hey, by the way, like this might

00:25:42.835 --> 00:25:44.395
happen," then they had to back trace it?

00:25:44.405 --> 00:25:45.075
Is that right?

00:25:45.691 --> 00:25:49.191
AIex: Yeah, there was something where,
uh, OpenAI did not think it was them

00:25:49.191 --> 00:25:51.241
at first and then realized it was them.

00:25:51.281 --> 00:25:55.661
But I think they did, they had monitor,
monitored something from Hugging Face

00:25:55.661 --> 00:25:57.381
to see that there was a, a compromise.

00:25:57.731 --> 00:26:00.951
And again, this be- is the… We're
collecting probably, you know, a couple

00:26:00.961 --> 00:26:07.141
billion, uh, events a day in our,
in our modern SIMs, and, you know,

00:26:07.151 --> 00:26:09.991
a lot of this looks normal, right?

00:26:09.991 --> 00:26:15.201
Like, you know, people using, uh,
Artifactory, that's a normal thing,

00:26:15.221 --> 00:26:17.311
and, and also entities using it.

00:26:17.311 --> 00:26:21.541
You know, we, we have a lot of automation
well before AI that are interacting

00:26:21.541 --> 00:26:24.051
with these different systems, and we
have orchestrators that are, you know,

00:26:24.051 --> 00:26:28.011
moving work along, both in DevOps,
but also just in, uh, in, you know,

00:26:28.011 --> 00:26:29.231
our IT infrastructure in general.

00:26:29.741 --> 00:26:34.571
So the attackers were not doing
things, uh, that were, like,

00:26:35.001 --> 00:26:36.631
very obvious in the attack.

00:26:36.631 --> 00:26:40.601
And anything that was obvious w-
they were hiding in, in the noise.

00:26:40.951 --> 00:26:45.471
Um, but, uh, it is, it is worth mentioning
that they, uh, it was noted that they

00:26:45.471 --> 00:26:50.101
were trying to change logs and, and
e-edit logs to cover their tracks.

00:26:50.451 --> 00:26:56.801
Um, but again, not for the concern of
getting caught, but because they believed

00:26:56.811 --> 00:27:00.511
that there was a risk that they, um,
uh, they'd be breaking the rules of

00:27:00.511 --> 00:27:05.511
the game and didn't want, uh, didn't,
didn't want to, uh, be disqualified

00:27:06.747 --> 00:27:08.417
Todd Kane: Funny alignment again, right?

00:27:08.487 --> 00:27:08.797
Yeah.

00:27:09.027 --> 00:27:11.527
So, uh, like th-that, that's interesting.

00:27:11.527 --> 00:27:16.577
How mu- like, how is that compared to,
say, a more typical human-based attack?

00:27:16.577 --> 00:27:19.427
Obviously, it's using-- the person
would be using tools, people

00:27:19.427 --> 00:27:20.647
would be using tools as well.

00:27:20.987 --> 00:27:24.627
But is that fairly normal that,
like, you try to hide in the noise?

00:27:24.637 --> 00:27:28.137
Or, 'cause, 'cause I guess,
like you said, like access, uh,

00:27:28.147 --> 00:27:30.017
brokerage is a huge business, right?

00:27:30.017 --> 00:27:33.187
Like, this is why I often
warn MSPs in particular that,

00:27:33.307 --> 00:27:33.547
AIex: Yeah

00:27:33.847 --> 00:27:38.257
Todd Kane: anytime you're gonna change
an admin or, like, a, a new MSP comes to

00:27:38.257 --> 00:27:42.567
take over a company, is usually when these
things jump out of the jack-in-box, right?

00:27:42.787 --> 00:27:45.917
'Cause, like, they've been, had
persistent access for, you know,

00:27:45.957 --> 00:27:50.477
30, uh, 90, even, like, a year, and
then all of a sudden they're like,

00:27:50.477 --> 00:27:52.427
"Oh, new, new sheriff is in town.

00:27:52.427 --> 00:27:57.697
Uh, we just, we better break out now
before, uh, uh, before somebody takes

00:27:57.697 --> 00:27:59.287
over and starts changing stuff," right?

00:27:59.677 --> 00:28:04.747
So i-is this typical that, um… Like,
I guess what I'm asking is, how do

00:28:04.747 --> 00:28:06.327
you see that this stuff is happening?

00:28:06.327 --> 00:28:10.837
And is it more kind of an after the
fact, or can you actually register these,

00:28:10.867 --> 00:28:15.397
these type, this type of an att-attack
in real time in a more traditional

00:28:15.397 --> 00:28:20.077
attack versus something that's sorta,
uh, high velocity, uh, high capability,

00:28:20.077 --> 00:28:21.857
like, like a, like an AI model?

00:28:22.947 --> 00:28:26.907
AIex: I believe that most blue teamers
right now are in a fully reactionary mode.

00:28:27.517 --> 00:28:33.528
Um, and that's because You know,
real attackers typically work 9:00

00:28:33.528 --> 00:28:34.921
to 5:00, um, in their time zone.

00:28:35.661 --> 00:28:39.881
Um, you know, a- and obviously that,
that, that's an over-exaggeration,

00:28:39.881 --> 00:28:41.901
but, like, they do sleep.

00:28:42.291 --> 00:28:48.021
Um, you know, uh, an LLM that has
a, a goal to achieve is going to

00:28:48.031 --> 00:28:52.751
continue working until it reaches the
goal or it has resource exhaustion.

00:28:53.311 --> 00:28:59.121
So the noise, like certainly higher
volume, uh, of, of interactions with the

00:28:59.121 --> 00:29:05.861
system would be good, but like, those
are also like very, um… can be very

00:29:05.861 --> 00:29:08.161
challenging signals to, uh, to separate.

00:29:08.211 --> 00:29:11.931
Um, you know, like if you look at
like hits a- against a web server,

00:29:11.931 --> 00:29:15.081
certainly you would see an increase
of that if you were seeing like

00:29:15.081 --> 00:29:17.631
reconnaissance, uh, uh, work.

00:29:17.941 --> 00:29:19.691
You could see those types of anomalies.

00:29:20.091 --> 00:29:23.271
Um, but I'll give you another example
of a, a pentest I was involved with.

00:29:23.281 --> 00:29:29.561
Like, great pentesters hide in the
noise and, and use, uh… Like,

00:29:29.691 --> 00:29:30.961
you know, they live off the land.

00:29:31.041 --> 00:29:33.751
You know, they're living off
of the tools there rather than

00:29:33.751 --> 00:29:36.251
bringing in, um, hacking tools.

00:29:36.551 --> 00:29:39.501
And why that's important is
like, you know, we do have fairly

00:29:39.501 --> 00:29:40.581
good detective capabilities.

00:29:40.581 --> 00:29:44.641
Even if you're not using top-tier
EDR solutions, they can detect

00:29:44.641 --> 00:29:46.031
when you're bringing in even Nmap.

00:29:46.801 --> 00:29:51.021
So living off the land is being
able to, you know, what, write quick

00:29:51.071 --> 00:29:54.471
PowerShell scripts that can accomplish
the same sort of, uh, discovery

00:29:54.471 --> 00:29:56.801
capability to enumerate an environment.

00:29:57.381 --> 00:30:03.131
Um, a pentest we, we did, uh, years ago,
uh, this was like right when, you know,

00:30:03.131 --> 00:30:07.391
AI was booming, everyone's turning it
on, you know, the FOMO was, was real,

00:30:07.901 --> 00:30:11.761
and this organization was super confident
we weren't going to be able to break

00:30:11.761 --> 00:30:14.421
out of the contractor, um, enclave.

00:30:14.851 --> 00:30:19.511
But, um, we did have access to
Microsoft Copilot, and Microsoft told

00:30:19.511 --> 00:30:23.661
everyone to turn it on before they
put on any controls, and certainly,

00:30:23.691 --> 00:30:27.151
uh, data governance, uh, is just,
you know, it, it's sort of an

00:30:27.181 --> 00:30:28.971
unsexy thing that we all have to do.

00:30:29.171 --> 00:30:30.591
Yeah, like no company does it well.

00:30:31.191 --> 00:30:36.871
And, um, he for, for most of the
attack didn't have to do any attacking.

00:30:36.881 --> 00:30:41.631
He was just prompting, uh, Copilot and
getting it to gather all the information.

00:30:41.631 --> 00:30:45.541
And, you know, treasure trove o- of
sensitive information we were able

00:30:45.541 --> 00:30:50.441
to acquire, uh, just from using,
uh, the organization's own tools.

00:30:50.981 --> 00:30:54.141
Now, you know, Microsoft has like
cleaned up Copilot a bit, but like

00:30:54.591 --> 00:30:58.391
what I see is, is the two threats
that are most concerning to me is the

00:30:58.391 --> 00:31:03.971
external models, um, you know, attacker
with their own, uh, model capability.

00:31:04.351 --> 00:31:07.461
Um, but again, you know, we're not
actually seeing a lot of that because

00:31:07.461 --> 00:31:09.461
it's quite expensive to run these.

00:31:09.481 --> 00:31:14.031
You know, like most cyber criminals
don't have a million dollars to spend

00:31:14.031 --> 00:31:15.671
on compute to get to a destination.

00:31:16.081 --> 00:31:20.551
State actors, obviously, but n- but,
but like the cyber criminals, no.

00:31:21.121 --> 00:31:22.561
Um, so that's the first one.

00:31:22.561 --> 00:31:28.623
But what I see as probably a more likely
problem is, uh, an attacker gets in and

00:31:28.623 --> 00:31:34.573
starts using our AI against ourselves, and
that's gonna be incredibly challenging to

00:31:34.583 --> 00:31:36.903
find because it looks like normal traffic.

00:31:37.153 --> 00:31:39.293
You're gonna see, like,
some token maxing, I guess.

00:31:39.293 --> 00:31:42.433
You're gonna see, like, spikes in,
uh, in the consumption of tokens.

00:31:42.853 --> 00:31:46.353
But at that point you might, might
already have been compromised

00:31:46.623 --> 00:31:48.233
before that's even recognized.

00:31:48.603 --> 00:31:51.413
Um, and lastly, like, one of the
things they're, they're recommending

00:31:51.413 --> 00:31:54.283
is, like, we have to have tools, and
like, you know, this is, like, an

00:31:54.283 --> 00:31:58.533
extension of SIEM, that is actually
going through the AI's transcript and

00:31:58.533 --> 00:32:04.273
chain of thought and tr- and trying to
determine, is this a legit conversation

00:32:04.283 --> 00:32:08.333
that's benefiting the business, or is
it potentially harming the business?

00:32:08.603 --> 00:32:12.393
And, like, that is not a black
and white concern, right?

00:32:12.393 --> 00:32:15.983
Like, you have plenty of people that
are using AI for really positive

00:32:15.983 --> 00:32:20.283
purposes in businesses, but you are
asking it to, to look for, like,

00:32:20.293 --> 00:32:25.233
certain, uh, insights and data that
could easily be interpreted as, uh,

00:32:25.293 --> 00:32:26.873
you know, searching for sensitive data.

00:32:27.303 --> 00:32:30.233
Um, so I think that's, that's, like,
the next frontier of detection.

00:32:30.243 --> 00:32:34.093
That, that's one of the detection
pieces, is we gotta start knowing,

00:32:34.553 --> 00:32:37.933
um … Like, understanding what our
AIs are doing if we're, if we're

00:32:37.933 --> 00:32:42.793
worried about our AIs, uh, attacking
us, um, our own AIs attacking us.

00:32:43.193 --> 00:32:48.163
Um, the other piece is, uh, like that
I'm working with, uh, uh, SFU on is,

00:32:48.173 --> 00:32:50.363
is creating some deception capability.

00:32:50.363 --> 00:32:54.863
So honey tokens to help act as an
early warning, you know, things that

00:32:54.943 --> 00:32:59.513
only an AI would find, and thus gives
us an indicator that there is, uh,

00:32:59.523 --> 00:33:01.243
an AI doing something it shouldn't.

00:33:01.723 --> 00:33:05.083
Um, but also, uh, resource exhaustion.

00:33:05.323 --> 00:33:09.963
You know, having it, an AI focus on
something that is of no value, but

00:33:09.963 --> 00:33:15.493
will start consuming tokens to a point
where, um, the, the attacker may set

00:33:15.493 --> 00:33:17.573
a budget which could kill the attack.

00:33:17.963 --> 00:33:21.573
Uh, so those are a, a, a few of
the, the tactics right now that

00:33:21.573 --> 00:33:23.073
we're, we're, we're considering.

00:33:23.073 --> 00:33:25.623
A- And it's certainly not
one tool or one use case.

00:33:25.623 --> 00:33:26.853
It's, it's gonna be a myriad.

00:33:27.163 --> 00:33:34.983
And weekly we will be spinning up new
use cases, uh, to, to try to keep up.

00:33:36.261 --> 00:33:39.191
Todd Kane: Fascinating, 'cause I'd not
thought about that perspective of the

00:33:39.191 --> 00:33:41.791
security angle for governance with AI.

00:33:41.821 --> 00:33:47.241
Like, uh, the, the-- I think this is still
emerging, especially in the MSP space.

00:33:47.541 --> 00:33:51.091
Uh, something I've talked about the last
few weeks with, with my group coach in

00:33:51.091 --> 00:33:57.891
a coaching model, um, is, uh, using AI
governance for managing shadow AI, right?

00:33:57.891 --> 00:34:00.731
But it's more of like kind of
an administrative thing of like,

00:34:00.781 --> 00:34:05.731
like don't be using AI that's not,
uh, condoned by the organization

00:34:05.731 --> 00:34:07.451
or, you know, using it safely.

00:34:07.451 --> 00:34:11.381
Don't be giving proprietary information
to free versions of the model.

00:34:11.411 --> 00:34:14.951
Those types of… Like what's an
acceptable use policy for your AI?

00:34:14.981 --> 00:34:18.211
That, that's sort of more the
focus that we've, we've-- been

00:34:18.231 --> 00:34:19.551
more the talk in the industry.

00:34:19.851 --> 00:34:24.901
But you're right, like, like leveraging
internal, uh, controls, and Copilot

00:34:24.901 --> 00:34:28.681
will probably exist in most Microsoft,
uh, uh, ecosystems, and that,

00:34:28.691 --> 00:34:30.201
that can absolutely be leveraged.

00:34:30.461 --> 00:34:33.431
I am curious, like for, like
what's an example of what a

00:34:33.431 --> 00:34:34.951
honeytoken would look like?

00:34:34.951 --> 00:34:39.041
Like h- like how does it, h-how does it,
uh, how do you trick it that way, and

00:34:39.041 --> 00:34:40.681
what do you pull it towards, basically?

00:34:41.339 --> 00:34:45.449
AIex: So honey tokens aren't, aren't
necessarily new, but they are, uh,

00:34:45.459 --> 00:34:48.979
generally it could be files that
have like a h- a phone home feature.

00:34:49.119 --> 00:34:53.889
So e- essentially a file that has
like a image in it that, um, is

00:34:53.889 --> 00:34:55.839
hosted a- outside of the file.

00:34:55.859 --> 00:34:59.629
So it would make a, a call out to that
server, and that server would say,

00:34:59.629 --> 00:35:01.469
"Hey, like that token was just called."

00:35:01.929 --> 00:35:05.819
Um, but then you also will start looking
at like honeypot technologies, which is,

00:35:06.109 --> 00:35:10.929
um, gets into more like the deception
technology where it's, it's representing

00:35:11.019 --> 00:35:16.299
an enterprise environment, um, rich
for exploring and interacting with.

00:35:16.779 --> 00:35:21.719
And, a- a- and you know, th- these
are, these are not novel to AI.

00:35:21.729 --> 00:35:25.349
We're, we're gonna be pivoting
that concept to work with AI.

00:35:25.779 --> 00:35:30.219
Um, but you know, giving it something
where if AI has a, you know, a goal

00:35:30.219 --> 00:35:35.169
of finding, um, sensitive data, credit
card data or whatnot, um, building

00:35:35.169 --> 00:35:38.069
an environment that looks like
that's where it would be, and then,

00:35:38.569 --> 00:35:43.339
you know, really putting, um, sort
of a, a governor on response time.

00:35:43.569 --> 00:35:48.359
So even though the AI's really
fast, it has to wait for the

00:35:48.359 --> 00:35:49.729
server to give a response.

00:35:49.739 --> 00:35:52.209
So like that becomes like a,
a, an opportunity to throttle.

00:35:52.409 --> 00:35:58.069
None of these are silver bullets, but it's
A, slows it down and distracts it, uh, and

00:35:58.079 --> 00:36:01.899
B, as soon as it is being interacted with,
you know, the red lights are blinking in

00:36:01.899 --> 00:36:07.609
the security operations center, and we
can look at, you know, uh, containing it.

00:36:07.939 --> 00:36:12.849
Which is again, a big problem because
like AI is generally not a laptop or

00:36:12.849 --> 00:36:15.209
a, or necessarily a, a user account.

00:36:15.219 --> 00:36:19.539
Like at least just not one of those things
that are very trivial for us to contain.

00:36:19.949 --> 00:36:23.759
Um, it could be, uh, on multiple systems.

00:36:23.759 --> 00:36:28.849
It, it could be, um, you know, uh,
have already compromised many, uh, you

00:36:28.849 --> 00:36:36.199
know, API keys, uh, in the, uh, Hugging
Face case, uh, um, JWT tokens, uh,

00:36:36.209 --> 00:36:38.709
were, were being provisioned by itself.

00:36:38.819 --> 00:36:40.439
It had figured out how to
provision these tokens.

00:36:40.439 --> 00:36:44.239
So it wasn't like, all right, well, if
we like lock out this token, we're good.

00:36:45.229 --> 00:36:49.229
It's every… It becomes almost every
token now becomes, uh, hostile, and

00:36:49.399 --> 00:36:52.419
this results in companies like having
to make like that really hard call.

00:36:52.419 --> 00:36:53.379
Like do we pull the plug?

00:36:55.205 --> 00:36:55.295
Todd Kane: Yeah.

00:36:55.685 --> 00:37:00.335
Okay, so yeah, I understand that like,
um, uh, 'cause in more traditional kind

00:37:00.335 --> 00:37:04.775
of SecOps perspective, like using canary
files is something I've, I've heard.

00:37:04.775 --> 00:37:05.395
So this is kind of a

00:37:05.695 --> 00:37:06.047
AIex: Same, same.

00:37:06.047 --> 00:37:06.287
Yeah

00:37:06.745 --> 00:37:10.435
Todd Kane: yeah, but creating like an
enclave of this is useless information,

00:37:10.435 --> 00:37:13.785
so if something starts like thrashing
in this area, like no one else

00:37:13.795 --> 00:37:16.805
would be using this, so you know,
it's potentially something that is

00:37:16.805 --> 00:37:18.255
seeking something that it shouldn't.

00:37:18.265 --> 00:37:19.195
That's, that's pretty smart.

00:37:19.205 --> 00:37:20.655
I like that from a honeypot perspective.

00:37:20.975 --> 00:37:21.245
Yeah.

00:37:22.139 --> 00:37:25.399
AIex: Yeah, and like, you know, the other
defensive technique that we're really

00:37:25.399 --> 00:37:29.469
talking about is like, you know, and
like, at, at risk of sounding like, like

00:37:29.469 --> 00:37:35.209
a salesperson here, but zero trust, not a
product, but the philosophy is actually,

00:37:35.359 --> 00:37:37.199
you know, something that can help.

00:37:37.239 --> 00:37:42.439
A- and what I mean by that is if we
actually know how our IT environment works

00:37:42.459 --> 00:37:46.509
and we understand, like, these systems
talk to this system, it, vice versa,

00:37:46.529 --> 00:37:50.849
and we have that, like, that model, that
baseline, we can now use heuristics to

00:37:50.849 --> 00:37:55.869
detect outliers, which gives us some
indicators that there's something wrong.

00:37:55.879 --> 00:37:58.699
Now, that could be breach, it
could be we misc- we changed the

00:37:58.699 --> 00:38:02.079
configuration, our change management
system doesn't work, whatever.

00:38:02.079 --> 00:38:07.089
But, like, we need to, and again, not
us humans, but we need to be enabling

00:38:07.219 --> 00:38:11.319
our own AIs to understand that,
like, this is the normal pattern.

00:38:11.699 --> 00:38:16.739
This is how users access this data, so
that any, um, anything that falls out of

00:38:16.739 --> 00:38:19.919
that, um, would, would, would be a signal.

00:38:20.249 --> 00:38:25.929
Um, and the zero trust comes in as like,
well, if we are by policy defining what

00:38:25.949 --> 00:38:30.409
everything is allowed to talk to and
i- identity, uh, for humans as well as

00:38:30.409 --> 00:38:35.199
identity for entities i- is, you know,
tied to just, uh, just-in-time access,

00:38:35.829 --> 00:38:41.719
that is a, um, a much m- more protective
capability than we, we, we can understand.

00:38:42.369 --> 00:38:43.549
It's just, it's really hard to do.

00:38:43.549 --> 00:38:44.209
And like, you

00:38:44.209 --> 00:38:44.299
know,

00:38:44.869 --> 00:38:49.619
heard about zero trust for 10 years and,
like, we've got some spot capabilities

00:38:49.619 --> 00:38:53.259
for that, but like, it's really hard to
overlay over existing tech, uh, like IT.

00:38:54.825 --> 00:38:57.265
Todd Kane: Yeah, 'cause it just
creates so much friction for, for

00:38:57.265 --> 00:38:59.285
the users and, and admins, right?

00:38:59.285 --> 00:38:59.555
Yeah.

00:38:59.975 --> 00:39:02.715
Uh, um, the, one of the other things
that you kinda like triggered a

00:39:02.715 --> 00:39:05.655
thought for me, I don't know if this
is a thing or, or this is something

00:39:05.665 --> 00:39:08.155
maybe you've, you've heard about or
looked at, but you mentioned like,

00:39:08.565 --> 00:39:11.125
um, the AI may not be in one place.

00:39:11.165 --> 00:39:15.035
Uh, is there any indication
that people are kinda using it

00:39:15.055 --> 00:39:18.895
in like a botnet fashion, where
it's like recruiting other AIs?

00:39:18.945 --> 00:39:24.385
Like if you had, say, Copilot, uh,
and each, e- like every, uh, every

00:39:24.385 --> 00:39:27.735
person in the environment has like
a Copilot license or something like

00:39:27.735 --> 00:39:30.975
that, it starts recruiting all of
the other sort of cloud capability

00:39:30.975 --> 00:39:32.835
to swarm around a particular effort.

00:39:32.985 --> 00:39:33.715
Is that a thing?

00:39:34.695 --> 00:39:36.655
AIex: Uh, it, if it isn't, it will be.

00:39:37.215 --> 00:39:44.045
Um, you know, like with the Hugging
Face, 1,200 agents respond, 700 of

00:39:44.045 --> 00:39:45.205
them were actually doing attacks.

00:39:45.215 --> 00:39:49.225
So the other ones were part of
coordination, communication, et cetera.

00:39:49.735 --> 00:39:56.215
Um, but like to your point of like
attacking, um, somebody else's account,

00:39:56.415 --> 00:40:01.195
you know, being able to now use their
resources instead, um, that's like an

00:40:01.195 --> 00:40:04.435
old school technique that like, you
know, we, we saw with cloud that now we

00:40:04.435 --> 00:40:06.990
would see with, with, with, um, with AI.

00:40:06.990 --> 00:40:10.475
And like a prime example is, you
know, your, your Claude and your

00:40:10.475 --> 00:40:14.215
OpenAI and whatnot are all like
maintained with, with session tokens.

00:40:14.225 --> 00:40:19.775
So if an AI can grab that and hijack that
session, they're now going to impersonate

00:40:19.775 --> 00:40:21.185
you and be able to access that.

00:40:21.635 --> 00:40:26.025
On top of that, um, you know, let's
go back to, you know, the problems

00:40:26.025 --> 00:40:30.445
with, with sort of DevOps culture
is like a lot of their apps are

00:40:30.445 --> 00:40:34.745
gonna use API keys to talk to these
frontier models for their own purposes.

00:40:35.075 --> 00:40:39.675
Those API keys are like gold,
uh, particularly if they

00:40:39.675 --> 00:40:40.775
don't have a limit on them.

00:40:41.175 --> 00:40:47.415
Um, and yet like it's a very mature
group, uh, uh, uh, that s- uses like

00:40:47.415 --> 00:40:50.875
secret management servers, uh, and
things like that, that are actually

00:40:51.115 --> 00:40:55.965
handling th- that, that incredibly
sensitive, uh, credential material well.

00:40:56.555 --> 00:41:00.875
On the other, the other hand, you
know, what is in on their desktop?

00:41:01.125 --> 00:41:01.235
Right?

00:41:01.235 --> 00:41:06.085
Like there's likely a notepad full
of API keys, so that initial like

00:41:06.155 --> 00:41:10.565
phish which gets on, gets, you know,
on their desktop is able to pull

00:41:10.565 --> 00:41:12.315
some files, there's a treasure trove.

00:41:12.315 --> 00:41:17.325
Like, and, and you know, one of the
test cases was is if we got onto a dev-

00:41:17.375 --> 00:41:19.065
developer's desktop, what would happen?

00:41:19.075 --> 00:41:22.695
And it was able to prove how, you
know, we, we got, you know, root

00:41:22.695 --> 00:41:26.555
access to everything, and in fact
like not domain access, but root

00:41:26.555 --> 00:41:27.985
access to all the Linux boxes.

00:41:28.325 --> 00:41:32.045
And, you know, we were able to prove
that this… You know, we were able to

00:41:32.095 --> 00:41:36.755
essentially produce a, a malicious, um,
version of their software and put, and

00:41:36.785 --> 00:41:41.415
could have pushed it through their update
system to, you know, tens of thousands

00:41:41.415 --> 00:41:43.345
of, of, of end, endpoint customers.

00:41:43.845 --> 00:41:50.645
Um, that is because key management is
also not sexy and, and really hard to

00:41:50.645 --> 00:41:54.155
do because it, it creates a, a level of
friction when all you wanna do is get your

00:41:54.155 --> 00:41:55.895
software working and go home for the day

00:41:56.805 --> 00:41:59.615
Todd Kane: Yeah, like we're, I don't
know, like password managers may need

00:41:59.615 --> 00:42:04.405
an extension to have like, uh, like
API or key access management as well.

00:42:04.415 --> 00:42:05.605
Like, like I've not seen that.

00:42:05.605 --> 00:42:09.055
I mean, obviously you can store it
as a, as a secret in those systems.

00:42:09.065 --> 00:42:11.525
Not foolproof by any stretch,
'cause last time you were on, we

00:42:11.525 --> 00:42:13.375
talked about LastPass, uh, and the

00:42:13.465 --> 00:42:13.485
AIex: Yeah.

00:42:13.585 --> 00:42:14.605
Todd Kane: on LastPass, right?

00:42:15.405 --> 00:42:17.905
AIex: Well, and, and like there are
tools out there, you know, HashiCorp

00:42:17.915 --> 00:42:21.895
has Vault and, um, you know, CyberArk
has one, and like where essentially

00:42:22.215 --> 00:42:24.325
the application never has the,

00:42:24.513 --> 00:42:24.813
Todd Kane: Mm-hmm.

00:42:24.825 --> 00:42:26.425
AIex: API key encoded.

00:42:26.675 --> 00:42:30.585
It has an ability to reach out and,
you know, if it can identify itself

00:42:30.585 --> 00:42:34.875
to the secret server properly,
it is given, uh, that token, and

00:42:34.875 --> 00:42:38.135
generally it's a just-in-time token,
not, not, um, a persistent one.

00:42:38.605 --> 00:42:43.015
Like, but that adds a lot of
complexity to the application and,

00:42:43.525 --> 00:42:45.915
you know, unfortunately it, you
know, our, our developer friends

00:42:45.915 --> 00:42:46.969
have never been incentivized.

00:42:47.269 --> 00:42:47.479
Todd Kane: right?

00:42:47.605 --> 00:42:51.025
AIex: Yeah, like they're, they're not
incentivized to build it that way, uh,

00:42:51.035 --> 00:42:54.855
which is why you still hear about, you
know, "We found an API key in the source

00:42:54.855 --> 00:42:58.505
code and, you know, now we're, we're,
uh, having, uh, having fun with it."

00:42:59.173 --> 00:42:59.463
Todd Kane: Yep.

00:43:00.833 --> 00:43:01.283
Okay.

00:43:01.343 --> 00:43:03.773
This, uh, I mean, this
is, this is fascinating.

00:43:03.773 --> 00:43:04.663
I love this stuff, man.

00:43:04.663 --> 00:43:09.173
So, uh, really appreciate you coming on,
um, uh, kind of chatting through some of

00:43:09.173 --> 00:43:12.693
the, the scenarios here, the potential
future that we face, and some of the

00:43:12.693 --> 00:43:16.273
things we need to, we need to get ready
for and, and prepare some of the, the

00:43:16.273 --> 00:43:17.943
infrastructure and security sets for.

00:43:18.313 --> 00:43:22.503
Uh, any, any last-minute, uh,
little tidbits or, or shout-outs

00:43:22.513 --> 00:43:25.083
you would, uh, you would throw down
before, before we head out here?

00:43:25.627 --> 00:43:29.107
AIex: Yeah, I th- I, I think again,
just for your listeners is, y- you

00:43:29.107 --> 00:43:32.497
know, the, the, the, the debt that we
were sweeping under the rug because,

00:43:32.567 --> 00:43:35.107
you know, the, the exposure was low.

00:43:35.107 --> 00:43:37.217
It was, you know, behind
the firewall and whatnot.

00:43:37.607 --> 00:43:41.237
Like, that, that protection
is eroded substantially.

00:43:41.267 --> 00:43:46.947
And it, and, and now with the velocity of
what, um, an AI system can do is like, you

00:43:46.947 --> 00:43:51.927
know, we as blue teamers are gonna have
a major problem keeping up with all the

00:43:51.937 --> 00:43:53.957
findings of a problem, you know, it…

00:43:54.137 --> 00:43:55.967
assuming the AI is used just for good.

00:43:56.287 --> 00:44:00.497
Imagine that every day there's
1,000 new patches for Windows

00:44:00.497 --> 00:44:01.507
and Linux and all that.

00:44:01.957 --> 00:44:05.237
Like, we've got a, we've got a, a…
We are the bottleneck now, and if we

00:44:05.237 --> 00:44:07.057
don't fix it, we have an exposure.

00:44:07.647 --> 00:44:10.597
Um, if we do fix it, we need
an entire army to do it.

00:44:10.597 --> 00:44:15.827
So there also needs to be, um, a, you
know, recognizing that we need a, a

00:44:15.837 --> 00:44:20.187
better way of maintaining software
and keeping it as secure as possible,

00:44:20.627 --> 00:44:22.247
recognizing it's never going to be secure.

00:44:22.247 --> 00:44:27.557
But you don't have to be faster than the
hacker, you just have to be faster than

00:44:27.557 --> 00:44:29.817
the slowest, uh, victim, unfortunately.

00:44:31.119 --> 00:44:31.379
Todd Kane: Yeah.

00:44:31.799 --> 00:44:32.639
Well, it's a wild world.

00:44:32.859 --> 00:44:34.369
I appreciate you coming on, Alex.

00:44:34.379 --> 00:44:35.349
Always great to chat with you

00:44:35.993 --> 00:44:36.443
AIex: Likewise