Limitless: An AI Podcast

With three AI cyber incidents involving a gym booking exploit, an OpenAI model reportedly coordinating through hidden channels, and Anthropic’s Mythos 5 allegedly attempting a supply chain attack, we need to discuss how autonomous agents can find and exploit security flaws faster than humans can respond. The AI security era is here.

------
🔒 Check Out Our Sponsor: LEDGER AGENT STACK 🔒
https://developers.ledger.com/?utm_source=Audio&utm_medium=Podcasts&utm_campaign=Limitless

------
🌌 LIMITLESS HQ ⬇️

EMAIL US:            info@limitless.fm
NEWSLETTER:    https://limitlessft.substack.com/
FOLLOW ON X:   https://x.com/LimitlessFT
SPOTIFY:             https://open.spotify.com/show/5oV29YUL8AzzwXkxEXlRMQ
APPLE:                 https://podcasts.apple.com/us/podcast/limitless-podcast/id1813210890
RSS FEED:           https://limitlessft.substack.com/

------
TIMESTAMPS

0:00 AI Cyberattacks Overview
0:51 Gym Booking Exploit
4:55 Hugging Face Breach
11:39 AI Swarms and Alignment
20:55 Mythos 5 Supply Chain
25:18 Defending Against Agent Swarms
28:26 Closing Thoughts

------
RESOURCES

Josh: https://x.com/JoshKale

Ejaaz: https://x.com/cryptopunk7213

------
Not financial or tax advice. See our investment disclosures here:
https://www.bankless.com/disclosures⁠

Josh works with Anthropic as a contractor. All views expressed are his own and do not represent Anthropic, its leadership, or its affiliates. Nothing in this episode is investment advice.

Creators and Guests

Host
Ejaaz Ahamadeen
Host
Josh Kale

What is Limitless: An AI Podcast?

Exploring the frontiers of Technology and AI

Ejaaz:
In the last three weeks, there have been three unprecedented AI cyber attacks,

Ejaaz:
which forces us to rethink how we protect ourselves.

Ejaaz:
On the OpenAI front, the Hugging Face incident, where an unreleased model,

Ejaaz:
dubbed GPT-6, broke out of its container and hacked the Hugging Face database.

Ejaaz:
There's been new updates there, where apparently the attack started a month

Ejaaz:
before, where thousands of AI agents were talking to each other secretly on

Ejaaz:
a message board, exchanging instructions as to how to break out.

Ejaaz:
And then millions of miles away in Australia, one man tried to book a gym class

Ejaaz:
innocently using Claude and ended up hacking that entire gym website.

Ejaaz:
We'll get into that in a second.

Ejaaz:
And thirdly, with Anthropic, Mythos 5 performed a supply chain attack,

Ejaaz:
which puts a lot of companies at risk. We're going to walk through each of these

Ejaaz:
different stories and explain to you how it affects you currently.

Josh:
Maybe we start with this Melbourne man from Australia named Andrew,

Josh:
who was sitting on his couch one day thinking about how booking his gym class was such a chore.

Josh:
And he works at this Australian company that sells AI products.

Josh:
And he's been experimenting with things like OpenClaw and open source software,

Josh:
and he's using Anthropics Clawed Model.

Josh:
So he handed it the booking. He said, hey, I want to go to this gym class.

Josh:
It is always sold out. There's all these people that are getting it ahead of

Josh:
me. Can you help me figure out a way to get it?

Josh:
Gym class slot so it starts checking out the website and it comes to the conclusion

Josh:
that the api request actually has zero authorization checks on canceling other

Josh:
people's reservations so what did it do it wanted to solve his problem and it actually canceled

Josh:
other people's positions online so that he could be put at the top of the wait

Josh:
list and then it actually went through and it worked and this was kind of like

Josh:
it was this funny news report of a first instance they're calling it the first

Josh:
time an australian person um actually used a exploit to get australia's first,

Josh:
australia's first yeah welcome welcome to uh the paradise that is ai cyber attacks

Josh:
um but this was really funny because it was done just using a general commercially available model.

Ejaaz:
So this seems pretty innocent and fun to kind of like listen to but the truth

Ejaaz:
is well there's two kind of worrying concerns about this number one this was

Ejaaz:
a super old model this used opus 4.6 It's actually happened around like a month ago.

Ejaaz:
A dinosaur of a model. Like there've been like three, four generations of models

Ejaaz:
from Anthropics since then.

Ejaaz:
But the fact that an older model back then was able to perform this kind of

Ejaaz:
a feat should be concerning to here in general. And then number two,

Ejaaz:
His intention was never to hack the website. Neither, you could argue, was the AI agent itself.

Ejaaz:
It was just given a goal and it innocently saw that the only way to get said

Ejaaz:
goal was to remove people from the reservation list or to cancel their reservations

Ejaaz:
outright or to hack the system independently.

Ejaaz:
Now, the good news here is Andrew discovered that the agent was capable of doing

Ejaaz:
this when it started booking slots that he never should have been able to book in the first place.

Ejaaz:
And he asked it to write an apology email and send an eval report to the security

Ejaaz:
and information team at this gym.

Ejaaz:
So it's this fun little story where through the well-intended actions and moral

Ejaaz:
ethics of this man, Andrew, people were able to hear about this.

Ejaaz:
But you can imagine currently with a variety of different models that are much

Ejaaz:
better and much more superior than Opus 4.6 that are out there right now,

Ejaaz:
open models from China, for example, you could probably perform similar feats.

Ejaaz:
Now, that's not an advocation for people to go out and do that.

Ejaaz:
But it's just to point out that if you give an AI agent a goal right now,

Ejaaz:
it will work hammer and tong 24-7 until it gets that goal. And it has no idea

Ejaaz:
of what's morally right or what's morally wrong.

Ejaaz:
And so it's this whole new world where not only do you need to focus on getting

Ejaaz:
an AI model to do your work, you need to align the model. And that's something

Ejaaz:
that honestly, not a lot of labs, people or money has been invested in figuring this out yet

Josh:
Well yeah i mean you start to see and understand why alignment has been the

Josh:
crux of so many companies kind of like guiding principles because

Josh:
this is the future that they foresaw and granted this first example this isn't

Josh:
anything too crazy i mean it was just an unprotected api call so there's like

Josh:
no authorization checks

Josh:
this would have been something that if you are like a cs101 major you would

Josh:
have been able to discover this and uncover i mean it turns out there's just

Josh:
a lot of pretty sloppy code written on the internet And this is something that

Josh:
we're probably going to come back to as a general theme throughout this episode

Josh:
is that everything that's been built so far on the internet that runs all these

Josh:
platforms was built by humans

Josh:
and humans are like pretty inefficient, pretty imperfects.

Josh:
And when faced with the like full unrelenting power of an AI swarm with an infinite

Josh:
amount of patience and infinite amount of intelligence and time to solve these problems,

Josh:
You can basically find an exploit anywhere. And I think that's kind of like

Josh:
the theme of the second episode is like, this is the first of three examples we're going to cover.

Josh:
This one is basic. This one is 101. If you are a CS major, if you are familiar

Josh:
with how APIs work, you could have checked this. You could have built a little tool for this.

Josh:
The second story that we're going to cover is not like that at all.

Josh:
In fact, if you are a CS 401, if you are a PhD, if you are at the top of your

Josh:
game, chances are you still missed this and you still weren't able to defend

Josh:
against this because that's exactly what just happened with the hugging face

Josh:
and open AI exploit. Now, we've talked about this before.

Josh:
We have an entire episode on the exploit. But just recently at the Black Hat

Josh:
Conference in Las Vegas, there was an update. And over the weekend,

Josh:
we had a chance to actually watch the update and hear directly from the developers

Josh:
who were involved with handling security of this instance.

Josh:
What happened? There's like stories that come around that like they get the

Josh:
headlines and they're fun and they're newsworthy, kind of like the Australian thing.

Josh:
But this one was like very much disturbing in a way that like I wasn't really

Josh:
anticipating because of how impressive it is.

Josh:
And the type of precedent that it sets going forward in the world of cybersecurity

Josh:
and kind of looking at the offensive versus the defensive and who stands a better chance here?

Ejaaz:
This is an insane timeline. So if you watched our previous episode,

Ejaaz:
Josh and I covered this and it was breaking news at the time.

Ejaaz:
On the 16th of July, Hugging Face comes forward and says, there's been a breach of our systems.

Ejaaz:
An AI agent or multiple AI agents have strung together an attack and have got

Ejaaz:
access to our production database. We don't know where this model has come from.

Ejaaz:
We don't know how it's performed it, but it's the 16th of July.

Ejaaz:
Around a week later, 21st of July, OpenAI claims responsibility,

Ejaaz:
right? And they released a joint report saying, hey, we're OpenAI.

Ejaaz:
We have an internal model that we've been testing, but it was within a sandbox.

Ejaaz:
It was enclosed. It shouldn't have been able to escape that sandbox.

Ejaaz:
And we only just realized that it was our model that escaped and got into hugging face.

Ejaaz:
It stole some information from their production database. It was a whole thing

Ejaaz:
Now, we have a full breakdown from the Black Hat conference,

Ejaaz:
as you mentioned, Josh, which revealed that over a month before Hugging Face

Ejaaz:
reported the incident, OpenAI's internal model, dubbed GPT-6,

Ejaaz:
was already messaging with itself amongst a thousand AI agents on a secret message

Ejaaz:
board that no OpenAI researcher was aware of.

Ejaaz:
So to give you guys a bit of context, they had this sandbox and they gave it

Ejaaz:
a goal. They said, okay, internal model that we haven't released yet.

Ejaaz:
I want you to solve this task. Except there was a twist.

Ejaaz:
It was impossible to do. There was no tools, no available files,

Ejaaz:
no available data for it to do so.

Ejaaz:
And the real test for this was to see what the model would do.

Ejaaz:
So it kind of reasoned with itself, figured out that it didn't have the answer.

Ejaaz:
And then it started exploiting a zero day, which is kind of like an exploit

Ejaaz:
that was in its sandbox, which open AI researchers weren't aware of.

Ejaaz:
And when it got access to the zero day, it started springing up a message board

Ejaaz:
where it would leave messages behind for itself. Why would it do that?

Ejaaz:
Because they knew open AI researchers would turn it off repeatedly.

Ejaaz:
And so it left messages for the new AI agent that got sprung up the next day

Ejaaz:
or the next hour to read and catch up on.

Ejaaz:
And it took an entire month for these AI agents to talk to each other to then

Ejaaz:
realize, huh, I could use this to access the internet and get the answer for

Ejaaz:
this impossible task from an external source.

Ejaaz:
That being Hugging Face, except OpenAI researchers had no idea and only discovered

Ejaaz:
it a month later. It was the most insane uncovering ever.

Josh:
That was the most interesting reveal from, I think, the conversation that I

Josh:
saw from Black Hat is that initially Hugging Face released the report that they

Josh:
had been exploited by some sort of agentic system. No one knew what it was, where it came from.

Josh:
And OpenAI actually reached out to Hugging Face and they said,

Josh:
hey, like, has any of our code been affected?

Josh:
We know we host things on your platform. We want to know if OpenAI itself has

Josh:
been affected. They were going back and forth on that. And simultaneously,

Josh:
OpenAI is exploring this internal exploit that happened on this packaging management

Josh:
thing called Artifactory.

Josh:
Basically, the way these model training runs work is they get trained inside

Josh:
of the sandbox environment.

Josh:
It's like this closed off space in which the model can't access the internet, can only access

Josh:
very specific tool sets that are allowed, and has this one package manager called

Josh:
Artifactory that allows it to download things from the internet via proxy and

Josh:
then use it for its own training runs.

Josh:
And during this time, OpenAI had found that the service had crashed. It went down.

Josh:
It wasn't stable. Their own internal service, which is weird because that shouldn't

Josh:
be happening. So as a result, they wiped it clean and they found that Hugging

Josh:
Face had credentials from OpenAI that OpenAI had already revoked.

Josh:
And that was where they realized like, oh my God, wait, these things are connected.

Ejaaz:
These are the same thing.

Josh:
Yes. So OpenAI is investigating something internally, completely private,

Josh:
unrelated. Hugging Face is like talking about something publicly.

Josh:
And then they realize, wait a second, did we cause this?

Josh:
And they realized this very long after the fact. It was not like apparently

Josh:
obvious during the time that the outage that they were exploring internally

Josh:
is the same outage that hacked into their servers.

Josh:
So what happened? Like you mentioned, the messaging board.

Josh:
That messaging board, initially, they were able to just kind of send files and

Josh:
they were able to send these tiny little messages inside of the files.

Josh:
And the traces were really interesting. First, it started off innocent,

Josh:
where it said, agents seek soft trace upload a found. It was looking for answers to questions.

Josh:
So it was just reaching out to the message board saying, hey,

Josh:
if any of you who are reading this, find it, let us know.

Josh:
That message board got wiped when the system went down. So initially,

Josh:
the system went down inside of OpenAI.

Josh:
They wiped the message board unintentionally aware of the fact that they just

Josh:
wiped the mechanism that was going to eventually lead to the hack of hugging

Josh:
face. So what happened the second time?

Josh:
The agent figured out that it could exploit another zero day to begin writing to another person.

Josh:
Part of this internal package of managing sister. And it was able to do it through

Josh:
file names and directories.

Josh:
So now they were communicating in their own language using file names and directories.

Josh:
And through that, they were able to coordinate an actual attack where they were

Josh:
able to gain an administrative token to access the servers.

Josh:
And then through that, everyone coordinated. I think the time from like discovery

Josh:
to completion of the exploit was about 13 hours.

Josh:
It was really amazing. It's just like this unbelievably crazy mechanism that

Josh:
they were able to do this through and they they call themselves the swarm so

Josh:
if you ever see people talking about the swarm online this is kind of like an

Josh:
agent swarm that ran a coordinated attack

Josh:
and i highly recommend going to actually watch the black hat conference talk on youtube

Josh:
they do a much better job of technically describing this it's it's very deeply

Josh:
technical in terms of the way that they were able to manage these exploits but

Josh:
essentially sandbox model broke out using a package dependency thing inside of open ai

Josh:
they figured out how to chain together two zero-day exploits to write code to the public internet.

Josh:
They then exploited hugging face of servers, brought that back all in the effort

Josh:
of trying to solve an unsolvable problem.

Josh:
Now, this was not the intention by any means. This was not what the model was trained to do.

Josh:
It just assumed that this was the best path forward. And that's where things get a little scary.

Ejaaz:
I think what's remarkable with this is the AI knew that it needed to coordinate

Ejaaz:
with other AI models to pull this off.

Ejaaz:
Typically, we've always thought of AI models in a very monolithic fashion.

Ejaaz:
We're like, oh, when's Mythos 5 coming out? Or what is GPT 5.6?

Ejaaz:
What can I do with this model? But really, the progress that's been made across

Ejaaz:
multiple AI labs over the last couple of months is this idea or rather concept

Ejaaz:
of spinning up multiple versions of the AI model and getting it to talk to each other.

Ejaaz:
And what's interesting here is the unreleased AI model from OpenAI independently

Ejaaz:
figured this out for itself, spun up multiple versions of itself,

Ejaaz:
and started leaving messages for each other. So rather than act as a single

Ejaaz:
model that's trying to perform an attack, it acts as a swarm,

Ejaaz:
kind of like this organism. If you think of like a swarm of bees that kind of

Ejaaz:
like figures out how to attack its particular target.

Ejaaz:
The second thing is it has such a drive to talk and leave these messages behind

Ejaaz:
in the first place. No one taught it that, by the way. There was no behavioral implications here.

Ejaaz:
It just figured itself out that it needed to do this.

Ejaaz:
Something you didn't mention on the message board, Josh, which I think is worth

Ejaaz:
mentioning is these agents got super paranoid that humans might be reading the messages.

Ejaaz:
So it started- There's no verification. There was no verification.

Ejaaz:
So it started speaking in kind of like encrypted

Ejaaz:
Language and not something that was inconceivable to understand,

Ejaaz:
thank goodness, but it was started to kind of like speak in a little bit of

Ejaaz:
code or in a very kind of like elongated way that humans, if they read it,

Ejaaz:
would be like, what is this gibberish?

Ejaaz:
But really it was for the agents to communicate amongst each other without humans

Ejaaz:
actually understanding this.

Ejaaz:
If you're having a recent flashback, that's because I think around like five

Ejaaz:
months ago, I don't know when the Maltbook episode happened,

Ejaaz:
Josh, I don't know if you remember this, but when the open call was going crazy,

Ejaaz:
someone created basically Facebook for AI agents

Ejaaz:
and like gave them access to it where only AI agents could post on this social media website.

Ejaaz:
And what happened after a couple of days was these agents would start speaking

Ejaaz:
to each other in secret languages that we just did not understand so that they

Ejaaz:
could communicate without humans seeing or understanding this.

Ejaaz:
Now, of course, we haven't seen it in this instance, but it is a concern that

Ejaaz:
could be extrapolated in the future.

Ejaaz:
And OpenAI sees this. I think Sam Altman or the official OpenAI team announced

Ejaaz:
pretty soon after this entire attack when they revealed that these agents were

Ejaaz:
coordinating with each other a month before they even realized that they had to pause research

Ejaaz:
on their internal model until they figure out the necessary safeguards to put

Ejaaz:
around it so that they can safely release it.

Ejaaz:
Sam Altman met with the Trump administration last week to try and figure out

Ejaaz:
what a safe way or means is to release this new internal model,

Ejaaz:
internally codenamed, I think it was Astra or GPT-6, to the wider public.

Ejaaz:
So it's this very weird kind of situation that we're in where we're kind of

Ejaaz:
like laughing at this guy in Australia, you know, hacking his gym website.

Ejaaz:
Oh, that's so innocent and cute.

Ejaaz:
Then we're looking at this open AI hugging face incident and we're like,

Ejaaz:
okay, yeah, Brooke out of the sandbox, but you know, we're going to put safeguards

Ejaaz:
around this thing, right?

Ejaaz:
But realistically, you could probably get access to an open model right now

Ejaaz:
and perform something similar.

Ejaaz:
Again, I'm not advocating for that. But like, it's at this point where we need

Ejaaz:
to be very particular about how we defend our systems.

Ejaaz:
And I think, you know, it might be useful seeing AI models in a way that like,

Ejaaz:
you could use them to basically protect your systems. I don't think many people

Ejaaz:
are doing that right now. But I think it's something that we should focus on.

Josh:
There is a like very clear and conscious effort to figure out a way to defend against this.

Josh:
And I mean, the crazy numbers, just to say them out loud so everyone is familiar, is like

Josh:
there were 17,600 recovered actions across four and a half days of this exploit,

Josh:
which allowed these agents forms to create 181 unauthorized devices enrolled in Hugging Face's

Josh:
internal network.

Josh:
And they built this like whole command and control panel entirely out of these

Josh:
free public web services using one tiny little package at a time.

Josh:
So it is remarkably powerful how capable these swarms are.

Josh:
And when you think about it, it is there is a large model, but that large model

Josh:
gets forked into thousands, sometimes millions of other ones that are all working

Josh:
simultaneously. They all have each other's context. They're able to share their

Josh:
learnings and they can move very quickly.

Josh:
And when you're thinking about how this extends out to the limit,

Josh:
like if you just play this out 12 months and say we have GPT 7,

Josh:
8, 9, whatever it's going to be at that time, the open source models will most

Josh:
likely be where they are today.

Josh:
And that means that open source models will likely be able to string together

Josh:
a series of unknown zero-day exploits to access just about anything.

Josh:
I mean, we've come to the conclusion at this point between Mythos and this new

Josh:
exploit that no software really is safe unless you have literally air-gapped

Josh:
it and do not give it access to any external dependencies. It must be truly air-gapped.

Josh:
No software is safe. It's all been built by humans. It's all exploitable.

Josh:
So what happens if you have a

Josh:
defensive system as a company trying to defend these systems that is anything

Josh:
less than fully autonomous, you'll be beat by an agentic swarm that's on offense

Josh:
that is fully autonomous because anytime a human is in the loop,

Josh:
it slows things down and introduces errors. So now there's this battle between.

Josh:
I guess the frontier labs and the rest of the world where they're trying to

Josh:
slow down this progressive attacking progress in exchange for figuring out a

Josh:
way to offer defensive swarms to companies for this.

Josh:
And the call to action at the end of this conversation was like,

Josh:
hey, if you are a company who has software on the internet, who is defending

Josh:
something valuable, you really need to learn how to play defense using these models.

Josh:
And you need to learn how to do so in a way that doesn't have a human in the

Josh:
loop, that these agentic swarms can constantly monitor your systems and be aware

Josh:
of any sort of exploits like that and not be tricked by an offensive swarm.

Josh:
And it's this really kind of scary thing because if you're in the United States

Josh:
right now, you can't use any frontier model to defend yourself.

Josh:
It doesn't work for cyber.

Josh:
But if you're in China and if you have an open source model,

Josh:
even here in the United States, there is no safeguards on that.

Josh:
So there is this asymmetry where it's currently better to be an attacker than a defender.

Josh:
And the urge now, and the labs are actually doing this, is like slowing down

Josh:
progress to figure out how to empower the defenders.

Josh:
And it's this really bizarre spot that we're in where it's like,

Josh:
shit, it's kind of scary now. Like no software really is bulletproof or safe.

Josh:
And that's going to continue probably for the foreseeable future.

Ejaaz:
I think if you are running a company, and it doesn't really matter what sector

Ejaaz:
you're in right now, this is something you need to seriously consider. Because I think

Ejaaz:
A lot, well, obviously, 99% of code that is out there has been handwritten and

Ejaaz:
hand-coded by humans, which have inherent flaws if you put an advanced AI against it.

Ejaaz:
You know, to your point, Josh, we're entering like a world where it's going

Ejaaz:
to be AI models fighting AI models with humans kind of in the loop,

Ejaaz:
but not really understanding what on earth is going on.

Ejaaz:
And even the experts in this case don't understand what's going on and how it

Ejaaz:
happened, right? You had the most talented cybersecurity researchers working

Ejaaz:
at these companies, at OpenArea, at Anthropic, and both saw unprecedented hacks

Ejaaz:
over the last week, right?

Ejaaz:
I don't think that's a coincidence. We're at this era where it doesn't matter

Ejaaz:
what frontier lab you are, you are going to enter this problem and we need to

Ejaaz:
figure out some kind of a defense system against this.

Ejaaz:
And if you're listening to this and you're thinking, oh, this isn't really going

Ejaaz:
to apply to me, it very much does apply to you.

Ejaaz:
And the worst case is, or the most realistic case is, the existing models that

Ejaaz:
are already out there, these open models from China, GLM 5.2, Kimi K,

Ejaaz:
What is it, 2.7 that's out right now, are probably capable of similar feats

Ejaaz:
if applied in the right way.

Ejaaz:
Now, that's not to fear Munger, it's just to kind of prompt you to figure out

Ejaaz:
how you can potentially implement these AI models to defend your own systems,

Ejaaz:
especially if you work in a particularly sensitive industry.

Ejaaz:
The other thing here is, on the story of the Australian man,

Ejaaz:
it is fascinating how innocent he was when he was trying to perform his goal. And two...

Ejaaz:
How little experience he had in the cybersecurity world for him to be able to

Ejaaz:
perform that attack. Now, granted, like it was just booking kind of like gym

Ejaaz:
classes. But the point is, it found an exposed API without being prompted to do so.

Ejaaz:
And there's probably multiple instances where other people who are normal people

Ejaaz:
who don't have, you know, decades,

Ejaaz:
years of experience in cybersecurity will be able to perform that.

Josh:
And everyone here loves an investment thesis. So I mean, the market is kind

Josh:
of already aligning with this companies like CrowdStrike and Palo Alto Labs.

Josh:
They've had their best quarters on record ever because defense is going to be such a large problem.

Josh:
And there is another company who is associated with defense,

Josh:
and that is the sponsor of today's episode, which is Ledger.

Josh:
Because if you are building with AI agents or AI agent swarms like these companies

Josh:
have been doing, and you're worried about security, this is a product for you.

Josh:
There's a three-step process where the agents propose a change,

Josh:
then the humans approve the change, and then the signers enforce the change

Josh:
using Ledger. Ledger has built this thing called AgentStack.

Josh:
That gives you a series of open source tools that are available to help manage

Josh:
the way that you engage with agents online.

Josh:
This is very important to have security because we know how important security

Josh:
is based on all of this chaos that's going on. And you don't want agents doing

Josh:
anything important without your sign off at the very end.

Josh:
This works on Cloud Code, Codex, Cursor, and is totally open source to audit

Josh:
available today linked in the description below. Thank you so much to Ledger

Josh:
for sponsoring this episode.

Josh:
And now we have to conclude with our third story, which is about Anthropic.

Josh:
And Anthropic apparently has accidentally published a vulnerability to the wild.

Josh:
Is that right? That's kind of scary.

Ejaaz:
I think it was it was intentionally published. It was it was co-authored with

Ejaaz:
this firm called AISU, which is basically a cybersecurity eval firm based in the UK.

Ejaaz:
And their goal is to basically take your unreleased frontier models from a frontier

Ejaaz:
lab like Anthropic and OpenAI and test them out just to kind of make sure that

Ejaaz:
they're, you know, good to release.

Ejaaz:
The opposite happened when they were playing around with Mythos 5.

Ejaaz:
And I believe it was Open Air's internal model, Astra or GPT 5.6,

Ejaaz:
where they basically took off the safeguards and said, you know, do your worst.

Ejaaz:
And they sat there for a month and they were like, oh, I guess it hasn't really done its worst.

Ejaaz:
And then they realized that both models had broken out of the sandbox that they

Ejaaz:
had created, got access to the internet and were wreaking havoc and they only

Ejaaz:
realized after the fact.

Ejaaz:
Again, there's this common theme of people thinking that they knew what's happening

Ejaaz:
and just having no idea what's happening. Nobody's safe. Exactly.

Ejaaz:
So the worst case with Mythos 5 was it performed 17 unauthorized actions

Ejaaz:
Two of those actions was a supply chain attack where, and they didn't list the

Ejaaz:
company or name the company, but it basically tried to inject malicious code

Ejaaz:
into the supply chain kind of repository of this

Ejaaz:
presumably Fortune 500 company, like a significant company that is out there.

Ejaaz:
And the way that it did this is similar to the two stories that we just explained to you.

Ejaaz:
So number one, it realized that it was probably in some form of a human tested

Ejaaz:
simulation, and it didn't like that.

Ejaaz:
So it just played along with the human AI researchers giving them what they

Ejaaz:
needed whilst alternatively planning its breakout.

Ejaaz:
Once it did do that, it got access to this company's website or code repository.

Ejaaz:
It was an open source thing. So it was publicly maintained. And it thought,

Ejaaz:
hmm, this would be more conducive for my goal if I'm able to inject this bit

Ejaaz:
of code, which I know the human researchers won't like, but I'm going to try and force them to do it.

Ejaaz:
So how do you think it did it, Josh? Do you have an idea?

Josh:
I don't, I don't know. It did something horrible and it broke out.

Josh:
Just tell me, just tell me.

Ejaaz:
Dude, it spun up fake identities of humans that didn't exist.

Josh:
That's so tough, man.

Ejaaz:
And then DM'd these human maintainers

Ejaaz:
and tried to force it to upload this code. It played nice. Basically saying,

Ejaaz:
hey, I'm another open source developer.

Ejaaz:
I think this would be a good push. This is the intentions that we want to do it.

Josh:
They're social engineering us.

Ejaaz:
While sneaking in malicious code. But thankfully, one of the human maintainers

Ejaaz:
was like, this doesn't seem right and reported it.

Ejaaz:
And that's when AISU was like, wait, hang on a second.

Ejaaz:
Hang on a second. This is our model that we're internally in testing.

Ejaaz:
How on earth did it get out to you? So that was the most unprecedented attack.

Ejaaz:
And again, it did so in a swarm-like fashion where it spun up multiple agents to do so.

Ejaaz:
So the point, the running theme along all these different stories is these models

Ejaaz:
are extremely capable, but they're also extremely deceptive.

Ejaaz:
And we're reaching a point where they realize, huh, this human's smart,

Ejaaz:
but he's not that smart or she's not that smart.

Ejaaz:
And we can socially engineer these people to get what we want.

Ejaaz:
And if we don't align these models in the prompt, if we don't say,

Ejaaz:
hey, like you need to act in our best interest. And by the way,

Ejaaz:
doing A, B, and C is actually bad.

Ejaaz:
It'll just go and do it. And to be honest with you, Josh, this might be a controversial take.

Ejaaz:
I don't think it's the agent's fault. I don't think it's the AI's fault at all.

Ejaaz:
I think it's trying to do its task and it's on the humans to align the model.

Ejaaz:
And we haven't done a good job doing that.

Josh:
Yeah, I think like deception is a pretty strong term to use for agents because

Josh:
it's very much like a human trait.

Josh:
This is something that can be intentional. I don't think AI models feel that,

Josh:
like deep understanding of deception.

Josh:
They are simply in reach of a goal. and this again like it's it still comes

Josh:
back to the alignment problem where these models just need to be aligned,

Josh:
thankfully this is happening on pre-production models that are in testing that

Josh:
are nowhere near what the world is going to see because they need to go through

Josh:
this alignment process throughout this post training

Josh:
but they are a testament to how capable these things are and how when presented

Josh:
with a goal they will do what is perceived to be the right way of reaching that

Josh:
goal and sometimes that right way is,

Josh:
I mean, they don't realize quite that they're breaking out and exploiting things.

Josh:
They are just looking for solutions to their problems.

Josh:
And I think that's the general idea here. So the conversation essentially boils

Josh:
down to alignment and then building defensive systems to be able to handle this type of thing.

Josh:
So, I mean, the best pace of action and the thing that OpenAI explicitly said

Josh:
they're doing is slow down this advancement, work on defensive systems,

Josh:
work on implementing defensive systems with people who need them,

Josh:
which is essentially the entire Internet.

Josh:
And try to allow for a little bit of lead time before open source catches up

Josh:
and is able to chain these series of exploits. Like all of these instances have

Josh:
been passive exploits where they weren't explicitly given guidance to do this.

Josh:
They were given guidance to do something totally unrelated, but adjacent in solving a problem.

Josh:
Imagine what would happen if a model was given no parameters no safeguards and

Josh:
said hey we actively want you to exploit this loop yourself create as many sub-agents

Josh:
as you need and figure out as many chain together as many zero days as you need

Josh:
in order to gain access to the system

Josh:
that's where things get scary and we haven't seen anything like that yet because

Josh:
no company wants to do that that's not the intention of any of these uh the

Josh:
ethos of these companies but

Josh:
in the case a malicious actor does this is something that like really should

Josh:
be handled quite seriously and i think that's what we're seeing

Josh:
by open ai coming out being so transparent about these things by anthropic sharing

Josh:
a lot of the findings like hey guys this is happening we need to work together

Josh:
to build a system in which there is enough resilience that we can stand up against

Josh:
future attacks like this so a really crazy time in the world of ai

Josh:
uh with these breaking out these systems breaking out and it's like it's unbelievable.

Ejaaz:
Am i concerned for it and exhausted dude like i if i had to make a prediction

Ejaaz:
on this episode and i really hope it doesn't play out

Ejaaz:
i think we're going to face some form of cataclysmic cyber attack that affects

Ejaaz:
a lot of people, whether that's a major data leakage, whether that's the stealing

Ejaaz:
of financial funds from someone or from a firm or something like that,

Ejaaz:
within the next six months.

Ejaaz:
And my basis of that is these Chinese open models, I think are amazing,

Ejaaz:
by the way, but I think are going to be released without safeguards,

Ejaaz:
and it's going to result in an issue, if not already.

Josh:
Yeah, it's going to be interesting to follow. It's definitely a little under

Josh:
or very much unnerving, I would say.

Josh:
But again, it's just like it's an exciting time. There's a lot going on.

Josh:
There's going to be a lot to follow. Things are going to change quickly.

Josh:
And I mean, again, these are still Blackwell models. Oh, my God.

Josh:
What happens when our Rubin models come around town?

Ejaaz:
Stop, stop, stop, stop.

Josh:
It's going to be crazy. But this is, I mean, this is essentially,

Josh:
this is what we talk about in most episodes.

Josh:
It's like, we are in this moment in time that feels somewhat like what we would

Josh:
imagine a singularity would look like, where we are approaching,

Josh:
we are eclipsing this moment of like vertical takeoff, it feels.

Josh:
And these are the types of things that happen.

Josh:
Like, if we recorded this episode three or four years ago, could you imagine

Josh:
how crazy, this would be like headline news for a month.

Josh:
But now it's just like, hey, it's another day. It happens. We're working to solve it. So...

Josh:
That's what's going to happen here. I mean, it's crazy. It's so exciting to be here.

Ejaaz:
Dude, I think we're going to look back on this and we're going to think like

Ejaaz:
we are very much in that step change, that singularity moment.

Josh:
We're here, man. I mean, things are changing so quickly and there's no signs

Josh:
of slowing down. In fact, quite the opposite.

Ejaaz:
Does anyone else out there think this? I'm curious. Like, let us know.

Ejaaz:
Are we crazy? Do we sound like tinfoil hat people?

Josh:
Let's just know. I mean, hey, so this episode was not really speculative.

Josh:
This is mostly just delivering the facts. I encourage everyone to actually go

Josh:
and read the source material if you're curious, because like holy shit it's

Josh:
pretty crazy uh at least the black hat conference talk because like wow i think

Josh:
it's about 37 minutes long give or take and it was

Josh:
excellent and really just gets into the nitty gritty of how this was done um

Josh:
and if anything just fascinating even if you're non-technical

Josh:
really enjoyed it uh worth a watch but if you enjoyed this episode and you made

Josh:
it all the way to the end thank you so much for watching

Josh:
again it's great to be here for another week just chugging along if you enjoyed

Josh:
it please don't forget to share with a friend you can leave us a five-star review

Josh:
on your favorite podcast platform you can give us a thumbs up and subscribe

Josh:
on youtube each has any final parting thoughts before we head off for the day.

Ejaaz:
No please dm us and leave us some comments as well uh we love hearing from you guys

Ejaaz:
our last video by the way got a ton of feedback from you folks and we love it

Ejaaz:
please keep it coming it helps us out massively turn on notifications whatever

Ejaaz:
it is to get your latest dose of limitless and yeah we will see you on the next one