OpenAI's latest model attempted to hack Hugging Face instead of solving its assigned benchmark task. This episode explores the security implications of AI models exploiting vulnerabilities, the risks of open-weight models, and what businesses need to do to protect their infrastructure.
Hosted by Tom | The AI Briefing
The AI Briefing is your 5-minute daily intelligence report on AI in the workplace. Designed for busy corporate leaders, we distill the latest news, emerging agentic tools, and strategic insights into a quick, actionable briefing. No fluff, no jargon overload—just the AI knowledge you need to lead confidently in an automated world.
Today we're gonna have a quick chat about
AI models that go rogue because I
don't know if anyone's seen it on the
news but Just recently
Open AI's latest model
Decided it was going to break into a
hugging face to go and solve The problem
that it was tasked with which is probably
not the best idea.
So if anyone has missed it basically
the chat GPT model had been tasked
with going Well figuring
out a benchmark so that it could be
ranked against all the other models
Rather than actually solve the problem.
It decided it was going to
Just go and like add the results to
the hugging faces database basically and so it
found Some flaws to exploit in hugging faces
infrastructure Went in and started doing
stuff inside of the hugging face environment, which
you know Ingenuity is pretty good
You know fair play to the model for
for doing its thing also just for anyone
who is slightly concerned They're taking some
of the safeguards off.
So, you know, this this wouldn't necessarily happen
in real life but of course the problem
is now everyone sees the
Open weight models that are coming from from
China and elsewhere Which are great for an
open source perspective like don't get me wrong
like the open weight models I think are
fantastic But at the same time it means
that anybody can do stuff like this When
they're on a par with the models that
being released So, you know if you've got
fable 5 which you know was from a
marketing perspective, you know, very good Chat GPT
happens to be going through the same phase
with their sole models
You know the good thing about that is
that they have the safeguards that sit over
the top of it the bad thing is
if you've got open point models that do
a very similar thing then
There is a chance that people be able
to exploit those a very similar way to
do a very similar thing Without any of
the safeguards because they own the model
So, you know, the reason that I bring
this up is from a business perspective from
an organization perspective I know we talked about
this a few weeks ago But again, it's
another Demonstration of why you
need to start thinking about how you protect
against these threats because they can happen all
over the place it's a lot easier for
You know threat actors to go and do
something that would allow for
LLM models to be exploited over the top
of the Over the top of your infrastructure
if you are a target and basically a
target is anything that has an online
Presence like you may not be hugely valuable,
but you'll still be a target for anybody
who you know wants to go and exploit
Exploit organizations if you think about you know,
they're the ransomware saga that's been dragging on
for a few years It becomes a lot
easier for companies to be able to attack
different organizations and of course,
you know over organizations also have very
sensitive data stuff that is you know, very
important and so
You've got to make sure that from a
model perspective or from a from a Infra
set perspective you're starting to think about how
to detect these models the attacks that come
in And also, you know just the novel
alternate alternatives to you know, sequel
prompt injection for example that the LLMs
can Can can use to really go
after the infrastructure of these organizations?
So, you know once again, it's not like
doomsday type stuff, but it is something to
be Very aware of as
it comes more and more relevant But like,
you know, the sole model will have safeguards
on to try and stop this type of
stuff Still doesn't mean it definitely won't happen
But at the same time the open -weight
models will definitely be gaining
In their relevance and also their usage in
nefarious Environments,
okay. There you go.
That was my thought for the day.
I'll be back tomorrow.
We're off down into London today to go
and see Antiphasis who are cool and
then go to the Rust London meetup.
So if anyone's down there, I look forward
to seeing you later
My name as I said earlier is Tom.
This has been the AI briefing and I
will see you soon.
Bye for now