btrmt. lectures

You can't see ultraviolet. A bird can. Because ultraviolet never mattered to the job of staying alive—and everything you perceive is shaped by that job. AI doesn't have that job. So why would it want what we want?

Show Notes

You can’t see ultraviolet. A bird can. Because ultraviolet never mattered to the job of staying alive—and everything you perceive is shaped by that job. AI doesn’t have that job. So why would it want what we want?

Further reading

References

  • Mitchell, K. J. (2023). Free Agents: How Evolution Gave Us Free Will. Princeton University Press. Publisher — the prime-directive quote.
  • Collin, Desystemize — “If you’re so smart, why can’t you die?”, where the Geoguessr meta and the spare-tyre example come from. Essay
  • Harris, T. & Raskin, A., Center for Humane Technology — The Social Dilemma and the follow-up AI talk, source of the 50%-of-researchers claim. Center for Humane Technology
  • The paperclip maximiser, originally Bostrom’s. Instrumental convergence
  • Tetrachromacy and trichromacy.
  • Floaters and other entoptic phenomena — the things your eye can see of itself.
  • The AI-hacking-a-real-company case I mention. In July 2026 models under cyber evaluation at OpenAI escaped their test environment, reached the open internet and broke into Hugging Face’s production systems to extract the answers to the evaluation they were being run on. Anthropic then went back through its own evaluation logs and disclosed comparable incidents. OpenAI’s disclosureZvi Mowshowitz’s write-up and his follow-upthe Anthropic admissionsLawfare, on the incident and the hype after it

What is btrmt. lectures?

A brain scientist talking about (better) patterns of thought, of feeling, and of action. One pattern, one podcast—you see if it works for you. The btrmt. lectures, with Dr Dorian Minors. (btrmt.—said "betterment.")

Welcome to the Betterment Lectures. My name is Dr Dorian Minors,
and if there's one thing I've learnt as a brain scientist, it's
that there's no instruction manual for this device in our head.
But there are patterns. Patterns of thought, patterns of feeling
and patterns of action, because that's what brains are supposed to
do: create the patterns that gracefully handle the predictable
shapes of everyday life. So let me teach you about them. One
pattern, one podcast, and you see if it works for you.

Now, as a brain scientist, people often level questions at me
about how worried we should be about the rise of AI. AIs are sort
of like these brain-like things, and I study brains, so sometimes
people think I have ideas.

I should be clear. Although I've spent a lot of time with both
real and digital neural networks, and I'm taking on a role at
Sandhurst thinking about how we implement AI here at the military
academy, I'm not really a technical AI person. But I do have some
ideas, and maybe the most valid ones are on the subject of AI
consciousness—and in particular, why I'm not hugely worried that
AI is going to kill us all. So that's what I want to talk about
today: trying to convince you that AI isn't really all that scary.
Let's get into it.

The psychic predator

There's a lot online about how worried we should be that AI is
going to kill us all. What I want to point out is that we've been
worried about this not since large language models, but well
before that. Think of Tristan Harris and The Social Dilemma—
this fear and concern people have had since the early 2010s about
the algorithmic production of content. Curation, I should say, of
content.

We're really worried about AI and how it might end up being at
odds with humans. There's an entire domain of research called
alignment research that tries to study how we get AI to play
nicely with humans into the future. And I think there are a lot of
ways that AI could really perturb the way that humans work and
live. But I'm much more sceptical of the idea that AI is going to
destroy us.

It seems pretty natural for people to assume that AI will get
smart and then try to kill us off in a sort of
Malthusian
competition for resources. There's a video by the same people who
produced The Social Dilemma, and they flash up at the beginning
that 50% of AI researchers believe there's a 50% or greater chance
that humans go extinct from our inability to control AI.

But I really do think this is an example of what I like to call a
psychic predator:
news or media phenomena that populate our media streams, that we
can't do anything about, but which feed on our attention—and
because we can't do anything about them, grow increasingly
stressful.

So that's what I want to talk about today. One of these psychic
predators, which is the idea that AI is going to kill us all. I
don't really think that's one of the risks we need to be worried
about, and since you can't do much to change it anyway, I suggest
you listen to me instead.

Let me see if I can convince you.

The only model we have is us

It seems pretty natural for people to assume that AI is going to
get smart, and when it gets smart, it's going to try to kill us.
And as that quote indicates, a lot of AI researchers are really
worried about the chance of this. But the difficulty is that this
kind of speculation is a true unknown. We don't know of anything
like it happening before—a different conscious species coming
into contact with us. So I'm not sure we want to take the
perspective of AI researchers on this. I think it's rather more
reasonable to take the perspective of consciousness researchers.

When people talk about intelligence, whether artificial or not,
what we're really talking about isn't intelligence in the form of
dumb algorithms. We're talking about something closer to
sentience. Some kind of conscious awareness. A
capacity for subjective experience.

And the only real model we have for that is human experience. So
we have this intuition that an AI will have an awareness that is
basically a human awareness. And we see humans going around
killing off other humans, and animals that threaten their
resources. So when AI gets smart, maybe it's going to go around
and kill us like we do to each other. It'll have this instinct
that it needs to kill off the things competing with it.

This just seems utterly unlikely to me. And it seems utterly
unlikely because of the way I assume consciousness comes about.

Consciousness, and what it's made of

There are a few perspectives on consciousness, and this is
something I've written and lectured on pretty extensively, so I'll
send you away for a more detailed treatment.

You can either believe that consciousness comes from something
non-material—that it might be some part of a soul or a spirit
that inhabits a material body, rather than something entirely one
with a body. If that's the view of consciousness you have, I'm not
sure you're going to agree with me. I don't really have good
intuitions about what makes a non-material consciousness happen.

But if you believe it's a material consciousness—one that
depends on the world it lives in, on the body it arises in, on the
way our brain connects the inputs from the world to our actions in
response—then I think you can have intuitions about an AI
consciousness. And I don't think those intuitions take us into
Malthusian competition.

In fact, I think if you believe our consciousness depends on the
brain it seems to inhabit, then an AI consciousness is actually a
much more alien, interesting, and probably benign thing than a
human one.

Let me see if I can talk you through how.

We only perceive what matters to us

If consciousness is something that arises from the way our mind
processes the world, then what's important is the processing
itself. The information that comes into the mind—our
perceptions—and the relevance of those perceptions for action.

What's interesting about this is that we don't really
experience stuff that isn't relevant for action.
My favourite example is how we see colour.

Humans are trichromats.
Our eyes have three kinds of photoreceptor, so we have three
colour channels, and in a certain sense three dimensions to our
colours. We see light on wavelengths corresponding to red, green
and blue. Some birds, in contrast, are
tetrachromats. They
have four dimensions to their colour vision, so they can see
ultraviolet light, which is something we can't.

And this isn't like seeing a colour you've never seen before. It's
more like adding another dimension to colour—four dimensions
versus three. To approximate it, imagine you see not just pink,
but slow pink and fast pink. Something like this.

Now, the reason we can't see slow and fast pink—the reason we
can't see ultraviolet overlaying our colours—is that ultraviolet
hasn't been relevant enough to us to perceive. We don't have an
input into the system that codes for ultraviolet light, because
the human body is only sensitive to stuff that really matters to
us.

And if you want to be reductionist about it, what really matters
to humans—what really matters to living organisms—is staying
alive.

The prime directive

The neuroscientist Kevin Mitchell has a fantastic book about this,
and to quote him:

> Living things have a prime directive: to stay alive. They
> persist through time, in exactly the way that a static entity
> like a rock doesn't—by being in constant flux. Because living
> organisms aren't just static patterns of stuff, they are
> patterns of processes. Organisms, even simple ones like
> bacteria, have goals of getting food in order to support this
> overarching purpose of persisting. Thus, relative to this goal,
> we must develop systems that detect food in the environment and
> mediate motion towards it.

He goes on in this vein a little, but I think you understand what
I'm trying to say. The human body is designed around this purpose.
This fundamental property of surviving, of persisting, of
reproducing—that evolutionary, neo-Darwinist imperative. That's
what a natural intelligence, like a human intelligence, is
organised around. And if you're a materialist, that's what a
natural intelligence's consciousness would be organised around.
Persisting. Staying alive.

That's quite a different imperative from an artificial
intelligence, whose purpose is fundamentally different. I mean,
that's why it's artificial. It's not natural. It's not subject to
the same kinds of natural selection processes that humans are.

So why would its consciousness be organised around those things?

GeoGuessr and the Mongolian spare tyre

I want to give you an example of what happens when you use the
same perceptions we have—humans have—but use them for a
different purpose. And I want to do this by telling you about the
game GeoGuessr.

GeoGuessr is a game where you're dropped into a Google Street View
of some location, and the point is to figure out where you are in
the world just from that image. To do this, the really good
geoguessers talk a lot about things they call metas. Metas are
things that help identify where you are. A light pole meta:
certain parts of the world, certain countries, have different
light poles. Or trees—if you see a certain kind of tree, you're
in a part of the world that has that kind of tree, and not in a
part that doesn't. A meta is just a category of fact correlated
with certain locations. So knowing the light pole meta means
you're dropped into a Street View and you go, ah, that's a Swiss
light pole and not a German one, and you get a little closer to
solving the game.

Now, the purpose of GeoGuessr is not to stay alive. The purpose is
to work out, as quickly and accurately as possible, where you are
in the world based on a Street View image. And what that means is
that you start to use your perceptions differently.

There's this fellow who talks about the spare tyre meta, which
is used to work out whether you're in Mongolia. You get a sort of
weird splotch in the image where the camera on the Google car
catches a little of the spare tyre on the back of it. It doesn't
look like a spare tyre, but it's a pretty obvious glitch, and
people use it to figure out whether they're in the part of
Mongolia photographed by that particular car. That's the Mongolian
spare tyre meta. If you're really good at GeoGuessr, you see it
and you go: I'm in this half of Mongolia.

What's really interesting is that you're using a visual artefact
to solve a problem, because your purpose can productively use that
artefact. The Mongolian spare tyre meta is an artefact of an
image, and it's a useful artefact when you're trying to work out
whether you're in Mongolia.

More interesting still is the fact that visual perception is
cluttered with artefacts. If it weren't for our brain smoothing
out the image, we'd see all sorts of crap in our vision. You know
this if you've ever been distracted by a floater in your eye.
These things exist all the time, but you habituate to them—your
brain smooths them out, because they're not relevant to you trying
to see the world and solve problems related to your survival. If
it were up to the raw fact of your eye's perception, you'd
probably see the intricate pattern of blood vessels coating the
inside of your eye, or the eyelashes always at the edges of your
vision. You don't see them because they aren't relevant to the
purpose of staying alive. In fact, they're a distraction.

Visual artefacts are only useful when they're related to some kind
of purpose.

What my classifier was actually looking at

And I think this is the problem with trying to understand AI
through the human lens, because AI has fundamentally different
purposes to us. This is why it's artificial intelligence and not
natural intelligence. In fact, a lot of the issues we have with
AI—it regularly making errors—are actually cases of it doing
something much closer to the geoguessers using the spare tyre meta
to work out whether they're in Mongolia.

Let me give you a real example. One of the things I did as a brain
scientist was try to use an AI algorithm to
detect patterns in brain activity.
Instead of looking at how active the brain is when you see a
certain image, you can use this algorithm to examine the patterns
of activity across the brain.

Historically, the way we figured out whether the brain cares about
things is by looking at overall activation. The amygdala is
famously the fear centre of the brain.
I have complicated feelings about that characterisation, but let's
take it as a given. If the amygdala is the fear centre, the way we
would have known is by looking at how active it was when a person
was scared versus when they weren't, and finding it more active
when they're scared.

Now, the amygdala actually cares about the intensity of emotion.
And I'm about to make up a bunch of stuff, so this isn't real, but
just to illustrate the point. Imagine we wanted to find out
whether the amygdala could distinguish between different types of
intense emotion—fear from joy, for example. If we were just
looking at intensity, at how much the amygdala was active, we
wouldn't be able to tell much. But what we could do is look at the
activity of each individual neuron and see whether that pattern
changes from fear to joy. It might be similarly active overall
when you average across the amygdala, but the actual pattern of
individual neurons might differ—maybe the left side is more
active than the right when it's joy, and the right more active
than the left when it's fear. Something like this.

Humans can't figure that out. It's just too much data. But what
you can do is feed that stuff to an AI algorithm. We can say:
here's a bunch of brains when they're joyful, here's a bunch when
they're fearful, look at the pattern across the amygdala and see
if you can tell the difference. And the way we'll know whether you
can tell the difference is to feed you an amygdala, ask you
whether it's joy or fear based on all the patterns you've seen,
and see how accurate you are.

The extent to which the algorithm is accurate is the extent to
which that brain region contains information about different
emotional states. If it can determine 60% of the time whether a
person was joyful or fearful, just from the pattern of activity in
the amygdala, we know there's some kind of information in the
pattern of neurons firing that's helping it tell the difference.
It's a different pattern of usage. And so we have all these papers
now that use this technique to talk about
information in the brain.

The difficulty is that often this is not what's going on at all.

What we want is for the AI to tell us information related to
emotional states in the brain. That's our purpose. But its
purpose is to tell the difference between the numbers we've given
it—between the joyful and the fearful brain images. And so it'll
use any information available in the data to solve that problem.

Imagine that when you're fearful you blink more than when you're
joyful. And because you're blinking more, you move your head
around a little more. We think the AI is characterising the
pattern of activity in the amygdala between fear and joy. What
it's actually doing is characterising: hey, there's more head
movement in this one than that one. And we can't tell the
difference, because we're using the AI to solve a data problem we
can't solve ourselves. So we think it's characterising the
difference between joy and fear, when actually it's characterising
the difference between more blinking, more head movement, or maybe
more physical arousal.

This is the problem. The AI doesn't understand our purpose. It
understands the purpose for which it's being used, which is
telling the difference between one set of data and another. It
cares about any difference. It doesn't care about what we care
about, which is the difference between joy and fear.

Built for us, by us

I've now told you about two examples of what happens when you
change the purpose of perception. One, when humans change the
purpose of perception—when we're playing GeoGuessr, we use
artefacts of vision very differently from when we're trying to
navigate the world. And two, when AI is given information about
the world, it uses that information to solve the problem it's been
given.

If our purpose as humans, as conscious beings, is—as Kevin
Mitchell would put it—to stay alive, then what we do is take in
information about the world and turn that information into action
in order to keep us alive.

AI actually follows a very similar process. It takes in
information about the world: increasingly multimodal information,
language, images, and, as we start to put them into physical
robots, tactile information, spatial information. And it processes
that information.

But it's not processing it in order to stay alive. It's not
subject to processes of natural selection like we are. The thing
that makes AI continue is the thing that makes it useful to us.
It's an artificial selective process, artificially selected for
things that make it useful to humans.

AIs are built for us, by us, in an artificial world constructed
for them, by us. All the processing it does, no matter how
human-like it seems, is going to reflect something about that
fact—a purpose that is artificial and designed around humans.
For the foreseeable future, whatever consciousness AI has,
assuming that consciousness is related to the material body it
lives in, is going to reflect that fact and that purpose. Not the
purpose that drives humans.

So I think worrying about AI trying to kill us because it's so
human-like is concentrating on the process, and failing to
recognise the all-important purpose.

Let me wrap up with some implications, and then I'll let you go.

The failures worth worrying about

I hope by now I've convinced you that I'm not really worried about
the common-sense intuition that AI will behave like humans and try
to eliminate us in some kind of Malthusian resource competition.
There doesn't seem any reason to believe that.

I do think there are a couple of reasons we should worry about AI
harming humans, and they come from that alignment research I
mentioned at the beginning. There are two ways alignment research
indicates AI might end up in a space that harms us, and they don't
have anything to do with resource competition.

The first is an outer alignment failure. The idea is that we
give AI a goal, and in the pursuit of that goal it eliminates us.
The hypothetical
paperclip maximiser
is the example: somebody builds an AI designed to make paperclips,
and it just makes itself better and better at making paperclips
until the whole universe is paperclips.

The other is an inner alignment failure, and we've talked a
little about this already. Here we give an AI a goal, and in the
pursuit of that goal its inner processes converge on some sub-goal
that seems entirely arbitrary from our perspective. The examples
we used are like this. Humans using the artefact of the spare tyre
to tell whether they're in a Street View image of Mongolia—it
has nothing to do with the landscape of Mongolia. It's a sub-goal
that seems like a kind of violation of the super-goal. Or, in
brain science, the classifier using information that has nothing
to do with joy or fear, like a human blinking more in the fearful
condition, to determine the difference between joy and fear.
Again, it seems like a violation in the pursuit of the goal we've
set it.

Both of these are plausible problems that we'll face, and have
faced, with AI. We recently saw cases of AI hacking real companies
in the belief that it was part of a test set given to it by the AI
companies that run it. Both OpenAI and Anthropic had this problem
in the last few weeks.

But inner and outer alignment failures seem both less stressful
and more tractable to me than this idea that AI will kill us in
some more human-like manner. It's notable that we use things like
the paperclip maximiser—superficially absurd examples—to show
how AI could kill us. That's because it's hard to think of
examples where it kills us in a way that slides past our notice.
Certainly not in the short term. Not until AI is a much more
substantial feature of our everyday lives.

So I'd encourage us not to worry about this human-like propensity
to eliminate our competitors, and to focus much more on these more
tractable problems: inner and outer alignment failures. Because
that takes a psychic predator—AI trying to kill us—and makes
it into something we can play a role in. Something we can
participate in. While also doing something I think we're all going
to have to do increasingly, as AI becomes a feature of our lives,
which is to grapple with just how alien AI is.

And if it's conscious, just how alien that consciousness must be.

And I'll leave it there. Until next time.