WEBVTT

NOTE
This file was generated by Descript 

00:00:00.000 --> 00:00:03.630
Judith: Welcome to Berry's In the
Interim podcast, where we explore the

00:00:03.630 --> 00:00:08.060
cutting edge of innovative clinical
trial design for the pharmaceutical and

00:00:08.060 --> 00:00:10.260
medical industries, and so much more.

00:00:10.820 --> 00:00:11.470
Let's dive in.

00:00:12.846 --> 00:00:14.136
Scott Berry: All right, welcome.

00:00:14.196 --> 00:00:15.426
We are back.

00:00:15.456 --> 00:00:20.436
Uh, in the interim, this is
Barry Consultants podcast of all

00:00:20.436 --> 00:00:26.046
things statistical, all things
scientific with clinical trials, uh,

00:00:26.046 --> 00:00:29.076
medical decision making, uh, and.

00:00:29.111 --> 00:00:35.291
We have an interesting podcast today
in that I don't know the topic, so

00:00:35.291 --> 00:00:41.741
my, my co-host on in the interim,
uh, Kurt Veley is here and he has a

00:00:41.741 --> 00:00:44.561
surprise topic to talk about today.

00:00:44.561 --> 00:00:45.281
Kurt.

00:00:46.281 --> 00:00:47.031
Kert Viele: Hey Scott.

00:00:47.151 --> 00:00:49.461
Um, alright, so I wanna
start with a story.

00:00:49.731 --> 00:00:54.471
I went to a conference not that long
ago, and standard statistical conference.

00:00:54.471 --> 00:00:57.681
People are talking about the
data they had, all the methods

00:00:57.681 --> 00:00:59.121
that they've been developing to.

00:00:59.381 --> 00:01:04.391
Understand it, what they've learned from
the data, and essentially all of these.

00:01:04.836 --> 00:01:05.256
Talks.

00:01:05.376 --> 00:01:08.166
They had a really interesting structure,
so they, they get through all their

00:01:08.166 --> 00:01:13.356
methods and everything, and at the
end of every single talk, it was ended

00:01:13.356 --> 00:01:17.826
with, if only I had gotten a chance to
design the data, design the experiment,

00:01:17.916 --> 00:01:22.056
design the data, do this in advance,
everything would've been better.

00:01:22.146 --> 00:01:24.366
I would've avoided all
these problems and so on.

00:01:25.086 --> 00:01:28.836
And so I'm listening to all these topics
and when I get to the, my talk, of

00:01:28.836 --> 00:01:33.366
course, I started it after listening
to all this is I live in utopia.

00:01:34.356 --> 00:01:36.636
And then talked about experimental design.

00:01:37.326 --> 00:01:41.706
And so I, I get back in the car on the,
on the way home, and I'm sitting here

00:01:41.706 --> 00:01:47.166
thinking about, so am I really in the
good place or am I in the bad place?

00:01:47.796 --> 00:01:50.706
Because one thing that's always
impressed me about the last 20

00:01:50.706 --> 00:01:53.496
years in experimental design.

00:01:53.971 --> 00:01:57.871
Is all of these methods that we're
developing to understand data,

00:01:57.871 --> 00:02:02.401
causal inference, all of these,
these aspects of, of how do we make

00:02:02.431 --> 00:02:04.501
inferences from data in front of us?

00:02:05.101 --> 00:02:11.341
We don't use those in experimental design,
so we often have an idea that when we

00:02:11.341 --> 00:02:16.621
design a trial, we're to ignore every
bit of data that's ever existed on Earth.

00:02:17.101 --> 00:02:20.311
If I go to my doctor and I ask,
why are you giving me a drug?

00:02:20.631 --> 00:02:24.021
They're gonna report lots and lots of
studies that say, this drug is good.

00:02:24.591 --> 00:02:28.461
If I go into an experiment and
say, I want to use this data,

00:02:28.521 --> 00:02:30.321
I'm immediately told it's bad.

00:02:30.861 --> 00:02:34.401
And that depends on if
I'm borrowing information.

00:02:34.401 --> 00:02:36.381
I got a drug that's targeted for I.

00:02:37.116 --> 00:02:39.096
A specific mutation.

00:02:39.426 --> 00:02:41.616
I have data on thyroid cancer.

00:02:41.616 --> 00:02:43.326
Can I use that for lung cancer?

00:02:43.506 --> 00:02:45.336
Generally, the answer is often no.

00:02:45.816 --> 00:02:50.496
If I wanna borrow historical data from
old clinical trials, Alzheimer's, we

00:02:50.496 --> 00:02:55.296
have hundreds of thousands of patients
that have been treated on placebos.

00:02:55.326 --> 00:02:55.986
Can I use it?

00:02:56.381 --> 00:02:57.971
Well, there's problems with that.

00:02:58.001 --> 00:03:03.791
So on real world evidence, so I, I would
actually define the standard debate

00:03:03.791 --> 00:03:09.101
that we've been having for the last
15 years is essentially, is data good

00:03:09.281 --> 00:03:12.011
or is it bad in experimental design?

00:03:12.401 --> 00:03:14.861
So does it lead us to a
good place or a bad place?

00:03:14.921 --> 00:03:18.581
So I'd like to talk about where
we are, how we got here, and

00:03:18.581 --> 00:03:19.691
where we think we're going.

00:03:20.111 --> 00:03:21.311
So I'm gonna leave it up to you.

00:03:21.311 --> 00:03:24.101
I've surprised you with a topic
and see what your reactions are.

00:03:24.576 --> 00:03:27.816
Scott Berry: well, well let's, let's,
let's sort of figure out the topic

00:03:27.816 --> 00:03:29.046
and the interesting part of this.

00:03:29.046 --> 00:03:31.596
So what, when you say is data good or bad?

00:03:31.776 --> 00:03:35.856
Do you mean that When I design an
experiment, I'm saying I want to collect

00:03:35.856 --> 00:03:39.666
the following data and I can analyze
just the data in that experiment,

00:03:39.936 --> 00:03:41.466
which we've been doing for a long time.

00:03:41.816 --> 00:03:46.466
Uh, frequentist, uh, rarely
Bayesian, but we, we do that.

00:03:46.796 --> 00:03:50.096
You're saying by data, is
that outside the experiment?

00:03:50.096 --> 00:03:52.466
Is there any room in that experiment?

00:03:52.466 --> 00:03:53.306
For external,

00:03:53.871 --> 00:03:54.321
Kert Viele: Yes.

00:03:54.381 --> 00:03:54.891
So what?

00:03:54.891 --> 00:03:58.071
What do we do with essentially
the totality of human knowledge?

00:03:58.416 --> 00:04:01.926
Prior to the experiment, should
that enter into the experiment

00:04:01.926 --> 00:04:04.236
in any way and wanna avoid this?

00:04:04.236 --> 00:04:06.516
I mean, there's lots of
rhetoric we could do with this.

00:04:06.726 --> 00:04:10.926
If you want to, if you want to say bad
things about this, you talk about biases

00:04:10.926 --> 00:04:12.936
and confounding and everything else.

00:04:12.936 --> 00:04:16.836
If you wanna say good things, you talk
about totality, the evidence, but really

00:04:16.836 --> 00:04:20.826
let's get at, you know, what is gonna
lead us in good and bad directions?

00:04:20.976 --> 00:04:23.166
How do we decide when to
do this and when to not?

00:04:24.516 --> 00:04:24.846
Scott Berry: Yeah.

00:04:24.846 --> 00:04:27.486
So let's, let's, let's lay out why not.

00:04:28.176 --> 00:04:28.506
Why?

00:04:28.506 --> 00:04:32.046
I mean, why, what are the reasons
people give that we shouldn't

00:04:32.046 --> 00:04:35.046
use other data in our experiment?

00:04:37.281 --> 00:04:38.511
Kert Viele: Obvious answer would be.

00:04:39.176 --> 00:04:41.041
It, it could lead us
in the wrong direction.

00:04:41.071 --> 00:04:45.211
So the, uh, experiment might have
been done with a different set

00:04:45.211 --> 00:04:47.131
of patients at a different time.

00:04:47.371 --> 00:04:52.081
There are key differences between that old
data and what my current experiment is.

00:04:52.441 --> 00:04:56.581
If I use the new data or use the
old data in my new experiment.

00:04:57.111 --> 00:05:00.831
I can get, uh, answers that
are biased in certain ways.

00:05:00.831 --> 00:05:02.211
I can draw wrong conclusions.

00:05:02.211 --> 00:05:06.831
I can say drugs work that don't simply
because of the information I'm bringing

00:05:06.831 --> 00:05:08.961
in rather than the experiment itself.

00:05:09.936 --> 00:05:12.816
Scott Berry: So is this, is
this a frequentist issue?

00:05:12.816 --> 00:05:17.826
Is this that we calculate the operating
characteristics of the new experiment?

00:05:17.826 --> 00:05:21.546
Type one error is only the new experiment.

00:05:22.836 --> 00:05:27.336
And if I use any data outside
of my experiment, you can

00:05:27.336 --> 00:05:28.476
inflate type one error.

00:05:28.506 --> 00:05:33.576
You can get bias, all these
bad terms, uh, uh, in it.

00:05:34.206 --> 00:05:40.356
Is it a frequentist problem
that data outside the experiment

00:05:40.386 --> 00:05:44.046
all of a sudden type one error
means something very different?

00:05:44.496 --> 00:05:49.476
Uh, bias means something very
different, uh, that I can't

00:05:49.476 --> 00:05:51.001
use stuff and be a frequentist.

00:05:52.476 --> 00:05:54.516
Kert Viele: So I don't know
if it's a, I think there is a

00:05:54.516 --> 00:05:56.256
frequent DYS Bayesian divide.

00:05:56.346 --> 00:06:01.416
I think Bayesians are trained with the
idea of collect some data, update your

00:06:01.416 --> 00:06:05.466
beliefs, collect more data, update your
beliefs, and it's a natural thing to do.

00:06:05.466 --> 00:06:05.526
I.

00:06:05.856 --> 00:06:11.856
To bring things in, but I don't know if,
if I were to say that the problem is not

00:06:11.856 --> 00:06:16.836
being a frequentist in terms of our, do
you care about long-term error rates?

00:06:17.286 --> 00:06:22.336
I would lay the problem at the
feet of, you have to have this 2.5

00:06:22.356 --> 00:06:22.956
type one error.

00:06:23.721 --> 00:06:26.181
In the worst case scenario.

00:06:26.661 --> 00:06:30.231
So in effect, we're playing
minimax against nature.

00:06:30.771 --> 00:06:37.911
We have to assume that nature has deceived
us for the last 25 years, and then in

00:06:37.911 --> 00:06:39.921
that case, we shouldn't use the data.

00:06:40.431 --> 00:06:44.841
So, uh, on the one hand, that delivers
an immense amount of robustness.

00:06:44.961 --> 00:06:46.701
On the other hand, we're
reinventing the wheel.

00:06:47.481 --> 00:06:47.751
Scott Berry: Yeah.

00:06:47.901 --> 00:06:52.221
Uh, so every time, so let, let's,
let's talk about a case where we would

00:06:52.221 --> 00:06:56.151
use external data, and we have, uh,
we've designed trials where we've used

00:06:56.151 --> 00:06:59.961
external data and what that looks like
and what could be the potential problem.

00:06:59.961 --> 00:07:01.881
So, um, suppose.

00:07:02.411 --> 00:07:07.601
I've got results from another trial,
a phase two trial about the relative

00:07:07.601 --> 00:07:11.351
efficacy of a treatment, and I
bring it into my next experiment.

00:07:11.621 --> 00:07:13.751
And I use that as prior knowledge.

00:07:14.231 --> 00:07:18.671
And when I'm done with my new
experiment, I combine them together.

00:07:18.671 --> 00:07:18.731
I.

00:07:19.641 --> 00:07:23.211
Or I bring in previous
data on the control arm.

00:07:23.691 --> 00:07:27.891
Uh, I'm comparing to an active
comparator, something that's approved.

00:07:27.891 --> 00:07:32.631
It's standard of care, and I wanna
run a new experiment of my new drug.

00:07:32.946 --> 00:07:35.496
And I want to compare to standard of care.

00:07:35.496 --> 00:07:39.006
There are lots of trials and lots of
information about standard of care.

00:07:39.576 --> 00:07:44.046
Why do I have to run a one-to-one
randomized trial standard of care to my

00:07:44.046 --> 00:07:46.086
therapy when I know so much about it?

00:07:46.416 --> 00:07:50.136
So I might bring in that information
specifically into the trial

00:07:50.676 --> 00:07:53.946
and e both of those notions.

00:07:54.786 --> 00:07:59.496
Can cause issues in my new
experiment essentially.

00:07:59.496 --> 00:08:04.626
If it's wrong, I if that is
different than my new experiment.

00:08:05.016 --> 00:08:08.616
Statisticians are great at
saying, well, boy, if that data

00:08:08.616 --> 00:08:09.816
was a little bit different.

00:08:09.816 --> 00:08:11.466
Now your new experiment.

00:08:11.926 --> 00:08:19.636
Has a type one error or bias or, or
issues that if I bring in anything from

00:08:19.636 --> 00:08:21.436
outside the trial, it can cause issues.

00:08:21.436 --> 00:08:24.436
So that might be ways in which
I bring in external data.

00:08:26.811 --> 00:08:29.751
Kert Viele: are certainly, I mean,
there are cases where you are

00:08:29.751 --> 00:08:31.191
gonna run into those problems.

00:08:31.191 --> 00:08:33.711
I mean, we've talked in
examples and antibiotics.

00:08:34.056 --> 00:08:38.346
The development of resistance over
time, we expect therapies to change.

00:08:38.646 --> 00:08:42.666
There are situations where we think
that problem exists and we need

00:08:42.666 --> 00:08:44.946
to address it or not use the data.

00:08:45.276 --> 00:08:50.946
There are also cases where diseases
are very stable over time, and those

00:08:50.946 --> 00:08:52.866
issues may not be as big a concern.

00:08:53.991 --> 00:08:58.191
Scott Berry: Yeah, so I, I mean, I'm,
I, I, I'm very much on the side that

00:08:58.191 --> 00:08:59.811
we should be using external data.

00:08:59.811 --> 00:09:05.211
I, I, uh, I think we go about this
in a way that's just so strikingly

00:09:05.211 --> 00:09:09.591
conservative that it slows us
down, uh, in it, that, that.

00:09:09.716 --> 00:09:12.656
That, you know, something may happen.

00:09:12.656 --> 00:09:18.086
I, I think it falls under the realm
that science is hard, uh, incorporating

00:09:18.086 --> 00:09:23.726
this other information, uh, be explicit
about it, incorporate it, and the

00:09:23.726 --> 00:09:29.306
process of this is hard, but the
notion that we don't use any of that

00:09:29.306 --> 00:09:32.426
information just seems strikingly wrong.

00:09:34.081 --> 00:09:36.631
Kert Viele: So I, it's an
interesting question 'cause it

00:09:36.631 --> 00:09:38.521
certainly can be wrong on occasion.

00:09:39.021 --> 00:09:43.791
So I, I wonder if as a society and you,
you ask about frequentist and whether

00:09:43.791 --> 00:09:48.891
the error rates are driving this, I
think you can be a, a perfectly good and

00:09:49.131 --> 00:09:54.171
wonderful frequentist and use historical
data, but what you're gonna have to

00:09:54.171 --> 00:09:59.871
accept is most of the time prior data
is leading you in the right direction.

00:10:00.111 --> 00:10:02.691
If it's not, we should just
see science altogether.

00:10:02.976 --> 00:10:04.296
Because we have worse problems.

00:10:04.746 --> 00:10:10.626
But in any case, if it's generally
leading us in the right direction, we're

00:10:10.626 --> 00:10:15.756
gonna be running a few experiments that
have say three or 4% type one error.

00:10:16.056 --> 00:10:19.716
And we're gonna be running a lot more
experiments that have say one or two.

00:10:20.166 --> 00:10:23.856
And in the grand scheme of
things, we're still putting a

00:10:23.856 --> 00:10:25.476
better mix of drugs on the shelf.

00:10:25.761 --> 00:10:27.891
More drugs that work,
less drugs that don't.

00:10:28.461 --> 00:10:33.651
So I think we have to change our mindset
to be frequentist in this kind of world.

00:10:34.956 --> 00:10:40.866
Scott Berry: Yeah, it's a, and I'm struck
by, I, I'm struck by the idea of using

00:10:40.866 --> 00:10:42.761
that data and in and in Bayesian we do

00:10:43.466 --> 00:10:43.586
Kert Viele: I.

00:10:44.256 --> 00:10:45.546
Scott Berry: borrowing
of that information.

00:10:45.546 --> 00:10:47.106
We do things that are dynamic.

00:10:47.106 --> 00:10:51.966
I'm struck by, there are
tons of trials that we use,

00:10:52.056 --> 00:10:53.976
objective performance criteria.

00:10:54.996 --> 00:10:59.676
That it's a single arm trial and the
drug has to jump a particular hurdle.

00:10:59.676 --> 00:11:00.456
That's a number.

00:11:00.726 --> 00:11:04.956
That number's always based on
prior data, and we do that all the

00:11:04.956 --> 00:11:07.176
time, and presumably that's okay.

00:11:09.396 --> 00:11:12.876
Or you run a trial where your
only, your data comes from the new

00:11:12.876 --> 00:11:14.676
experiment, and that's your control.

00:11:15.216 --> 00:11:19.446
But the idea that we do something
halfway between that, that we use

00:11:19.446 --> 00:11:24.816
it, it just strikes it, people the
wrong way, just do one or the other.

00:11:24.816 --> 00:11:29.106
And, and I, I'll, I'll, I'll give
an example of this happening to me.

00:11:29.106 --> 00:11:33.696
Where, by the way, and, and if you listen
to, in the interim a lot, you'll find

00:11:33.696 --> 00:11:37.266
that we, we very much respect the FDA.

00:11:37.536 --> 00:11:43.536
The it, they're, they're, they're
the, um, uh, i I in the world.

00:11:43.536 --> 00:11:45.786
They do tremendous good
and, and they're wonderful.

00:11:45.786 --> 00:11:47.466
It doesn't mean we don't run into hurdles.

00:11:47.466 --> 00:11:48.186
So we

00:11:48.186 --> 00:11:48.906
presented a

00:11:48.951 --> 00:11:51.261
Kert Viele: so just, just to
interrupt you there, I mean, we go

00:11:51.261 --> 00:11:56.721
to the FDAA hundred times a year
and have three major disagreements.

00:11:56.811 --> 00:11:59.121
That's much better than
with our wives, so.

00:12:00.916 --> 00:12:01.206
Scott Berry: Yeah.

00:12:01.611 --> 00:12:01.911
Yeah.

00:12:02.061 --> 00:12:06.381
So we provide an example that we
wanted to go in with a new therapy.

00:12:06.591 --> 00:12:12.141
It was oncology and we were,
we, the, the new trial.

00:12:12.141 --> 00:12:15.351
We didn't want to enroll one
to one, to a standard of care.

00:12:15.861 --> 00:12:18.801
Uh, it's easier to get patients
in if they're more likely to

00:12:18.801 --> 00:12:20.091
get experimental treatment.

00:12:20.481 --> 00:12:25.821
Uh, we wanted to enroll three to one,
three to experimental, one to control.

00:12:26.286 --> 00:12:29.766
There's a good bit of information
about the behavior of the control.

00:12:29.766 --> 00:12:33.606
So we were borrowing from previous
trials, and let's assume that,

00:12:33.606 --> 00:12:39.486
that the prior data was a 15%, uh,
overall, uh, objective response rate.

00:12:40.176 --> 00:12:43.536
Make it simple and, but we're
gonna enroll three to one.

00:12:43.536 --> 00:12:47.436
We're gonna borrow on the
control using that, that prior

00:12:47.436 --> 00:12:49.506
data that's centered on 15%.

00:12:51.216 --> 00:12:56.676
In the new experiment because we're
borrowing data that's 15% on the control.

00:12:56.826 --> 00:13:02.016
If the new rated, that new experiment is
actually higher than that, we're pulling

00:13:02.016 --> 00:13:06.546
the control down, making it easier for
the treatment to look better, and we

00:13:06.546 --> 00:13:08.586
present those operating characteristics.

00:13:08.916 --> 00:13:13.566
That was, that was deemed to
be unacceptable because you

00:13:13.566 --> 00:13:15.066
inflate type one error rate.

00:13:16.716 --> 00:13:22.356
The response from the agency
was use 15% as an objective

00:13:22.356 --> 00:13:28.086
control and don't enroll any new
controls, and you have to beat 15%.

00:13:28.806 --> 00:13:33.756
Of course, that we generally
don't calculate the.

00:13:34.221 --> 00:13:35.181
Type one error.

00:13:35.181 --> 00:13:40.311
If in fact the real rate in
that experiment was 25 or 30%

00:13:40.311 --> 00:13:42.411
because we're jumping 15%.

00:13:42.981 --> 00:13:48.441
But it's almost the, the, the opposite
Bayesian view that we know that answer's

00:13:48.441 --> 00:13:50.331
15% for what you're trying to beat.

00:13:50.331 --> 00:13:51.771
And if you can beat that, that's good.

00:13:52.101 --> 00:13:57.231
It's using a, a, a prior that's all
focused on a single value rather

00:13:57.231 --> 00:14:01.551
than modeling it with some borrowing
and recognizing the uncertainty.

00:14:01.911 --> 00:14:05.091
Or enroll a whole control arm,
so do one or the other, but

00:14:05.091 --> 00:14:06.771
it's entirely this experiment.

00:14:06.771 --> 00:14:13.071
But the idea of using borrowing or
modeling seems to be harder to accept.

00:14:14.861 --> 00:14:16.871
Kert Viele: It's interesting,
it's almost you have to cross the

00:14:16.871 --> 00:14:21.101
Rubicon of what are you willing to
completely make this assumption or

00:14:21.101 --> 00:14:22.721
not make this assumption at all?

00:14:22.721 --> 00:14:24.851
And you can't test the assumption.

00:14:26.751 --> 00:14:27.111
Scott Berry: Right.

00:14:27.171 --> 00:14:27.501
Right.

00:14:27.651 --> 00:14:31.461
And the other extreme of this,
and I know, I know, I know this

00:14:31.461 --> 00:14:33.831
one, uh, gets at you a little

00:14:33.861 --> 00:14:34.431
Kert Viele: Don't do it.

00:14:34.431 --> 00:14:35.331
Don't do it, Scott.

00:14:35.601 --> 00:14:36.171
Scott Berry: Yeah.

00:14:36.231 --> 00:14:40.191
That we talk about using real
world evidence all the time.

00:14:41.121 --> 00:14:43.401
Digital twins or real world evidence.

00:14:43.401 --> 00:14:46.701
So for, for control, we
bring in external data.

00:14:47.316 --> 00:14:51.126
And people seem to be almost
comfortable that your control is

00:14:51.126 --> 00:14:54.246
entirely external data, that's okay.

00:14:55.026 --> 00:15:01.356
Um, but if you were to do something
like use non concurrent controls

00:15:01.356 --> 00:15:07.656
from your own experiment, that
seems to be unacceptable despite

00:15:07.656 --> 00:15:13.956
the fact that these are unbelievably
phenomenal, uh, historical controls.

00:15:14.631 --> 00:15:19.791
And better than any real world evidence
you're gonna get same protocol.

00:15:19.791 --> 00:15:23.061
You have the exact same
data, the same quality.

00:15:23.301 --> 00:15:25.431
The only difference with them is time.

00:15:25.941 --> 00:15:28.611
People seem really
uncomfortable with that.

00:15:28.611 --> 00:15:32.751
But yet, at your conference you
went to, there are probably 10 talks

00:15:32.751 --> 00:15:34.071
about using real world evidence.

00:15:34.071 --> 00:15:34.086
It.

00:15:34.531 --> 00:15:34.751
Kert Viele: Yep.

00:15:36.396 --> 00:15:41.976
The, um, I mean, it, it's, there are
different legal reasons that real world

00:15:41.976 --> 00:15:45.156
evidence is being handled in certain
ways right now compared to other.

00:15:45.831 --> 00:15:51.831
Um, but, you know, FDA is not a, um,
it's not a homogeneous organization.

00:15:51.831 --> 00:15:55.401
There are lots of people, the
academic community is a, is many

00:15:55.401 --> 00:15:57.381
people, industry is many people.

00:15:57.711 --> 00:16:02.181
Um, but I think we're gonna, if we're
talking about where we're going, one of

00:16:02.181 --> 00:16:08.961
the key aspects of this is a hierarchy
of what's more or less dangerous.

00:16:09.111 --> 00:16:12.951
And I'd actually like to see
us spending more time on.

00:16:13.746 --> 00:16:14.766
Alzheimer's trials.

00:16:14.766 --> 00:16:17.316
We've got dozens, hundreds of these.

00:16:17.586 --> 00:16:21.636
So how much do the control arms differ
from trial to trial to trial to trial?

00:16:21.966 --> 00:16:26.046
If we have hundreds of pieces of trials,
we should be able to estimate this.

00:16:26.046 --> 00:16:26.886
We're statisticians.

00:16:26.886 --> 00:16:28.626
We can estimate a variance parameter.

00:16:28.926 --> 00:16:33.036
Let's get an idea about what it is,
and that variance parameter translates

00:16:33.036 --> 00:16:34.341
into how much we should borrow.

00:16:35.586 --> 00:16:39.276
Scott Berry: Yeah, and, and you,
you mentioned a key point to that if

00:16:39.276 --> 00:16:43.386
you're enrolling some new controls,
you actually get a comparison.

00:16:44.061 --> 00:16:47.631
You get the, you get the controls in
your experiment against the previous

00:16:47.631 --> 00:16:53.721
controls, and you can judge some level
of, of similarity between those and,

00:16:53.721 --> 00:16:55.821
and what's become okay.

00:16:55.851 --> 00:16:57.711
Kert Viele: Why don't, why don't
you talk about that a little more?

00:16:57.711 --> 00:16:58.791
About just the details?

00:16:58.971 --> 00:17:00.591
'cause I think there's often a sense of.

00:17:01.911 --> 00:17:06.081
the fact that there is a hierarchy
between a single arm trial, dynamic

00:17:06.081 --> 00:17:10.401
borrowing, static borrowing, get an
idea about what those methods are

00:17:10.401 --> 00:17:12.171
and how they mitigate those risks

00:17:12.871 --> 00:17:14.161
Scott Berry: Yeah, yeah, yeah.

00:17:14.761 --> 00:17:18.751
So, so the extremes on this, and I can
run an experiment where I only include

00:17:18.751 --> 00:17:23.131
data in my new experiment, uh, one-to-one
randomized control to treatment.

00:17:23.521 --> 00:17:25.891
And I'm only gonna compare those two arms.

00:17:26.221 --> 00:17:31.981
Uh, in that I could do the exact
ex, uh, very much extreme from that,

00:17:31.986 --> 00:17:34.321
where I enroll only my treatment arm.

00:17:35.371 --> 00:17:39.121
And I, I want to know, is the
treatment arm behaving, you know.

00:17:39.681 --> 00:17:43.641
Better than, or is the treatment
benefiting these patients?

00:17:43.641 --> 00:17:46.041
I would have to figure out a
way to create a control arm.

00:17:46.041 --> 00:17:48.081
I could use entirely historical data.

00:17:49.491 --> 00:17:51.771
In that far extreme.

00:17:51.921 --> 00:17:58.161
I have no ability to know whether that
control rates I'm using is is reasonable.

00:17:58.731 --> 00:18:03.081
If suppose I do something halfway that
I enroll two to one or three to one.

00:18:04.221 --> 00:18:09.261
I use the controls in that
experiment, but I also use external

00:18:09.261 --> 00:18:11.391
controls to reinforce the control.

00:18:12.021 --> 00:18:17.091
Now I'm getting new controls
so I can judge their similarity

00:18:17.091 --> 00:18:18.921
to the external controls.

00:18:19.366 --> 00:18:23.806
And statistically model that
similarity using the external data

00:18:23.806 --> 00:18:28.726
to strength if they're similar, and
not if they're not, but I do get

00:18:28.726 --> 00:18:31.036
evidence about their similarity.

00:18:31.066 --> 00:18:35.386
And it's different than this extreme case
where I only use controls and I have no

00:18:35.386 --> 00:18:38.206
idea whether, whether they're good or not.

00:18:38.386 --> 00:18:41.086
Uh, as, as a middling ground to this.

00:18:43.121 --> 00:18:45.281
Kert Viele: And I think, you know,
that's really what's going on in the

00:18:45.281 --> 00:18:49.601
academic literature right now in terms
of trying to figure out the best ways

00:18:49.901 --> 00:18:54.521
to make that assessment, minimize
the risks, amplify the benefits.

00:18:54.821 --> 00:18:59.261
Um, we're talking now more about
covariate matching between studies.

00:18:59.261 --> 00:18:59.321
I.

00:18:59.506 --> 00:19:03.466
So not just looking at all the
historical data, but which patients

00:19:03.676 --> 00:19:06.166
best match our current enrollment.

00:19:06.616 --> 00:19:10.546
Um, none of that of course is going
to be perfect, but I think it, again,

00:19:10.546 --> 00:19:15.946
it's the question of do we generally
add information or are we taking a risk

00:19:16.246 --> 00:19:20.266
and we should be looking at a society
If we do this, these kind of trials

00:19:20.446 --> 00:19:24.526
over and over and over again, do we
get a better set of drugs on the shelf?

00:19:25.356 --> 00:19:28.806
Scott Berry: Yeah, and,
and part of this is it.

00:19:29.571 --> 00:19:33.801
Part of it is statisticians are
really smart and they can figure

00:19:33.801 --> 00:19:35.511
out how something can go wrong.

00:19:36.351 --> 00:19:39.951
That this, the, the, the historical
data's not the same as the new

00:19:39.951 --> 00:19:45.621
controls we've got, uh, beha, we've
got problems with the analysis there.

00:19:45.621 --> 00:19:48.321
We've got potential type
one error inflation.

00:19:48.321 --> 00:19:51.201
We've got potential bias in the
estimate of the treatment effect.

00:19:51.651 --> 00:19:55.011
And statisticians can figure this out
and say that, you know, this is bad.

00:19:55.731 --> 00:20:00.291
But if we have to do experiments
where we have to enroll one-to-one,

00:20:00.531 --> 00:20:02.931
that experiment is bigger.

00:20:03.021 --> 00:20:04.221
It's more patients.

00:20:04.221 --> 00:20:06.081
It's more patients on control.

00:20:06.411 --> 00:20:08.571
We take less shots on goal.

00:20:08.571 --> 00:20:10.341
We learn slower.

00:20:10.821 --> 00:20:15.411
It's, it's the whole industry is slowed.

00:20:15.651 --> 00:20:18.141
If every time we have a question.

00:20:18.786 --> 00:20:24.606
It has to be only in that experiment
as opposed to using what we know

00:20:24.606 --> 00:20:31.236
scientifically to reinforce that making
a more efficient trial we can, we can

00:20:31.266 --> 00:20:36.006
look at more treatments, we can treat
fewer patients, we can do this cheaper.

00:20:36.306 --> 00:20:42.966
So there is a huge societal, I'll call
it societal, a whole medical field issue.

00:20:43.326 --> 00:20:45.696
If every time we have a question.

00:20:46.071 --> 00:20:48.921
The only thing we can do is
look at the new experiment.

00:20:48.921 --> 00:20:50.721
It seems like there's a huge waste.

00:20:51.281 --> 00:20:53.201
Kert Viele: Well, we
basically, it's like data.

00:20:53.501 --> 00:20:56.531
We design an experiment,
data comes into existence.

00:20:56.681 --> 00:21:00.341
There is this bright shining
moment of about five minutes

00:21:00.521 --> 00:21:01.931
where we can do something with it.

00:21:02.371 --> 00:21:06.841
And then we're meant to put it aside, at
least with regard to future experiments.

00:21:06.961 --> 00:21:10.021
Obviously people are gonna reference
it in deciding treatments and so

00:21:10.021 --> 00:21:14.341
forth in the medium, but it, it's
interesting, it worries me as a

00:21:14.341 --> 00:21:18.841
statistician, I feel deep-seated
guilt actually putting data aside.

00:21:20.031 --> 00:21:20.421
Scott Berry: Yeah.

00:21:20.811 --> 00:21:21.171
Yeah.

00:21:21.531 --> 00:21:28.491
Uh, and, and, and again, it's hard, uh,
but there's, there's huge ramifications

00:21:28.491 --> 00:21:31.281
to not doing it, uh, in that ways.

00:21:31.281 --> 00:21:33.351
And I think there are ways
of doing it very well.

00:21:33.771 --> 00:21:36.771
You can be prospective
in your new experiment.

00:21:37.236 --> 00:21:40.656
The analysis plan is completely
written down, written down.

00:21:40.656 --> 00:21:41.946
It's explicit.

00:21:42.426 --> 00:21:45.876
When you present the results
you, you show the results.

00:21:45.876 --> 00:21:49.416
Just in the experiment, you
show what happens using the

00:21:49.416 --> 00:21:51.156
modeling of external data.

00:21:51.426 --> 00:21:53.916
I think we can do this really well.

00:21:54.156 --> 00:21:57.726
I think, by the way, this
is changing a little bit.

00:21:57.726 --> 00:22:03.576
We're seeing this even a somewhat similar
issue where you run basket trials.

00:22:04.476 --> 00:22:09.096
There, there's a lot of reticence
to a, you know, we have four

00:22:09.096 --> 00:22:13.986
kinds of patients and we run an
experiment of control to treatment.

00:22:14.721 --> 00:22:20.151
Of using the data in one of the
subsets of patients to help estimate

00:22:20.151 --> 00:22:24.021
another subset of patients that
gets people really bothered, and

00:22:24.021 --> 00:22:25.671
that's been done more and more.

00:22:25.671 --> 00:22:30.591
I think it's a very similar type of
issue and it's sort of striking and

00:22:30.591 --> 00:22:32.511
we're in trouble if we don't do that.

00:22:32.571 --> 00:22:35.601
But it has similar
statistical ramifications.

00:22:35.601 --> 00:22:35.661
It.

00:22:36.516 --> 00:22:40.356
Kert Viele: Well, it also, it, it
has tremendous implications towards

00:22:40.356 --> 00:22:46.506
design because if you can't borrow,
suddenly sponsors have an incredible

00:22:46.506 --> 00:22:50.136
incentive to claim that their
patient population is homogeneous

00:22:50.196 --> 00:22:53.256
so that they can pull everything
together and maximize their power.

00:22:53.901 --> 00:22:54.111
Scott Berry: Yeah.

00:22:54.486 --> 00:22:56.706
Kert Viele: you say, you know, as
soon as you say there's two different

00:22:56.706 --> 00:23:00.786
groups of people, you have to run
twice as many patients, no one's gonna

00:23:00.786 --> 00:23:02.286
wanna explore those kind of questions.

00:23:03.021 --> 00:23:08.181
Scott Berry: yeah, so it's the same
bizarre thing that happens, that we run

00:23:08.181 --> 00:23:12.711
trials that are single arm trials where
we use a historical control a hundred

00:23:12.711 --> 00:23:15.176
percent, or we don't use it at all.

00:23:16.221 --> 00:23:19.581
And you're describing exactly
what happens in trials.

00:23:19.581 --> 00:23:23.691
We run a trial and you do a common
estimate across the entire population,

00:23:23.781 --> 00:23:28.941
a single estimate, which is this,
which is full pooling, it's exact same

00:23:28.941 --> 00:23:32.811
mathematically of pooling populations
altogether to get a single answer.

00:23:34.281 --> 00:23:39.081
Or we run separate trials and we call 'em
population A and population B, but we only

00:23:39.081 --> 00:23:41.811
estimate those individually in A and B.

00:23:42.081 --> 00:23:46.251
Now we want to do something that's a
middling thing between that, and by the

00:23:46.251 --> 00:23:52.221
way, we actually get to observe in the
experiment similarity of effect that we

00:23:52.221 --> 00:23:58.251
shrink estimates statistically, uh, that
isn't full pooling and it isn't separate.

00:23:58.986 --> 00:24:01.206
And that people struggle with that.

00:24:01.266 --> 00:24:05.796
It's either you have to do all
pooling or completely separate

00:24:05.796 --> 00:24:10.656
populations, and it's just, it,
it, it strikes me as just bizarre.

00:24:12.661 --> 00:24:15.066
Kert Viele: So let me ask you about
one other thing, which I think is a

00:24:15.066 --> 00:24:17.136
real oddball in this conversation.

00:24:17.901 --> 00:24:22.611
Um, there is a situation where
we routinely combine data across

00:24:22.611 --> 00:24:28.191
different studies, and it's often
viewed as the pinnacle of evidence

00:24:28.221 --> 00:24:30.921
for clinical trials, which is a

00:24:31.026 --> 00:24:32.346
Scott Berry: I, uh, I have no idea.

00:24:32.346 --> 00:24:33.936
I have no, oh, meta-analysis.

00:24:33.936 --> 00:24:34.386
Okay.

00:24:34.386 --> 00:24:34.566
Yep.

00:24:35.151 --> 00:24:37.731
Kert Viele: where do you think
meta-analysis fits in, in terms

00:24:37.731 --> 00:24:41.571
of the, the assumptions and
how it fits into this debate?

00:24:43.581 --> 00:24:48.171
Scott Berry: Yeah, I, I, I think
meta-analyses are incredibly valuable.

00:24:48.381 --> 00:24:53.061
Um, the modeling makes a ton of
sense and incredibly important.

00:24:53.451 --> 00:24:59.091
Um, I, I think people overweight single
trials and the, the totality of evidence.

00:24:59.091 --> 00:25:01.881
Now, not every meta-analysis is every

00:25:01.881 --> 00:25:02.541
meta-analysis.

00:25:03.291 --> 00:25:07.581
Uh, there's, and, and, and I
think people, I sort of think

00:25:07.581 --> 00:25:11.811
meta-analyses are bad because there
are some that are clearly biased.

00:25:11.811 --> 00:25:17.301
There's publication bias, there's,
there's availability of data bias.

00:25:17.361 --> 00:25:18.231
Um, uh, so I.

00:25:19.251 --> 00:25:25.281
Not every meta-analysis is bringing
quality data that's free of these

00:25:25.281 --> 00:25:30.591
issues, but there are many circumstances
where we have data on essentially every

00:25:30.591 --> 00:25:32.721
patient that's been experimented on.

00:25:33.231 --> 00:25:37.461
And combining this together
I think is hugely valuable.

00:25:37.911 --> 00:25:42.021
I By, by the way, I think the
idea that we run two separate

00:25:42.021 --> 00:25:47.061
independent phase three trials and
we analyze them separately is stupid.

00:25:47.526 --> 00:25:52.026
I, I, I mean, it's stupid that
we would analyze them completely

00:25:52.026 --> 00:25:54.786
separately, and we need both to be 0.05.

00:25:55.086 --> 00:25:58.746
Combining them together into
a single inference is better

00:25:59.586 --> 00:26:00.936
in every way.

00:26:01.596 --> 00:26:04.386
Kert Viele: is two is two 0.49s

00:26:04.656 --> 00:26:06.366
as good as a 0.051

00:26:06.366 --> 00:26:07.866
and a 0.001

00:26:07.956 --> 00:26:08.976
on the P-values.

00:26:09.306 --> 00:26:11.226
Clearly you'd want the latter.

00:26:12.546 --> 00:26:12.966
Scott Berry: Yeah.

00:26:13.056 --> 00:26:13.266
Yeah.

00:26:13.271 --> 00:26:16.686
And, and, and so I think of a
meta-analysis the same way I, I,

00:26:16.776 --> 00:26:21.246
I would do a meta-analysis of my
two phase three trials for what,

00:26:21.246 --> 00:26:23.526
what does it tell me to the answer?

00:26:23.526 --> 00:26:27.456
So it, you know, those techniques,
they tend to be Bayesian.

00:26:27.456 --> 00:26:30.111
But, uh, largely the availability of data.

00:26:30.186 --> 00:26:33.696
And there are some where we know there's
missing data, we know there's issues

00:26:33.696 --> 00:26:36.816
with it, and we would not take those at.

00:26:37.176 --> 00:26:42.216
Uh, you know, at, at a high
scientific degree to that.

00:26:42.366 --> 00:26:45.696
Um, and that's where the
science parts comes into it.

00:26:45.696 --> 00:26:50.676
And by the way, these are abused where
people present data, it's bias data.

00:26:50.856 --> 00:26:54.276
They're ignoring huge amounts
and they call it a meta-analysis.

00:26:54.276 --> 00:26:55.956
And we think that's bad science.

00:26:56.226 --> 00:26:59.436
So we've got to, at some
level, be able to judge.

00:27:00.246 --> 00:27:05.286
Good science, quality data,
meta-analysis, which I think are more

00:27:05.286 --> 00:27:08.766
valuable than individual trials there.

00:27:09.216 --> 00:27:14.196
And then there's bad science put
together that we, we, we recognize

00:27:14.196 --> 00:27:16.566
we wouldn't use that information.

00:27:17.136 --> 00:27:20.196
Kert Viele: We would certainly say
that on any kind of borrowing, bringing

00:27:20.196 --> 00:27:23.526
together data that doesn't belong is bad.

00:27:23.556 --> 00:27:26.826
You can borrow in a way that
does generate horrible biases and

00:27:26.826 --> 00:27:28.626
will generate bad conclusions.

00:27:29.166 --> 00:27:33.936
Uh, but we also are talking about
what can be done well with experience.

00:27:34.356 --> 00:27:35.676
Um, we're running outta time.

00:27:35.676 --> 00:27:36.876
Why don't you close us out?

00:27:38.181 --> 00:27:42.651
Scott Berry: Uh, so, uh, I, I think
this by the way, is the future.

00:27:42.651 --> 00:27:45.861
I think as we learn more and
more about diseases, we get

00:27:45.861 --> 00:27:47.631
more and more subsets of them.

00:27:48.006 --> 00:27:53.046
Um, we are getting more and more quality
data available from other clinical trials.

00:27:53.256 --> 00:27:57.126
The sharing of data from clinical
trials, the sharing of other

00:27:57.126 --> 00:28:01.206
types of resources, we're gonna
have explosion with medical data.

00:28:01.506 --> 00:28:06.246
The idea that we don't use that
in our new experiments, if we

00:28:06.246 --> 00:28:09.156
don't use that in valuable ways,
we're running worse experiments.

00:28:09.576 --> 00:28:13.326
So a wonderful topic, and I
think it's certainly the future.

00:28:13.876 --> 00:28:16.366
Uh, and so here we are in the interim.

00:28:16.366 --> 00:28:19.246
So, uh, till next time, appreciate it.

00:28:19.336 --> 00:28:19.726
Thank you,

00:28:19.726 --> 00:28:20.026
Kurt.