WEBVTT

00:00:00.000 --> 00:00:04.428 align:middle line:84%
[SQUEAKING]
[RUSTLING] [CLICKING]

00:00:04.428 --> 00:00:09.820 align:middle line:90%


00:00:09.820 --> 00:00:11.820 align:middle line:84%
SCOTT KEMP: All right,
why don't we get started?

00:00:11.820 --> 00:00:17.720 align:middle line:84%
So today, we're going to talk
about the linear no-threshold

00:00:17.720 --> 00:00:18.720 align:middle line:90%
model.

00:00:18.720 --> 00:00:23.120 align:middle line:84%
And some of you, probably most
of you, have heard about it.

00:00:23.120 --> 00:00:25.720 align:middle line:84%
In the last class
on Wednesday, we

00:00:25.720 --> 00:00:29.360 align:middle line:84%
went through this, the very
basics of radiation and dose.

00:00:29.360 --> 00:00:32.400 align:middle line:84%
And we came up to
the point where

00:00:32.400 --> 00:00:35.900 align:middle line:84%
we talked about these
units of sieverts or REM,

00:00:35.900 --> 00:00:38.440 align:middle line:84%
which are so-called
equivalent dose.

00:00:38.440 --> 00:00:41.560 align:middle line:84%
They have units of
energy per unit mass,

00:00:41.560 --> 00:00:44.320 align:middle line:84%
but they don't really
mean that because they've

00:00:44.320 --> 00:00:48.960 align:middle line:84%
been multiplied by these
unitless adjustment factors.

00:00:48.960 --> 00:00:50.880 align:middle line:84%
But the problem is
that we don't really

00:00:50.880 --> 00:00:53.160 align:middle line:90%
know what to do with this unit.

00:00:53.160 --> 00:00:57.440 align:middle line:84%
It doesn't tell us anything
about how much cancer that

00:00:57.440 --> 00:00:58.800 align:middle line:90%
gives us.

00:00:58.800 --> 00:01:01.650 align:middle line:84%
It just tells us how
much energy is deposited,

00:01:01.650 --> 00:01:05.030 align:middle line:84%
but the amount of energy that's
actually deposited by radiation

00:01:05.030 --> 00:01:06.810 align:middle line:84%
is trivially small
compared to, say,

00:01:06.810 --> 00:01:10.430 align:middle line:84%
taking a bath in warm water,
which deposits a lot more energy

00:01:10.430 --> 00:01:12.150 align:middle line:90%
into your body.

00:01:12.150 --> 00:01:14.270 align:middle line:84%
And it has to do
with, as we mentioned,

00:01:14.270 --> 00:01:17.870 align:middle line:84%
how that interaction happens,
where the energy is going.

00:01:17.870 --> 00:01:23.310 align:middle line:84%
So we need some way of
converting this otherwise

00:01:23.310 --> 00:01:27.590 align:middle line:84%
useless unit into something that
tells us about health effects.

00:01:27.590 --> 00:01:31.030 align:middle line:84%
And so that is the so-called
dose-response model.

00:01:31.030 --> 00:01:35.430 align:middle line:84%
Now there are lots of different
kinds of health effects

00:01:35.430 --> 00:01:38.870 align:middle line:84%
that we could build
dose-response models for.

00:01:38.870 --> 00:01:43.430 align:middle line:84%
So we typically break these down
into two different categories--

00:01:43.430 --> 00:01:46.270 align:middle line:84%
what are called the
deterministic diseases

00:01:46.270 --> 00:01:48.390 align:middle line:90%
and the stochastic diseases.

00:01:48.390 --> 00:01:52.150 align:middle line:84%
So deterministic diseases
have the property

00:01:52.150 --> 00:01:54.710 align:middle line:84%
that the intensity
of the illness

00:01:54.710 --> 00:01:59.140 align:middle line:84%
is directly proportional
to the amount of radiation

00:01:59.140 --> 00:02:01.020 align:middle line:90%
that you get.

00:02:01.020 --> 00:02:04.900 align:middle line:84%
And sometimes there's
a threshold effect that

00:02:04.900 --> 00:02:07.260 align:middle line:90%
might exist in those diseases.

00:02:07.260 --> 00:02:13.140 align:middle line:84%
So the classic examples here
are like skin burns, cataracts,

00:02:13.140 --> 00:02:16.900 align:middle line:84%
hematological disorders
of various kinds.

00:02:16.900 --> 00:02:20.260 align:middle line:84%
And it's just basically,
if you put enough energy

00:02:20.260 --> 00:02:22.980 align:middle line:84%
onto the surface of your
skin from radiation,

00:02:22.980 --> 00:02:26.033 align:middle line:84%
more than capillary dilation
is able to take away.

00:02:26.033 --> 00:02:28.200 align:middle line:84%
So what happens is when you
put energy on your skin,

00:02:28.200 --> 00:02:29.617 align:middle line:84%
your blood vessels
instantaneously

00:02:29.617 --> 00:02:32.100 align:middle line:84%
expand to carry the
extra heat away.

00:02:32.100 --> 00:02:35.220 align:middle line:84%
If you can't do that,
then you get a burn.

00:02:35.220 --> 00:02:41.040 align:middle line:84%
And that's going to be worse if
the energy goes up in value--

00:02:41.040 --> 00:02:42.780 align:middle line:90%
very straightforward.

00:02:42.780 --> 00:02:45.860 align:middle line:84%
Stochastic diseases
are different.

00:02:45.860 --> 00:02:48.020 align:middle line:84%
The intensity of
the disease is not

00:02:48.020 --> 00:02:51.380 align:middle line:84%
proportional to the amount
of radiation you get.

00:02:51.380 --> 00:02:53.540 align:middle line:84%
The probability of
getting the disease

00:02:53.540 --> 00:02:56.200 align:middle line:84%
is what is proportional to the
amount of radiation you get.

00:02:56.200 --> 00:02:59.370 align:middle line:84%
And so there are also
lots of different things

00:02:59.370 --> 00:03:00.622 align:middle line:90%
in the stochastic category.

00:03:00.622 --> 00:03:02.330 align:middle line:84%
But the big one that
we always talk about

00:03:02.330 --> 00:03:05.650 align:middle line:90%
is cancer, genomic mutations.

00:03:05.650 --> 00:03:07.610 align:middle line:84%
And some genomic
mutations do things

00:03:07.610 --> 00:03:09.650 align:middle line:90%
other than causing cancer.

00:03:09.650 --> 00:03:11.930 align:middle line:84%
So for example, you
can have mutations

00:03:11.930 --> 00:03:14.130 align:middle line:84%
that cause all kinds
of classical disease

00:03:14.130 --> 00:03:19.330 align:middle line:84%
by either leading to
maybe a single nucleotide

00:03:19.330 --> 00:03:24.210 align:middle line:84%
variant in some protein that
then is miscoded and then

00:03:24.210 --> 00:03:25.910 align:middle line:84%
doesn't operate
correctly in your body,

00:03:25.910 --> 00:03:29.370 align:middle line:84%
and you get some kind
of functional disorder.

00:03:29.370 --> 00:03:35.410 align:middle line:84%
But the one that kills us
if we are otherwise healthy

00:03:35.410 --> 00:03:37.230 align:middle line:84%
and then kills us in
mid-life is cancer.

00:03:37.230 --> 00:03:39.430 align:middle line:84%
And so that's the one
that we tend to focus on.

00:03:39.430 --> 00:03:41.770 align:middle line:84%
So we're going to just
talk about cancer only.

00:03:41.770 --> 00:03:44.010 align:middle line:84%
But just to remind you that
there are lots of things

00:03:44.010 --> 00:03:46.090 align:middle line:90%
that can happen.

00:03:46.090 --> 00:03:46.590 align:middle line:90%
All right.

00:03:46.590 --> 00:03:51.890 align:middle line:84%
So how do we construct
this dose-response model?

00:03:51.890 --> 00:03:55.880 align:middle line:84%
So there's some amount
of dose, and then there's

00:03:55.880 --> 00:04:00.320 align:middle line:84%
this black box, which
is a blue box, which

00:04:00.320 --> 00:04:02.540 align:middle line:84%
has some mysterious
biological process.

00:04:02.540 --> 00:04:05.240 align:middle line:90%
And out pops some cancer rate.

00:04:05.240 --> 00:04:08.040 align:middle line:84%
And we need to understand
what's in that blue box.

00:04:08.040 --> 00:04:11.180 align:middle line:84%
So one approach is that
we say, you know what?

00:04:11.180 --> 00:04:13.640 align:middle line:84%
We don't bother
with what's inside.

00:04:13.640 --> 00:04:17.079 align:middle line:84%
We'll just build a
statistical model

00:04:17.079 --> 00:04:20.640 align:middle line:84%
that seems to describe
what we observe

00:04:20.640 --> 00:04:22.840 align:middle line:90%
happening on the outside.

00:04:22.840 --> 00:04:25.120 align:middle line:84%
You put a certain
amount of radiation in,

00:04:25.120 --> 00:04:28.280 align:middle line:84%
and maybe we get some cancer
rate out, some estimate

00:04:28.280 --> 00:04:29.480 align:middle line:90%
of the cancer rate.

00:04:29.480 --> 00:04:32.080 align:middle line:84%
So we'll talk a lot about
this statistical model.

00:04:32.080 --> 00:04:35.560 align:middle line:84%
The prevailing statistical
model is called

00:04:35.560 --> 00:04:38.160 align:middle line:90%
the linear no-threshold model.

00:04:38.160 --> 00:04:41.800 align:middle line:84%
And we'll get into
that a little bit more.

00:04:41.800 --> 00:04:43.520 align:middle line:84%
The problem with
statistical models

00:04:43.520 --> 00:04:46.680 align:middle line:84%
is a good model requires
good statistics.

00:04:46.680 --> 00:04:50.060 align:middle line:84%
And it's actually rather
hard to get good statistics.

00:04:50.060 --> 00:04:54.510 align:middle line:84%
And I'll show you why in class
a little bit later today.

00:04:54.510 --> 00:04:58.390 align:middle line:84%
The other way we could
try to build a model

00:04:58.390 --> 00:05:00.030 align:middle line:90%
would be mechanistically.

00:05:00.030 --> 00:05:02.350 align:middle line:84%
We say, OK, well, there's
some Compton scattering,

00:05:02.350 --> 00:05:03.770 align:middle line:84%
and there's
photoelectric effect.

00:05:03.770 --> 00:05:06.103 align:middle line:84%
And then this is going to do
this to the genes, and then

00:05:06.103 --> 00:05:07.570 align:middle line:90%
this, and et cetera, et cetera.

00:05:07.570 --> 00:05:09.710 align:middle line:84%
But you can see
what the problem is.

00:05:09.710 --> 00:05:13.030 align:middle line:84%
We don't understand
how the body works.

00:05:13.030 --> 00:05:16.010 align:middle line:90%
We have some understanding.

00:05:16.010 --> 00:05:17.970 align:middle line:84%
When you get down to
the biochemistry of it,

00:05:17.970 --> 00:05:20.670 align:middle line:84%
there's a lot we don't
really understand.

00:05:20.670 --> 00:05:25.550 align:middle line:84%
And so this model is not a model
that we can actually build.

00:05:25.550 --> 00:05:29.930 align:middle line:84%
So we are stuck with
the statistical model.

00:05:29.930 --> 00:05:34.710 align:middle line:84%
But we still have insights
from the mechanistic world

00:05:34.710 --> 00:05:39.110 align:middle line:84%
that we can use to bound how we
build these statistical models.

00:05:39.110 --> 00:05:40.910 align:middle line:90%
And that's what we want to do.

00:05:40.910 --> 00:05:44.750 align:middle line:84%
So my goal today is to
walk you through how

00:05:44.750 --> 00:05:49.630 align:middle line:84%
we got to the statistical model
we have and go from there.

00:05:49.630 --> 00:05:50.170 align:middle line:90%
All right.

00:05:50.170 --> 00:05:52.650 align:middle line:84%
So if we want to
build such a model,

00:05:52.650 --> 00:05:57.710 align:middle line:84%
we need first to decide
what our metrics are.

00:05:57.710 --> 00:05:58.650 align:middle line:90%
What is our input?

00:05:58.650 --> 00:05:59.770 align:middle line:90%
What is our output?

00:05:59.770 --> 00:06:02.130 align:middle line:90%
Is the input going to be dose?

00:06:02.130 --> 00:06:07.738 align:middle line:84%
Is it going to be
the effective dose?

00:06:07.738 --> 00:06:08.530 align:middle line:90%
What is the output?

00:06:08.530 --> 00:06:11.490 align:middle line:84%
Is it going to be the
cancer rate, et cetera?

00:06:11.490 --> 00:06:14.950 align:middle line:90%
So here's what we have chosen.

00:06:14.950 --> 00:06:18.990 align:middle line:90%
We have these definitions.

00:06:18.990 --> 00:06:20.770 align:middle line:90%
So the first is what is risk?

00:06:20.770 --> 00:06:23.190 align:middle line:84%
It's the baseline
probability that a person

00:06:23.190 --> 00:06:25.490 align:middle line:84%
in a defined population
has the disease.

00:06:25.490 --> 00:06:27.950 align:middle line:84%
So you have a baseline
risk of cancer, which

00:06:27.950 --> 00:06:30.190 align:middle line:90%
turns out to be about 20%.

00:06:30.190 --> 00:06:33.510 align:middle line:84%
It's about one in five of
you will probably get cancer.

00:06:33.510 --> 00:06:36.190 align:middle line:84%
You have a relative
risk, which is

00:06:36.190 --> 00:06:38.230 align:middle line:84%
to say the exposed,
in this case,

00:06:38.230 --> 00:06:40.750 align:middle line:84%
to radiation, the exposed
population divided

00:06:40.750 --> 00:06:43.390 align:middle line:84%
by the risk of the
unexposed, or control group.

00:06:43.390 --> 00:06:45.110 align:middle line:90%
This is a function of dose.

00:06:45.110 --> 00:06:49.950 align:middle line:84%
So here is our first
mechanistic choice.

00:06:49.950 --> 00:06:52.680 align:middle line:84%
Just in writing this
definition, it's hidden.

00:06:52.680 --> 00:06:54.940 align:middle line:84%
But there is a choice
about how cancer

00:06:54.940 --> 00:07:00.700 align:middle line:90%
works in writing relative risk.

00:07:00.700 --> 00:07:04.100 align:middle line:84%
What this means is if
a person is a smoker

00:07:04.100 --> 00:07:07.340 align:middle line:84%
and they have a higher
baseline risk then

00:07:07.340 --> 00:07:13.660 align:middle line:84%
for a given unit of radiation,
their relative risk is the same,

00:07:13.660 --> 00:07:17.740 align:middle line:84%
but their absolute
risk is now higher.

00:07:17.740 --> 00:07:21.460 align:middle line:84%
A unit of radiation for a
person who has a higher baseline

00:07:21.460 --> 00:07:26.240 align:middle line:84%
risk of cancer causes a higher
probability of getting cancer.

00:07:26.240 --> 00:07:28.700 align:middle line:90%
Does that make sense?

00:07:28.700 --> 00:07:31.680 align:middle line:84%
Because it's relative
to your baseline risk.

00:07:31.680 --> 00:07:33.980 align:middle line:84%
And if your baseline
risk is higher,

00:07:33.980 --> 00:07:37.140 align:middle line:84%
then your absolute
probability is higher.

00:07:37.140 --> 00:07:40.060 align:middle line:84%
And so that is a choice
that we are making

00:07:40.060 --> 00:07:42.660 align:middle line:90%
when we build this model.

00:07:42.660 --> 00:07:45.580 align:middle line:84%
And it's not one that we
actually talk about a lot.

00:07:45.580 --> 00:07:48.420 align:middle line:84%
Now strictly speaking, we
could have various controls

00:07:48.420 --> 00:07:51.530 align:middle line:84%
to interact with different
kinds of baseline risks

00:07:51.530 --> 00:07:53.090 align:middle line:90%
and try to control for it.

00:07:53.090 --> 00:07:55.870 align:middle line:84%
We should ask the question,
what is the truth of the matter?

00:07:55.870 --> 00:07:57.110 align:middle line:90%
Has someone looked into this?

00:07:57.110 --> 00:07:58.027 align:middle line:90%
And the answer is yes.

00:07:58.027 --> 00:07:59.170 align:middle line:90%
People look into this.

00:07:59.170 --> 00:08:03.250 align:middle line:84%
And as best as they can
tell, this is about correct.

00:08:03.250 --> 00:08:07.650 align:middle line:84%
So for example, a famous
one is radon exposure,

00:08:07.650 --> 00:08:10.170 align:middle line:84%
which is an inhaled
form of radiation,

00:08:10.170 --> 00:08:16.690 align:middle line:84%
and smoking, another
inhaled kind of carcinogen.

00:08:16.690 --> 00:08:18.630 align:middle line:90%
Does of relationship hold?

00:08:18.630 --> 00:08:22.010 align:middle line:84%
The answer is for light
smokers, it turns out

00:08:22.010 --> 00:08:25.910 align:middle line:90%
to be submultiplicative.

00:08:25.910 --> 00:08:28.590 align:middle line:84%
Sorry, light smokers,
it's super multiplicative.

00:08:28.590 --> 00:08:32.049 align:middle line:84%
And for heavy smokers,
it's submultiplicative.

00:08:32.049 --> 00:08:36.490 align:middle line:90%
So it's about multiplicative.

00:08:36.490 --> 00:08:41.049 align:middle line:84%
Another classic one is arsenic
exposure and radon in miners.

00:08:41.049 --> 00:08:44.370 align:middle line:84%
And that one is also found to
be basically multiplicative.

00:08:44.370 --> 00:08:49.300 align:middle line:84%
So this definition is basically
supported by the evidence,

00:08:49.300 --> 00:08:52.920 align:middle line:84%
but there are places where the
interactions are not strictly

00:08:52.920 --> 00:08:54.080 align:middle line:90%
multiplicative.

00:08:54.080 --> 00:08:56.080 align:middle line:84%
This is just a
simplification we have

00:08:56.080 --> 00:09:01.240 align:middle line:84%
to make in order to figure
out how to build the model.

00:09:01.240 --> 00:09:03.360 align:middle line:84%
Then we have this
definition, which we will

00:09:03.360 --> 00:09:05.300 align:middle line:90%
tend to use when presenting.

00:09:05.300 --> 00:09:08.740 align:middle line:84%
We typically take the relative
risk, and we subtract 1.

00:09:08.740 --> 00:09:11.080 align:middle line:84%
Or we take one and
subtract, subtract

00:09:11.080 --> 00:09:14.640 align:middle line:84%
it to get what we call
the excess relative risk.

00:09:14.640 --> 00:09:18.960 align:middle line:84%
So if you have a 20%
chance of getting cancer,

00:09:18.960 --> 00:09:23.440 align:middle line:84%
and then a 21% chance
of getting cancer,

00:09:23.440 --> 00:09:25.420 align:middle line:84%
if you also are
exposed to radiation,

00:09:25.420 --> 00:09:29.880 align:middle line:84%
then the excess relative
risk is just that extra 1%.

00:09:29.880 --> 00:09:34.983 align:middle line:84%
So let's look at some
dose responses that

00:09:34.983 --> 00:09:35.900 align:middle line:90%
are in the literature.

00:09:35.900 --> 00:09:38.360 align:middle line:84%
So here's the excess relative
risk of a solid cancer.

00:09:38.360 --> 00:09:40.660 align:middle line:84%
A solid cancer
means not leukemia.

00:09:40.660 --> 00:09:43.800 align:middle line:90%


00:09:43.800 --> 00:09:48.210 align:middle line:84%
By the dose to colon-- the
colon is, as it turns out,

00:09:48.210 --> 00:09:49.930 align:middle line:90%
a standard organ.

00:09:49.930 --> 00:09:51.970 align:middle line:84%
Every organ has a slightly
different response,

00:09:51.970 --> 00:09:55.230 align:middle line:84%
and people tend to use the
colon for whatever reasons

00:09:55.230 --> 00:09:56.530 align:middle line:90%
as the standard.

00:09:56.530 --> 00:09:59.390 align:middle line:84%
And so what they typically do
is they adjust all the exposures

00:09:59.390 --> 00:10:02.430 align:middle line:90%
to colon-weighted exposure.

00:10:02.430 --> 00:10:05.550 align:middle line:84%
And this is for the
Japanese nuclear bomb

00:10:05.550 --> 00:10:06.730 align:middle line:90%
survivors in Hiroshima.

00:10:06.730 --> 00:10:09.150 align:middle line:84%
It's called the
life span survey.

00:10:09.150 --> 00:10:13.190 align:middle line:84%
And you see the
data here in blue.

00:10:13.190 --> 00:10:17.870 align:middle line:84%
And they've binned it into
different neutron-rated doses.

00:10:17.870 --> 00:10:21.710 align:middle line:90%
And there's a model there.

00:10:21.710 --> 00:10:27.650 align:middle line:84%
So this model is a linear
no-threshold model in red.

00:10:27.650 --> 00:10:30.630 align:middle line:84%
The black are the error
bars of the model.

00:10:30.630 --> 00:10:33.330 align:middle line:84%
And let's just look
at what it means.

00:10:33.330 --> 00:10:35.510 align:middle line:90%
Linear means, well, it's linear.

00:10:35.510 --> 00:10:40.590 align:middle line:84%
No threshold means that
the intercept is at 0, 0.

00:10:40.590 --> 00:10:45.700 align:middle line:84%
There is no threshold at which
radiation produces zero effect.

00:10:45.700 --> 00:10:47.580 align:middle line:90%
That's what that simply means.

00:10:47.580 --> 00:10:48.840 align:middle line:90%
And what do we see?

00:10:48.840 --> 00:10:52.420 align:middle line:84%
It's pretty clear it
doesn't fit the data.

00:10:52.420 --> 00:10:56.300 align:middle line:84%
The data rolls
over at high doses.

00:10:56.300 --> 00:10:59.520 align:middle line:84%
So there's no actual
controversy here because no,

00:10:59.520 --> 00:11:03.500 align:middle line:84%
if you are, in fact, looking
at a population of people

00:11:03.500 --> 00:11:07.420 align:middle line:84%
who are getting doses
of 3 and 4 gray,

00:11:07.420 --> 00:11:09.340 align:middle line:84%
which is equivalent to
basically three or four

00:11:09.340 --> 00:11:13.180 align:middle line:84%
sieverts, something
like that, typically.

00:11:13.180 --> 00:11:15.300 align:middle line:84%
Well, this is
neutron-weighted, so that

00:11:15.300 --> 00:11:16.800 align:middle line:90%
wouldn't be exactly equivalent.

00:11:16.800 --> 00:11:21.420 align:middle line:84%
But anyway, if you're getting
people at these high doses,

00:11:21.420 --> 00:11:22.880 align:middle line:90%
you don't use the linear model.

00:11:22.880 --> 00:11:25.460 align:middle line:84%
You use one of these
nonlinear models, which

00:11:25.460 --> 00:11:27.220 align:middle line:90%
we do have in the arsenal.

00:11:27.220 --> 00:11:32.140 align:middle line:84%
Where we use the linear models,
only really down here below 1.

00:11:32.140 --> 00:11:35.340 align:middle line:84%
And the reason why
it's controversial

00:11:35.340 --> 00:11:39.740 align:middle line:84%
is because when you look
at nuclear accident events,

00:11:39.740 --> 00:11:41.800 align:middle line:84%
this is in fact where
most people lined up.

00:11:41.800 --> 00:11:45.090 align:middle line:84%
Most people get almost
no dose from an accident.

00:11:45.090 --> 00:11:47.150 align:middle line:84%
And they're right down
here in the linear regime.

00:11:47.150 --> 00:11:49.770 align:middle line:84%
And the question is, what
exactly is the shape down here?

00:11:49.770 --> 00:11:53.450 align:middle line:84%
And it looks like it
kind of is linear.

00:11:53.450 --> 00:11:56.250 align:middle line:84%
But we really
would like to know.

00:11:56.250 --> 00:12:00.210 align:middle line:90%
So that's the story.

00:12:00.210 --> 00:12:02.930 align:middle line:90%
So let's look at some more data.

00:12:02.930 --> 00:12:07.830 align:middle line:84%
This is DNA mutations in thyroid
tumors of Chernobyl survivors.

00:12:07.830 --> 00:12:10.610 align:middle line:90%


00:12:10.610 --> 00:12:13.690 align:middle line:84%
The maximum recorded dose
here is 1000 milligray,

00:12:13.690 --> 00:12:14.590 align:middle line:90%
which is one gray.

00:12:14.590 --> 00:12:16.950 align:middle line:84%
So we are completely
in the low-dose regime.

00:12:16.950 --> 00:12:26.415 align:middle line:84%
And at first, let's just
look at what we have here.

00:12:26.415 --> 00:12:28.290 align:middle line:84%
The first thing to notice
is there's actually

00:12:28.290 --> 00:12:31.490 align:middle line:84%
considerable background
population at 0 dose.

00:12:31.490 --> 00:12:34.470 align:middle line:84%
This is the distribution
of observations.

00:12:34.470 --> 00:12:37.410 align:middle line:84%
And you see that on
average is something

00:12:37.410 --> 00:12:44.640 align:middle line:84%
like 25 small deletions in
this genome, even at no dose.

00:12:44.640 --> 00:12:48.000 align:middle line:84%
We can draw a
straight line, and we

00:12:48.000 --> 00:12:51.980 align:middle line:84%
can call this the baseline
risk of a mutation.

00:12:51.980 --> 00:12:54.340 align:middle line:84%
This is like the baseline
risk of cancer, if you will.

00:12:54.340 --> 00:12:56.520 align:middle line:90%
It's just a different measure.

00:12:56.520 --> 00:13:02.120 align:middle line:84%
And then we can look at what
the effect is after that.

00:13:02.120 --> 00:13:05.160 align:middle line:84%
The other thing you can observe
in this data is there seems

00:13:05.160 --> 00:13:08.080 align:middle line:84%
to be a lot of people
clustered here at 1,000.

00:13:08.080 --> 00:13:11.320 align:middle line:84%
So clearly, something happened
in how they were doing the dose

00:13:11.320 --> 00:13:12.060 align:middle line:90%
reconstruction.

00:13:12.060 --> 00:13:13.977 align:middle line:84%
They're basically saying
oh, anyone over here,

00:13:13.977 --> 00:13:15.880 align:middle line:84%
we're just going to
put them into this bin,

00:13:15.880 --> 00:13:18.060 align:middle line:84%
which now we have to
be careful of that bin

00:13:18.060 --> 00:13:23.600 align:middle line:84%
and how that affects
the measurements.

00:13:23.600 --> 00:13:26.640 align:middle line:84%
So there are some systematic
errors, which are a problem.

00:13:26.640 --> 00:13:28.660 align:middle line:84%
And indeed, if you
look carefully,

00:13:28.660 --> 00:13:30.520 align:middle line:84%
you'll see that
there appears to be

00:13:30.520 --> 00:13:33.160 align:middle line:84%
a discontinuity
between the zeroth bin

00:13:33.160 --> 00:13:35.880 align:middle line:84%
and the first bin, at
which there's actually dose

00:13:35.880 --> 00:13:37.760 align:middle line:90%
in terms of the distribution.

00:13:37.760 --> 00:13:40.030 align:middle line:84%
So there's something else
going on there as well

00:13:40.030 --> 00:13:41.810 align:middle line:90%
between the two control groups.

00:13:41.810 --> 00:13:44.550 align:middle line:90%
So lots of issues.

00:13:44.550 --> 00:13:47.110 align:middle line:84%
Another issue is
that you might just

00:13:47.110 --> 00:13:50.470 align:middle line:84%
say, oh, I'm going to throw
this into a regression

00:13:50.470 --> 00:13:52.670 align:middle line:90%
to fit that line.

00:13:52.670 --> 00:13:53.590 align:middle line:90%
Seems reasonable.

00:13:53.590 --> 00:13:55.110 align:middle line:84%
I would just throw
this into Excel,

00:13:55.110 --> 00:13:58.110 align:middle line:84%
and it would perform
a least squares fit.

00:13:58.110 --> 00:14:00.230 align:middle line:84%
Would that be a
good choice here?

00:14:00.230 --> 00:14:02.110 align:middle line:90%
Why?

00:14:02.110 --> 00:14:04.295 align:middle line:84%
AUDIENCE: The data
is really wide.

00:14:04.295 --> 00:14:05.670 align:middle line:84%
SCOTT KEMP: That
in and of itself

00:14:05.670 --> 00:14:07.850 align:middle line:84%
is not a problem
doing least squares.

00:14:07.850 --> 00:14:11.630 align:middle line:90%


00:14:11.630 --> 00:14:15.350 align:middle line:84%
A least squares fit is a
maximum likelihood estimate

00:14:15.350 --> 00:14:21.050 align:middle line:84%
only if the underlying data
have a Gaussian distribution.

00:14:21.050 --> 00:14:22.550 align:middle line:84%
Gaussian distribution
has tails that

00:14:22.550 --> 00:14:24.910 align:middle line:84%
go off to infinity
in both directions,

00:14:24.910 --> 00:14:28.310 align:middle line:84%
but you can't have fewer
than zero deletions.

00:14:28.310 --> 00:14:30.610 align:middle line:84%
So this is not
Gaussian-distributed.

00:14:30.610 --> 00:14:34.690 align:middle line:84%
And if you now fit something
using least squares,

00:14:34.690 --> 00:14:36.463 align:middle line:84%
you're not going to
get the correct fit.

00:14:36.463 --> 00:14:38.380 align:middle line:84%
You're not going to get
the maximum likelihood

00:14:38.380 --> 00:14:40.560 align:middle line:90%
estimation of the data.

00:14:40.560 --> 00:14:45.740 align:middle line:84%
So this, turns out, is more
of a binomial distribution.

00:14:45.740 --> 00:14:49.780 align:middle line:84%
And so you would need to use an
appropriate regression model.

00:14:49.780 --> 00:14:52.940 align:middle line:84%
So all this is to say that
this is not a tidy problem.

00:14:52.940 --> 00:14:54.700 align:middle line:84%
And there's lots of
places where people

00:14:54.700 --> 00:14:59.500 align:middle line:84%
can go wrong in these types of
analyzing these types of things.

00:14:59.500 --> 00:15:02.020 align:middle line:84%
So should we take
this fit, which

00:15:02.020 --> 00:15:04.260 align:middle line:84%
is the one provided to
us by the authors, which

00:15:04.260 --> 00:15:05.820 align:middle line:90%
is a linear model?

00:15:05.820 --> 00:15:07.700 align:middle line:90%
Well, it looks all right.

00:15:07.700 --> 00:15:10.300 align:middle line:84%
It does look like there's
some low region here and maybe

00:15:10.300 --> 00:15:12.500 align:middle line:90%
some higher region over here.

00:15:12.500 --> 00:15:15.780 align:middle line:90%
Maybe we should take this.

00:15:15.780 --> 00:15:19.760 align:middle line:84%
Maybe we should do put some
lowness, maybe make that green.

00:15:19.760 --> 00:15:21.860 align:middle line:84%
Maybe that green one is
more convincing to us,

00:15:21.860 --> 00:15:25.140 align:middle line:84%
looks like actually
fits the data better.

00:15:25.140 --> 00:15:28.280 align:middle line:90%
Or maybe we should do that.

00:15:28.280 --> 00:15:30.700 align:middle line:84%
maybe that actually
fits the data better.

00:15:30.700 --> 00:15:31.380 align:middle line:90%
I don't know.

00:15:31.380 --> 00:15:35.940 align:middle line:84%
I mean, the data do
definitely seem to go up,

00:15:35.940 --> 00:15:37.670 align:middle line:84%
and then they definitely
seem to go down.

00:15:37.670 --> 00:15:40.230 align:middle line:84%
And then there's this big
outlier point over here.

00:15:40.230 --> 00:15:42.410 align:middle line:90%
And what do we do?

00:15:42.410 --> 00:15:43.710 align:middle line:90%
How do we choose these models?

00:15:43.710 --> 00:15:50.590 align:middle line:84%
So here's where we're going
to have to make a choice.

00:15:50.590 --> 00:15:57.890 align:middle line:84%
And this is the practice
of model selection.

00:15:57.890 --> 00:16:00.770 align:middle line:84%
And I don't think
people really--

00:16:00.770 --> 00:16:03.845 align:middle line:84%
people know about overfitting
data as a problem.

00:16:03.845 --> 00:16:05.970 align:middle line:84%
But I don't think people
are really given the tools

00:16:05.970 --> 00:16:08.210 align:middle line:90%
to talk about model selection.

00:16:08.210 --> 00:16:17.130 align:middle line:84%
So as a general rule, we
want to choose simple models.

00:16:17.130 --> 00:16:19.610 align:middle line:84%
But there's a good
reason behind why we

00:16:19.610 --> 00:16:21.850 align:middle line:90%
want to choose simple models.

00:16:21.850 --> 00:16:23.750 align:middle line:84%
We don't need to choose
the simplest model.

00:16:23.750 --> 00:16:25.810 align:middle line:84%
We still want to
describe the data well.

00:16:25.810 --> 00:16:27.910 align:middle line:84%
But there's a preference
for simple models.

00:16:27.910 --> 00:16:30.930 align:middle line:84%
And that comes from Occam's
razor, which is to say,

00:16:30.930 --> 00:16:37.290 align:middle line:84%
basically, simple things
make the most sense.

00:16:37.290 --> 00:16:39.070 align:middle line:90%
Let's say I have a picture.

00:16:39.070 --> 00:16:41.850 align:middle line:90%


00:16:41.850 --> 00:16:46.930 align:middle line:84%
And there's a picture
of the NW14 building.

00:16:46.930 --> 00:16:50.490 align:middle line:84%
And remember, there's a
road here in front of it,

00:16:50.490 --> 00:16:51.710 align:middle line:90%
and then there's a sidewalk.

00:16:51.710 --> 00:16:58.090 align:middle line:84%
And then there are these
trees planted on the sidewalk

00:16:58.090 --> 00:17:00.570 align:middle line:90%
along the front of NW14.

00:17:00.570 --> 00:17:03.050 align:middle line:90%
There's a little door here.

00:17:03.050 --> 00:17:12.430 align:middle line:84%
And then there are some
windows in the NW14 building.

00:17:12.430 --> 00:17:15.690 align:middle line:90%


00:17:15.690 --> 00:17:16.670 align:middle line:90%
We're behind the trees?

00:17:16.670 --> 00:17:22.329 align:middle line:90%
Maybe I should make the windows.

00:17:22.329 --> 00:17:24.270 align:middle line:84%
I don't really have
any different colors.

00:17:24.270 --> 00:17:27.569 align:middle line:90%


00:17:27.569 --> 00:17:28.270 align:middle line:90%
There we go.

00:17:28.270 --> 00:17:30.970 align:middle line:90%


00:17:30.970 --> 00:17:31.470 align:middle line:90%
All right.

00:17:31.470 --> 00:17:35.400 align:middle line:84%
How many windows
are in this picture?

00:17:35.400 --> 00:17:36.640 align:middle line:90%
Four.

00:17:36.640 --> 00:17:37.140 align:middle line:90%
All right.

00:17:37.140 --> 00:17:39.515 align:middle line:84%
But how do you know there's
not like a little tiny window

00:17:39.515 --> 00:17:40.060 align:middle line:90%
right here?

00:17:40.060 --> 00:17:43.240 align:middle line:90%


00:17:43.240 --> 00:17:48.800 align:middle line:84%
OK, maybe, but it's
strange credulity.

00:17:48.800 --> 00:17:52.400 align:middle line:90%


00:17:52.400 --> 00:17:55.680 align:middle line:84%
Maybe there's actually a secret
little window right here.

00:17:55.680 --> 00:17:59.720 align:middle line:84%
OK, It gets obviously the
more complex your model space

00:17:59.720 --> 00:18:02.700 align:middle line:84%
where you're having to consider
all kinds of possibilities,

00:18:02.700 --> 00:18:05.100 align:middle line:90%
the less likely it is.

00:18:05.100 --> 00:18:07.780 align:middle line:84%
So that's the kind of
intuitive explanation.

00:18:07.780 --> 00:18:11.120 align:middle line:84%
But we can actually show
this mathematically.

00:18:11.120 --> 00:18:13.980 align:middle line:84%
So we can go back and
derive Bayes' rule.

00:18:13.980 --> 00:18:18.000 align:middle line:84%
So let's say we have the
probability of A and B--

00:18:18.000 --> 00:18:23.120 align:middle line:90%
A and B. Well, what is that?

00:18:23.120 --> 00:18:30.760 align:middle line:84%
Well, it's the probability of
A times the probability of B

00:18:30.760 --> 00:18:34.250 align:middle line:84%
if A happens given A.
This is straightforward.

00:18:34.250 --> 00:18:37.090 align:middle line:84%
And then we can just flip the
variables and write the reverse.

00:18:37.090 --> 00:18:45.910 align:middle line:90%


00:18:45.910 --> 00:18:47.910 align:middle line:90%
Everyone is happy with that?

00:18:47.910 --> 00:18:50.430 align:middle line:84%
And does everyone agree that
the probability of A and B

00:18:50.430 --> 00:18:53.230 align:middle line:84%
is the same thing as the
probability of B and A?

00:18:53.230 --> 00:18:56.910 align:middle line:84%
Yes, so then we can
just create these.

00:18:56.910 --> 00:19:01.270 align:middle line:84%
Probability of A times the
probability of B given A

00:19:01.270 --> 00:19:07.670 align:middle line:84%
is equal to the probability of B
times the probability of A given

00:19:07.670 --> 00:19:11.850 align:middle line:84%
B. And then I can just
divide both sides by this.

00:19:11.850 --> 00:19:15.470 align:middle line:90%


00:19:15.470 --> 00:19:17.750 align:middle line:90%
And that's Bayes' law.

00:19:17.750 --> 00:19:20.890 align:middle line:84%
So that's a very useful way
to derive Bayes' equation

00:19:20.890 --> 00:19:23.090 align:middle line:90%
if you don't really remember.

00:19:23.090 --> 00:19:26.550 align:middle line:90%


00:19:26.550 --> 00:19:28.890 align:middle line:84%
This actually embodies
Occam's razor.

00:19:28.890 --> 00:19:30.710 align:middle line:90%
And let me just show you why.

00:19:30.710 --> 00:19:32.560 align:middle line:90%
So let's make a substitution.

00:19:32.560 --> 00:19:36.500 align:middle line:90%


00:19:36.500 --> 00:19:44.080 align:middle line:84%
Let's call B here a hypothesis,
and let's call A the data.

00:19:44.080 --> 00:19:47.020 align:middle line:90%


00:19:47.020 --> 00:19:48.660 align:middle line:84%
If I write over here,
people are going

00:19:48.660 --> 00:19:50.520 align:middle line:84%
to have to turn
their heads too much?

00:19:50.520 --> 00:19:51.020 align:middle line:90%
No?

00:19:51.020 --> 00:19:51.560 align:middle line:90%
OK.

00:19:51.560 --> 00:19:52.380 align:middle line:90%
All right.

00:19:52.380 --> 00:19:54.280 align:middle line:84%
So we can write
something like this.

00:19:54.280 --> 00:19:58.460 align:middle line:84%
The probability of hypothesis
one, given some data,

00:19:58.460 --> 00:20:03.520 align:middle line:84%
is equal to the probability of
the data given hypothesis one.

00:20:03.520 --> 00:20:05.680 align:middle line:84%
Or let's actually
call this model one.

00:20:05.680 --> 00:20:08.540 align:middle line:90%
It's a better model one.

00:20:08.540 --> 00:20:15.300 align:middle line:84%
Model one, given the data, times
the probability of the model

00:20:15.300 --> 00:20:19.860 align:middle line:84%
divided by the
probability of the data.

00:20:19.860 --> 00:20:20.360 align:middle line:90%
All right.

00:20:20.360 --> 00:20:23.660 align:middle line:84%
So this is the probability
that the model is true

00:20:23.660 --> 00:20:28.940 align:middle line:84%
given the observed data, which
is what we would like to know.

00:20:28.940 --> 00:20:31.295 align:middle line:84%
And so we could also
have a second model.

00:20:31.295 --> 00:20:33.170 align:middle line:84%
We say, OK, well, we
can have the probability

00:20:33.170 --> 00:20:37.110 align:middle line:84%
of model 2 given the observed
data, which are unchanging.

00:20:37.110 --> 00:20:47.450 align:middle line:90%


00:20:47.450 --> 00:20:48.770 align:middle line:90%
All right.

00:20:48.770 --> 00:20:51.610 align:middle line:84%
And we can just divide
these two equations.

00:20:51.610 --> 00:20:55.650 align:middle line:84%
So we get the probability
of this minus this.

00:20:55.650 --> 00:20:58.170 align:middle line:90%
These terms cancel.

00:20:58.170 --> 00:21:00.990 align:middle line:90%
And so we have this times this.

00:21:00.990 --> 00:21:04.010 align:middle line:84%
Let me just rewrite this here,
make it a little bit cleaner.

00:21:04.010 --> 00:21:08.370 align:middle line:84%
This is equal to the probability
of the data given model

00:21:08.370 --> 00:21:13.770 align:middle line:84%
1 times the probability
of model 1 divided

00:21:13.770 --> 00:21:18.570 align:middle line:84%
by the probability of
the data model 2 times

00:21:18.570 --> 00:21:21.650 align:middle line:90%
probability of model 2.

00:21:21.650 --> 00:21:23.970 align:middle line:90%
All right.

00:21:23.970 --> 00:21:30.880 align:middle line:84%
So maybe it will make more
sense if I just do that.

00:21:30.880 --> 00:21:34.000 align:middle line:90%
So what is this term here?

00:21:34.000 --> 00:21:35.277 align:middle line:90%
So this is your prior.

00:21:35.277 --> 00:21:37.360 align:middle line:84%
This is the probability
that you believe the model

00:21:37.360 --> 00:21:39.000 align:middle line:90%
without the data.

00:21:39.000 --> 00:21:41.720 align:middle line:84%
So is my prior belief in
the probability of model

00:21:41.720 --> 00:21:44.300 align:middle line:84%
1 versus my prior
belief in model 2.

00:21:44.300 --> 00:21:47.320 align:middle line:84%
And if I'm agnostic
between models,

00:21:47.320 --> 00:21:49.540 align:middle line:90%
then this is basically unity.

00:21:49.540 --> 00:21:50.500 align:middle line:90%
This goes to 1.

00:21:50.500 --> 00:21:53.360 align:middle line:90%


00:21:53.360 --> 00:21:54.400 align:middle line:90%
All right.

00:21:54.400 --> 00:21:55.940 align:middle line:90%
So now we simply have--

00:21:55.940 --> 00:21:59.280 align:middle line:90%


00:21:59.280 --> 00:22:02.360 align:middle line:90%
this looks a bit sloppy.

00:22:02.360 --> 00:22:03.980 align:middle line:90%
Let me just rewrite over here.

00:22:03.980 --> 00:22:11.360 align:middle line:84%
Probability of model 1 given
the data probability of model 2.

00:22:11.360 --> 00:22:12.960 align:middle line:84%
So now we have
this, which is what

00:22:12.960 --> 00:22:16.240 align:middle line:84%
is the ratio of that we
believe in view of the data,

00:22:16.240 --> 00:22:18.120 align:middle line:90%
model 1 over model 2.

00:22:18.120 --> 00:22:20.720 align:middle line:84%
And it's just the ratio
of the likelihoods.

00:22:20.720 --> 00:22:23.240 align:middle line:84%
This is what's called the
likelihood-- the probability

00:22:23.240 --> 00:22:26.710 align:middle line:84%
that I see the data
given the model.

00:22:26.710 --> 00:22:28.950 align:middle line:90%
So here's the insight.

00:22:28.950 --> 00:22:33.030 align:middle line:84%
If model one is
the simple model,

00:22:33.030 --> 00:22:36.390 align:middle line:84%
it has a small number
of possible states.

00:22:36.390 --> 00:22:38.310 align:middle line:84%
If model 2 is a
complex model, it

00:22:38.310 --> 00:22:41.230 align:middle line:90%
has many more possible states.

00:22:41.230 --> 00:22:46.270 align:middle line:84%
So therefore, the probability
of seeing the given data

00:22:46.270 --> 00:22:50.550 align:middle line:84%
for a complex model is
always going to be smaller.

00:22:50.550 --> 00:23:03.150 align:middle line:84%
So for model 2
complex relative to 1,

00:23:03.150 --> 00:23:09.070 align:middle line:84%
this ratio here is going
to be greater than 1

00:23:09.070 --> 00:23:12.590 align:middle line:84%
because this will
be less than this.

00:23:12.590 --> 00:23:13.910 align:middle line:90%
Does that make sense?

00:23:13.910 --> 00:23:17.510 align:middle line:84%
And so that means we will
prefer by some amount model

00:23:17.510 --> 00:23:18.510 align:middle line:90%
1 over model 2.

00:23:18.510 --> 00:23:19.790 align:middle line:90%
I can just take model 2.

00:23:19.790 --> 00:23:22.870 align:middle line:84%
I can multiply it by both
sides, put it over here,

00:23:22.870 --> 00:23:25.580 align:middle line:84%
and it has 1 times model
2, which means something

00:23:25.580 --> 00:23:29.420 align:middle line:84%
greater than 1 times model 2
is the probability of model 1.

00:23:29.420 --> 00:23:33.420 align:middle line:84%
So that just tells us that we
have a preference for models

00:23:33.420 --> 00:23:38.100 align:middle line:84%
which have more simple
descriptions, which

00:23:38.100 --> 00:23:43.020 align:middle line:84%
is essentially this,
just mathematically.

00:23:43.020 --> 00:23:55.420 align:middle line:84%
So it takes a lot of data
or some kind of prior reason

00:23:55.420 --> 00:23:59.620 align:middle line:84%
to support choosing
a different model.

00:23:59.620 --> 00:24:03.160 align:middle line:84%
And so let's look
at how we decide.

00:24:03.160 --> 00:24:07.620 align:middle line:84%
So one of the things that
comes up often in radiation

00:24:07.620 --> 00:24:10.140 align:middle line:90%
is should there be a threshold.

00:24:10.140 --> 00:24:15.940 align:middle line:84%
And so one question might
be, is this model actually

00:24:15.940 --> 00:24:18.240 align:middle line:84%
really more complicated
than this model?

00:24:18.240 --> 00:24:21.260 align:middle line:90%


00:24:21.260 --> 00:24:22.600 align:middle line:90%
And the answer is yes.

00:24:22.600 --> 00:24:27.290 align:middle line:84%
So in the no-threshold model,
our model is something like y

00:24:27.290 --> 00:24:29.930 align:middle line:90%
equals alpha x.

00:24:29.930 --> 00:24:34.450 align:middle line:84%
And in the threshold model,
our model is piecewise,

00:24:34.450 --> 00:24:39.170 align:middle line:84%
and it looks like y
equals 0 if x less than

00:24:39.170 --> 00:24:41.510 align:middle line:90%
or equal to some threshold.

00:24:41.510 --> 00:24:44.690 align:middle line:84%
Otherwise, x greater
than threshold

00:24:44.690 --> 00:24:49.970 align:middle line:90%
is beta x minus some intercept.

00:24:49.970 --> 00:24:56.810 align:middle line:84%
So we have now extra parameters
in the threshold model.

00:24:56.810 --> 00:25:00.210 align:middle line:84%
And the threshold model
thus requires more data

00:25:00.210 --> 00:25:03.250 align:middle line:84%
to defend it than the
no-threshold model.

00:25:03.250 --> 00:25:05.290 align:middle line:90%
All right.

00:25:05.290 --> 00:25:08.850 align:middle line:90%
So how is model selection done?

00:25:08.850 --> 00:25:11.170 align:middle line:84%
So the problem is
that when we look

00:25:11.170 --> 00:25:13.290 align:middle line:84%
at this going back
to the Bayes rule

00:25:13.290 --> 00:25:17.510 align:middle line:84%
here, when we look at the
probability of the data which.

00:25:17.510 --> 00:25:20.130 align:middle line:90%


00:25:20.130 --> 00:25:22.840 align:middle line:90%
Well I deleted.

00:25:22.840 --> 00:25:26.660 align:middle line:84%
And this, I was able to delete
it here and not worry about it.

00:25:26.660 --> 00:25:29.440 align:middle line:84%
But when you want to
calculate it directly,

00:25:29.440 --> 00:25:32.800 align:middle line:84%
you need to do this
bit of math here.

00:25:32.800 --> 00:25:34.760 align:middle line:84%
And the problem is
that this bit of math

00:25:34.760 --> 00:25:39.580 align:middle line:84%
is often impossible to do
or very difficult to do.

00:25:39.580 --> 00:25:41.720 align:middle line:84%
The way you traditionally
do it numerically

00:25:41.720 --> 00:25:45.360 align:middle line:84%
is through a Markov
chain Monte Carlo.

00:25:45.360 --> 00:25:50.080 align:middle line:84%
But analytically, it's a
really difficult thing to do.

00:25:50.080 --> 00:25:53.160 align:middle line:84%
So instead, what we
tend to do is come up

00:25:53.160 --> 00:25:56.280 align:middle line:84%
with some kind of
approximation method.

00:25:56.280 --> 00:26:03.240 align:middle line:84%
So the best way to do this
if you have a large data set

00:26:03.240 --> 00:26:05.840 align:middle line:90%
is called cross-validation.

00:26:05.840 --> 00:26:13.760 align:middle line:84%
So what you do is you have a
lot of samples in your data.

00:26:13.760 --> 00:26:17.000 align:middle line:84%
Here's a whole bunch of
people in my data set.

00:26:17.000 --> 00:26:20.990 align:middle line:84%
And I train my model
on this population.

00:26:20.990 --> 00:26:25.870 align:middle line:84%
And then I use these people who
didn't go into the trained model

00:26:25.870 --> 00:26:29.250 align:middle line:84%
to test how well the
model works to see, OK,

00:26:29.250 --> 00:26:33.590 align:middle line:84%
what is the prediction error
for each of these observations?

00:26:33.590 --> 00:26:36.490 align:middle line:84%
And then I do that again, but
I choose a different subset.

00:26:36.490 --> 00:26:38.910 align:middle line:84%
And I go through all
the possible subsets.

00:26:38.910 --> 00:26:44.750 align:middle line:84%
And I basically measure
how badly the model

00:26:44.750 --> 00:26:48.470 align:middle line:90%
predicts the outliers.

00:26:48.470 --> 00:26:51.830 align:middle line:84%
That only works if you
have a lot of data.

00:26:51.830 --> 00:26:54.770 align:middle line:84%
And you can actually train
the model on a small subset.

00:26:54.770 --> 00:26:59.430 align:middle line:84%
If you don't have a lot of data,
then you're kind of in trouble.

00:26:59.430 --> 00:27:03.310 align:middle line:84%
The other way to do this is that
you can use a scoring function.

00:27:03.310 --> 00:27:06.690 align:middle line:84%
And there are two good
scoring functions.

00:27:06.690 --> 00:27:09.310 align:middle line:84%
One is called the Bayesian
information Criterion,

00:27:09.310 --> 00:27:14.350 align:middle line:84%
and the other one is the
Akaike Information Criterion.

00:27:14.350 --> 00:27:21.270 align:middle line:84%
If you have reason to believe
that the true model that

00:27:21.270 --> 00:27:24.430 align:middle line:84%
actually contains
the real physics

00:27:24.430 --> 00:27:26.710 align:middle line:84%
is somehow in the set
of models that you're

00:27:26.710 --> 00:27:30.870 align:middle line:84%
evaluating, which is
almost never the case,

00:27:30.870 --> 00:27:36.270 align:middle line:84%
then BIC can be shown to be
the correct way to approach it.

00:27:36.270 --> 00:27:39.910 align:middle line:84%
Otherwise, the Akaike
Information Criterion

00:27:39.910 --> 00:27:41.870 align:middle line:84%
is usually a better
way to do it.

00:27:41.870 --> 00:27:43.310 align:middle line:84%
Here it is-- the
number of samples

00:27:43.310 --> 00:27:46.030 align:middle line:84%
times the log of the
residual sums of squares,

00:27:46.030 --> 00:27:49.310 align:middle line:84%
plus the number of
free parameters.

00:27:49.310 --> 00:27:54.790 align:middle line:84%
Basically, if you run if you
run any data through NumPy

00:27:54.790 --> 00:27:59.510 align:middle line:84%
or something, it has a way of
just calling for this value.

00:27:59.510 --> 00:28:01.110 align:middle line:90%
And you can get out the value.

00:28:01.110 --> 00:28:05.230 align:middle line:84%
And basically, what you do
is you get this number out

00:28:05.230 --> 00:28:07.390 align:middle line:84%
for every model, and you
just choose the model

00:28:07.390 --> 00:28:09.210 align:middle line:90%
that has the lowest AIC.

00:28:09.210 --> 00:28:12.790 align:middle line:84%
That is the model that contains
the most possible information.

00:28:12.790 --> 00:28:14.310 align:middle line:84%
This is a very
simple result which

00:28:14.310 --> 00:28:18.140 align:middle line:84%
comes from a long
information-theoretic

00:28:18.140 --> 00:28:21.440 align:middle line:84%
derivation, and you can go look
it up if you're interested.

00:28:21.440 --> 00:28:24.020 align:middle line:84%
But it's basically
simply a way to penalize

00:28:24.020 --> 00:28:28.020 align:middle line:84%
overfitting-- that is, at least
has a foundation in information

00:28:28.020 --> 00:28:28.740 align:middle line:90%
theory.

00:28:28.740 --> 00:28:30.995 align:middle line:84%
So this is how you would
do model selection.

00:28:30.995 --> 00:28:32.620 align:middle line:84%
And you should know
this because you're

00:28:32.620 --> 00:28:34.980 align:middle line:84%
going to have to do
model selection if you're

00:28:34.980 --> 00:28:37.780 align:middle line:84%
a researcher, especially
if you're a grad student.

00:28:37.780 --> 00:28:42.140 align:middle line:84%
And people frequently say, oh,
here's some experimental data,

00:28:42.140 --> 00:28:43.280 align:middle line:90%
and it looks like this.

00:28:43.280 --> 00:28:44.780 align:middle line:84%
And you fit something,
and you maybe

00:28:44.780 --> 00:28:47.180 align:middle line:90%
don't give it a second thought.

00:28:47.180 --> 00:28:51.420 align:middle line:84%
And really, you should be trying
multiple models that possibly

00:28:51.420 --> 00:28:54.220 align:middle line:84%
make sense and scoring
them either using

00:28:54.220 --> 00:28:59.780 align:middle line:84%
k-fold cross-validation or using
AIC before just choosing one

00:28:59.780 --> 00:29:01.420 align:middle line:90%
for your thesis.

00:29:01.420 --> 00:29:02.860 align:middle line:90%
Don't go on R-squared.

00:29:02.860 --> 00:29:07.840 align:middle line:84%
R-squared only tells you how
well it matches the data points.

00:29:07.840 --> 00:29:10.680 align:middle line:84%
It doesn't tell you how much
information is in the model.

00:29:10.680 --> 00:29:13.540 align:middle line:90%


00:29:13.540 --> 00:29:18.310 align:middle line:84%
This is elementary, but if I
have these data points like this

00:29:18.310 --> 00:29:20.650 align:middle line:84%
and I have a model
that does this,

00:29:20.650 --> 00:29:22.750 align:middle line:84%
it fits the fits
the data perfectly,

00:29:22.750 --> 00:29:27.210 align:middle line:84%
but probably has almost
no predictive value.

00:29:27.210 --> 00:29:27.970 align:middle line:90%
OK.

00:29:27.970 --> 00:29:32.810 align:middle line:84%
So this is where linear
no threshold comes from.

00:29:32.810 --> 00:29:36.410 align:middle line:84%
If you actually go through this
process rigorously and you start

00:29:36.410 --> 00:29:42.090 align:middle line:84%
testing a bunch of models,
as I have been doing lately,

00:29:42.090 --> 00:29:45.470 align:middle line:84%
you will find that
the LNT model,

00:29:45.470 --> 00:29:47.690 align:middle line:84%
this is a k-fold
cross-validation result

00:29:47.690 --> 00:29:50.490 align:middle line:90%
for a 20-fold cross-validation.

00:29:50.490 --> 00:29:52.715 align:middle line:84%
This is for the
lifespan survey data.

00:29:52.715 --> 00:29:54.090 align:middle line:84%
These are the
people in Hiroshima

00:29:54.090 --> 00:29:56.530 align:middle line:84%
who were exposed to radiation
from the atom bomb, which

00:29:56.530 --> 00:29:59.470 align:middle line:84%
is the largest single data
set for low-dose exposure.

00:29:59.470 --> 00:30:02.770 align:middle line:84%
I think I have some information
about how many people

00:30:02.770 --> 00:30:05.770 align:middle line:90%
are in a subsequent slide.

00:30:05.770 --> 00:30:09.650 align:middle line:84%
Only below 1 sievert-- so we
are in the potentially linear

00:30:09.650 --> 00:30:10.490 align:middle line:90%
regime.

00:30:10.490 --> 00:30:12.470 align:middle line:84%
And these are different
models-- linear model,

00:30:12.470 --> 00:30:14.370 align:middle line:84%
linear-quadratic model,
a threshold model

00:30:14.370 --> 00:30:16.680 align:middle line:84%
where we allow the
threshold to exist

00:30:16.680 --> 00:30:19.840 align:middle line:84%
at the optimal possible
location that give the best

00:30:19.840 --> 00:30:21.840 align:middle line:90%
possible fit to the data.

00:30:21.840 --> 00:30:23.460 align:middle line:90%
Quadratic, this is a Gaussian.

00:30:23.460 --> 00:30:25.700 align:middle line:84%
These are non-parametric
models, Gaussian process,

00:30:25.700 --> 00:30:27.840 align:middle line:90%
Linear Gaussian process other.

00:30:27.840 --> 00:30:31.960 align:middle line:84%
And you see that basically
the linear model just

00:30:31.960 --> 00:30:33.440 align:middle line:90%
constantly wins.

00:30:33.440 --> 00:30:38.960 align:middle line:84%
And this is frustrating
because I would like to--

00:30:38.960 --> 00:30:43.400 align:middle line:84%
I would like a model that
explains why this happens.

00:30:43.400 --> 00:30:47.320 align:middle line:84%
But that shape, you'll
see it over and over.

00:30:47.320 --> 00:30:52.280 align:middle line:84%
But statistically, the only
thing we can really defend

00:30:52.280 --> 00:30:53.640 align:middle line:90%
is the linear model.

00:30:53.640 --> 00:30:54.140 align:middle line:90%
All right.

00:30:54.140 --> 00:30:55.807 align:middle line:84%
Now I have to go
through all this again.

00:30:55.807 --> 00:30:58.360 align:middle line:90%


00:30:58.360 --> 00:31:05.240 align:middle line:84%
So there are some
other pitfalls that you

00:31:05.240 --> 00:31:09.040 align:middle line:84%
should be aware of when looking
at radiation dose-response

00:31:09.040 --> 00:31:12.840 align:middle line:84%
models that sometimes
people fall into.

00:31:12.840 --> 00:31:15.830 align:middle line:84%
And here's an example
of one of them.

00:31:15.830 --> 00:31:19.910 align:middle line:84%
So this is the
excess relative risk

00:31:19.910 --> 00:31:29.310 align:middle line:84%
for cancer for exposures
per sievert of radiation

00:31:29.310 --> 00:31:31.730 align:middle line:90%
to different organs.

00:31:31.730 --> 00:31:34.970 align:middle line:84%
And so remember we talked
about colon-weighted dose.

00:31:34.970 --> 00:31:39.510 align:middle line:84%
You can see here that it's not
uniform depending on the tissue.

00:31:39.510 --> 00:31:40.230 align:middle line:90%
All right.

00:31:40.230 --> 00:31:40.970 align:middle line:90%
So let's see.

00:31:40.970 --> 00:31:45.990 align:middle line:84%
So 0 excess risk means that
that's your baseline risk

00:31:45.990 --> 00:31:47.550 align:middle line:90%
for no radiation.

00:31:47.550 --> 00:31:51.750 align:middle line:84%
And here, we see that
for uterus exposure,

00:31:51.750 --> 00:31:55.750 align:middle line:84%
the probability of
cancer is actually lower.

00:31:55.750 --> 00:32:00.170 align:middle line:84%
So should we go around
irradiating women's uteruses?

00:32:00.170 --> 00:32:04.670 align:middle line:90%


00:32:04.670 --> 00:32:05.750 align:middle line:90%
AUDIENCE: Just for fun?

00:32:05.750 --> 00:32:09.430 align:middle line:84%
SCOTT KEMP: No,
for their benefit.

00:32:09.430 --> 00:32:12.460 align:middle line:84%
Because we will reduce the
risk of uterine cancer.

00:32:12.460 --> 00:32:13.880 align:middle line:90%
Should we go and do that?

00:32:13.880 --> 00:32:17.260 align:middle line:90%


00:32:17.260 --> 00:32:18.720 align:middle line:90%
This data is correct.

00:32:18.720 --> 00:32:20.540 align:middle line:90%
I mean, the real data.

00:32:20.540 --> 00:32:22.020 align:middle line:90%
AUDIENCE: How much data is it?

00:32:22.020 --> 00:32:24.040 align:middle line:84%
SCOTT KEMP: Well,
that's a good question.

00:32:24.040 --> 00:32:25.100 align:middle line:90%
There's the error bar.

00:32:25.100 --> 00:32:31.403 align:middle line:90%


00:32:31.403 --> 00:32:33.320 align:middle line:84%
AUDIENCE: The gist is
it doesn't cause cancer,

00:32:33.320 --> 00:32:35.230 align:middle line:84%
but then maybe
there's other effects

00:32:35.230 --> 00:32:36.480 align:middle line:90%
that we're not accounting for.

00:32:36.480 --> 00:32:39.263 align:middle line:90%


00:32:39.263 --> 00:32:40.680 align:middle line:84%
SCOTT KEMP: There's
other effects.

00:32:40.680 --> 00:32:41.960 align:middle line:90%
So you're being precautionary.

00:32:41.960 --> 00:32:43.780 align:middle line:84%
But on the basis of
cancer, you don't

00:32:43.780 --> 00:32:45.627 align:middle line:90%
have an argument against it.

00:32:45.627 --> 00:32:46.960 align:middle line:90%
AUDIENCE: Cancer for the mother?

00:32:46.960 --> 00:32:48.380 align:middle line:90%
SCOTT KEMP: For the uterus.

00:32:48.380 --> 00:32:50.300 align:middle line:90%
AUDIENCE: Yeah.

00:32:50.300 --> 00:32:52.420 align:middle line:84%
AUDIENCE: You just
have to look at how

00:32:52.420 --> 00:32:54.200 align:middle line:84%
cancer affects the
quality of life,

00:32:54.200 --> 00:32:56.617 align:middle line:90%
not just the incidence of some.

00:32:56.617 --> 00:32:58.200 align:middle line:84%
SCOTT KEMP: Well,
if you don't get it,

00:32:58.200 --> 00:32:59.325 align:middle line:90%
that's better than nothing.

00:32:59.325 --> 00:33:02.940 align:middle line:90%


00:33:02.940 --> 00:33:05.380 align:middle line:84%
This is called the
look-again effect.

00:33:05.380 --> 00:33:07.580 align:middle line:84%
Have you heard of the
look-again effect.

00:33:07.580 --> 00:33:08.160 align:middle line:90%
All right.

00:33:08.160 --> 00:33:13.690 align:middle line:84%
So if these error bars are 95%
confidence intervals, which

00:33:13.690 --> 00:33:16.690 align:middle line:84%
means 95% of the
observations are estimated

00:33:16.690 --> 00:33:20.490 align:middle line:84%
to fall within the error bar,
which means one out of every 20

00:33:20.490 --> 00:33:22.450 align:middle line:84%
is going to fall
out of the error.

00:33:22.450 --> 00:33:25.230 align:middle line:84%
The estimation will be wrong
by the amount of the error bar.

00:33:25.230 --> 00:33:29.150 align:middle line:84%
So there's about
20 things up here.

00:33:29.150 --> 00:33:31.590 align:middle line:90%
And one of them is wrong.

00:33:31.590 --> 00:33:33.090 align:middle line:90%
And that's what it is.

00:33:33.090 --> 00:33:37.250 align:middle line:84%
But if you look at
all of these things

00:33:37.250 --> 00:33:42.930 align:middle line:84%
and then just picked this
result out and wrote a paper,

00:33:42.930 --> 00:33:45.270 align:middle line:90%
then you would have a problem.

00:33:45.270 --> 00:33:49.490 align:middle line:84%
So in fact, this is easily shown
if we also plot the mortality

00:33:49.490 --> 00:33:50.670 align:middle line:90%
from radiation risk.

00:33:50.670 --> 00:33:53.870 align:middle line:84%
And you see, it was just
a statistical fluke.

00:33:53.870 --> 00:33:57.290 align:middle line:84%
The mortality shows that,
in fact, radiation exposure

00:33:57.290 --> 00:34:00.170 align:middle line:84%
does increase your
mortality of uterine cancer,

00:34:00.170 --> 00:34:02.170 align:middle line:84%
even though it appears
to reduce the probability

00:34:02.170 --> 00:34:03.370 align:middle line:90%
of uterine cancer.

00:34:03.370 --> 00:34:06.810 align:middle line:84%
And here you see here the
effect has shown up here

00:34:06.810 --> 00:34:10.260 align:middle line:84%
for oral and pancreas in
the opposite direction.

00:34:10.260 --> 00:34:11.780 align:middle line:90%
It's just this effect.

00:34:11.780 --> 00:34:15.480 align:middle line:84%
So there was a famous
XKCD cartoon about this.

00:34:15.480 --> 00:34:17.139 align:middle line:90%
Jelly beans cause acne.

00:34:17.139 --> 00:34:18.400 align:middle line:90%
Scientists investigate.

00:34:18.400 --> 00:34:21.360 align:middle line:84%
And there's this claim
that they found this link.

00:34:21.360 --> 00:34:24.739 align:middle line:90%
And then here is this thing.

00:34:24.739 --> 00:34:27.120 align:middle line:84%
It appears to be
a certain color.

00:34:27.120 --> 00:34:29.199 align:middle line:84%
And so they go through
all the colors.

00:34:29.199 --> 00:34:34.120 align:middle line:84%
Found no link between
purple and brown and pink

00:34:34.120 --> 00:34:37.280 align:middle line:84%
and blue and teal
and salmon color,

00:34:37.280 --> 00:34:40.400 align:middle line:84%
red color, turquoise,
magenta, yellow.

00:34:40.400 --> 00:34:44.840 align:middle line:84%
The p-value is always too
large to defend any link.

00:34:44.840 --> 00:34:47.699 align:middle line:84%
And then, lo and behold,
they get to green.

00:34:47.699 --> 00:34:54.880 align:middle line:84%
And the p-value appears to
justify that there is acne

00:34:54.880 --> 00:34:55.900 align:middle line:90%
caused by jelly beans.

00:34:55.900 --> 00:34:57.700 align:middle line:84%
And then you have
this publication.

00:34:57.700 --> 00:35:00.680 align:middle line:90%


00:35:00.680 --> 00:35:05.260 align:middle line:84%
If you look at the data and
you start ignoring things,

00:35:05.260 --> 00:35:10.510 align:middle line:84%
you actually have to adjust
how you calculate the p-value,

00:35:10.510 --> 00:35:12.730 align:middle line:84%
because every time
you pull a sample,

00:35:12.730 --> 00:35:16.630 align:middle line:84%
every time you run a test,
you are actually adjusting.

00:35:16.630 --> 00:35:19.990 align:middle line:84%
You are making choices
and filtering the data,

00:35:19.990 --> 00:35:22.550 align:middle line:84%
and therefore you have
to actually change

00:35:22.550 --> 00:35:24.090 align:middle line:90%
how you do that calculation.

00:35:24.090 --> 00:35:27.070 align:middle line:84%
This is frequently
forgotten about.

00:35:27.070 --> 00:35:30.190 align:middle line:84%
And people do this, and they
run their science experiment

00:35:30.190 --> 00:35:32.610 align:middle line:84%
on their instrument until
they get the result they want.

00:35:32.610 --> 00:35:35.230 align:middle line:84%
Now finally, the
instrument worked right.

00:35:35.230 --> 00:35:36.830 align:middle line:90%
And then they publish.

00:35:36.830 --> 00:35:39.710 align:middle line:84%
It's like, well,
actually, you have

00:35:39.710 --> 00:35:42.790 align:middle line:84%
to take into account all
of those unsuccessful runs,

00:35:42.790 --> 00:35:45.470 align:middle line:90%
and maybe it's just noise.

00:35:45.470 --> 00:35:46.850 align:middle line:90%
So this happens all the time.

00:35:46.850 --> 00:35:48.270 align:middle line:84%
And one of the
ways this manifests

00:35:48.270 --> 00:35:51.910 align:middle line:84%
is that people get these results
out of small little laboratory

00:35:51.910 --> 00:35:53.670 align:middle line:84%
experiments about
radiation, and they

00:35:53.670 --> 00:35:56.010 align:middle line:84%
show that some radiation
is good for you.

00:35:56.010 --> 00:35:58.590 align:middle line:84%
And then they say, Lo, and
we have discovered hormesis,

00:35:58.590 --> 00:36:03.310 align:middle line:84%
this idea that a little bit
of radiation is good for you,

00:36:03.310 --> 00:36:05.380 align:middle line:90%
and it's not real.

00:36:05.380 --> 00:36:11.940 align:middle line:90%
So you've been warned.

00:36:11.940 --> 00:36:13.180 align:middle line:90%
All right.

00:36:13.180 --> 00:36:19.140 align:middle line:84%
This is kind of like the grand
plot from the lifespan survey.

00:36:19.140 --> 00:36:21.512 align:middle line:84%
This is from ICRP
Publication 99, which

00:36:21.512 --> 00:36:22.720 align:middle line:90%
is actually assigned reading.

00:36:22.720 --> 00:36:24.620 align:middle line:90%
I hope you looked at it.

00:36:24.620 --> 00:36:28.103 align:middle line:84%
And here is up to
two sieverts of dose.

00:36:28.103 --> 00:36:29.520 align:middle line:84%
And you can see
the linear regime.

00:36:29.520 --> 00:36:31.620 align:middle line:84%
And then you see in
the low-dose regime,

00:36:31.620 --> 00:36:35.360 align:middle line:84%
there does appear to be this
like super linear region.

00:36:35.360 --> 00:36:40.820 align:middle line:90%


00:36:40.820 --> 00:36:42.940 align:middle line:84%
You could say, well,
maybe that's just it does

00:36:42.940 --> 00:36:45.200 align:middle line:84%
appear consistently
between data points.

00:36:45.200 --> 00:36:47.580 align:middle line:90%
So that could be.

00:36:47.580 --> 00:36:49.900 align:middle line:84%
It looks like it's
not just noise.

00:36:49.900 --> 00:36:53.260 align:middle line:84%
But I have tried as I could
to find a non-linear model

00:36:53.260 --> 00:36:55.300 align:middle line:90%
to describe this.

00:36:55.300 --> 00:36:56.580 align:middle line:90%
You can't do it.

00:36:56.580 --> 00:36:57.860 align:middle line:90%
It does not.

00:36:57.860 --> 00:37:01.500 align:middle line:84%
Nothing is statistically
justifiable.

00:37:01.500 --> 00:37:05.730 align:middle line:84%
So I mentioned this
data set earlier.

00:37:05.730 --> 00:37:09.310 align:middle line:84%
It's the single
largest data set.

00:37:09.310 --> 00:37:13.970 align:middle line:84%
It tracked 80,000 individuals
who were exposed to radiation

00:37:13.970 --> 00:37:16.850 align:middle line:90%
because of the atom bomb.

00:37:16.850 --> 00:37:21.010 align:middle line:84%
A lot of people think, oh,
well, that's a lot of radiation.

00:37:21.010 --> 00:37:25.890 align:middle line:84%
Well, it turns out that 64,000
of those 80,000 were exposed

00:37:25.890 --> 00:37:28.690 align:middle line:90%
to less than 100 millisieverts.

00:37:28.690 --> 00:37:33.670 align:middle line:84%
So most of these people are
really in the low-dose regime.

00:37:33.670 --> 00:37:36.290 align:middle line:84%
To put that into perspective,
100 million millisieverts is

00:37:36.290 --> 00:37:39.730 align:middle line:84%
about the same amount of dose a
radiation worker in the United

00:37:39.730 --> 00:37:43.050 align:middle line:84%
States would be allowed to
receive over a 20-year career.

00:37:43.050 --> 00:37:47.810 align:middle line:90%
So it's not that much radiation.

00:37:47.810 --> 00:37:53.730 align:middle line:84%
So we do have 64,000
observations in here.

00:37:53.730 --> 00:37:56.930 align:middle line:84%
And it still doesn't
support anything other

00:37:56.930 --> 00:37:58.550 align:middle line:90%
than the linear no threshold.

00:37:58.550 --> 00:38:01.290 align:middle line:90%


00:38:01.290 --> 00:38:03.870 align:middle line:84%
One of the things I
want to talk about,

00:38:03.870 --> 00:38:07.822 align:middle line:84%
though, is just to look
at the slopes here.

00:38:07.822 --> 00:38:11.010 align:middle line:84%
Right now, the ICRP, the
International Committee

00:38:11.010 --> 00:38:14.250 align:middle line:84%
on Radiation Protection,
which is one of several bodies

00:38:14.250 --> 00:38:17.850 align:middle line:84%
that recommend safety
standards for radiation,

00:38:17.850 --> 00:38:22.170 align:middle line:84%
they need to publish a number
for the slope of what this dose

00:38:22.170 --> 00:38:23.370 align:middle line:90%
response is--

00:38:23.370 --> 00:38:27.210 align:middle line:84%
how much cancer you're going
to get for a certain dose.

00:38:27.210 --> 00:38:31.930 align:middle line:90%
So the ICRP recommends 0.53.

00:38:31.930 --> 00:38:39.450 align:middle line:84%
But if you read the slope of
that chart, it's actually 0.64.

00:38:39.450 --> 00:38:43.310 align:middle line:84%
And if you really limit
yourself to the low-dose regime,

00:38:43.310 --> 00:38:47.290 align:middle line:84%
you would actually read
one, about one per sievert,

00:38:47.290 --> 00:38:50.610 align:middle line:84%
excess relative risk
of one per sievert.

00:38:50.610 --> 00:38:55.430 align:middle line:84%
So should we conclude that the
ICRP is being conservative,

00:38:55.430 --> 00:38:58.270 align:middle line:84%
and that the linear no-threshold
model is conservative,

00:38:58.270 --> 00:38:59.570 align:middle line:90%
as a lot of people say?

00:38:59.570 --> 00:39:01.840 align:middle line:90%
No, in fact, the opposite--

00:39:01.840 --> 00:39:04.960 align:middle line:84%
the data show that it
is not conservative.

00:39:04.960 --> 00:39:08.720 align:middle line:84%
They are being generous in
allowing more radiation out

00:39:08.720 --> 00:39:13.360 align:middle line:84%
there than is defendable
and as opposed to the idea

00:39:13.360 --> 00:39:16.400 align:middle line:84%
that LNT is somehow
a conservative model.

00:39:16.400 --> 00:39:19.180 align:middle line:84%
It is also not supported
because of the statistics.

00:39:19.180 --> 00:39:24.960 align:middle line:84%
The threshold, the threshold
model is not really supported.

00:39:24.960 --> 00:39:32.000 align:middle line:84%
So let's say, we really
liked the idea of a threshold

00:39:32.000 --> 00:39:35.080 align:middle line:84%
because most people in
the nuclear community

00:39:35.080 --> 00:39:39.160 align:middle line:90%
would love that to be true.

00:39:39.160 --> 00:39:43.600 align:middle line:84%
Someone here was doing
a thesis on this.

00:39:43.600 --> 00:39:49.000 align:middle line:84%
And so we could ask,
OK, what would it

00:39:49.000 --> 00:39:54.120 align:middle line:84%
take to actually establish
that a threshold does

00:39:54.120 --> 00:39:56.800 align:middle line:90%
exist at some point?

00:39:56.800 --> 00:39:58.920 align:middle line:84%
And so we can do
that calculation.

00:39:58.920 --> 00:40:02.210 align:middle line:84%
So let's just do it
really quickly here.

00:40:02.210 --> 00:40:04.950 align:middle line:90%


00:40:04.950 --> 00:40:07.950 align:middle line:84%
We set up a little
thought experiment.

00:40:07.950 --> 00:40:11.010 align:middle line:84%
And this is by the way, borrowed
from ICRP Publication 99--

00:40:11.010 --> 00:40:15.590 align:middle line:84%
section 2.42 if you want to
go read their version of it.

00:40:15.590 --> 00:40:17.630 align:middle line:84%
And what we're
going to do is ask

00:40:17.630 --> 00:40:19.950 align:middle line:84%
how many observations or
basically, what does it

00:40:19.950 --> 00:40:24.630 align:middle line:84%
take to observe a threshold
at one millisievert?

00:40:24.630 --> 00:40:30.230 align:middle line:84%
Now, one millisievert
is not all that low,

00:40:30.230 --> 00:40:35.110 align:middle line:84%
but I think it's like just
50% of natural background.

00:40:35.110 --> 00:40:37.890 align:middle line:84%
So that's not really
super low dose.

00:40:37.890 --> 00:40:39.670 align:middle line:84%
People who talk about
threshold models

00:40:39.670 --> 00:40:42.950 align:middle line:84%
often talk about even
lower thresholds.

00:40:42.950 --> 00:40:46.435 align:middle line:84%
But it seems like a reasonable
place to put your test.

00:40:46.435 --> 00:40:47.810 align:middle line:84%
If you put it
really, really low,

00:40:47.810 --> 00:40:50.150 align:middle line:84%
it gets really, really
hard to see because you

00:40:50.150 --> 00:40:51.977 align:middle line:90%
have very, very small effect.

00:40:51.977 --> 00:40:54.310 align:middle line:84%
So we're going to be a little
generous, and we'll put it

00:40:54.310 --> 00:40:57.150 align:middle line:84%
we'll put it at
one millisievert--

00:40:57.150 --> 00:40:59.140 align:middle line:90%
half of background.

00:40:59.140 --> 00:40:59.640 align:middle line:90%
All right.

00:40:59.640 --> 00:41:01.300 align:middle line:84%
Now we need to make
some assumptions

00:41:01.300 --> 00:41:05.520 align:middle line:84%
about how the world works
in order to run this model.

00:41:05.520 --> 00:41:07.020 align:middle line:84%
That first assumption
is we're going

00:41:07.020 --> 00:41:11.540 align:middle line:84%
to assume that the baseline
lifetime lethal cancer risk

00:41:11.540 --> 00:41:13.620 align:middle line:90%
is 10%.

00:41:13.620 --> 00:41:17.420 align:middle line:84%
Now, earlier, I said it
was 20%, which is true.

00:41:17.420 --> 00:41:19.420 align:middle line:84%
But what we're going
to do is assume

00:41:19.420 --> 00:41:23.260 align:middle line:84%
it's 10% in this little
thought experiment.

00:41:23.260 --> 00:41:24.540 align:middle line:90%
Why?

00:41:24.540 --> 00:41:27.920 align:middle line:84%
Well, if the background
rate is lower,

00:41:27.920 --> 00:41:32.260 align:middle line:84%
it's easier to see the
effect of the radiation.

00:41:32.260 --> 00:41:37.060 align:middle line:84%
So it's going to bias
the result of our model

00:41:37.060 --> 00:41:40.060 align:middle line:84%
to make it appear to be
easy to see the threshold,

00:41:40.060 --> 00:41:42.060 align:middle line:90%
easier than reality.

00:41:42.060 --> 00:41:44.900 align:middle line:90%
Does that make sense?

00:41:44.900 --> 00:41:47.180 align:middle line:90%
All right.

00:41:47.180 --> 00:41:48.320 align:middle line:90%
We need another assumption.

00:41:48.320 --> 00:41:53.180 align:middle line:84%
We need an assumption about
what of excess risk of cancer

00:41:53.180 --> 00:41:55.660 align:middle line:90%
is per unit radiation exposure.

00:41:55.660 --> 00:41:59.850 align:middle line:84%
So let's adopt the excess
relative risk of one

00:41:59.850 --> 00:42:02.250 align:middle line:90%
per sievert.

00:42:02.250 --> 00:42:07.610 align:middle line:84%
Now remember that the
ICRP recommended 0.53.

00:42:07.610 --> 00:42:10.290 align:middle line:84%
But we are going
to take one, which

00:42:10.290 --> 00:42:14.290 align:middle line:84%
is to assume that the
effective radiation is stronger

00:42:14.290 --> 00:42:17.730 align:middle line:84%
than the consensus view,
which makes it easier

00:42:17.730 --> 00:42:20.190 align:middle line:84%
to see the effects of a little
bit of dose of radiation

00:42:20.190 --> 00:42:22.810 align:middle line:84%
and makes it artificially
easy to estimate

00:42:22.810 --> 00:42:25.110 align:middle line:90%
the existence of a threshold.

00:42:25.110 --> 00:42:28.110 align:middle line:84%
So both of these are biased
in the same direction.

00:42:28.110 --> 00:42:31.690 align:middle line:84%
It's making the experiment
easier than it really is.

00:42:31.690 --> 00:42:35.410 align:middle line:84%
And three, we're going to assume
perfect experimental conditions.

00:42:35.410 --> 00:42:37.930 align:middle line:84%
You say we can have
any population.

00:42:37.930 --> 00:42:40.630 align:middle line:84%
We'll know the dose that
everyone gets perfectly.

00:42:40.630 --> 00:42:43.850 align:middle line:84%
There's not going to be
any uncertainties there.

00:42:43.850 --> 00:42:46.730 align:middle line:84%
And we'll know all the
control parameters exactly.

00:42:46.730 --> 00:42:51.410 align:middle line:84%
We don't need a
fancy control group.

00:42:51.410 --> 00:42:54.490 align:middle line:84%
We don't have confounders from
people smoking and other stuff

00:42:54.490 --> 00:42:55.860 align:middle line:90%
noising up the data.

00:42:55.860 --> 00:42:59.783 align:middle line:84%
We have exactly the right
number of identical individuals,

00:42:59.783 --> 00:43:01.200 align:middle line:84%
and we're going
to all expose them

00:43:01.200 --> 00:43:04.480 align:middle line:84%
to the exact unambiguous
amount of dose,

00:43:04.480 --> 00:43:08.000 align:middle line:84%
which, again, is making the
experiment artificially easy.

00:43:08.000 --> 00:43:11.880 align:middle line:84%
So we have made all of these
generous assumptions that

00:43:11.880 --> 00:43:16.000 align:middle line:84%
will cause us to underestimate
the difficulty of actually

00:43:16.000 --> 00:43:18.040 align:middle line:90%
looking for a threshold.

00:43:18.040 --> 00:43:18.540 align:middle line:90%
All right.

00:43:18.540 --> 00:43:21.200 align:middle line:90%
So now we have some hypotheses.

00:43:21.200 --> 00:43:25.780 align:middle line:84%
And we're going to set the null
hypothesis to the threshold.

00:43:25.780 --> 00:43:28.800 align:middle line:84%
We'll say the null hypothesis
is that at one millisievert,

00:43:28.800 --> 00:43:30.040 align:middle line:90%
the dose is zero.

00:43:30.040 --> 00:43:31.600 align:middle line:90%
The effect is zero.

00:43:31.600 --> 00:43:33.620 align:middle line:84%
At 1 millisievert dose,
the effect is zero.

00:43:33.620 --> 00:43:34.840 align:middle line:90%
That's our null.

00:43:34.840 --> 00:43:36.562 align:middle line:90%
That's the threshold model.

00:43:36.562 --> 00:43:38.520 align:middle line:84%
And then we'll have this
alternative hypothesis

00:43:38.520 --> 00:43:41.960 align:middle line:84%
that says that the relative
risk at one millisievert

00:43:41.960 --> 00:43:47.040 align:middle line:84%
will be as predicted by the
linear model, one per sievert.

00:43:47.040 --> 00:43:49.600 align:middle line:90%
So those are our two hypotheses.

00:43:49.600 --> 00:43:50.280 align:middle line:90%
OK.

00:43:50.280 --> 00:43:53.400 align:middle line:84%
So let's see what
we would expect

00:43:53.400 --> 00:43:58.010 align:middle line:84%
will assume to make the
math easy incorrectly,

00:43:58.010 --> 00:44:03.790 align:middle line:84%
that these are Gaussians
which enlarge n is

00:44:03.790 --> 00:44:06.630 align:middle line:90%
a reasonable approximation.

00:44:06.630 --> 00:44:11.730 align:middle line:84%
And so here we have
the null hypothesis,

00:44:11.730 --> 00:44:13.910 align:middle line:90%
which is the threshold model.

00:44:13.910 --> 00:44:17.470 align:middle line:84%
We expect that the
population sample

00:44:17.470 --> 00:44:23.950 align:middle line:84%
will have a mean at the
baseline cancer risk of 10%.

00:44:23.950 --> 00:44:25.630 align:middle line:84%
Because at one
millisievert, there's

00:44:25.630 --> 00:44:26.930 align:middle line:90%
no effect from radiation.

00:44:26.930 --> 00:44:29.190 align:middle line:84%
So you just get
the baseline risk.

00:44:29.190 --> 00:44:32.950 align:middle line:84%
And then there will be
some statistical variation

00:44:32.950 --> 00:44:36.030 align:middle line:90%
around that expected mean.

00:44:36.030 --> 00:44:40.230 align:middle line:84%
If hypothesis one is correct,
then we expect the background

00:44:40.230 --> 00:44:45.070 align:middle line:84%
rate plus the excess risk,
which is the background rate--

00:44:45.070 --> 00:44:46.830 align:middle line:90%
remember because it's relative--

00:44:46.830 --> 00:44:51.030 align:middle line:84%
times 1 per sievert
times 1 millisievert.

00:44:51.030 --> 00:44:54.420 align:middle line:90%
So it's going to be 0.1001.

00:44:54.420 --> 00:44:57.360 align:middle line:84%
So that's where we expect the
mean of the distribution to be.

00:44:57.360 --> 00:45:01.060 align:middle line:84%
And again, we'll have
some variance around that.

00:45:01.060 --> 00:45:03.700 align:middle line:84%
So the question is,
how big of a sample

00:45:03.700 --> 00:45:08.580 align:middle line:84%
do we need so that the
uncertainty is squished

00:45:08.580 --> 00:45:11.740 align:middle line:84%
sufficiently small
that we can tell

00:45:11.740 --> 00:45:14.620 align:middle line:84%
the difference between
these two distributions.

00:45:14.620 --> 00:45:17.700 align:middle line:90%
That's the test of what we need.

00:45:17.700 --> 00:45:20.140 align:middle line:90%
So we need to choose two things.

00:45:20.140 --> 00:45:23.860 align:middle line:84%
We need to choose this critical
value, which will be our cutoff.

00:45:23.860 --> 00:45:28.820 align:middle line:84%
And we need to choose the sample
size, which will determine

00:45:28.820 --> 00:45:31.380 align:middle line:90%
the width of this distribution.

00:45:31.380 --> 00:45:36.140 align:middle line:84%
So the cutoff is basically where
we decide to accept or reject

00:45:36.140 --> 00:45:39.780 align:middle line:90%
the null hypothesis.

00:45:39.780 --> 00:45:43.660 align:middle line:84%
If we see a cancer
rate above this level,

00:45:43.660 --> 00:45:45.100 align:middle line:90%
we'll reject the null.

00:45:45.100 --> 00:45:49.260 align:middle line:84%
Otherwise, we're going to assume
that the null hypothesis, which

00:45:49.260 --> 00:45:52.810 align:middle line:84%
is in this case is the
threshold model is correct.

00:45:52.810 --> 00:45:54.930 align:middle line:84%
So we are being very
generous because what

00:45:54.930 --> 00:46:00.090 align:middle line:84%
we've done is we've used
standard 95% confidence here

00:46:00.090 --> 00:46:05.690 align:middle line:84%
and 80% power, which is to say
that we will accidentally reject

00:46:05.690 --> 00:46:11.010 align:middle line:90%
the null only 5% of the time.

00:46:11.010 --> 00:46:16.250 align:middle line:84%
And by setting 80% power
for the other test,

00:46:16.250 --> 00:46:23.210 align:middle line:84%
it's OK if we reject the LNT
theory at least 20% of the time

00:46:23.210 --> 00:46:25.250 align:middle line:90%
that it was correct.

00:46:25.250 --> 00:46:28.170 align:middle line:84%
We don't really care about
getting the LNT part wrong,

00:46:28.170 --> 00:46:32.690 align:middle line:84%
but we want to be very careful
not to accidentally reject

00:46:32.690 --> 00:46:34.703 align:middle line:84%
the threshold model
as being wrong.

00:46:34.703 --> 00:46:36.870 align:middle line:84%
So we're being very generous
to the threshold model.

00:46:36.870 --> 00:46:38.790 align:middle line:84%
Normally, you would do
this in the reverse.

00:46:38.790 --> 00:46:41.770 align:middle line:84%
You would have a higher
standard for the threshold model

00:46:41.770 --> 00:46:44.130 align:middle line:84%
because it's a more
complicated model.

00:46:44.130 --> 00:46:46.290 align:middle line:84%
But here we're going to
be generous and assume

00:46:46.290 --> 00:46:50.080 align:middle line:84%
the threshold model
is the notional model.

00:46:50.080 --> 00:46:54.040 align:middle line:84%
So now we need to
basically write

00:46:54.040 --> 00:46:57.280 align:middle line:84%
what I've drawn here
in some equations

00:46:57.280 --> 00:47:01.360 align:middle line:84%
with these type I
and type II error.

00:47:01.360 --> 00:47:06.000 align:middle line:84%
And so this is just the
z-test, as you recall.

00:47:06.000 --> 00:47:08.640 align:middle line:90%
Here's the definition.

00:47:08.640 --> 00:47:11.400 align:middle line:84%
Z is the number of standard
deviations away from the mean.

00:47:11.400 --> 00:47:16.080 align:middle line:90%
X-bar is your sample mean.

00:47:16.080 --> 00:47:17.960 align:middle line:90%
Mu is what you expect.

00:47:17.960 --> 00:47:22.520 align:middle line:84%
Sigma over root n is
your standard deviation

00:47:22.520 --> 00:47:25.880 align:middle line:90%
of your sample.

00:47:25.880 --> 00:47:28.240 align:middle line:84%
Just some criteria--
this technically

00:47:28.240 --> 00:47:31.220 align:middle line:90%
requires a normal distribution.

00:47:31.220 --> 00:47:32.720 align:middle line:84%
We're using the
central limit theory

00:47:32.720 --> 00:47:37.280 align:middle line:84%
to basically say that as long
as the sample size is large,

00:47:37.280 --> 00:47:40.120 align:middle line:84%
the number of people that
we're going to test is large.

00:47:40.120 --> 00:47:41.952 align:middle line:90%
Over 30-ish, this will hold.

00:47:41.952 --> 00:47:43.160 align:middle line:90%
We'll go back and check this.

00:47:43.160 --> 00:47:46.338 align:middle line:84%
When we do the math, if we get a
number less than 30 individuals,

00:47:46.338 --> 00:47:47.880 align:middle line:84%
then we'll say this
was a bad choice.

00:47:47.880 --> 00:47:51.720 align:middle line:90%
And we'll see what it is.

00:47:51.720 --> 00:47:54.600 align:middle line:90%
So let's just put in the math.

00:47:54.600 --> 00:47:57.120 align:middle line:84%
For type I error,
you can look this up.

00:47:57.120 --> 00:48:01.280 align:middle line:84%
If you want a 95% confidence
of accepting the model,

00:48:01.280 --> 00:48:06.020 align:middle line:84%
or only a 5% chance of wrongly
rejecting it is 1.5 sigma out.

00:48:06.020 --> 00:48:07.380 align:middle line:90%
So there's your equation.

00:48:07.380 --> 00:48:10.120 align:middle line:84%
There's the equation
for type II error

00:48:10.120 --> 00:48:15.400 align:middle line:84%
on these other
distribution -0.84.

00:48:15.400 --> 00:48:18.620 align:middle line:84%
And we have two equations
and two unknowns.

00:48:18.620 --> 00:48:20.500 align:middle line:84%
And so we do the
math, and we solve it.

00:48:20.500 --> 00:48:24.200 align:middle line:84%
And we see the
threshold is 1.0007.

00:48:24.200 --> 00:48:31.400 align:middle line:84%
Anything above this, we'll
reject the linear model

00:48:31.400 --> 00:48:34.320 align:middle line:90%
at 56 million people.

00:48:34.320 --> 00:48:35.900 align:middle line:84%
And actually, to
be more precise,

00:48:35.900 --> 00:48:39.460 align:middle line:84%
56 million people
is the minimum size.

00:48:39.460 --> 00:48:45.540 align:middle line:84%
If your observed incidence of
cancer rate deviates from this,

00:48:45.540 --> 00:48:50.110 align:middle line:84%
you will need actually
potentially a larger population.

00:48:50.110 --> 00:48:55.710 align:middle line:90%
So here's the issue.

00:48:55.710 --> 00:49:01.110 align:middle line:84%
Everything we've done favors,
underestimates the difficulty

00:49:01.110 --> 00:49:04.870 align:middle line:90%
of seeing the threshold.

00:49:04.870 --> 00:49:06.990 align:middle line:84%
And the question is,
can you do an experiment

00:49:06.990 --> 00:49:13.030 align:middle line:84%
with 56 identically
humans and give them

00:49:13.030 --> 00:49:16.320 align:middle line:84%
all one millisievert
of radiation

00:49:16.320 --> 00:49:17.570 align:middle line:90%
in order to see the threshold?

00:49:17.570 --> 00:49:21.950 align:middle line:84%
That is an experiment you
probably cannot practically do.

00:49:21.950 --> 00:49:23.550 align:middle line:90%
That's the point.

00:49:23.550 --> 00:49:25.910 align:middle line:84%
We'll never have
this experiment.

00:49:25.910 --> 00:49:28.290 align:middle line:90%
And this is really optimistic.

00:49:28.290 --> 00:49:31.910 align:middle line:90%


00:49:31.910 --> 00:49:36.190 align:middle line:84%
If the threshold does exist,
it is extremely difficult

00:49:36.190 --> 00:49:39.070 align:middle line:90%
to show that it exists.

00:49:39.070 --> 00:49:43.910 align:middle line:84%
So since we can't
prove or disprove

00:49:43.910 --> 00:49:47.540 align:middle line:84%
the existence of a threshold
with any reasonable experiment,

00:49:47.540 --> 00:49:48.440 align:middle line:90%
what happens?

00:49:48.440 --> 00:49:51.300 align:middle line:90%


00:49:51.300 --> 00:49:55.140 align:middle line:84%
People just choose to believe
what they want to believe,

00:49:55.140 --> 00:49:56.740 align:middle line:90%
but they're not supposed to.

00:49:56.740 --> 00:50:00.220 align:middle line:84%
They're supposed to
follow Occam's razor

00:50:00.220 --> 00:50:04.420 align:middle line:84%
and choose the model that has
the best predictive power given

00:50:04.420 --> 00:50:05.860 align:middle line:90%
what we know.

00:50:05.860 --> 00:50:09.820 align:middle line:84%
So that's the story
between threshold models

00:50:09.820 --> 00:50:14.100 align:middle line:84%
and no-threshold models, is
that the no-threshold is not

00:50:14.100 --> 00:50:19.420 align:middle line:84%
defensible because it is
a higher complexity model,

00:50:19.420 --> 00:50:23.128 align:middle line:84%
but because you can't prove
the threshold people wrong,

00:50:23.128 --> 00:50:25.420 align:middle line:84%
there will be people who will
insist on using threshold

00:50:25.420 --> 00:50:26.240 align:middle line:90%
no matter what.

00:50:26.240 --> 00:50:28.940 align:middle line:90%


00:50:28.940 --> 00:50:29.600 align:middle line:90%
All right.

00:50:29.600 --> 00:50:33.780 align:middle line:84%
Now, could we actually
do better than this?

00:50:33.780 --> 00:50:35.280 align:middle line:90%
And the answer is yes.

00:50:35.280 --> 00:50:38.460 align:middle line:84%
If we take a
mechanistic perspective,

00:50:38.460 --> 00:50:42.460 align:middle line:84%
we can learn something about
how radiation biology works

00:50:42.460 --> 00:50:47.090 align:middle line:84%
and inform us as to whether we
think there might be a threshold

00:50:47.090 --> 00:50:47.890 align:middle line:90%
or not.

00:50:47.890 --> 00:50:51.370 align:middle line:84%
So this is basically
saying, now we're

00:50:51.370 --> 00:50:58.130 align:middle line:84%
going to allow this term here to
be something other than unity.

00:50:58.130 --> 00:51:03.210 align:middle line:84%
We're going to inject a
prior into the model based

00:51:03.210 --> 00:51:06.370 align:middle line:90%
on some additional information.

00:51:06.370 --> 00:51:11.010 align:middle line:84%
So ready to learn
how cancer happens?

00:51:11.010 --> 00:51:11.590 align:middle line:90%
All right.

00:51:11.590 --> 00:51:15.390 align:middle line:84%
So you have some amount
of radiation coming in.

00:51:15.390 --> 00:51:18.530 align:middle line:84%
And one of the things
it does is it hits water

00:51:18.530 --> 00:51:23.890 align:middle line:84%
because you are mostly water,
as usually what happens.

00:51:23.890 --> 00:51:24.990 align:middle line:90%
And so what does it do?

00:51:24.990 --> 00:51:27.890 align:middle line:84%
It ionizes the water,
kicks off an electron,

00:51:27.890 --> 00:51:31.490 align:middle line:90%
and it produces hydronium.

00:51:31.490 --> 00:51:37.530 align:middle line:84%
And eventually, you get
OH hydronium and OH dot.

00:51:37.530 --> 00:51:39.570 align:middle line:90%
And this is a free radical.

00:51:39.570 --> 00:51:46.680 align:middle line:84%
And it goes and attacks one
of the sugars on your DNA

00:51:46.680 --> 00:51:49.000 align:middle line:90%
and causes a problem.

00:51:49.000 --> 00:51:51.140 align:middle line:90%
So this is usually what happens.

00:51:51.140 --> 00:51:54.440 align:middle line:84%
And this is happening all
the time in your body,

00:51:54.440 --> 00:51:55.620 align:middle line:90%
not just from radiation.

00:51:55.620 --> 00:51:58.760 align:middle line:84%
Also, chemicals cause
free radicals and do this,

00:51:58.760 --> 00:52:00.780 align:middle line:84%
and your body has
to deal with it.

00:52:00.780 --> 00:52:04.000 align:middle line:90%


00:52:04.000 --> 00:52:05.640 align:middle line:84%
Another thing that
you can have is

00:52:05.640 --> 00:52:09.200 align:middle line:84%
that you could have the
radiation, either the photon

00:52:09.200 --> 00:52:12.640 align:middle line:84%
or an electron,
directly interact

00:52:12.640 --> 00:52:20.280 align:middle line:84%
with an atom in the chain in
your DNA, the non-water atom,

00:52:20.280 --> 00:52:24.240 align:middle line:84%
and cause that molecule
to break apart.

00:52:24.240 --> 00:52:26.720 align:middle line:84%
And this is where
things get interesting.

00:52:26.720 --> 00:52:31.500 align:middle line:84%
So recall that it's not really
like the total energy deposit.

00:52:31.500 --> 00:52:33.560 align:middle line:84%
It's how and where
it's deposited.

00:52:33.560 --> 00:52:36.900 align:middle line:84%
So I showed you
this on Wednesday.

00:52:36.900 --> 00:52:39.600 align:middle line:84%
And what I've done now is I've
superimposed in approximately

00:52:39.600 --> 00:52:42.670 align:middle line:90%
to scale the size of a cell.

00:52:42.670 --> 00:52:47.310 align:middle line:84%
And you can see here that
for low LET radiation,

00:52:47.310 --> 00:52:49.870 align:middle line:84%
you can have lots
of interactions

00:52:49.870 --> 00:52:52.270 align:middle line:90%
on the scale of the nucleus.

00:52:52.270 --> 00:52:57.790 align:middle line:84%
And so it is possible
to have radiation.

00:52:57.790 --> 00:53:01.470 align:middle line:84%
A single radioactive particle, a
single electron or single photon

00:53:01.470 --> 00:53:04.950 align:middle line:84%
enter the body and undergo
photoelectric effect and Compton

00:53:04.950 --> 00:53:08.230 align:middle line:84%
scattering and cause all
of these subsequent changes

00:53:08.230 --> 00:53:10.830 align:middle line:90%
multiple times within a cell.

00:53:10.830 --> 00:53:14.290 align:middle line:84%
And this may generate multiple
free radicals from water,

00:53:14.290 --> 00:53:18.470 align:middle line:84%
these OH dot ions, or it
may directly attack the DNA.

00:53:18.470 --> 00:53:25.830 align:middle line:84%
And so the probability of
getting multiple breaks

00:53:25.830 --> 00:53:31.650 align:middle line:84%
to the genome goes up because
of the nature of radiation,

00:53:31.650 --> 00:53:34.690 align:middle line:84%
because of this
low LET radiation.

00:53:34.690 --> 00:53:37.230 align:middle line:84%
So this is looking
at different kinds

00:53:37.230 --> 00:53:38.410 align:middle line:90%
of breaks that might happen.

00:53:38.410 --> 00:53:42.460 align:middle line:84%
The top one is what we get when
we have that OH dot showing

00:53:42.460 --> 00:53:44.060 align:middle line:84%
up just at some
random place, and it

00:53:44.060 --> 00:53:45.400 align:middle line:90%
breaks one side of the DNA.

00:53:45.400 --> 00:53:46.900 align:middle line:84%
And then we have
this process called

00:53:46.900 --> 00:53:48.580 align:middle line:90%
non-homologous end joining.

00:53:48.580 --> 00:53:50.180 align:middle line:84%
And it's like a
little thing that

00:53:50.180 --> 00:53:53.460 align:middle line:84%
comes along and detects
the break and glues

00:53:53.460 --> 00:53:55.460 align:middle line:90%
it back together.

00:53:55.460 --> 00:53:59.020 align:middle line:84%
And it turns out we
have one of these

00:53:59.020 --> 00:54:05.500 align:middle line:84%
happen every half-second
per cell in our body.

00:54:05.500 --> 00:54:06.875 align:middle line:90%
It's happening all the time.

00:54:06.875 --> 00:54:09.000 align:middle line:84%
All the cells breaking,
breaking, fixing, breaking,

00:54:09.000 --> 00:54:11.020 align:middle line:84%
fixing, breaking,
fixing all the time.

00:54:11.020 --> 00:54:13.980 align:middle line:90%
It's amazing.

00:54:13.980 --> 00:54:20.700 align:middle line:84%
But this kind of break where
we have a double-strand break

00:54:20.700 --> 00:54:24.380 align:middle line:90%
is pretty rare.

00:54:24.380 --> 00:54:27.980 align:middle line:84%
That just two happen to
happen in close proximity

00:54:27.980 --> 00:54:29.060 align:middle line:90%
to each other.

00:54:29.060 --> 00:54:32.500 align:middle line:84%
It's something like
once every-- turns out

00:54:32.500 --> 00:54:35.460 align:middle line:90%
this is like every two hours.

00:54:35.460 --> 00:54:38.130 align:middle line:90%
So it still happens.

00:54:38.130 --> 00:54:39.890 align:middle line:84%
And you still have
non-homologous end

00:54:39.890 --> 00:54:42.050 align:middle line:84%
joining coming and
gluing this together.

00:54:42.050 --> 00:54:46.530 align:middle line:84%
But sometimes you get
something like this.

00:54:46.530 --> 00:54:47.830 align:middle line:90%
And what is the end?

00:54:47.830 --> 00:54:52.107 align:middle line:84%
This thing floats away, and the
enzyme comes along and goes,

00:54:52.107 --> 00:54:53.690 align:middle line:84%
uhh. and then it
cuts these things off

00:54:53.690 --> 00:54:58.250 align:middle line:84%
and it glues it together and
just takes out a few base pairs.

00:54:58.250 --> 00:55:00.610 align:middle line:84%
And this happens all
the time in your genome.

00:55:00.610 --> 00:55:06.850 align:middle line:84%
And sometimes it's benign,
and sometimes it's not.

00:55:06.850 --> 00:55:11.610 align:middle line:84%
And when it's not,
you get cancer.

00:55:11.610 --> 00:55:15.130 align:middle line:90%
So here's the key takeaway.

00:55:15.130 --> 00:55:19.450 align:middle line:84%
The repair process is what's
giving you the cancer.

00:55:19.450 --> 00:55:23.145 align:middle line:84%
If you simply break the genome
and you don't repair it,

00:55:23.145 --> 00:55:24.770 align:middle line:84%
then generally what
happens is the cell

00:55:24.770 --> 00:55:28.450 align:middle line:84%
becomes senescent and
doesn't continue to operate

00:55:28.450 --> 00:55:31.770 align:middle line:90%
and just dies generally.

00:55:31.770 --> 00:55:34.890 align:middle line:84%
But if you go and you
introduce these genomic errors

00:55:34.890 --> 00:55:36.880 align:middle line:90%
from this repair process.

00:55:36.880 --> 00:55:38.680 align:middle line:90%
That's when you get cancer.

00:55:38.680 --> 00:55:45.080 align:middle line:84%
So sometimes you hear people who
have seen the idea of hormesis,

00:55:45.080 --> 00:55:47.320 align:middle line:84%
that a little bit of
radiation is good for you,

00:55:47.320 --> 00:55:49.860 align:middle line:84%
and they are looking for
ways to justify that.

00:55:49.860 --> 00:55:54.720 align:middle line:84%
Remember, most likely that is
actually the look again effect

00:55:54.720 --> 00:55:56.220 align:middle line:84%
that they've picked
up on initially.

00:55:56.220 --> 00:55:58.060 align:middle line:84%
And they say, well what
would describe this?

00:55:58.060 --> 00:56:00.160 align:middle line:84%
And they'll say,
oh, it turns out

00:56:00.160 --> 00:56:01.940 align:middle line:84%
that if you give
a cell radiation,

00:56:01.940 --> 00:56:05.352 align:middle line:84%
we can observe that the
repair process is upregulated.

00:56:05.352 --> 00:56:06.060 align:middle line:90%
Of course, it is.

00:56:06.060 --> 00:56:07.435 align:middle line:84%
You've done damage
to the genome.

00:56:07.435 --> 00:56:09.960 align:middle line:84%
It's going to have
more repair to do.

00:56:09.960 --> 00:56:12.240 align:middle line:84%
But then they think oh, but
maybe that repair process

00:56:12.240 --> 00:56:14.740 align:middle line:84%
is preventing
background cancers.

00:56:14.740 --> 00:56:17.120 align:middle line:90%
It's like no.

00:56:17.120 --> 00:56:20.600 align:middle line:84%
That repair process is
causing the cancers.

00:56:20.600 --> 00:56:25.880 align:middle line:84%
And so this is a
key insight here.

00:56:25.880 --> 00:56:28.200 align:middle line:84%
The other insight I want
you to take away this

00:56:28.200 --> 00:56:32.360 align:middle line:84%
is possible from
a single photon--

00:56:32.360 --> 00:56:33.760 align:middle line:90%
comes in, it comes through.

00:56:33.760 --> 00:56:35.090 align:middle line:90%
You get Compton scatters.

00:56:35.090 --> 00:56:36.010 align:middle line:90%
You get an electron.

00:56:36.010 --> 00:56:38.090 align:middle line:84%
You get another photon,
another Compton.

00:56:38.090 --> 00:56:39.130 align:middle line:90%
Two electrons.

00:56:39.130 --> 00:56:40.190 align:middle line:90%
Boom, boom, boom.

00:56:40.190 --> 00:56:44.390 align:middle line:90%
This is technically possible.

00:56:44.390 --> 00:56:51.190 align:middle line:84%
Therefore, with some very,
very, very small probability,

00:56:51.190 --> 00:56:55.150 align:middle line:84%
a single photon could
give you cancer--

00:56:55.150 --> 00:56:57.172 align:middle line:84%
very very, very
small, not normal.

00:56:57.172 --> 00:56:58.630 align:middle line:84%
Normally, it requires
multiple hits

00:56:58.630 --> 00:57:01.310 align:middle line:84%
to the genome and multiple
places, et cetera.

00:57:01.310 --> 00:57:05.750 align:middle line:84%
But if there's a very
small non-zero probability,

00:57:05.750 --> 00:57:07.410 align:middle line:90%
a threshold cannot exist.

00:57:07.410 --> 00:57:10.710 align:middle line:90%


00:57:10.710 --> 00:57:14.590 align:middle line:84%
So this is the
mechanistic argument

00:57:14.590 --> 00:57:17.070 align:middle line:84%
for why we don't have
a threshold model.

00:57:17.070 --> 00:57:19.200 align:middle line:90%
Yes.

00:57:19.200 --> 00:57:20.950 align:middle line:84%
AUDIENCE: What energy
photons are required

00:57:20.950 --> 00:57:22.050 align:middle line:90%
to do a damage like that?

00:57:22.050 --> 00:57:24.790 align:middle line:90%
Is there a threshold?

00:57:24.790 --> 00:57:26.710 align:middle line:84%
SCOTT KEMP: You need
to deposit at least

00:57:26.710 --> 00:57:32.580 align:middle line:84%
basically an EV in the
interaction to ionize something.

00:57:32.580 --> 00:57:36.580 align:middle line:84%
And typically, these photons
have hundreds of hundreds

00:57:36.580 --> 00:57:38.460 align:middle line:90%
of keV of energy--

00:57:38.460 --> 00:57:39.720 align:middle line:90%
so plenty of energy.

00:57:39.720 --> 00:57:40.220 align:middle line:90%
Yeah.

00:57:40.220 --> 00:57:44.220 align:middle line:90%


00:57:44.220 --> 00:57:45.180 align:middle line:90%
OK.

00:57:45.180 --> 00:57:49.920 align:middle line:84%
So we do live in a
radioactive environment.

00:57:49.920 --> 00:57:53.220 align:middle line:84%
Our bodies do have
these repair mechanisms.

00:57:53.220 --> 00:57:56.020 align:middle line:84%
There are error-proof
repair mechanisms.

00:57:56.020 --> 00:58:03.580 align:middle line:84%
There's something called, sorry,
homologous recombination, which

00:58:03.580 --> 00:58:06.580 align:middle line:84%
is a way of
error-checking the genome.

00:58:06.580 --> 00:58:12.620 align:middle line:84%
But in human cells, it only
exists during cell division.

00:58:12.620 --> 00:58:15.900 align:middle line:84%
In the normal growth
phase of the cell,

00:58:15.900 --> 00:58:17.640 align:middle line:84%
we only have L1
copy of the genome,

00:58:17.640 --> 00:58:20.460 align:middle line:84%
so we have nothing to
fact-check it against.

00:58:20.460 --> 00:58:23.660 align:middle line:84%
Certain bacteria, like
Deinococcus radiodurans,

00:58:23.660 --> 00:58:25.400 align:middle line:84%
has multiple copies
of its genome.

00:58:25.400 --> 00:58:27.340 align:middle line:84%
And every time it fixes
it, it checks repair

00:58:27.340 --> 00:58:28.720 align:middle line:90%
against the other copies.

00:58:28.720 --> 00:58:32.220 align:middle line:84%
And so it makes it extremely
robust against large doses

00:58:32.220 --> 00:58:33.280 align:middle line:90%
of radiation.

00:58:33.280 --> 00:58:35.380 align:middle line:90%
But human cells, only once--

00:58:35.380 --> 00:58:37.740 align:middle line:84%
only copy, except
during cell division.

00:58:37.740 --> 00:58:41.320 align:middle line:84%
And so we have no
way of enduring

00:58:41.320 --> 00:58:43.080 align:middle line:90%
those radioactive assaults.

00:58:43.080 --> 00:58:48.360 align:middle line:84%
So yes, we are evolved to live
in the radioactive environment.

00:58:48.360 --> 00:58:52.520 align:middle line:84%
We have repair processes
that keep the cells going,

00:58:52.520 --> 00:58:53.700 align:middle line:90%
keep them alive.

00:58:53.700 --> 00:58:57.460 align:middle line:84%
But it doesn't guarantee us
that the cells don't eventually

00:58:57.460 --> 00:58:58.500 align:middle line:90%
get cancer.

00:58:58.500 --> 00:59:03.980 align:middle line:84%
It only guarantees that on
average, most people will

00:59:03.980 --> 00:59:06.900 align:middle line:90%
live long enough to procreate.

00:59:06.900 --> 00:59:10.300 align:middle line:84%
And that's all
evolution has given us.

00:59:10.300 --> 00:59:16.180 align:middle line:84%
So little bits of radiation
do matter very likely.

00:59:16.180 --> 00:59:19.340 align:middle line:84%
And we shouldn't
be seduced by any

00:59:19.340 --> 00:59:22.780 align:middle line:84%
of the ideas that
radioactive environment,

00:59:22.780 --> 00:59:25.440 align:middle line:84%
and we're accustomed to
this, cetera, cetera.

00:59:25.440 --> 00:59:28.860 align:middle line:90%


00:59:28.860 --> 00:59:30.410 align:middle line:90%
That's it.

00:59:30.410 --> 00:59:33.330 align:middle line:84%
That's why LNT is
right, or very likely.

00:59:33.330 --> 00:59:34.680 align:middle line:90%
Yep.

00:59:34.680 --> 00:59:36.230 align:middle line:90%
AUDIENCE: Here's a question.

00:59:36.230 --> 00:59:40.730 align:middle line:84%
So if the cluster
double-strand breaks,

00:59:40.730 --> 00:59:44.930 align:middle line:84%
if the repair mechanism
happens to repair

00:59:44.930 --> 00:59:50.170 align:middle line:84%
in a way that's
malignant, is there

00:59:50.170 --> 00:59:55.170 align:middle line:84%
a threshold for a certain amount
of those dangerous repairs

00:59:55.170 --> 00:59:55.830 align:middle line:90%
have to happen?

00:59:55.830 --> 00:59:58.390 align:middle line:84%
Because I'm assuming
if probabilistically,

00:59:58.390 --> 01:00:01.130 align:middle line:90%
there's been cancerous repairs.

01:00:01.130 --> 01:00:02.850 align:middle line:90%
SCOTT KEMP: Yep.

01:00:02.850 --> 01:00:03.350 align:middle line:90%
Yeah.

01:00:03.350 --> 01:00:07.610 align:middle line:84%
So they call them
promoters or potentiators.

01:00:07.610 --> 01:00:10.070 align:middle line:84%
So I believe-- and this
is going from memory,

01:00:10.070 --> 01:00:11.570 align:middle line:84%
so you should
double-check this--

01:00:11.570 --> 01:00:13.370 align:middle line:84%
but I believe the
general consensus

01:00:13.370 --> 01:00:22.170 align:middle line:84%
is that you need on average
about six or so of these things

01:00:22.170 --> 01:00:24.730 align:middle line:84%
to get the cell really going
into the cancerous mode.

01:00:24.730 --> 01:00:27.505 align:middle line:90%
AUDIENCE: In a specific cell?

01:00:27.505 --> 01:00:29.140 align:middle line:84%
SCOTT KEMP: In a
specific genome.

01:00:29.140 --> 01:00:29.680 align:middle line:90%
Yeah.

01:00:29.680 --> 01:00:32.500 align:middle line:84%
That doesn't mean you
actually need six.

01:00:32.500 --> 01:00:37.200 align:middle line:84%
If the damage happens in the
right region of the genome,

01:00:37.200 --> 01:00:40.840 align:middle line:84%
you in fact can cause
problems right away.

01:00:40.840 --> 01:00:42.963 align:middle line:84%
You can also, if
the damage happens

01:00:42.963 --> 01:00:44.380 align:middle line:84%
in the right region
of the genome,

01:00:44.380 --> 01:00:46.320 align:middle line:84%
you can also just
instantly kill the cell,

01:00:46.320 --> 01:00:47.420 align:middle line:90%
which is not a problem.

01:00:47.420 --> 01:00:49.160 align:middle line:90%
Cell die is fine.

01:00:49.160 --> 01:00:53.960 align:middle line:84%
But on average, it does need
to happen more than once.

01:00:53.960 --> 01:00:56.840 align:middle line:84%
But remember, it is happening
all the time in your body.

01:00:56.840 --> 01:01:01.520 align:middle line:84%
So you are accumulating
these potentiated cells

01:01:01.520 --> 01:01:03.560 align:middle line:84%
over your lifetime, which
is why you typically

01:01:03.560 --> 01:01:07.320 align:middle line:90%
get cancer when you're old.

01:01:07.320 --> 01:01:10.940 align:middle line:84%
And this is also why the
relative risk model makes sense.

01:01:10.940 --> 01:01:14.560 align:middle line:84%
The idea that a smoker
has a higher probability

01:01:14.560 --> 01:01:18.460 align:middle line:84%
of getting cancer from
radiation makes sense,

01:01:18.460 --> 01:01:21.240 align:middle line:84%
because they've already
trashed their genome

01:01:21.240 --> 01:01:22.800 align:middle line:90%
from cigarette smoking.

01:01:22.800 --> 01:01:26.600 align:middle line:84%
So they've accumulated
these potentiated genomes.

01:01:26.600 --> 01:01:29.030 align:middle line:84%
So an additional
damage from radiation

01:01:29.030 --> 01:01:32.550 align:middle line:84%
is more likely to
actually cause the cancer.

01:01:32.550 --> 01:01:37.730 align:middle line:84%
So yeah, that is what's
likely going on in most cells.

01:01:37.730 --> 01:01:38.230 align:middle line:90%
Yeah.

01:01:38.230 --> 01:01:39.330 align:middle line:90%
AUDIENCE: That's a good point.

01:01:39.330 --> 01:01:40.240 align:middle line:90%
It's like you have these--

01:01:40.240 --> 01:01:41.907 align:middle line:84%
SCOTT KEMP: Just
depends on where it is.

01:01:41.907 --> 01:01:44.250 align:middle line:90%
It's a stochastic process.

01:01:44.250 --> 01:01:46.190 align:middle line:90%
It's not a threshold phenomenon.

01:01:46.190 --> 01:01:47.870 align:middle line:90%
It's stochastic.

01:01:47.870 --> 01:01:48.730 align:middle line:90%
Yeah.

01:01:48.730 --> 01:01:52.150 align:middle line:84%
AUDIENCE: What is the effect of
how long you get the radiation

01:01:52.150 --> 01:01:53.870 align:middle line:90%
dose over [INAUDIBLE]?

01:01:53.870 --> 01:01:55.630 align:middle line:90%
SCOTT KEMP: The dose rate.

01:01:55.630 --> 01:02:01.932 align:middle line:84%
Yeah, so there is evidence
of dose-rate effects.

01:02:01.932 --> 01:02:03.390 align:middle line:84%
If you want to read
about that, you

01:02:03.390 --> 01:02:06.710 align:middle line:84%
can go to our ICRP,
which is in the reading.

01:02:06.710 --> 01:02:10.470 align:middle line:84%
Generally, what the dose
rate effects show is like

01:02:10.470 --> 01:02:15.670 align:middle line:84%
if you increase the dose
rate, then it tips over,

01:02:15.670 --> 01:02:19.190 align:middle line:84%
and then it's to say that
the additional units of dose

01:02:19.190 --> 01:02:22.202 align:middle line:84%
at a high rate don't
have as big of an effect.

01:02:22.202 --> 01:02:23.410 align:middle line:90%
And that kind of makes sense.

01:02:23.410 --> 01:02:29.660 align:middle line:84%
You have a lot of
damages occurring,

01:02:29.660 --> 01:02:34.020 align:middle line:84%
and they're all happening before
the repair process happens

01:02:34.020 --> 01:02:35.220 align:middle line:90%
kind of a thing.

01:02:35.220 --> 01:02:40.280 align:middle line:84%
So that is the extent of
my understanding of it.

01:02:40.280 --> 01:02:41.960 align:middle line:90%
And you can go and look more.

01:02:41.960 --> 01:02:42.460 align:middle line:90%
Yeah.

01:02:42.460 --> 01:02:47.700 align:middle line:90%


01:02:47.700 --> 01:02:48.380 align:middle line:90%
Yep.

01:02:48.380 --> 01:02:50.500 align:middle line:84%
AUDIENCE: Any of
the nonlinear models

01:02:50.500 --> 01:02:54.580 align:middle line:84%
end up kind of converging
on setting those nonlinear

01:02:54.580 --> 01:02:55.900 align:middle line:90%
weights to 0?

01:02:55.900 --> 01:02:58.700 align:middle line:84%
I guess, do you see those models
converging on a linear model

01:02:58.700 --> 01:02:59.420 align:middle line:90%
still?

01:02:59.420 --> 01:03:03.460 align:middle line:84%
SCOTT KEMP: The
nonlinear models are

01:03:03.460 --> 01:03:10.100 align:middle line:84%
defensible once you allow the
high-dose regime like this.

01:03:10.100 --> 01:03:14.340 align:middle line:84%
Then they do converge
for very high doses.

01:03:14.340 --> 01:03:16.500 align:middle line:90%
Yeah.

01:03:16.500 --> 01:03:19.880 align:middle line:84%
For what we would still permit
in the lower-dose regime.

01:03:19.880 --> 01:03:22.500 align:middle line:84%
I think we now have
enough data on leukemia

01:03:22.500 --> 01:03:28.410 align:middle line:84%
to support a nonlinear
model on leukemia,

01:03:28.410 --> 01:03:32.170 align:middle line:84%
but they're basically all linear
in the extreme low-- and so

01:03:32.170 --> 01:03:34.610 align:middle line:90%
in the limit of ultra-low doses.

01:03:34.610 --> 01:03:36.030 align:middle line:90%
They all kind go like that.

01:03:36.030 --> 01:03:39.810 align:middle line:90%


01:03:39.810 --> 01:03:42.450 align:middle line:90%
Yep.

01:03:42.450 --> 01:03:44.430 align:middle line:84%
AUDIENCE: [INAUDIBLE]
why would it be?

01:03:44.430 --> 01:03:47.650 align:middle line:84%
You said earlier that
people in industry

01:03:47.650 --> 01:03:49.090 align:middle line:90%
would want this to be true.

01:03:49.090 --> 01:03:50.870 align:middle line:84%
Why is there such
a small effect?

01:03:50.870 --> 01:03:51.785 align:middle line:90%
Why would it be--

01:03:51.785 --> 01:03:52.410 align:middle line:90%
SCOTT KEMP: Oh.

01:03:52.410 --> 01:03:55.890 align:middle line:84%
Well, maybe we can have an
answer from the audience.

01:03:55.890 --> 01:03:57.810 align:middle line:84%
AUDIENCE: I mean, I'm
not really trying to--

01:03:57.810 --> 01:04:00.510 align:middle line:84%
I haven't really looked into
threshold models or not.

01:04:00.510 --> 01:04:03.090 align:middle line:84%
Just what would the
cost difference be?

01:04:03.090 --> 01:04:10.330 align:middle line:90%
And I'm going to tell you.

01:04:10.330 --> 01:04:11.550 align:middle line:90%
It's going to be small.

01:04:11.550 --> 01:04:14.070 align:middle line:90%


01:04:14.070 --> 01:04:18.710 align:middle line:84%
The purpose of the research is
to say that for nuclear power,

01:04:18.710 --> 01:04:20.130 align:middle line:90%
it does not matter.

01:04:20.130 --> 01:04:23.840 align:middle line:84%
Just focus on why the
containment is so damn expensive

01:04:23.840 --> 01:04:25.760 align:middle line:90%
and all these things.

01:04:25.760 --> 01:04:28.480 align:middle line:90%
That's essentially what I do.

01:04:28.480 --> 01:04:32.020 align:middle line:84%
I'm not looking
into why threshold.

01:04:32.020 --> 01:04:33.060 align:middle line:90%
Why not threshold?

01:04:33.060 --> 01:04:35.480 align:middle line:84%
Obviously, I've done
my [INAUDIBLE] review.

01:04:35.480 --> 01:04:38.480 align:middle line:90%
But it's [INAUDIBLE]

01:04:38.480 --> 01:04:42.140 align:middle line:84%
SCOTT KEMP: But what's
motivating your study?

01:04:42.140 --> 01:04:45.180 align:middle line:84%
So you believe that it's likely
to be or have a small impact?

01:04:45.180 --> 01:04:47.200 align:middle line:84%
Bongiorno believes likely
to have a small impact

01:04:47.200 --> 01:04:48.900 align:middle line:90%
on the total cost of reactors.

01:04:48.900 --> 01:04:50.783 align:middle line:90%
I agree with that.

01:04:50.783 --> 01:04:52.200 align:middle line:84%
But there are
people out there who

01:04:52.200 --> 01:04:54.920 align:middle line:84%
don't believe that,
which is why you're doing

01:04:54.920 --> 01:04:56.420 align:middle line:90%
the study in the first place.

01:04:56.420 --> 01:04:57.420 align:middle line:90%
AUDIENCE: Yeah, exactly.

01:04:57.420 --> 01:05:00.820 align:middle line:84%
SCOTT KEMP: Yeah, and so,
just to drive the point home,

01:05:00.820 --> 01:05:04.600 align:middle line:84%
in May of this year, the White
House issued Executive Order

01:05:04.600 --> 01:05:09.040 align:middle line:84%
14300, which instructed the
Nuclear Regulatory Commission

01:05:09.040 --> 01:05:14.120 align:middle line:84%
to quote, "specifically consider
adopting a radiation limit,"

01:05:14.120 --> 01:05:18.400 align:middle line:84%
which is, again, the
politicization of our regulation

01:05:18.400 --> 01:05:20.780 align:middle line:90%
with nonscientific views.

01:05:20.780 --> 01:05:25.670 align:middle line:84%
There are people out there
pushing hard for this

01:05:25.670 --> 01:05:30.150 align:middle line:84%
because they believe,
without evidence, that this

01:05:30.150 --> 01:05:33.750 align:middle line:84%
is going to make it cheap
to build nuclear power.

01:05:33.750 --> 01:05:37.510 align:middle line:90%
So yeah.

01:05:37.510 --> 01:05:39.210 align:middle line:84%
AUDIENCE: A more
general question.

01:05:39.210 --> 01:05:45.110 align:middle line:84%
So when you're not familiar with
the subject-- so for example,

01:05:45.110 --> 01:05:47.690 align:middle line:84%
this is going back to
the look-again effect.

01:05:47.690 --> 01:05:54.190 align:middle line:84%
Yeah, so using the example of
the jelly beans or whatever.

01:05:54.190 --> 01:05:57.390 align:middle line:90%
Say I was familiar.

01:05:57.390 --> 01:06:01.050 align:middle line:84%
How do you tune your radar to be
aware of those types of things,

01:06:01.050 --> 01:06:03.190 align:middle line:84%
besides knowing the
subject and being like,

01:06:03.190 --> 01:06:05.970 align:middle line:84%
that's very contrary to
all the existing evidence?

01:06:05.970 --> 01:06:10.705 align:middle line:90%


01:06:10.705 --> 01:06:12.330 align:middle line:84%
SCOTT KEMP: How do
you tune your radar?

01:06:12.330 --> 01:06:16.190 align:middle line:90%


01:06:16.190 --> 01:06:17.690 align:middle line:90%
I'm not sure.

01:06:17.690 --> 01:06:20.787 align:middle line:90%


01:06:20.787 --> 01:06:22.620 align:middle line:84%
Are you asking the
question, if you had only

01:06:22.620 --> 01:06:25.300 align:middle line:84%
looked at the green
jelly bean scenario?

01:06:25.300 --> 01:06:28.215 align:middle line:84%
AUDIENCE: I think I
mean, as a reader who

01:06:28.215 --> 01:06:29.340 align:middle line:90%
looks at various articles--

01:06:29.340 --> 01:06:31.882 align:middle line:84%
SCOTT KEMP: Oh, how do they're
not doing the green jelly bean

01:06:31.882 --> 01:06:32.580 align:middle line:90%
scenario?

01:06:32.580 --> 01:06:33.260 align:middle line:90%
Yeah.

01:06:33.260 --> 01:06:34.380 align:middle line:90%
Yeah.

01:06:34.380 --> 01:06:37.300 align:middle line:90%
You don't often.

01:06:37.300 --> 01:06:38.060 align:middle line:90%
Yeah.

01:06:38.060 --> 01:06:41.540 align:middle line:84%
And if you go read the
medical literature,

01:06:41.540 --> 01:06:43.900 align:middle line:84%
you will find all kinds of
articles where you're like?

01:06:43.900 --> 01:06:45.200 align:middle line:90%
Why did you look at this?

01:06:45.200 --> 01:06:46.940 align:middle line:90%
Why did you study this?

01:06:46.940 --> 01:06:50.860 align:middle line:84%
And I think a lot
of that is oh, look.

01:06:50.860 --> 01:06:54.020 align:middle line:84%
And then they write
a paper about it.

01:06:54.020 --> 01:06:56.940 align:middle line:84%
So that's why you,
you want to avoid

01:06:56.940 --> 01:07:01.420 align:middle line:84%
what I call single-study
syndrome, where you just

01:07:01.420 --> 01:07:04.120 align:middle line:84%
find one study that
supports something.

01:07:04.120 --> 01:07:08.360 align:middle line:84%
If that study has not been
replicated, you shouldn't do it.

01:07:08.360 --> 01:07:10.860 align:middle line:84%
And so we try to teach
people not to do this.

01:07:10.860 --> 01:07:13.700 align:middle line:84%
One of the things
is something called

01:07:13.700 --> 01:07:18.715 align:middle line:84%
hypothesizing after the
fact, which is to say,

01:07:18.715 --> 01:07:19.590 align:middle line:90%
oh, I saw this thing.

01:07:19.590 --> 01:07:20.350 align:middle line:90%
I wonder why it happened.

01:07:20.350 --> 01:07:21.350 align:middle line:90%
Oh, maybe this is the reason.

01:07:21.350 --> 01:07:22.310 align:middle line:90%
This is my hypothesis.

01:07:22.310 --> 01:07:24.050 align:middle line:90%
Here's the data, publish.

01:07:24.050 --> 01:07:26.690 align:middle line:90%
You're not supposed to do that.

01:07:26.690 --> 01:07:28.530 align:middle line:84%
What you're doing is
you're exacerbating

01:07:28.530 --> 01:07:29.750 align:middle line:90%
the look-again effect.

01:07:29.750 --> 01:07:31.490 align:middle line:84%
You're supposed to
have the hypothesis do

01:07:31.490 --> 01:07:33.490 align:middle line:90%
the test of your hypothesis.

01:07:33.490 --> 01:07:36.490 align:middle line:90%
And then if it fails, it fails.

01:07:36.490 --> 01:07:40.090 align:middle line:84%
But unfortunately,
this is a problem

01:07:40.090 --> 01:07:44.150 align:middle line:90%
with how we promote people.

01:07:44.150 --> 01:07:46.650 align:middle line:84%
It's a social problem
because we say

01:07:46.650 --> 01:07:51.690 align:middle line:84%
you need publications to show
your impact as a scientist.

01:07:51.690 --> 01:07:54.090 align:middle line:84%
And therefore people
are incentivized to have

01:07:54.090 --> 01:07:57.690 align:middle line:84%
bad behavior and to
publish everything they can

01:07:57.690 --> 01:07:59.610 align:middle line:90%
when it ought not be published.

01:07:59.610 --> 01:08:02.050 align:middle line:84%
And then journals
also don't want

01:08:02.050 --> 01:08:05.130 align:middle line:84%
to publish null results
because they're not exciting.

01:08:05.130 --> 01:08:06.970 align:middle line:84%
And so there are
lots of experiments

01:08:06.970 --> 01:08:11.450 align:middle line:84%
that are done that are never
reported in the literature.

01:08:11.450 --> 01:08:14.570 align:middle line:84%
And all kinds of
people are coming up

01:08:14.570 --> 01:08:16.840 align:middle line:84%
with the same hypothesis,
doing the same experiment,

01:08:16.840 --> 01:08:18.720 align:middle line:84%
getting the same null
result, and not knowing

01:08:18.720 --> 01:08:21.479 align:middle line:84%
that the repeating people's
experiments because we

01:08:21.479 --> 01:08:23.359 align:middle line:90%
don't publish null results.

01:08:23.359 --> 01:08:25.220 align:middle line:90%
It's a real big problem.

01:08:25.220 --> 01:08:25.720 align:middle line:90%
Yeah.

01:08:25.720 --> 01:08:26.220 align:middle line:90%
Yeah.

01:08:26.220 --> 01:08:27.840 align:middle line:84%
AUDIENCE: Follow-up
question of mine--

01:08:27.840 --> 01:08:31.600 align:middle line:84%
since an important part
of the scientific process

01:08:31.600 --> 01:08:35.240 align:middle line:90%
is repeatability.

01:08:35.240 --> 01:08:38.200 align:middle line:84%
This is also coming from
a lack of knowledge.

01:08:38.200 --> 01:08:40.160 align:middle line:84%
Is the system set
up so that it is

01:08:40.160 --> 01:08:43.240 align:middle line:84%
hard to get funding to just
prove or disprove someone else's

01:08:43.240 --> 01:08:45.279 align:middle line:90%
stuff, because someone would?

01:08:45.279 --> 01:08:46.772 align:middle line:84%
We want you to
make breakthroughs.

01:08:46.772 --> 01:08:47.939 align:middle line:90%
SCOTT KEMP: Something novel.

01:08:47.939 --> 01:08:49.200 align:middle line:90%
Yeah, sure.

01:08:49.200 --> 01:08:50.899 align:middle line:90%
Absolutely, it's really hard.

01:08:50.899 --> 01:08:53.760 align:middle line:84%
It would be really hard to
justify a replication study

01:08:53.760 --> 01:08:56.779 align:middle line:84%
unless the initial result
was, for some reason,

01:08:56.779 --> 01:09:00.220 align:middle line:84%
very controversial
or very important.

01:09:00.220 --> 01:09:03.399 align:middle line:84%
Someone say, oh, it's
already been done.

01:09:03.399 --> 01:09:07.000 align:middle line:84%
AUDIENCE: So does that make it
such that single studies set

01:09:07.000 --> 01:09:12.319 align:middle line:84%
the tone of the body of
literature, when in reality, it

01:09:12.319 --> 01:09:14.140 align:middle line:84%
should be, a
collection of studies?

01:09:14.140 --> 01:09:15.740 align:middle line:84%
SCOTT KEMP: Yes,
they absolutely do.

01:09:15.740 --> 01:09:20.080 align:middle line:84%
And I mean, in the
best of all worlds,

01:09:20.080 --> 01:09:24.140 align:middle line:84%
the first study changes
the direction of research,

01:09:24.140 --> 01:09:25.720 align:middle line:90%
but research goes and checks.

01:09:25.720 --> 01:09:28.800 align:middle line:90%


01:09:28.800 --> 01:09:32.800 align:middle line:84%
It is the case that sometimes
we go down the wrong path,

01:09:32.800 --> 01:09:35.040 align:middle line:90%
sometimes for a very long time.

01:09:35.040 --> 01:09:38.220 align:middle line:84%
So I mean, maybe
the classic example,

01:09:38.220 --> 01:09:40.840 align:middle line:84%
which is not really relevant
to how modern research is done

01:09:40.840 --> 01:09:45.160 align:middle line:84%
and published, but would
be Ptolemaic epicycles

01:09:45.160 --> 01:09:52.279 align:middle line:84%
to explain how the sun and
the planets orbit the Earth.

01:09:52.279 --> 01:09:58.720 align:middle line:84%
And it took hundreds of years
before that view was refuted

01:09:58.720 --> 01:10:02.220 align:middle line:90%
because of the firmness.

01:10:02.220 --> 01:10:05.000 align:middle line:84%
There are other
phenomena which I

01:10:05.000 --> 01:10:07.320 align:middle line:84%
can't bring to mind,
which have existed

01:10:07.320 --> 01:10:11.460 align:middle line:90%
for various periods of time.

01:10:11.460 --> 01:10:14.710 align:middle line:90%
Does anyone remember the--

01:10:14.710 --> 01:10:20.190 align:middle line:84%
during the COVID pandemic,
there was a big push for--

01:10:20.190 --> 01:10:23.202 align:middle line:84%
what is the name of
the anti-malarial drug?

01:10:23.202 --> 01:10:24.410 align:middle line:90%
AUDIENCE: Hydroxychloroquine?

01:10:24.410 --> 01:10:25.702 align:middle line:90%
SCOTT KEMP: Hydroxychloroquine.

01:10:25.702 --> 01:10:28.190 align:middle line:84%
Do you remember
hydroxychloroquine?

01:10:28.190 --> 01:10:34.990 align:middle line:84%
So that didn't just come
from some wacky person.

01:10:34.990 --> 01:10:38.070 align:middle line:84%
I was about to Google
it, but you got it.

01:10:38.070 --> 01:10:41.950 align:middle line:84%
That actually came from a
scientific study published

01:10:41.950 --> 01:10:44.950 align:middle line:90%
in Nature, top-quality journal.

01:10:44.950 --> 01:10:47.315 align:middle line:84%
And what they did
is they found--

01:10:47.315 --> 01:10:49.190 align:middle line:84%
people who want to go
and don't want to learn

01:10:49.190 --> 01:10:53.830 align:middle line:90%
about rate biology, you can go.

01:10:53.830 --> 01:10:59.910 align:middle line:84%
So what they did is they
used Vero cells, which

01:10:59.910 --> 01:11:01.470 align:middle line:90%
are derived from liver.

01:11:01.470 --> 01:11:05.870 align:middle line:84%
And it's a standard test cell
to see if different drugs would

01:11:05.870 --> 01:11:07.550 align:middle line:90%
prevent COVID infection.

01:11:07.550 --> 01:11:13.700 align:middle line:84%
So they put the cell inside a
solution of hydroxychloroquine.

01:11:13.700 --> 01:11:19.620 align:middle line:84%
They infect the supernatant
with little virus of COVID.

01:11:19.620 --> 01:11:25.420 align:middle line:84%
And lo and behold, it reduced
the probability of infection.

01:11:25.420 --> 01:11:28.940 align:middle line:84%
So everyone assumed
that this would work.

01:11:28.940 --> 01:11:33.700 align:middle line:84%
The thing is that it turns out
that viruses can infect cells

01:11:33.700 --> 01:11:35.380 align:middle line:90%
in a couple of different ways.

01:11:35.380 --> 01:11:37.380 align:middle line:84%
So one of the classic
ways is the virus lands

01:11:37.380 --> 01:11:41.820 align:middle line:84%
on the top of the
cell, like here.

01:11:41.820 --> 01:11:45.460 align:middle line:84%
And the cell detects
something is wrong,

01:11:45.460 --> 01:11:48.900 align:middle line:84%
and it endocytosis,
creates a vesicle,

01:11:48.900 --> 01:11:51.460 align:middle line:90%
brings the virus inside.

01:11:51.460 --> 01:11:55.235 align:middle line:84%
And then what hydroxychloroquine
does is it acidifies.

01:11:55.235 --> 01:11:56.360 align:middle line:90%
This is called an endosome.

01:11:56.360 --> 01:12:00.460 align:middle line:84%
And it acidifies this endosome
and destroys the virus.

01:12:00.460 --> 01:12:04.300 align:middle line:84%
And so it prevents the viral,
which is being endocytosed

01:12:04.300 --> 01:12:08.540 align:middle line:84%
is the term, from
infecting the cell.

01:12:08.540 --> 01:12:12.490 align:middle line:84%
So it does indeed prevent
infection via this mechanism.

01:12:12.490 --> 01:12:15.490 align:middle line:84%
But this mechanism is
the dominant mechanism

01:12:15.490 --> 01:12:17.490 align:middle line:84%
in liver cells, but not
the dominant mechanism

01:12:17.490 --> 01:12:20.410 align:middle line:84%
in your respiratory tract,
which happens to express

01:12:20.410 --> 01:12:23.650 align:middle line:90%
a protein here called ACE2.

01:12:23.650 --> 01:12:29.130 align:middle line:84%
And most of the viruses land
on ACE2 and get in this way.

01:12:29.130 --> 01:12:31.050 align:middle line:84%
And so hydroxychloroquine
does nothing

01:12:31.050 --> 01:12:34.010 align:middle line:90%
to prevent infection via ACE2.

01:12:34.010 --> 01:12:37.290 align:middle line:84%
So it was good science,
but they forgot

01:12:37.290 --> 01:12:38.970 align:middle line:84%
that this cell type
did not express

01:12:38.970 --> 01:12:43.810 align:middle line:84%
this protein, which is expressed
in your airway epithelial.

01:12:43.810 --> 01:12:46.150 align:middle line:90%
And so the result was wrong.

01:12:46.150 --> 01:12:48.570 align:middle line:84%
But you will not
believe how hard

01:12:48.570 --> 01:12:52.850 align:middle line:84%
it was to get hydroxychloroquine
for a large number of months

01:12:52.850 --> 01:12:56.230 align:middle line:84%
because of this one paper,
which had never been validated.

01:12:56.230 --> 01:12:58.510 align:middle line:84%
And that wasn't even
the look-again effect.

01:12:58.510 --> 01:13:01.210 align:middle line:84%
It was just that they
used the wrong cell type.

01:13:01.210 --> 01:13:05.330 align:middle line:84%
So it was good science,
but it wasn't complete.

01:13:05.330 --> 01:13:06.410 align:middle line:90%
So yeah.

01:13:06.410 --> 01:13:08.890 align:middle line:84%
AUDIENCE: Continuing unless
someone else has a question.

01:13:08.890 --> 01:13:10.307 align:middle line:84%
AUDIENCE: I was
just going to say,

01:13:10.307 --> 01:13:15.160 align:middle line:84%
I feel like there'd be a need
for a scientific magazine called

01:13:15.160 --> 01:13:19.060 align:middle line:84%
Null, Not Void, which
is just null results.

01:13:19.060 --> 01:13:23.000 align:middle line:84%
SCOTT KEMP: I think there is
a journal of null results.

01:13:23.000 --> 01:13:25.880 align:middle line:84%
But I don't think most people
want to bother taking the time

01:13:25.880 --> 01:13:28.440 align:middle line:90%
to write up their work.

01:13:28.440 --> 01:13:31.560 align:middle line:90%
So yeah.

01:13:31.560 --> 01:13:34.060 align:middle line:90%
Anyway, this is a problem.

01:13:34.060 --> 01:13:36.400 align:middle line:84%
We'll talk more about
this in the final class.

01:13:36.400 --> 01:13:38.920 align:middle line:84%
No other questions
about radiation biology?

01:13:38.920 --> 01:13:41.680 align:middle line:90%
All right, good.

01:13:41.680 --> 01:13:47.560 align:middle line:84%
So on Wednesday, we will
calculate how many people die

01:13:47.560 --> 01:13:49.640 align:middle line:90%
from nuclear reactor accidents.

01:13:49.640 --> 01:13:53.120 align:middle line:84%
And that will tell us something
about whether nuclear power is

01:13:53.120 --> 01:13:53.980 align:middle line:90%
safe or not.

01:13:53.980 --> 01:13:55.980 align:middle line:84%
And we'll use the linear
no-threshold model,

01:13:55.980 --> 01:14:00.430 align:middle line:84%
which I hope you will believe
now that I've defended it.

01:14:00.430 --> 01:14:09.000 align:middle line:90%