WEBVTT

00:00:00.000 --> 00:00:00.040 align:middle line:90%


00:00:00.040 --> 00:00:02.460 align:middle line:84%
The following content is
provided under a Creative

00:00:02.460 --> 00:00:03.870 align:middle line:90%
Commons license.

00:00:03.870 --> 00:00:06.910 align:middle line:84%
Your support will help MIT
OpenCourseWare continue to

00:00:06.910 --> 00:00:10.560 align:middle line:84%
offer high quality educational
resources for free.

00:00:10.560 --> 00:00:13.460 align:middle line:84%
To make a donation or view
additional materials from

00:00:13.460 --> 00:00:19.290 align:middle line:84%
hundreds of MIT courses, visit
MIT OpenCourseWare at

00:00:19.290 --> 00:00:20.540 align:middle line:90%
ocw.mit.edu.

00:00:20.540 --> 00:00:22.648 align:middle line:90%


00:00:22.648 --> 00:00:25.410 align:middle line:84%
JOHN TSITSIKLIS: So today we're
going to finish with the

00:00:25.410 --> 00:00:28.240 align:middle line:90%
core material of this class.

00:00:28.240 --> 00:00:30.980 align:middle line:84%
That is the material that has to
do with probability theory

00:00:30.980 --> 00:00:31.690 align:middle line:90%
in general.

00:00:31.690 --> 00:00:34.240 align:middle line:84%
And then for the rest of the
semester we're going to look

00:00:34.240 --> 00:00:38.290 align:middle line:84%
at some special types of models,
talk about inference.

00:00:38.290 --> 00:00:40.970 align:middle line:84%
Well, there's also going to
be a small module of core

00:00:40.970 --> 00:00:42.840 align:middle line:90%
material coming later.

00:00:42.840 --> 00:00:46.940 align:middle line:84%
But today we're basically
finishing chapter four.

00:00:46.940 --> 00:00:50.800 align:middle line:84%
And what we're going to do is
we're going to look at a

00:00:50.800 --> 00:00:53.720 align:middle line:84%
somewhat familiar concept, the
concept of the conditional

00:00:53.720 --> 00:00:55.000 align:middle line:90%
expectation.

00:00:55.000 --> 00:00:58.690 align:middle line:84%
But we're going to look at it
from a slightly different

00:00:58.690 --> 00:01:02.840 align:middle line:84%
angle, from a slightly more
sophisticated angle.

00:01:02.840 --> 00:01:05.370 align:middle line:84%
And together with the
conditional expectation we

00:01:05.370 --> 00:01:08.445 align:middle line:84%
will also talk about conditional
variances.

00:01:08.445 --> 00:01:11.840 align:middle line:84%
It's something that we're going
to denote this way.

00:01:11.840 --> 00:01:15.180 align:middle line:84%
And we're going to see what they
are, and there are some

00:01:15.180 --> 00:01:17.820 align:middle line:84%
subtle concepts that
are involved here.

00:01:17.820 --> 00:01:20.780 align:middle line:84%
And we're going to apply some
of the tools we're going to

00:01:20.780 --> 00:01:24.390 align:middle line:84%
develop to deal with a special
type of situation in which

00:01:24.390 --> 00:01:26.660 align:middle line:90%
we're adding random variables.

00:01:26.660 --> 00:01:31.860 align:middle line:84%
But we're adding a random number
of random variables.

00:01:31.860 --> 00:01:34.720 align:middle line:84%
OK, so let's start talking
about conditional

00:01:34.720 --> 00:01:37.410 align:middle line:90%
expectations.

00:01:37.410 --> 00:01:39.970 align:middle line:84%
I guess you know
what they are.

00:01:39.970 --> 00:01:43.660 align:middle line:84%
Suppose we are in the discrete
the world. xy, or discrete

00:01:43.660 --> 00:01:45.590 align:middle line:90%
random variables.

00:01:45.590 --> 00:01:49.340 align:middle line:84%
We defined the conditional
expectation of x given that I

00:01:49.340 --> 00:01:52.480 align:middle line:84%
told you the value of the
random variable y.

00:01:52.480 --> 00:01:56.800 align:middle line:84%
And the way we define it is the
same way as an ordinary

00:01:56.800 --> 00:02:01.020 align:middle line:84%
expectation, except that we're
using the conditional PMF.

00:02:01.020 --> 00:02:03.440 align:middle line:84%
So we're using the probabilities
that apply to

00:02:03.440 --> 00:02:06.910 align:middle line:84%
the new universe where we are
told the value of the random

00:02:06.910 --> 00:02:08.289 align:middle line:90%
variable y.

00:02:08.289 --> 00:02:12.050 align:middle line:84%
So this is still a familiar
concept so far.

00:02:12.050 --> 00:02:14.720 align:middle line:84%
If we're dealing with the
continuous random variable x

00:02:14.720 --> 00:02:17.170 align:middle line:84%
the formula is the same, except
that here we have an

00:02:17.170 --> 00:02:21.450 align:middle line:84%
integral, and we have to use
the conditional density

00:02:21.450 --> 00:02:25.020 align:middle line:90%
function of x.

00:02:25.020 --> 00:02:28.770 align:middle line:84%
Now what I'm going to do, I want
to introduce it gently

00:02:28.770 --> 00:02:32.200 align:middle line:84%
through the example that we
talked about last time.

00:02:32.200 --> 00:02:35.290 align:middle line:84%
So last time we talked about
having a stick that has a

00:02:35.290 --> 00:02:36.770 align:middle line:90%
certain length.

00:02:36.770 --> 00:02:41.950 align:middle line:84%
And we take that stick, and we
break it at some point that we

00:02:41.950 --> 00:02:43.790 align:middle line:90%
choose uniformly at random.

00:02:43.790 --> 00:02:49.390 align:middle line:84%
And let's denote why the place
where we chose to break it.

00:02:49.390 --> 00:02:52.750 align:middle line:84%
Having chosen y, then
we're left with a

00:02:52.750 --> 00:02:53.930 align:middle line:90%
piece of the stick.

00:02:53.930 --> 00:02:57.750 align:middle line:84%
And I'm going to choose a place
to break it once more

00:02:57.750 --> 00:03:01.330 align:middle line:84%
uniformly at random
between 0 and y.

00:03:01.330 --> 00:03:04.170 align:middle line:84%
So this is the second place at
which we are going to break

00:03:04.170 --> 00:03:07.900 align:middle line:90%
it, and we call that place x.

00:03:07.900 --> 00:03:12.040 align:middle line:84%
OK, so what's the conditional
expectation of x if I tell you

00:03:12.040 --> 00:03:13.630 align:middle line:90%
the value of y?

00:03:13.630 --> 00:03:16.740 align:middle line:84%
I tell you that capital Y
happens to take a specific

00:03:16.740 --> 00:03:18.800 align:middle line:90%
numerical value.

00:03:18.800 --> 00:03:22.770 align:middle line:84%
So this capital Y is now a
specific numerical value, x is

00:03:22.770 --> 00:03:25.280 align:middle line:84%
chosen uniformly over
this range.

00:03:25.280 --> 00:03:29.780 align:middle line:84%
So the expected value of x is
going to be half of this range

00:03:29.780 --> 00:03:30.810 align:middle line:90%
between 0 and y.

00:03:30.810 --> 00:03:36.850 align:middle line:84%
So the conditional expectation
is little y over 2.

00:03:36.850 --> 00:03:40.210 align:middle line:84%
The important thing to realize
here is that this

00:03:40.210 --> 00:03:42.170 align:middle line:90%
quantity is a number.

00:03:42.170 --> 00:03:45.570 align:middle line:84%
I told you that the random
variable took a certain

00:03:45.570 --> 00:03:49.170 align:middle line:84%
numerical value,
let's say 3.5.

00:03:49.170 --> 00:03:52.780 align:middle line:84%
And then you tell me given that
the random variable took

00:03:52.780 --> 00:03:59.940 align:middle line:84%
the numerical value 3.5 the
expected value of x is 1.75.

00:03:59.940 --> 00:04:04.080 align:middle line:84%
So this is an equality
between numbers.

00:04:04.080 --> 00:04:08.110 align:middle line:84%
On the other hand, before you
do the experiment you don't

00:04:08.110 --> 00:04:12.160 align:middle line:84%
know what y is going
to turn out to be.

00:04:12.160 --> 00:04:15.680 align:middle line:84%
So this little y is the
numerical value that has been

00:04:15.680 --> 00:04:18.990 align:middle line:84%
observed when you start doing
the experiments and you

00:04:18.990 --> 00:04:22.700 align:middle line:84%
observe the value of capital
Y. So in some sense this

00:04:22.700 --> 00:04:27.770 align:middle line:84%
quantity is not known ahead of
time, it is random itself.

00:04:27.770 --> 00:04:33.670 align:middle line:84%
So maybe we can start thinking
of it as a random variable.

00:04:33.670 --> 00:04:37.010 align:middle line:84%
So to put it differently, before
we do the experiment I

00:04:37.010 --> 00:04:41.030 align:middle line:84%
ask you what's the expected
value of x given y?

00:04:41.030 --> 00:04:44.740 align:middle line:84%
You're going to answer me well
I don't know, it depends on

00:04:44.740 --> 00:04:47.580 align:middle line:84%
what y is going to
turn out to be.

00:04:47.580 --> 00:04:52.690 align:middle line:84%
So the expected value of x given
y itself can be viewed

00:04:52.690 --> 00:04:56.540 align:middle line:84%
as a random variable, because
it depends on the random

00:04:56.540 --> 00:04:58.330 align:middle line:90%
variable capital Y.

00:04:58.330 --> 00:05:02.080 align:middle line:84%
So hidden here there's some kind
of statement about random

00:05:02.080 --> 00:05:04.810 align:middle line:90%
variables instead of numbers.

00:05:04.810 --> 00:05:07.770 align:middle line:84%
And that statement about
random variables, we

00:05:07.770 --> 00:05:09.660 align:middle line:90%
write it this way.

00:05:09.660 --> 00:05:12.770 align:middle line:84%
By thinking of the expected
value, the conditional

00:05:12.770 --> 00:05:17.330 align:middle line:84%
expectation, as a random
variable instead of a number.

00:05:17.330 --> 00:05:20.380 align:middle line:84%
It's a random variable when we
do not specify a specific

00:05:20.380 --> 00:05:23.410 align:middle line:84%
number, but we think of it
as an abstract object.

00:05:23.410 --> 00:05:29.560 align:middle line:84%
The expected value of x given
the random variable y is the

00:05:29.560 --> 00:05:34.390 align:middle line:84%
random variable y over 2 no
matter what capital Y

00:05:34.390 --> 00:05:37.090 align:middle line:90%
turns out to be.

00:05:37.090 --> 00:05:39.530 align:middle line:84%
So we turn and take a statement
that deals with

00:05:39.530 --> 00:05:43.460 align:middle line:84%
equality of two numbers, and we
make it a statement that's

00:05:43.460 --> 00:05:46.740 align:middle line:84%
an equality between two
random variables.

00:05:46.740 --> 00:05:49.910 align:middle line:84%
OK so this is clearly a random
variable because

00:05:49.910 --> 00:05:52.330 align:middle line:90%
capital Y is random.

00:05:52.330 --> 00:05:54.170 align:middle line:90%
What exactly is this object?

00:05:54.170 --> 00:05:57.130 align:middle line:84%
I didn't yet define it
for you formally.

00:05:57.130 --> 00:06:02.150 align:middle line:84%
So let's now give the formal
definition of this object

00:06:02.150 --> 00:06:04.570 align:middle line:84%
that's going to be
denoted this way.

00:06:04.570 --> 00:06:09.400 align:middle line:84%
The conditional expectation of
x given the random variable y

00:06:09.400 --> 00:06:12.830 align:middle line:90%
is a random variable.

00:06:12.830 --> 00:06:14.900 align:middle line:90%
Which random variable is it?

00:06:14.900 --> 00:06:19.670 align:middle line:84%
It's the random variable that
takes this specific numerical

00:06:19.670 --> 00:06:24.330 align:middle line:84%
value whenever capital Y happens
to take the specific

00:06:24.330 --> 00:06:26.480 align:middle line:90%
numerical value little y.

00:06:26.480 --> 00:06:30.010 align:middle line:84%
In particular, this is a random
variable, which is a

00:06:30.010 --> 00:06:33.840 align:middle line:84%
function of the random variable
capital Y. In this

00:06:33.840 --> 00:06:36.680 align:middle line:84%
instance, it's given by a simple
formula in terms of

00:06:36.680 --> 00:06:39.540 align:middle line:84%
capital Y. In other situations
it might be a

00:06:39.540 --> 00:06:41.520 align:middle line:90%
more complicated formula.

00:06:41.520 --> 00:06:44.680 align:middle line:84%
So again, to summarize,
it's a random.

00:06:44.680 --> 00:06:48.530 align:middle line:84%
The conditional expectation can
be thought of as a random

00:06:48.530 --> 00:06:53.040 align:middle line:84%
variable instead of something
that's just a number.

00:06:53.040 --> 00:06:55.940 align:middle line:84%
So in any specific context when
you're given the value of

00:06:55.940 --> 00:06:59.110 align:middle line:84%
capital Y the conditional
expectation becomes a number.

00:06:59.110 --> 00:07:02.890 align:middle line:84%
This is the realized value
of this random variable.

00:07:02.890 --> 00:07:06.260 align:middle line:84%
But before the experiment
starts, before you know what

00:07:06.260 --> 00:07:10.140 align:middle line:84%
capital Y is going to be, all
that you can say is that the

00:07:10.140 --> 00:07:14.320 align:middle line:84%
conditional expectation is going
to be 1/2 of whatever

00:07:14.320 --> 00:07:16.840 align:middle line:90%
capital Y turns out to be.

00:07:16.840 --> 00:07:20.270 align:middle line:84%
This is a pretty subtle concept,
it's an abstraction,

00:07:20.270 --> 00:07:22.990 align:middle line:90%
but it's a useful abstraction.

00:07:22.990 --> 00:07:29.440 align:middle line:84%
And we're going to see
today how to use it.

00:07:29.440 --> 00:07:32.940 align:middle line:84%
All right, I have made the point
that the conditional

00:07:32.940 --> 00:07:37.200 align:middle line:84%
expectation, the random variable
that takes these

00:07:37.200 --> 00:07:40.490 align:middle line:84%
numerical values is
a random variable.

00:07:40.490 --> 00:07:43.090 align:middle line:84%
If it is a random variable
this means that it has an

00:07:43.090 --> 00:07:45.710 align:middle line:90%
expectation of its own.

00:07:45.710 --> 00:07:48.590 align:middle line:84%
So let's start thinking what
the expectation of the

00:07:48.590 --> 00:07:53.432 align:middle line:84%
conditional expectation is
going to turn out to be.

00:07:53.432 --> 00:07:59.210 align:middle line:84%
OK, so the conditional
expectation is a random

00:07:59.210 --> 00:08:03.030 align:middle line:84%
variable, and in general it's
some function of the random

00:08:03.030 --> 00:08:05.465 align:middle line:84%
variable y that we
are observing.

00:08:05.465 --> 00:08:07.970 align:middle line:90%


00:08:07.970 --> 00:08:13.910 align:middle line:84%
In terms of numerical values if
capital Y happens to take a

00:08:13.910 --> 00:08:17.490 align:middle line:84%
specific numerical value then
the conditional expectation

00:08:17.490 --> 00:08:20.830 align:middle line:84%
also takes a specific numerical
value, and we use

00:08:20.830 --> 00:08:22.630 align:middle line:84%
the same function
to evaluate it.

00:08:22.630 --> 00:08:25.770 align:middle line:84%
The difference here is that this
is an equality of random

00:08:25.770 --> 00:08:29.440 align:middle line:84%
variables, this is an equality
between numbers.

00:08:29.440 --> 00:08:33.120 align:middle line:84%
Now if we want to calculate
the expected value of the

00:08:33.120 --> 00:08:38.539 align:middle line:84%
conditional expectation we're
basically talking about the

00:08:38.539 --> 00:08:44.080 align:middle line:84%
expected value of a function
of a random variable.

00:08:44.080 --> 00:08:48.620 align:middle line:84%
And we know how to calculate
expected values of a function.

00:08:48.620 --> 00:08:54.330 align:middle line:84%
If we are in the discrete case,
for example, this would

00:08:54.330 --> 00:09:02.690 align:middle line:84%
be a sum over all y's of the
function who's expected value

00:09:02.690 --> 00:09:09.580 align:middle line:84%
we're taking times the
probability that y takes on a

00:09:09.580 --> 00:09:11.940 align:middle line:90%
specific numerical value.

00:09:11.940 --> 00:09:16.360 align:middle line:84%
OK, but let's remember
what g is.

00:09:16.360 --> 00:09:22.690 align:middle line:84%
So g is the numerical value
of the conditional

00:09:22.690 --> 00:09:25.300 align:middle line:90%
expectation of x with y.

00:09:25.300 --> 00:09:29.530 align:middle line:90%


00:09:29.530 --> 00:09:33.450 align:middle line:84%
And now when you see this
expression you recognize it.

00:09:33.450 --> 00:09:35.630 align:middle line:84%
This is the expression
that we get in the

00:09:35.630 --> 00:09:37.190 align:middle line:90%
total expectation theorem.

00:09:37.190 --> 00:09:41.300 align:middle line:90%


00:09:41.300 --> 00:09:42.795 align:middle line:90%
Did I miss something?

00:09:42.795 --> 00:09:45.570 align:middle line:90%


00:09:45.570 --> 00:09:48.700 align:middle line:84%
Yes, in the total expectation
theorem to find the expected

00:09:48.700 --> 00:09:52.720 align:middle line:84%
value of x, we divide the world
into different scenarios

00:09:52.720 --> 00:09:55.970 align:middle line:90%
depending on what y happens.

00:09:55.970 --> 00:09:59.110 align:middle line:84%
We calculate the expectation
in each one of the possible

00:09:59.110 --> 00:10:01.750 align:middle line:84%
worlds, and we take the
weighted average.

00:10:01.750 --> 00:10:04.770 align:middle line:84%
So this is a formula that you
have seen before, and you

00:10:04.770 --> 00:10:08.610 align:middle line:84%
recognize that this is the
expected value of x.

00:10:08.610 --> 00:10:13.280 align:middle line:84%
So this is a longer, more
detailed derivation of what I

00:10:13.280 --> 00:10:17.770 align:middle line:84%
had written up here, but the
important thing to keep in

00:10:17.770 --> 00:10:22.790 align:middle line:84%
mind is the moral of the
story, the punchline.

00:10:22.790 --> 00:10:26.640 align:middle line:84%
The expected value of the
conditional expectation is the

00:10:26.640 --> 00:10:27.890 align:middle line:90%
expectation itself.

00:10:27.890 --> 00:10:30.710 align:middle line:90%


00:10:30.710 --> 00:10:35.030 align:middle line:84%
So this is just our total
expectation theorem, but

00:10:35.030 --> 00:10:37.700 align:middle line:84%
written in more abstract
notation.

00:10:37.700 --> 00:10:40.035 align:middle line:84%
And it comes handy to have this
more abstract notation,

00:10:40.035 --> 00:10:43.570 align:middle line:84%
as as we're going to
see in a while.

00:10:43.570 --> 00:10:47.320 align:middle line:84%
OK, we can apply this to
our stick example.

00:10:47.320 --> 00:10:50.220 align:middle line:84%
If we want to find the expected
value of x how much

00:10:50.220 --> 00:10:53.110 align:middle line:84%
of the stick is left
at the end?

00:10:53.110 --> 00:10:57.370 align:middle line:84%
We can calculate it using this
law of iterated expectations.

00:10:57.370 --> 00:11:00.190 align:middle line:84%
It's the expected value of the
conditional expectation.

00:11:00.190 --> 00:11:03.790 align:middle line:84%
We know that the conditional
expectation is y over 2.

00:11:03.790 --> 00:11:10.730 align:middle line:84%
So expected value of y is l over
2, because y is uniform

00:11:10.730 --> 00:11:12.830 align:middle line:90%
so we get l over 4.

00:11:12.830 --> 00:11:15.440 align:middle line:84%
So this gives us the same answer
that we derived last

00:11:15.440 --> 00:11:18.210 align:middle line:90%
time in a rather long way.

00:11:18.210 --> 00:11:24.470 align:middle line:90%


00:11:24.470 --> 00:11:27.750 align:middle line:84%
All right, now that we have
mastered conditional

00:11:27.750 --> 00:11:33.100 align:middle line:84%
expectations, let's raise the
bar a little more and talk

00:11:33.100 --> 00:11:35.590 align:middle line:90%
about conditional variances.

00:11:35.590 --> 00:11:38.750 align:middle line:84%
So the conditional expectation
is the mean value, or the

00:11:38.750 --> 00:11:41.380 align:middle line:84%
expected value, in a conditional
universe where

00:11:41.380 --> 00:11:43.450 align:middle line:90%
you're told the value of y.

00:11:43.450 --> 00:11:47.270 align:middle line:84%
In that same conditional
universe you can talk about

00:11:47.270 --> 00:11:51.360 align:middle line:84%
the conditional distribution
of x, which has a mean--

00:11:51.360 --> 00:11:52.810 align:middle line:90%
the conditional expectation--

00:11:52.810 --> 00:11:54.140 align:middle line:84%
but the conditional
distribution of

00:11:54.140 --> 00:11:56.130 align:middle line:90%
x also has a variance.

00:11:56.130 --> 00:11:58.730 align:middle line:84%
So we can talk about the
variance of x in that

00:11:58.730 --> 00:12:01.500 align:middle line:90%
conditional universe.

00:12:01.500 --> 00:12:07.390 align:middle line:84%
The conditional variance as a
number is the natural thing.

00:12:07.390 --> 00:12:11.680 align:middle line:84%
It's the variance of x, except
that all the calculations are

00:12:11.680 --> 00:12:13.790 align:middle line:84%
done in the conditional
universe.

00:12:13.790 --> 00:12:19.940 align:middle line:84%
In the conditional universe the
expected value of x is the

00:12:19.940 --> 00:12:21.740 align:middle line:90%
conditional expectation.

00:12:21.740 --> 00:12:24.530 align:middle line:84%
This is the distance from the
mean in the conditional

00:12:24.530 --> 00:12:26.310 align:middle line:90%
universe squared.

00:12:26.310 --> 00:12:30.080 align:middle line:84%
And we take the average value
of the squared distance, but

00:12:30.080 --> 00:12:32.660 align:middle line:84%
calculate it again using the
probabilities that apply in

00:12:32.660 --> 00:12:35.240 align:middle line:90%
the conditional universe.

00:12:35.240 --> 00:12:38.020 align:middle line:84%
This is an equality
between numbers.

00:12:38.020 --> 00:12:43.720 align:middle line:84%
I tell you the value of y, once
you know that value for y

00:12:43.720 --> 00:12:47.730 align:middle line:84%
you can go ahead and plot the
conditional distribution of x.

00:12:47.730 --> 00:12:50.090 align:middle line:84%
And for that conditional
distribution you can calculate

00:12:50.090 --> 00:12:52.890 align:middle line:84%
the number which is the
variance of x in that

00:12:52.890 --> 00:12:54.650 align:middle line:90%
conditional universe.

00:12:54.650 --> 00:12:57.820 align:middle line:84%
So now let's repeat the mental
gymnastics from the previous

00:12:57.820 --> 00:13:03.560 align:middle line:84%
slide, and abstract things, and
define a random variable--

00:13:03.560 --> 00:13:06.080 align:middle line:90%
the conditional variance.

00:13:06.080 --> 00:13:08.900 align:middle line:84%
And it's going to be a random
variable because we leave the

00:13:08.900 --> 00:13:12.010 align:middle line:84%
numerical value of capital
Y unspecified.

00:13:12.010 --> 00:13:15.670 align:middle line:84%
So ahead of time we don't know
what capital Y is going to be,

00:13:15.670 --> 00:13:18.860 align:middle line:84%
and because of that we don't
know ahead of time what the

00:13:18.860 --> 00:13:20.870 align:middle line:84%
conditional variance
is going to be.

00:13:20.870 --> 00:13:24.500 align:middle line:84%
So before the experiment starts
if I ask you what's the

00:13:24.500 --> 00:13:26.060 align:middle line:90%
conditional variance of x?

00:13:26.060 --> 00:13:28.440 align:middle line:84%
You're going to tell me well I
don't know, It depends on what

00:13:28.440 --> 00:13:30.300 align:middle line:90%
y is going to turn out to be.

00:13:30.300 --> 00:13:32.770 align:middle line:84%
It's going to be something
that depends on y.

00:13:32.770 --> 00:13:36.210 align:middle line:84%
So it's a random variable,
which is a function of y.

00:13:36.210 --> 00:13:38.980 align:middle line:84%
So more precisely, the
conditional variance when

00:13:38.980 --> 00:13:42.480 align:middle line:84%
written in this notation just
with capital letters, is a

00:13:42.480 --> 00:13:43.730 align:middle line:90%
random variable.

00:13:43.730 --> 00:13:47.560 align:middle line:84%
It's a random variable whose
value is completely determined

00:13:47.560 --> 00:13:52.330 align:middle line:84%
once you learned the value of
capital Y. And it takes a

00:13:52.330 --> 00:13:55.070 align:middle line:90%
specific numerical value.

00:13:55.070 --> 00:13:58.700 align:middle line:84%
If capital Y happens to get a
realization that's a specific

00:13:58.700 --> 00:14:03.130 align:middle line:84%
number, then the variance also
becomes a specific number.

00:14:03.130 --> 00:14:05.390 align:middle line:84%
And it's just a conditional
variance of y

00:14:05.390 --> 00:14:09.420 align:middle line:90%
over x in that universe.

00:14:09.420 --> 00:14:12.390 align:middle line:84%
All right, OK, so let's continue
what we did in the

00:14:12.390 --> 00:14:13.620 align:middle line:90%
previous slide.

00:14:13.620 --> 00:14:15.960 align:middle line:84%
We had the law of iterated
expectations.

00:14:15.960 --> 00:14:18.350 align:middle line:84%
That told us that expected
value of a conditional

00:14:18.350 --> 00:14:21.360 align:middle line:84%
expectation is the unconditional
expectation.

00:14:21.360 --> 00:14:26.140 align:middle line:84%
Is there a similar rule that
might apply in this context?

00:14:26.140 --> 00:14:29.810 align:middle line:84%
So you might guess that the
variance of x could be found

00:14:29.810 --> 00:14:33.590 align:middle line:84%
by taking the expected value of
the conditional variance.

00:14:33.590 --> 00:14:35.680 align:middle line:84%
It turns out that this
is not true.

00:14:35.680 --> 00:14:38.480 align:middle line:84%
There is a formula for the
variance in terms of

00:14:38.480 --> 00:14:40.060 align:middle line:90%
conditional quantities.

00:14:40.060 --> 00:14:42.280 align:middle line:84%
But the formula is a little
more complicated.

00:14:42.280 --> 00:14:46.200 align:middle line:84%
If involves two terms
instead of one.

00:14:46.200 --> 00:14:50.010 align:middle line:84%
So we're going to go
quickly through the

00:14:50.010 --> 00:14:52.470 align:middle line:90%
derivation of this formula.

00:14:52.470 --> 00:14:55.260 align:middle line:84%
And then, through examples
we'll try to get some

00:14:55.260 --> 00:14:58.480 align:middle line:84%
interpretation of what the
different terms here

00:14:58.480 --> 00:15:01.440 align:middle line:90%
correspond to.

00:15:01.440 --> 00:15:04.800 align:middle line:84%
All right, so let's try
to prove this formula.

00:15:04.800 --> 00:15:08.940 align:middle line:84%
And the proof is sort of a
useful exercise to make sure

00:15:08.940 --> 00:15:11.860 align:middle line:84%
you understand all the symbols
that are involved in here.

00:15:11.860 --> 00:15:14.850 align:middle line:84%
So the proof is not difficult,
it's 4 and 1/2 lines of

00:15:14.850 --> 00:15:18.220 align:middle line:84%
algebra, of just writing
down formulas.

00:15:18.220 --> 00:15:21.710 align:middle line:84%
But the challenge is to make
sure that at each point you

00:15:21.710 --> 00:15:25.070 align:middle line:84%
understand what each one
of the objects is.

00:15:25.070 --> 00:15:27.880 align:middle line:84%
So we go into formula for
the variance affects.

00:15:27.880 --> 00:15:32.480 align:middle line:84%
We know in general that the
variance of x has this nice

00:15:32.480 --> 00:15:34.590 align:middle line:84%
expression that we often
use to calculate it.

00:15:34.590 --> 00:15:37.340 align:middle line:84%
The expected value of the
squared of the random variable

00:15:37.340 --> 00:15:41.220 align:middle line:90%
minus the mean squared.

00:15:41.220 --> 00:15:45.290 align:middle line:84%
This formula, for the variances,
of course it should

00:15:45.290 --> 00:15:48.380 align:middle line:84%
apply to conditional
universes.

00:15:48.380 --> 00:15:50.430 align:middle line:84%
I mean it's a general formula
about variances.

00:15:50.430 --> 00:15:53.650 align:middle line:84%
If we put ourselves in a
conditional universe where the

00:15:53.650 --> 00:15:58.380 align:middle line:84%
random variable y is given to us
the same math should work.

00:15:58.380 --> 00:16:01.220 align:middle line:84%
So we should have a similar
formula for

00:16:01.220 --> 00:16:02.900 align:middle line:90%
the conditional variances.

00:16:02.900 --> 00:16:05.430 align:middle line:84%
It's just the same formula,
but applied to

00:16:05.430 --> 00:16:07.370 align:middle line:90%
the conditional universe.

00:16:07.370 --> 00:16:10.130 align:middle line:84%
The variance of x in the
conditional universe is the

00:16:10.130 --> 00:16:12.050 align:middle line:90%
expected value of x squared--

00:16:12.050 --> 00:16:13.770 align:middle line:90%
in the conditional universe--

00:16:13.770 --> 00:16:16.700 align:middle line:84%
minus the mean of x-- in the
conditional universe--

00:16:16.700 --> 00:16:17.730 align:middle line:90%
squared.

00:16:17.730 --> 00:16:20.350 align:middle line:90%
So this formula looks fine.

00:16:20.350 --> 00:16:23.620 align:middle line:84%
Now let's take expected
values of both sides.

00:16:23.620 --> 00:16:27.470 align:middle line:84%
Remember the conditional
variance is a random variable,

00:16:27.470 --> 00:16:30.600 align:middle line:84%
because its value depends on
whatever realization we get

00:16:30.600 --> 00:16:33.860 align:middle line:84%
for capital Y. So we can
take expectations here.

00:16:33.860 --> 00:16:36.320 align:middle line:84%
We get the expected value
of the variance.

00:16:36.320 --> 00:16:39.380 align:middle line:84%
Then we have the expected
value of a conditional

00:16:39.380 --> 00:16:40.740 align:middle line:90%
expectation.

00:16:40.740 --> 00:16:44.420 align:middle line:84%
Here we use the fact that
we discussed before.

00:16:44.420 --> 00:16:48.020 align:middle line:84%
The expected value of a
conditional expectation is the

00:16:48.020 --> 00:16:50.560 align:middle line:84%
same as the unconditional
expectation.

00:16:50.560 --> 00:16:52.780 align:middle line:90%
So this term becomes this.

00:16:52.780 --> 00:16:57.240 align:middle line:84%
And finally, here we just have
some weird looking random

00:16:57.240 --> 00:17:02.360 align:middle line:84%
variable, and we take the
expected value of it.

00:17:02.360 --> 00:17:06.210 align:middle line:84%
All right, now we need to do
something about this term.

00:17:06.210 --> 00:17:10.130 align:middle line:84%
Let's use the same
rule up here to

00:17:10.130 --> 00:17:14.030 align:middle line:90%
write down this variance.

00:17:14.030 --> 00:17:17.810 align:middle line:84%
So variance of an expectation,
that's kind of strange, but

00:17:17.810 --> 00:17:21.460 align:middle line:84%
you remember that the
conditional expectation is

00:17:21.460 --> 00:17:23.790 align:middle line:90%
random, because y is random.

00:17:23.790 --> 00:17:26.099 align:middle line:84%
So this thing is a random
variable, so

00:17:26.099 --> 00:17:28.390 align:middle line:90%
this thing has a variance.

00:17:28.390 --> 00:17:30.310 align:middle line:84%
What is the variance
of this thing?

00:17:30.310 --> 00:17:37.740 align:middle line:84%
It's the expected value of the
thing squared minus the square

00:17:37.740 --> 00:17:40.590 align:middle line:84%
of the expected value
of the thing.

00:17:40.590 --> 00:17:43.340 align:middle line:84%
Now what's the expected
value of that thing?

00:17:43.340 --> 00:17:47.230 align:middle line:84%
By the law of iterated
expectations, once more, the

00:17:47.230 --> 00:17:49.990 align:middle line:84%
expected value of this thing
is the unconditional

00:17:49.990 --> 00:17:51.090 align:middle line:90%
expectation.

00:17:51.090 --> 00:17:54.560 align:middle line:84%
And that's why here I put the
unconditional expectation.

00:17:54.560 --> 00:17:58.040 align:middle line:84%
So I'm using again this general
rule about how to

00:17:58.040 --> 00:18:01.510 align:middle line:84%
calculate variances, and I'm
applying it to calculate the

00:18:01.510 --> 00:18:05.680 align:middle line:84%
variance of the conditional
expectation.

00:18:05.680 --> 00:18:10.030 align:middle line:84%
And now you notice that if you
add these two expressions c

00:18:10.030 --> 00:18:15.040 align:middle line:84%
and d we get this plus
that, which is this.

00:18:15.040 --> 00:18:17.220 align:middle line:90%
It's equal to--

00:18:17.220 --> 00:18:22.360 align:middle line:84%
these two terms cancel, we're
left with this minus that,

00:18:22.360 --> 00:18:24.810 align:middle line:90%
which is the variance of x.

00:18:24.810 --> 00:18:27.430 align:middle line:84%
And that's the end
of the proof.

00:18:27.430 --> 00:18:31.105 align:middle line:84%
This one of those proofs that
do not convey any intuition.

00:18:31.105 --> 00:18:34.310 align:middle line:90%


00:18:34.310 --> 00:18:37.880 align:middle line:84%
This, as I said, it's a useful
proof to go through just to

00:18:37.880 --> 00:18:40.250 align:middle line:84%
make sure you understand
the symbols.

00:18:40.250 --> 00:18:44.020 align:middle line:84%
It starts to get pretty
confusing, and a little bit on

00:18:44.020 --> 00:18:45.490 align:middle line:90%
the abstract side.

00:18:45.490 --> 00:18:48.010 align:middle line:84%
So it's good to understand
what's going on.

00:18:48.010 --> 00:18:52.610 align:middle line:84%
Now there is intuition behind
this formula, some of which is

00:18:52.610 --> 00:18:54.780 align:middle line:84%
better left for later
in the class when

00:18:54.780 --> 00:18:56.680 align:middle line:90%
we talk about inference.

00:18:56.680 --> 00:19:01.380 align:middle line:84%
The idea is that the conditional
expectation you

00:19:01.380 --> 00:19:04.110 align:middle line:84%
can interpret it as an estimate
of the random

00:19:04.110 --> 00:19:06.700 align:middle line:84%
variable that you
are trying to--

00:19:06.700 --> 00:19:10.240 align:middle line:84%
an estimate of x based on
measurements of y, you can

00:19:10.240 --> 00:19:14.090 align:middle line:84%
think of these variances as
having something to do with an

00:19:14.090 --> 00:19:15.650 align:middle line:90%
estimation error.

00:19:15.650 --> 00:19:19.040 align:middle line:84%
And once you start thinking in
those terms an interpretation

00:19:19.040 --> 00:19:20.060 align:middle line:90%
will come about.

00:19:20.060 --> 00:19:23.750 align:middle line:84%
But again as I said this is
better left for when we start

00:19:23.750 --> 00:19:25.320 align:middle line:90%
talking about inference.

00:19:25.320 --> 00:19:28.080 align:middle line:84%
Nevertheless, we're going to get
some intuition about all

00:19:28.080 --> 00:19:33.010 align:middle line:84%
these formulas by considering
a baby example where we're

00:19:33.010 --> 00:19:35.900 align:middle line:84%
going to apply the law of
iterated expectations, and the

00:19:35.900 --> 00:19:38.060 align:middle line:90%
law of total variance.

00:19:38.060 --> 00:19:42.360 align:middle line:84%
So the baby example is that we
do this beautiful experiment

00:19:42.360 --> 00:19:47.190 align:middle line:84%
of giving a quiz to a class
consisting of many sections.

00:19:47.190 --> 00:19:49.325 align:middle line:84%
And we're interested in
two random variables.

00:19:49.325 --> 00:19:52.440 align:middle line:90%


00:19:52.440 --> 00:19:54.590 align:middle line:84%
So we have a number of students,
and they're all

00:19:54.590 --> 00:19:55.980 align:middle line:90%
allocated to sections.

00:19:55.980 --> 00:19:59.890 align:middle line:84%
The experiment is that I pick
a student at random, and I

00:19:59.890 --> 00:20:01.180 align:middle line:90%
look at two random variables.

00:20:01.180 --> 00:20:05.880 align:middle line:84%
One is the quiz score of the
randomly selected student, and

00:20:05.880 --> 00:20:09.960 align:middle line:84%
the other random variable is
the section number of the

00:20:09.960 --> 00:20:13.040 align:middle line:90%
student that I have selected.

00:20:13.040 --> 00:20:17.010 align:middle line:84%
We're given some statistics
about the two sections.

00:20:17.010 --> 00:20:19.960 align:middle line:84%
Section one has 10 students,
section two has 20 students.

00:20:19.960 --> 00:20:22.430 align:middle line:84%
The quiz average in section
one was 90.

00:20:22.430 --> 00:20:25.860 align:middle line:84%
Quiz average in section
two was 60.

00:20:25.860 --> 00:20:28.320 align:middle line:84%
What's the expected
value of x?

00:20:28.320 --> 00:20:32.990 align:middle line:84%
What's the expected quiz score
if I pick a student at random?

00:20:32.990 --> 00:20:34.420 align:middle line:90%
Well, each student has the same

00:20:34.420 --> 00:20:35.930 align:middle line:90%
probability of being selected.

00:20:35.930 --> 00:20:38.740 align:middle line:84%
I'm making that assumption
out of the 30 students.

00:20:38.740 --> 00:20:43.520 align:middle line:84%
I need to add the quiz scores
of all of the students.

00:20:43.520 --> 00:20:47.210 align:middle line:84%
So I need to add the quiz scores
in section one, which

00:20:47.210 --> 00:20:48.860 align:middle line:90%
is 90 times 10.

00:20:48.860 --> 00:20:51.030 align:middle line:84%
I need to add the quiz scores
in that section,

00:20:51.030 --> 00:20:52.720 align:middle line:90%
which is 60 times 20.

00:20:52.720 --> 00:20:55.220 align:middle line:84%
And we find that the overall
average was 70.

00:20:55.220 --> 00:20:58.310 align:middle line:84%
So this is the usual
unconditional expectation.

00:20:58.310 --> 00:21:00.990 align:middle line:84%
Let's look at the conditional
expectation, and let's look at

00:21:00.990 --> 00:21:03.000 align:middle line:84%
the elementary version
where we're talking

00:21:03.000 --> 00:21:04.690 align:middle line:90%
about numerical values.

00:21:04.690 --> 00:21:07.330 align:middle line:84%
If I tell you that the randomly
selected student was

00:21:07.330 --> 00:21:10.780 align:middle line:84%
in section one what's the
expected value of the quiz

00:21:10.780 --> 00:21:12.490 align:middle line:90%
score of that student?

00:21:12.490 --> 00:21:16.900 align:middle line:84%
Well, given this information,
we're picking a random student

00:21:16.900 --> 00:21:20.820 align:middle line:84%
uniformly from that section in
which the average was 90.

00:21:20.820 --> 00:21:23.070 align:middle line:84%
The expected value of the
score of that student

00:21:23.070 --> 00:21:24.580 align:middle line:90%
is going to be 90.

00:21:24.580 --> 00:21:28.800 align:middle line:84%
So given the specific value of
y, the specific section, the

00:21:28.800 --> 00:21:31.280 align:middle line:84%
conditional expectation or the
expected value of the quiz

00:21:31.280 --> 00:21:34.470 align:middle line:84%
score is a specific number,
the number 90.

00:21:34.470 --> 00:21:37.900 align:middle line:84%
Similarly for the second section
the expected value is

00:21:37.900 --> 00:21:41.480 align:middle line:84%
60, that's the average score
in the second section.

00:21:41.480 --> 00:21:42.940 align:middle line:84%
This is the elementary
version.

00:21:42.940 --> 00:21:45.000 align:middle line:84%
What about the abstract
version?

00:21:45.000 --> 00:21:48.350 align:middle line:84%
In the abstract version the
conditional expectation is a

00:21:48.350 --> 00:21:52.540 align:middle line:84%
random variable because
it depends.

00:21:52.540 --> 00:21:57.220 align:middle line:84%
In which section is the
student that I picked?

00:21:57.220 --> 00:22:01.680 align:middle line:84%
And with probability 1/3, I'm
going to pick a student in the

00:22:01.680 --> 00:22:04.890 align:middle line:84%
first section, in which case
the conditional expectation

00:22:04.890 --> 00:22:08.180 align:middle line:84%
will be 90, and with probability
2/3 I'm going to

00:22:08.180 --> 00:22:10.260 align:middle line:84%
pick a student in the
second section.

00:22:10.260 --> 00:22:12.450 align:middle line:84%
And in that case the conditional
expectation will

00:22:12.450 --> 00:22:14.220 align:middle line:90%
take the value of 60.

00:22:14.220 --> 00:22:17.020 align:middle line:84%
So this illustrates the idea
that the conditional

00:22:17.020 --> 00:22:19.300 align:middle line:84%
expectation is a random
variable.

00:22:19.300 --> 00:22:21.760 align:middle line:84%
Depending on what y is going
to be, the conditional

00:22:21.760 --> 00:22:25.320 align:middle line:84%
expectation is going to be one
or the other value with

00:22:25.320 --> 00:22:27.260 align:middle line:90%
certain probabilities.

00:22:27.260 --> 00:22:29.230 align:middle line:84%
Now that we have the
distribution of the

00:22:29.230 --> 00:22:31.610 align:middle line:84%
conditional expectation
we can calculate the

00:22:31.610 --> 00:22:33.560 align:middle line:90%
expected value of it.

00:22:33.560 --> 00:22:37.220 align:middle line:84%
And the expected value of such a
random variable is 1/3 times

00:22:37.220 --> 00:22:44.000 align:middle line:84%
90, plus 2/3 times 60, and
it comes out to equal 70.

00:22:44.000 --> 00:22:49.020 align:middle line:84%
Which miraculously is the same
number that we got up there.

00:22:49.020 --> 00:22:53.060 align:middle line:84%
So this tells you that you can
calculate the overall average

00:22:53.060 --> 00:22:58.180 align:middle line:84%
in a large class by taking the
averages in each one of the

00:22:58.180 --> 00:23:02.900 align:middle line:84%
sections and weighing each one
of the sections according to

00:23:02.900 --> 00:23:06.320 align:middle line:84%
the number of students
that it has.

00:23:06.320 --> 00:23:10.560 align:middle line:84%
So this section had 90 students
but only 1/3 of the

00:23:10.560 --> 00:23:13.850 align:middle line:84%
students, so it gets
a weight of 1/3.

00:23:13.850 --> 00:23:16.520 align:middle line:84%
So the law of iterated
expectations, once more, is

00:23:16.520 --> 00:23:18.540 align:middle line:90%
nothing too complicated.

00:23:18.540 --> 00:23:20.770 align:middle line:84%
It's just that you can calculate
overall class

00:23:20.770 --> 00:23:22.780 align:middle line:84%
average by looking
at the section

00:23:22.780 --> 00:23:26.330 align:middle line:90%
averages and combine them.

00:23:26.330 --> 00:23:28.680 align:middle line:84%
Now since the conditional
expectation is a random

00:23:28.680 --> 00:23:31.860 align:middle line:84%
variable, of course it has
a variance of it's own.

00:23:31.860 --> 00:23:34.080 align:middle line:84%
So let's calculate
the variance.

00:23:34.080 --> 00:23:36.060 align:middle line:90%
How do we calculate variances?

00:23:36.060 --> 00:23:38.960 align:middle line:84%
We look at all the possible
numerical values of this

00:23:38.960 --> 00:23:42.270 align:middle line:84%
random variable, which
are 90 and 60.

00:23:42.270 --> 00:23:45.620 align:middle line:84%
We look at the difference of
those possible numerical

00:23:45.620 --> 00:23:49.910 align:middle line:84%
values from the mean of this
random variable, and the mean

00:23:49.910 --> 00:23:53.770 align:middle line:84%
of that random variable, we
found that's it's 70.

00:23:53.770 --> 00:23:57.480 align:middle line:84%
And then we weight the different
possible numerical

00:23:57.480 --> 00:23:59.960 align:middle line:84%
values according to their
probabilities.

00:23:59.960 --> 00:24:03.930 align:middle line:84%
So with probability 1/3 the
conditional expectation is 90,

00:24:03.930 --> 00:24:06.940 align:middle line:84%
which is 20 away
from the mean.

00:24:06.940 --> 00:24:08.470 align:middle line:84%
And we get this squared
distance.

00:24:08.470 --> 00:24:11.750 align:middle line:84%
With probability 2/3 the
conditional expectation is 60,

00:24:11.750 --> 00:24:14.400 align:middle line:84%
which is 10 away from the
mean, has this squared

00:24:14.400 --> 00:24:16.910 align:middle line:84%
distance and gets weighed
by 2/3, which is the

00:24:16.910 --> 00:24:18.470 align:middle line:90%
probability of 60.

00:24:18.470 --> 00:24:21.130 align:middle line:84%
So you do the numbers, and you
get the value for the variance

00:24:21.130 --> 00:24:26.800 align:middle line:90%
equal to 200.

00:24:26.800 --> 00:24:30.250 align:middle line:84%
All right, so now we want to
move towards using that more

00:24:30.250 --> 00:24:33.770 align:middle line:84%
complicated formula involving
the conditional variances.

00:24:33.770 --> 00:24:36.650 align:middle line:90%


00:24:36.650 --> 00:24:40.470 align:middle line:84%
OK, suppose someone goes and
calculates the variance of the

00:24:40.470 --> 00:24:44.060 align:middle line:84%
quiz scores inside each
one of the sections.

00:24:44.060 --> 00:24:47.680 align:middle line:84%
So someone gives us these two
pieces of information.

00:24:47.680 --> 00:24:53.230 align:middle line:84%
In section one we take the
differences from the mean in

00:24:53.230 --> 00:24:57.900 align:middle line:84%
that section, and let's say that
the various turns out to

00:24:57.900 --> 00:25:00.240 align:middle line:84%
be a number equal
to 10 similarly

00:25:00.240 --> 00:25:01.410 align:middle line:90%
in the second section.

00:25:01.410 --> 00:25:05.280 align:middle line:84%
So these are the variances
of the quiz scores inside

00:25:05.280 --> 00:25:07.520 align:middle line:90%
individual sections.

00:25:07.520 --> 00:25:09.850 align:middle line:84%
The variance in one conditional
universe, the

00:25:09.850 --> 00:25:13.290 align:middle line:84%
variance in the other
conditional universe.

00:25:13.290 --> 00:25:18.860 align:middle line:84%
So if I pick a student in
section one and I don't tell

00:25:18.860 --> 00:25:21.400 align:middle line:84%
you anything more about the
student, what's the variance

00:25:21.400 --> 00:25:23.530 align:middle line:84%
of the random score
of that student?

00:25:23.530 --> 00:25:25.810 align:middle line:90%
The variance is 10.

00:25:25.810 --> 00:25:28.210 align:middle line:84%
I know why, but I don't
know the student.

00:25:28.210 --> 00:25:31.260 align:middle line:84%
So the score is still a random
variable in that universe.

00:25:31.260 --> 00:25:33.860 align:middle line:84%
It has a variance, and
that's the variance.

00:25:33.860 --> 00:25:36.330 align:middle line:84%
Similarly, in the other
universe, the variance of the

00:25:36.330 --> 00:25:39.110 align:middle line:84%
quiz scores is this
number, 20.

00:25:39.110 --> 00:25:42.650 align:middle line:84%
Once more, this is an equality
between numbers.

00:25:42.650 --> 00:25:44.920 align:middle line:84%
I have fixed the specific
value of y.

00:25:44.920 --> 00:25:48.440 align:middle line:84%
So I put myself in a specific
universe, I can calculate the

00:25:48.440 --> 00:25:51.430 align:middle line:84%
variance in that specific
universe.

00:25:51.430 --> 00:25:55.150 align:middle line:84%
If I don't specify a numerical
value for capital Y, and say I

00:25:55.150 --> 00:25:58.390 align:middle line:84%
don't know what Y is going to
be, it's going to be random.

00:25:58.390 --> 00:26:02.510 align:middle line:84%
Then what kind of section
variance I'm going to get

00:26:02.510 --> 00:26:04.500 align:middle line:90%
itself will be random.

00:26:04.500 --> 00:26:09.530 align:middle line:84%
With probability 1/3, I pick a
student in the first section

00:26:09.530 --> 00:26:14.740 align:middle line:84%
in which case the conditional
variance given what I have

00:26:14.740 --> 00:26:16.630 align:middle line:90%
picked is going to be 10.

00:26:16.630 --> 00:26:20.990 align:middle line:84%
Or with probability 2/3 I pick
y equal to 2, and I place

00:26:20.990 --> 00:26:22.690 align:middle line:90%
myself in that universe.

00:26:22.690 --> 00:26:25.790 align:middle line:84%
And in that universe the
conditional variance is 20.

00:26:25.790 --> 00:26:28.320 align:middle line:84%
So you see again from here that
the conditional variance

00:26:28.320 --> 00:26:32.410 align:middle line:84%
is a random variable that takes
different values with

00:26:32.410 --> 00:26:33.920 align:middle line:90%
certain probabilities.

00:26:33.920 --> 00:26:37.830 align:middle line:84%
And which value it takes depends
on the realization of

00:26:37.830 --> 00:26:41.670 align:middle line:84%
the random variable capital Y.
So this happens if capital Y

00:26:41.670 --> 00:26:45.970 align:middle line:84%
is one, this happens if capital
Y is equal to 2.

00:26:45.970 --> 00:26:50.000 align:middle line:84%
Once you have something
of this form--

00:26:50.000 --> 00:26:52.040 align:middle line:84%
a random variable that takes
values with certain

00:26:52.040 --> 00:26:53.150 align:middle line:90%
probabilities--

00:26:53.150 --> 00:26:55.690 align:middle line:84%
then you can certainly calculate
the expected value

00:26:55.690 --> 00:26:57.320 align:middle line:90%
of that random variable.

00:26:57.320 --> 00:27:00.110 align:middle line:84%
Don't get intimidated by the
fact that this random

00:27:00.110 --> 00:27:03.555 align:middle line:84%
variable, it's something that's
described by a string

00:27:03.555 --> 00:27:07.850 align:middle line:84%
of eight symbols, or
seven, instead of

00:27:07.850 --> 00:27:09.440 align:middle line:90%
just a single letter.

00:27:09.440 --> 00:27:15.290 align:middle line:84%
Think of this whole string of
symbols there as just being a

00:27:15.290 --> 00:27:16.940 align:middle line:90%
random variable.

00:27:16.940 --> 00:27:21.790 align:middle line:84%
You could call it z for example,
use one letter.

00:27:21.790 --> 00:27:25.990 align:middle line:84%
So z is a random variable that
takes these two values with

00:27:25.990 --> 00:27:27.990 align:middle line:84%
these corresponding
probabilities.

00:27:27.990 --> 00:27:31.210 align:middle line:84%
So we can talk about the
expected value of Z, which is

00:27:31.210 --> 00:27:35.560 align:middle line:84%
going to be 1/3 times 10, 2/3
times 20, and we get a certain

00:27:35.560 --> 00:27:38.260 align:middle line:90%
number from here.

00:27:38.260 --> 00:27:41.620 align:middle line:84%
And now we have all the pieces
to calculate the overall

00:27:41.620 --> 00:27:43.620 align:middle line:90%
variance of x.

00:27:43.620 --> 00:27:49.330 align:middle line:84%
The formula from the previous
slide tells us this.

00:27:49.330 --> 00:27:51.310 align:middle line:90%
Do we have all the pieces?

00:27:51.310 --> 00:27:53.190 align:middle line:84%
The expected value of
the variance, we

00:27:53.190 --> 00:27:55.160 align:middle line:90%
just calculated it.

00:27:55.160 --> 00:27:58.710 align:middle line:84%
The variance of the expected
value, this was the last

00:27:58.710 --> 00:28:00.410 align:middle line:84%
calculation in the
previous slide.

00:28:00.410 --> 00:28:03.490 align:middle line:84%
We did get a number for
it, it was 200.

00:28:03.490 --> 00:28:05.765 align:middle line:84%
You add the two, you find
the total variance.

00:28:05.765 --> 00:28:09.050 align:middle line:90%


00:28:09.050 --> 00:28:12.350 align:middle line:84%
Now the useful piece of this
exercise is to try to

00:28:12.350 --> 00:28:16.490 align:middle line:84%
interpret these two numbers,
and see what they mean.

00:28:16.490 --> 00:28:20.350 align:middle line:90%


00:28:20.350 --> 00:28:26.670 align:middle line:84%
The variance of x given y for
a specific y is the variance

00:28:26.670 --> 00:28:28.850 align:middle line:90%
inside section one.

00:28:28.850 --> 00:28:31.820 align:middle line:84%
This is the variance
inside section two.

00:28:31.820 --> 00:28:34.940 align:middle line:84%
The expected value is some
kind of average of the

00:28:34.940 --> 00:28:38.440 align:middle line:84%
variances inside individual
sections.

00:28:38.440 --> 00:28:41.770 align:middle line:84%
So this term tells us
something about the

00:28:41.770 --> 00:28:46.010 align:middle line:84%
variability of this course,
how widely spread they are

00:28:46.010 --> 00:28:47.856 align:middle line:90%
within individual sections.

00:28:47.856 --> 00:28:50.580 align:middle line:90%


00:28:50.580 --> 00:28:57.870 align:middle line:84%
So we have three sections, and
this course happens to be--

00:28:57.870 --> 00:29:01.180 align:middle line:84%
OK, let's say the sections
are really different.

00:29:01.180 --> 00:29:03.190 align:middle line:84%
So here you have undergraduates
and here you

00:29:03.190 --> 00:29:05.860 align:middle line:90%
have post-doctoral students.

00:29:05.860 --> 00:29:08.590 align:middle line:84%
And these are the quiz scores,
that's section one, section

00:29:08.590 --> 00:29:09.960 align:middle line:90%
two, section three.

00:29:09.960 --> 00:29:13.360 align:middle line:84%
Here's the mean of the
first section.

00:29:13.360 --> 00:29:16.200 align:middle line:84%
And the variance has something
to do with the spread.

00:29:16.200 --> 00:29:18.430 align:middle line:84%
The variance in the second
section has something to do

00:29:18.430 --> 00:29:21.830 align:middle line:84%
with the spread, similarly
with the third spread.

00:29:21.830 --> 00:29:28.220 align:middle line:84%
And the expected value of the
conditional variances is some

00:29:28.220 --> 00:29:31.690 align:middle line:84%
weighted average of the three
variances that we get from

00:29:31.690 --> 00:29:33.720 align:middle line:90%
individual sections.

00:29:33.720 --> 00:29:37.060 align:middle line:84%
So variability within sections
definitely contributes

00:29:37.060 --> 00:29:40.000 align:middle line:84%
something to the overall
variability of this course.

00:29:40.000 --> 00:29:45.340 align:middle line:84%
But if you ask me about the
variability over the entire

00:29:45.340 --> 00:29:47.740 align:middle line:90%
class there's a second effect.

00:29:47.740 --> 00:29:50.470 align:middle line:84%
That has to do with the fact
that different sections are

00:29:50.470 --> 00:29:52.660 align:middle line:84%
very different from
each other.

00:29:52.660 --> 00:29:59.440 align:middle line:84%
That these courses here are far
away from those scores.

00:29:59.440 --> 00:30:02.490 align:middle line:84%
And this term is the one
that does the job.

00:30:02.490 --> 00:30:08.410 align:middle line:84%
This one looks at the expected
values inside each section,

00:30:08.410 --> 00:30:12.840 align:middle line:84%
and these expected values are
this, this, and that.

00:30:12.840 --> 00:30:18.230 align:middle line:84%
And asks a question how widely
spread are they?

00:30:18.230 --> 00:30:23.000 align:middle line:84%
It asks how different from
each other are the means

00:30:23.000 --> 00:30:25.400 align:middle line:90%
inside individual sections?

00:30:25.400 --> 00:30:28.280 align:middle line:84%
And in this picture it would be
a large number because the

00:30:28.280 --> 00:30:31.980 align:middle line:84%
difference section means
are quite different.

00:30:31.980 --> 00:30:35.890 align:middle line:84%
So the story that this formula
is telling us is that the

00:30:35.890 --> 00:30:40.810 align:middle line:84%
overall variability of the quiz
scores consists of two

00:30:40.810 --> 00:30:44.720 align:middle line:84%
factors that can be quantified
and added.

00:30:44.720 --> 00:30:49.580 align:middle line:84%
One factor is how much
variability is there inside

00:30:49.580 --> 00:30:51.420 align:middle line:90%
individual sections?

00:30:51.420 --> 00:30:54.990 align:middle line:84%
And the other factor is how
different are the sections

00:30:54.990 --> 00:30:56.100 align:middle line:90%
from each other?

00:30:56.100 --> 00:30:58.620 align:middle line:84%
Both effects contribute
to the overall

00:30:58.620 --> 00:30:59.885 align:middle line:90%
variability of this course.

00:30:59.885 --> 00:31:03.920 align:middle line:90%


00:31:03.920 --> 00:31:08.290 align:middle line:84%
Let's continue with just one
more numerical example.

00:31:08.290 --> 00:31:11.730 align:middle line:84%
Just to get the hang of doing
these kinds of calculations,

00:31:11.730 --> 00:31:15.810 align:middle line:84%
and apply this formula to do a
divide and conquer calculation

00:31:15.810 --> 00:31:18.270 align:middle line:84%
of the variance of a
random variable.

00:31:18.270 --> 00:31:20.830 align:middle line:84%
Just for variety now we're going
to take a continuous

00:31:20.830 --> 00:31:22.140 align:middle line:90%
random variable.

00:31:22.140 --> 00:31:25.890 align:middle line:84%
Somebody gives you a PDF if this
form, and they ask you

00:31:25.890 --> 00:31:26.640 align:middle line:90%
for the variance.

00:31:26.640 --> 00:31:29.490 align:middle line:84%
And you say oh that's too
complicated, I don't want to

00:31:29.490 --> 00:31:30.350 align:middle line:90%
do integrals.

00:31:30.350 --> 00:31:32.480 align:middle line:90%
Can I divide and conquer?

00:31:32.480 --> 00:31:35.210 align:middle line:84%
And you say OK, let me do
the following trick.

00:31:35.210 --> 00:31:37.830 align:middle line:84%
Let me define a random
variable, y.

00:31:37.830 --> 00:31:43.450 align:middle line:84%
Which takes the value 1 if x
falls in here, and takes the

00:31:43.450 --> 00:31:47.080 align:middle line:84%
value 2 if x falls in
the second interval.

00:31:47.080 --> 00:31:51.340 align:middle line:84%
And let me try to work in the
conditional world where things

00:31:51.340 --> 00:31:54.340 align:middle line:84%
might be easier, and then
add things up to

00:31:54.340 --> 00:31:57.540 align:middle line:90%
get the overall variance.

00:31:57.540 --> 00:32:01.500 align:middle line:84%
So I have defined y this
particular way.

00:32:01.500 --> 00:32:04.562 align:middle line:84%
In this example y becomes
a function of x.

00:32:04.562 --> 00:32:07.370 align:middle line:84%
y is completely determined
by x.

00:32:07.370 --> 00:32:11.230 align:middle line:84%
And I'm going to calculate the
overall variance by trying to

00:32:11.230 --> 00:32:14.420 align:middle line:84%
calculate all of the terms
that are involved here.

00:32:14.420 --> 00:32:16.430 align:middle line:90%
So let's start calculating.

00:32:16.430 --> 00:32:21.690 align:middle line:84%
First observation is that this
event has probability 1/3, and

00:32:21.690 --> 00:32:24.390 align:middle line:84%
this event has probability
2/3.

00:32:24.390 --> 00:32:28.480 align:middle line:84%
The expected value of x given
that we are in this universe

00:32:28.480 --> 00:32:31.260 align:middle line:84%
is 1/2, because we
have a uniform

00:32:31.260 --> 00:32:33.350 align:middle line:90%
distribution from 0 to 1.

00:32:33.350 --> 00:32:36.630 align:middle line:84%
Here we have a uniform
distribution from 1 to 2, so

00:32:36.630 --> 00:32:40.820 align:middle line:84%
the conditional expectation of
x in that universe is 3/2.

00:32:40.820 --> 00:32:43.200 align:middle line:84%
How about conditional
variances?

00:32:43.200 --> 00:32:48.920 align:middle line:84%
In the world who are y is equal
to 1 x has a uniform

00:32:48.920 --> 00:32:50.770 align:middle line:84%
distribution on a
unit interval.

00:32:50.770 --> 00:32:53.090 align:middle line:90%
What's the variance of x?

00:32:53.090 --> 00:32:57.480 align:middle line:84%
By now you've probably seen that
formula, it's 1 over 12.

00:32:57.480 --> 00:33:00.580 align:middle line:84%
1 over 12 is the variance of a
uniform distribution over a

00:33:00.580 --> 00:33:01.880 align:middle line:90%
unit interval.

00:33:01.880 --> 00:33:07.120 align:middle line:84%
When y is equal to 2 the
variance is again 1 over 12.

00:33:07.120 --> 00:33:10.850 align:middle line:84%
Because in this instance again
x has a uniform distribution

00:33:10.850 --> 00:33:13.360 align:middle line:84%
over an interval
of unit length.

00:33:13.360 --> 00:33:16.010 align:middle line:84%
What's the overall expected
value of x?

00:33:16.010 --> 00:33:19.080 align:middle line:84%
The way you find the overall
expected value is to consider

00:33:19.080 --> 00:33:21.370 align:middle line:84%
the different numerical values
of the conditional

00:33:21.370 --> 00:33:22.450 align:middle line:90%
expectation.

00:33:22.450 --> 00:33:25.570 align:middle line:84%
And weigh them according
to their probabilities.

00:33:25.570 --> 00:33:28.770 align:middle line:84%
So with probability 1/3
the conditional

00:33:28.770 --> 00:33:30.830 align:middle line:90%
expectation is 1/2.

00:33:30.830 --> 00:33:34.170 align:middle line:84%
And with probability
2/3 the conditional

00:33:34.170 --> 00:33:36.460 align:middle line:90%
expectation is 3 over 2.

00:33:36.460 --> 00:33:39.555 align:middle line:84%
And this turns out
to be 7 over 6.

00:33:39.555 --> 00:33:45.080 align:middle line:90%


00:33:45.080 --> 00:33:48.450 align:middle line:84%
So this is the advance work
we need to do, now let's

00:33:48.450 --> 00:33:50.660 align:middle line:90%
calculate a few things here.

00:33:50.660 --> 00:33:56.660 align:middle line:84%
What's the variance of the
expected value of x given y?

00:33:56.660 --> 00:34:00.800 align:middle line:84%
Expected value of x given y is
a random variable that takes

00:34:00.800 --> 00:34:06.600 align:middle line:84%
these two values with
these probabilities.

00:34:06.600 --> 00:34:10.610 align:middle line:84%
So to find the variance we
consider the probability that

00:34:10.610 --> 00:34:18.730 align:middle line:84%
the expected value takes the
numerical value of 1/2 minus

00:34:18.730 --> 00:34:23.659 align:middle line:84%
the mean of the conditional
expectation.

00:34:23.659 --> 00:34:26.820 align:middle line:84%
What's the mean of the
conditional expectation?

00:34:26.820 --> 00:34:28.560 align:middle line:84%
It's the unconditional
expectation.

00:34:28.560 --> 00:34:30.980 align:middle line:90%
So it's 7 over 6.

00:34:30.980 --> 00:34:32.889 align:middle line:90%
We just did that calculation.

00:34:32.889 --> 00:34:38.050 align:middle line:84%
So I'm putting here that number,
7 over 6 squared.

00:34:38.050 --> 00:34:41.830 align:middle line:84%
And then there's a second term
with probability 2/3, the

00:34:41.830 --> 00:34:48.760 align:middle line:84%
conditional expectation takes
this value of 3 over 2, which

00:34:48.760 --> 00:34:54.380 align:middle line:84%
is so much away from the mean,
and we get this contribution.

00:34:54.380 --> 00:34:57.800 align:middle line:84%
So this way we have calculated
the variance of the

00:34:57.800 --> 00:35:01.590 align:middle line:84%
conditional expectation,
this is this term.

00:35:01.590 --> 00:35:04.000 align:middle line:90%
What is this?

00:35:04.000 --> 00:35:05.940 align:middle line:84%
Any guesses what
this number is?

00:35:05.940 --> 00:35:09.900 align:middle line:90%


00:35:09.900 --> 00:35:11.740 align:middle line:90%
It's 1 over 12, why?

00:35:11.740 --> 00:35:15.740 align:middle line:84%
The conditional variance just
happened in this example to be

00:35:15.740 --> 00:35:18.550 align:middle line:90%
1 over 12 no matter what.

00:35:18.550 --> 00:35:21.240 align:middle line:84%
So the conditional variance
is a deterministic random

00:35:21.240 --> 00:35:23.530 align:middle line:84%
variable that takes
a constant value.

00:35:23.530 --> 00:35:27.110 align:middle line:84%
So the expected value of
this random variable

00:35:27.110 --> 00:35:29.490 align:middle line:90%
is just 1 over 12.

00:35:29.490 --> 00:35:35.460 align:middle line:84%
So we got the two pieces that we
need, and so we do have the

00:35:35.460 --> 00:35:39.515 align:middle line:84%
overall variance of the
random variable x.

00:35:39.515 --> 00:35:45.680 align:middle line:90%


00:35:45.680 --> 00:35:50.750 align:middle line:84%
So this was just an academic
example in order to get the

00:35:50.750 --> 00:35:56.660 align:middle line:84%
hang of how to manipulate
various quantities.

00:35:56.660 --> 00:36:00.480 align:middle line:84%
Now let's use what we have
learned and the tools that we

00:36:00.480 --> 00:36:04.410 align:middle line:84%
have to do something a little
more interesting.

00:36:04.410 --> 00:36:07.820 align:middle line:84%
OK, so by now you're all in
love with probabilities.

00:36:07.820 --> 00:36:11.590 align:middle line:84%
So over the weekend you're going
to bookstores to buy

00:36:11.590 --> 00:36:13.540 align:middle line:90%
probability books.

00:36:13.540 --> 00:36:19.110 align:middle line:84%
So you're going to visit a
random number bookstores, and

00:36:19.110 --> 00:36:23.900 align:middle line:84%
at each one of the bookstores
you're going to spend a random

00:36:23.900 --> 00:36:26.420 align:middle line:90%
amount of money.

00:36:26.420 --> 00:36:31.060 align:middle line:84%
So let n be the number of stores
that you are visiting.

00:36:31.060 --> 00:36:32.890 align:middle line:90%
So n is an integer--

00:36:32.890 --> 00:36:34.870 align:middle line:90%
non-negative random variable--

00:36:34.870 --> 00:36:37.050 align:middle line:84%
and perhaps you know
the distribution

00:36:37.050 --> 00:36:39.230 align:middle line:90%
of that random variable.

00:36:39.230 --> 00:36:44.080 align:middle line:84%
Each time that you walk into a
store your mind is clear from

00:36:44.080 --> 00:36:48.580 align:middle line:84%
whatever you did before, and you
just buy a random number

00:36:48.580 --> 00:36:51.530 align:middle line:84%
of books that has nothing to
do with how many books you

00:36:51.530 --> 00:36:53.650 align:middle line:90%
bought earlier on the day.

00:36:53.650 --> 00:36:55.890 align:middle line:84%
It has nothing to do with
how many stores you are

00:36:55.890 --> 00:36:57.490 align:middle line:90%
visiting, and so on.

00:36:57.490 --> 00:37:00.760 align:middle line:84%
So each time you enter as a
brand new person, and buy a

00:37:00.760 --> 00:37:02.180 align:middle line:84%
random number of books,
and spend a

00:37:02.180 --> 00:37:03.580 align:middle line:90%
random amount of money.

00:37:03.580 --> 00:37:07.160 align:middle line:84%
So what I'm saying, more
precisely, is that I'm making

00:37:07.160 --> 00:37:08.760 align:middle line:90%
the following assumptions.

00:37:08.760 --> 00:37:11.130 align:middle line:90%
That for each store i--

00:37:11.130 --> 00:37:14.360 align:middle line:84%
if you end up visiting
the i-th store--

00:37:14.360 --> 00:37:17.480 align:middle line:84%
the amount of money that you
spend is a random variable

00:37:17.480 --> 00:37:19.090 align:middle line:84%
that has a certain
distribution.

00:37:19.090 --> 00:37:23.410 align:middle line:84%
That distribution is the same
for each store, and the xi's

00:37:23.410 --> 00:37:26.890 align:middle line:84%
from store to store are
independent from each other.

00:37:26.890 --> 00:37:30.800 align:middle line:84%
And furthermore, the xi's are
all independent of n.

00:37:30.800 --> 00:37:34.130 align:middle line:84%
So how much I'm spending at the
store-- once I get in--

00:37:34.130 --> 00:37:37.280 align:middle line:84%
has nothing to do with how
many stores I'm visiting.

00:37:37.280 --> 00:37:40.700 align:middle line:84%
So this is the setting that
we're going to look at.

00:37:40.700 --> 00:37:45.470 align:middle line:84%
y is the total amount of money
that you did spend.

00:37:45.470 --> 00:37:48.790 align:middle line:84%
It's the sum of how much you
spent in the stores, but the

00:37:48.790 --> 00:37:53.980 align:middle line:84%
index goes up to capital N.
And what's the twist here?

00:37:53.980 --> 00:37:57.460 align:middle line:84%
It's that we're dealing with the
sum of independent random

00:37:57.460 --> 00:38:02.690 align:middle line:84%
variables except that how many
random variables we have is

00:38:02.690 --> 00:38:07.470 align:middle line:84%
not given to us ahead of time,
but it is chosen at random.

00:38:07.470 --> 00:38:12.480 align:middle line:84%
So it's a sum of a random number
of random variables.

00:38:12.480 --> 00:38:15.360 align:middle line:84%
We would like to calculate some
quantities that have to

00:38:15.360 --> 00:38:19.690 align:middle line:84%
do with y, in particular the
expected value of y, or the

00:38:19.690 --> 00:38:21.930 align:middle line:90%
variance of y.

00:38:21.930 --> 00:38:23.540 align:middle line:90%
How do we go about it?

00:38:23.540 --> 00:38:26.950 align:middle line:84%
OK, we know something about the
linearity of expectations.

00:38:26.950 --> 00:38:31.890 align:middle line:84%
That expectation of a sum is the
sum of the expectations.

00:38:31.890 --> 00:38:37.180 align:middle line:84%
But we have used that rule only
in the case where it's

00:38:37.180 --> 00:38:39.850 align:middle line:84%
the sum of a fixed number
of random variables.

00:38:39.850 --> 00:38:43.670 align:middle line:84%
So expected value of x plus y
plus z is expectation of x,

00:38:43.670 --> 00:38:46.390 align:middle line:84%
plus expectation of y, plus
expectation of z.

00:38:46.390 --> 00:38:48.960 align:middle line:84%
We know this for a fixed number
of random variables.

00:38:48.960 --> 00:38:53.140 align:middle line:84%
We don't know it, or how it
would work for the case of a

00:38:53.140 --> 00:38:54.430 align:middle line:90%
random number.

00:38:54.430 --> 00:38:57.870 align:middle line:84%
Well, if we know something
about the case for fixed

00:38:57.870 --> 00:39:01.730 align:middle line:84%
random variables let's transport
ourselves to a

00:39:01.730 --> 00:39:05.310 align:middle line:84%
conditional universe where the
number of random variables

00:39:05.310 --> 00:39:07.570 align:middle line:90%
we're summing is fixed.

00:39:07.570 --> 00:39:11.640 align:middle line:84%
So let's try to break the
problem divide and conquer by

00:39:11.640 --> 00:39:15.300 align:middle line:84%
conditioning on the different
possible values of the number

00:39:15.300 --> 00:39:17.290 align:middle line:84%
of bookstores that
we're visiting.

00:39:17.290 --> 00:39:19.860 align:middle line:84%
So let's work in the conditional
universe, find the

00:39:19.860 --> 00:39:24.950 align:middle line:84%
conditional expectation in this
universe, and then use

00:39:24.950 --> 00:39:29.630 align:middle line:84%
our law of iterated expectations
to see what

00:39:29.630 --> 00:39:32.840 align:middle line:90%
happens more generally.

00:39:32.840 --> 00:39:37.120 align:middle line:84%
If I told you that I visited
exactly little n stores, where

00:39:37.120 --> 00:39:40.420 align:middle line:84%
little n now is a number,
let's say 10.

00:39:40.420 --> 00:39:44.840 align:middle line:84%
Then the amount of money you're
spending is x1 plus x2

00:39:44.840 --> 00:39:51.060 align:middle line:84%
all the way up to x10 given
that we visited 10 stores.

00:39:51.060 --> 00:39:54.640 align:middle line:84%
So what I have done here is that
I've replaced the capital

00:39:54.640 --> 00:39:59.370 align:middle line:84%
N with little n, and I can do
this because I'm now in the

00:39:59.370 --> 00:40:01.160 align:middle line:84%
conditional universe
where I know that

00:40:01.160 --> 00:40:04.160 align:middle line:90%
capital N is little n.

00:40:04.160 --> 00:40:06.840 align:middle line:90%
Now little n is fixed.

00:40:06.840 --> 00:40:10.810 align:middle line:84%
We have assumed that n is
independent from the xi's.

00:40:10.810 --> 00:40:15.900 align:middle line:84%
So in this universe of a fixed
n this information here

00:40:15.900 --> 00:40:20.400 align:middle line:84%
doesn't tell me anything new
about the values of the x's.

00:40:20.400 --> 00:40:24.600 align:middle line:84%
If you're conditioning random
variables that are independent

00:40:24.600 --> 00:40:27.220 align:middle line:84%
from the random variables you
are interested in, the

00:40:27.220 --> 00:40:30.630 align:middle line:84%
conditioning has no effect,
and so it can be dropped.

00:40:30.630 --> 00:40:33.000 align:middle line:84%
So in this conditional universe
where you visit

00:40:33.000 --> 00:40:35.720 align:middle line:84%
exactly 10 stores the expected
amount of money you're

00:40:35.720 --> 00:40:40.840 align:middle line:84%
spending is the expectation of
the amount of money spent in

00:40:40.840 --> 00:40:44.350 align:middle line:84%
10 stores, which is the sum of
the expected amount of money

00:40:44.350 --> 00:40:45.880 align:middle line:90%
in each store.

00:40:45.880 --> 00:40:48.760 align:middle line:84%
Each one of these is the same
number, because the random

00:40:48.760 --> 00:40:50.960 align:middle line:84%
variables have identical
distributions.

00:40:50.960 --> 00:40:54.130 align:middle line:84%
So it's n times the expected
value of money you spent in a

00:40:54.130 --> 00:40:57.140 align:middle line:90%
typical store.

00:40:57.140 --> 00:41:02.240 align:middle line:84%
This is almost obvious without
doing it formally.

00:41:02.240 --> 00:41:05.010 align:middle line:84%
If I'm telling you that you're
visiting 10 stores, what you

00:41:05.010 --> 00:41:09.220 align:middle line:84%
expect to spend is 10 times the
amount you expect to spend

00:41:09.220 --> 00:41:12.180 align:middle line:90%
in each store individually.

00:41:12.180 --> 00:41:16.480 align:middle line:84%
Now let's take this equality
here and rewrite it in our

00:41:16.480 --> 00:41:20.030 align:middle line:84%
abstract notation, in terms
of random variables.

00:41:20.030 --> 00:41:22.170 align:middle line:84%
This is an equality
between numbers.

00:41:22.170 --> 00:41:25.440 align:middle line:84%
Expected value of y given that
you visit 10 stores is 10

00:41:25.440 --> 00:41:28.220 align:middle line:90%
times this particular number.

00:41:28.220 --> 00:41:30.345 align:middle line:84%
Let's translate it into
random variables.

00:41:30.345 --> 00:41:36.290 align:middle line:84%
In random variable notation,
the expected value of money

00:41:36.290 --> 00:41:39.610 align:middle line:84%
you're spending given the
number of stores--

00:41:39.610 --> 00:41:42.480 align:middle line:84%
but without telling you
a specific number--

00:41:42.480 --> 00:41:46.720 align:middle line:84%
is whatever that number of
stores turns out to be times

00:41:46.720 --> 00:41:49.300 align:middle line:90%
the expected value of x.

00:41:49.300 --> 00:41:55.110 align:middle line:84%
So this is a random variable
that takes this as a numerical

00:41:55.110 --> 00:41:58.150 align:middle line:84%
value whenever capital
N happens to be

00:41:58.150 --> 00:42:00.030 align:middle line:90%
equal to little n.

00:42:00.030 --> 00:42:04.570 align:middle line:84%
This is a random variable, which
by definition takes this

00:42:04.570 --> 00:42:07.450 align:middle line:84%
numerical value whenever
capital N is

00:42:07.450 --> 00:42:09.520 align:middle line:90%
equal to little n.

00:42:09.520 --> 00:42:14.960 align:middle line:84%
So no matter what capital N
happens to be what specific

00:42:14.960 --> 00:42:18.870 align:middle line:84%
value, little n, it takes
this is equal to that.

00:42:18.870 --> 00:42:21.590 align:middle line:84%
Therefore the value of this
random variable is going to be

00:42:21.590 --> 00:42:23.350 align:middle line:90%
equal to that random variable.

00:42:23.350 --> 00:42:26.750 align:middle line:84%
So as random variables, these
two random variables are equal

00:42:26.750 --> 00:42:28.000 align:middle line:90%
to each other.

00:42:28.000 --> 00:42:29.940 align:middle line:90%


00:42:29.940 --> 00:42:33.200 align:middle line:84%
And now we use the law of
iterated expectations.

00:42:33.200 --> 00:42:35.750 align:middle line:84%
The law of iterated expectations
tells us that the

00:42:35.750 --> 00:42:39.530 align:middle line:84%
overall expected value of y is
the expected value of the

00:42:39.530 --> 00:42:41.270 align:middle line:90%
conditional expectation.

00:42:41.270 --> 00:42:43.650 align:middle line:84%
We have a formula for the
conditional expectation.

00:42:43.650 --> 00:42:46.580 align:middle line:84%
It's n times expected
value of x.

00:42:46.580 --> 00:42:50.390 align:middle line:84%
Now the expected value
of x is a number.

00:42:50.390 --> 00:42:54.970 align:middle line:84%
Expected value of something
random times a number is

00:42:54.970 --> 00:42:58.320 align:middle line:84%
expected value of the
random variable

00:42:58.320 --> 00:42:59.820 align:middle line:90%
times the number itself.

00:42:59.820 --> 00:43:02.880 align:middle line:84%
We can take a number outside
the expectation.

00:43:02.880 --> 00:43:06.060 align:middle line:84%
So expected value of
x gets pulled out.

00:43:06.060 --> 00:43:09.790 align:middle line:84%
And that's the conclusion,
that overall the expected

00:43:09.790 --> 00:43:13.340 align:middle line:84%
amount of money you're going to
spend is equal to how many

00:43:13.340 --> 00:43:16.670 align:middle line:84%
stores you expect to visit on
the average, and how much

00:43:16.670 --> 00:43:22.050 align:middle line:84%
money you expect to spend on
each one on the average.

00:43:22.050 --> 00:43:24.890 align:middle line:84%
You might have guessed that
this is the answer.

00:43:24.890 --> 00:43:30.400 align:middle line:84%
If you expect to visit 10
stores, and you expect to

00:43:30.400 --> 00:43:34.460 align:middle line:84%
spend $100 on each store, then
yes, you expect to spend

00:43:34.460 --> 00:43:36.150 align:middle line:90%
$1,000 today.

00:43:36.150 --> 00:43:39.050 align:middle line:84%
You're not going to impress your
Harvard friends if you

00:43:39.050 --> 00:43:40.300 align:middle line:90%
tell them that story.

00:43:40.300 --> 00:43:42.900 align:middle line:90%


00:43:42.900 --> 00:43:46.410 align:middle line:84%
It's one of the cases where
reasoning, on the average,

00:43:46.410 --> 00:43:50.160 align:middle line:84%
does give you the plausible
answer.

00:43:50.160 --> 00:43:54.290 align:middle line:84%
But you will be able to impress
your Harvard friends

00:43:54.290 --> 00:43:56.940 align:middle line:84%
if you tell them that I can
actually calculate the

00:43:56.940 --> 00:44:01.510 align:middle line:84%
variance of how much
I can spend.

00:44:01.510 --> 00:44:05.500 align:middle line:84%
And we're going to work by
applying this formula that we

00:44:05.500 --> 00:44:09.710 align:middle line:84%
have, and the difficulty is
basically sorting out all

00:44:09.710 --> 00:44:14.360 align:middle line:84%
those terms here, and
what they mean.

00:44:14.360 --> 00:44:20.630 align:middle line:90%
So let's start with this term.

00:44:20.630 --> 00:44:23.460 align:middle line:84%
So the expected value of y given
that you're visiting n

00:44:23.460 --> 00:44:26.280 align:middle line:84%
stores is n times the
expected value of x.

00:44:26.280 --> 00:44:28.250 align:middle line:84%
That's what we did in
the previous slide.

00:44:28.250 --> 00:44:32.540 align:middle line:84%
So this thing is a random
variable, it has a variance.

00:44:32.540 --> 00:44:34.300 align:middle line:90%
What is the variance?

00:44:34.300 --> 00:44:39.240 align:middle line:84%
Is the variance of n times
the expected value of x.

00:44:39.240 --> 00:44:42.010 align:middle line:84%
Remember expected value
of x is a number.

00:44:42.010 --> 00:44:46.180 align:middle line:84%
So we're dealing with the
variance of n times a number.

00:44:46.180 --> 00:44:48.330 align:middle line:84%
What happens when you
multiply a random

00:44:48.330 --> 00:44:50.800 align:middle line:90%
variable by a constant?

00:44:50.800 --> 00:44:55.020 align:middle line:84%
The variance becomes the
previous variance times the

00:44:55.020 --> 00:44:56.650 align:middle line:90%
constant squared.

00:44:56.650 --> 00:45:01.900 align:middle line:84%
So the variance of this is the
variance of n times the square

00:45:01.900 --> 00:45:04.300 align:middle line:84%
of that constant that
we had here.

00:45:04.300 --> 00:45:08.570 align:middle line:84%
So this tells us the variance
of the expected

00:45:08.570 --> 00:45:10.290 align:middle line:90%
value of y given n.

00:45:10.290 --> 00:45:13.380 align:middle line:84%
This is the part of the
variability of how much money

00:45:13.380 --> 00:45:16.950 align:middle line:84%
you're spending, which is
attributed to the randomness,

00:45:16.950 --> 00:45:19.650 align:middle line:84%
or the variability, in
the number of stores

00:45:19.650 --> 00:45:21.380 align:middle line:90%
that you are visiting.

00:45:21.380 --> 00:45:24.450 align:middle line:84%
So the interpretation of the
two terms is there's

00:45:24.450 --> 00:45:27.760 align:middle line:84%
randomness in how much you're
going to spend, and this is

00:45:27.760 --> 00:45:32.480 align:middle line:84%
attributed to the randomness
in the number of stores

00:45:32.480 --> 00:45:36.660 align:middle line:84%
together with the randomness
inside individual stores.

00:45:36.660 --> 00:45:40.110 align:middle line:84%
Well, after I tell you how many
stores you're visiting.

00:45:40.110 --> 00:45:42.570 align:middle line:84%
So now let's deal with this
term-- the variance inside

00:45:42.570 --> 00:45:45.020 align:middle line:90%
individual stores.

00:45:45.020 --> 00:45:47.070 align:middle line:90%
Let's take it slow.

00:45:47.070 --> 00:45:50.490 align:middle line:84%
If I tell you that you're
visiting exactly little n

00:45:50.490 --> 00:45:54.220 align:middle line:84%
stores, then y is how much
money you spent in those

00:45:54.220 --> 00:45:55.490 align:middle line:90%
little n stores.

00:45:55.490 --> 00:45:59.480 align:middle line:84%
You're dealing with the sum of
little n random variables.

00:45:59.480 --> 00:46:01.290 align:middle line:84%
What is the variance
of the sum of

00:46:01.290 --> 00:46:03.120 align:middle line:90%
little n random variables?

00:46:03.120 --> 00:46:05.880 align:middle line:84%
It's the sum of their
variances.

00:46:05.880 --> 00:46:10.590 align:middle line:84%
So each store contributes a
variance of x, and you're

00:46:10.590 --> 00:46:12.600 align:middle line:90%
adding over little n stores.

00:46:12.600 --> 00:46:16.520 align:middle line:84%
That's the variance of money
spent if I tell you

00:46:16.520 --> 00:46:18.040 align:middle line:90%
the number of stores.

00:46:18.040 --> 00:46:26.430 align:middle line:84%
Now let's translate this into
random variable notation.

00:46:26.430 --> 00:46:30.310 align:middle line:84%
This is a random variable that
takes this numerical value

00:46:30.310 --> 00:46:33.630 align:middle line:84%
whenever capital N is
equal to little n.

00:46:33.630 --> 00:46:37.020 align:middle line:84%
This is a random variable that
takes this numerical value

00:46:37.020 --> 00:46:39.250 align:middle line:84%
whenever capital N is
equal to little n.

00:46:39.250 --> 00:46:40.760 align:middle line:90%
This is equal to that.

00:46:40.760 --> 00:46:43.960 align:middle line:84%
Therefore, these two are always
equal, no matter what

00:46:43.960 --> 00:46:45.400 align:middle line:90%
happens to y.

00:46:45.400 --> 00:46:49.100 align:middle line:84%
So we have an equality here
between random variables.

00:46:49.100 --> 00:46:51.620 align:middle line:84%
Now we take expectations
of both.

00:46:51.620 --> 00:46:56.160 align:middle line:84%
Expected value of the variance
is expected value of this.

00:46:56.160 --> 00:46:59.890 align:middle line:84%
OK it may look confusing to
think of the expected value of

00:46:59.890 --> 00:47:05.740 align:middle line:84%
the variance here, but the
variance of x is a number, not

00:47:05.740 --> 00:47:06.650 align:middle line:90%
a random variable.

00:47:06.650 --> 00:47:08.480 align:middle line:90%
You think of it as a constant.

00:47:08.480 --> 00:47:12.580 align:middle line:84%
So its expected value of n times
a constant gives us the

00:47:12.580 --> 00:47:16.420 align:middle line:84%
expected value of n times
the constant itself.

00:47:16.420 --> 00:47:20.840 align:middle line:84%
So now we got the second term
as well, and now we put

00:47:20.840 --> 00:47:24.900 align:middle line:84%
everything together, this plus
that to get an expression for

00:47:24.900 --> 00:47:28.050 align:middle line:90%
the overall variance of y.

00:47:28.050 --> 00:47:32.380 align:middle line:84%
Which again, as I said before,
the overall variability in y

00:47:32.380 --> 00:47:36.790 align:middle line:84%
has to do with the variability
of how much you spent inside

00:47:36.790 --> 00:47:39.210 align:middle line:90%
the typical store.

00:47:39.210 --> 00:47:43.000 align:middle line:84%
And the variability in
the number of stores

00:47:43.000 --> 00:47:45.510 align:middle line:90%
that you are visiting.

00:47:45.510 --> 00:47:48.820 align:middle line:90%
OK, so this is it for today.

00:47:48.820 --> 00:47:52.600 align:middle line:84%
We'll change subjects quite
radically from next time.

00:47:52.600 --> 00:47:53.850 align:middle line:90%