WEBVTT

00:00:00.000 --> 00:00:00.040 align:middle line:90%


00:00:00.040 --> 00:00:02.460 align:middle line:84%
The following content is
provided under a Creative

00:00:02.460 --> 00:00:03.870 align:middle line:90%
Commons license.

00:00:03.870 --> 00:00:06.910 align:middle line:84%
Your support will help MIT
OpenCourseWare continue to

00:00:06.910 --> 00:00:10.560 align:middle line:84%
offer high-quality educational
resources for free.

00:00:10.560 --> 00:00:13.460 align:middle line:84%
To make a donation or view
additional materials from

00:00:13.460 --> 00:00:19.290 align:middle line:84%
hundreds of MIT courses, visit
MIT OpenCourseWare at

00:00:19.290 --> 00:00:20.540 align:middle line:90%
ocw.mit.edu.

00:00:20.540 --> 00:00:22.460 align:middle line:90%


00:00:22.460 --> 00:00:26.290 align:middle line:90%
PROFESSOR: OK so let's start.

00:00:26.290 --> 00:00:28.120 align:middle line:84%
So today, we're going
to continue the

00:00:28.120 --> 00:00:29.550 align:middle line:90%
subject from last time.

00:00:29.550 --> 00:00:32.619 align:middle line:84%
And the subject is
random variables.

00:00:32.619 --> 00:00:35.890 align:middle line:84%
As we discussed, random
variables basically associate

00:00:35.890 --> 00:00:39.550 align:middle line:84%
numerical values with the
outcomes of an experiment.

00:00:39.550 --> 00:00:42.690 align:middle line:84%
And we want to learn how
to manipulate them.

00:00:42.690 --> 00:00:45.890 align:middle line:84%
Now to a large extent, what's
going to happen, what's

00:00:45.890 --> 00:00:50.850 align:middle line:84%
happening during this chapter,
is that we are revisiting the

00:00:50.850 --> 00:00:53.220 align:middle line:84%
same concepts we have
seen in chapter one.

00:00:53.220 --> 00:00:57.340 align:middle line:84%
But we're going to introduce
a lot of new notation, but

00:00:57.340 --> 00:01:00.210 align:middle line:84%
really dealing with the
same kind of stuff.

00:01:00.210 --> 00:01:05.500 align:middle line:84%
The only difference where we
go beyond the new notation,

00:01:05.500 --> 00:01:08.470 align:middle line:84%
the new concept in this chapter
is the concept of the

00:01:08.470 --> 00:01:10.470 align:middle line:84%
expectation or expected
values.

00:01:10.470 --> 00:01:14.010 align:middle line:84%
And we're going to learn how
to manipulate expectations.

00:01:14.010 --> 00:01:17.280 align:middle line:84%
So let us start with a quick
review of what we

00:01:17.280 --> 00:01:19.610 align:middle line:90%
discussed last time.

00:01:19.610 --> 00:01:22.830 align:middle line:84%
We talked about random
variables.

00:01:22.830 --> 00:01:25.860 align:middle line:84%
Loosely speaking, random
variables are random

00:01:25.860 --> 00:01:28.670 align:middle line:84%
quantities that result
from an experiment.

00:01:28.670 --> 00:01:31.950 align:middle line:84%
More precisely speaking,
mathematically speaking, a

00:01:31.950 --> 00:01:35.310 align:middle line:84%
random variable is a function
from the sample space to the

00:01:35.310 --> 00:01:35.900 align:middle line:90%
real numbers.

00:01:35.900 --> 00:01:39.370 align:middle line:84%
That is, you give me an outcome,
and based on that

00:01:39.370 --> 00:01:42.940 align:middle line:84%
outcome, I can tell you the
value of the random variable.

00:01:42.940 --> 00:01:45.750 align:middle line:84%
So the value of the random
variable is a function of the

00:01:45.750 --> 00:01:47.870 align:middle line:90%
outcome that we have.

00:01:47.870 --> 00:01:51.170 align:middle line:84%
Now given a random variable,
some of the numerical outcomes

00:01:51.170 --> 00:01:52.850 align:middle line:90%
are more likely than others.

00:01:52.850 --> 00:01:55.600 align:middle line:84%
And we want to say which ones
are more likely and

00:01:55.600 --> 00:01:57.420 align:middle line:90%
how likely they are.

00:01:57.420 --> 00:02:00.900 align:middle line:84%
And the way we do that is by
writing down the probabilities

00:02:00.900 --> 00:02:04.030 align:middle line:84%
of the different possible
numerical outcomes.

00:02:04.030 --> 00:02:05.530 align:middle line:90%
Notice here, the notation.

00:02:05.530 --> 00:02:08.660 align:middle line:84%
We use uppercase to denote
the random variable.

00:02:08.660 --> 00:02:11.840 align:middle line:84%
We use lowercase to denote
real numbers.

00:02:11.840 --> 00:02:15.880 align:middle line:84%
So the way you read this, this
is the probability that the

00:02:15.880 --> 00:02:20.160 align:middle line:84%
random variable, capital X,
happens to take the numerical

00:02:20.160 --> 00:02:21.840 align:middle line:90%
value, little x.

00:02:21.840 --> 00:02:25.780 align:middle line:84%
This is a concept that's
familiar from chapter one.

00:02:25.780 --> 00:02:28.810 align:middle line:84%
And this is just the new
notation we will be using for

00:02:28.810 --> 00:02:30.080 align:middle line:90%
that concept.

00:02:30.080 --> 00:02:33.690 align:middle line:84%
It's the Probability Mass
Function of the random

00:02:33.690 --> 00:02:37.620 align:middle line:84%
variable, capital X. So the
subscript just indicates which

00:02:37.620 --> 00:02:40.190 align:middle line:84%
random variable we're
talking about.

00:02:40.190 --> 00:02:44.070 align:middle line:84%
And it's the probability
assigned to

00:02:44.070 --> 00:02:45.300 align:middle line:90%
a particular outcome.

00:02:45.300 --> 00:02:48.060 align:middle line:84%
And we want to assign such
probabilities for all possibly

00:02:48.060 --> 00:02:49.400 align:middle line:90%
numerical values.

00:02:49.400 --> 00:02:52.690 align:middle line:84%
So you can think of this as
being a function of little x.

00:02:52.690 --> 00:02:56.910 align:middle line:84%
And it tells you how likely
every little x is going to be.

00:02:56.910 --> 00:02:59.310 align:middle line:84%
Now the new concept we
introduced last time is the

00:02:59.310 --> 00:03:02.520 align:middle line:84%
concept of the expected value
for random variable, which is

00:03:02.520 --> 00:03:04.160 align:middle line:90%
defined this way.

00:03:04.160 --> 00:03:06.870 align:middle line:84%
You look at all the
possible outcomes.

00:03:06.870 --> 00:03:10.900 align:middle line:84%
And you form some kind of
average of all the possible

00:03:10.900 --> 00:03:15.000 align:middle line:84%
numerical values over the random
variable capital X. You

00:03:15.000 --> 00:03:17.830 align:middle line:84%
consider all the possible
numerical values, and you form

00:03:17.830 --> 00:03:18.560 align:middle line:90%
an average.

00:03:18.560 --> 00:03:23.160 align:middle line:84%
In fact, it's a weighted average
where, to every little

00:03:23.160 --> 00:03:26.750 align:middle line:84%
x, you assign a weight equal to
the probability that that

00:03:26.750 --> 00:03:30.190 align:middle line:84%
particular little x is
going to be realized.

00:03:30.190 --> 00:03:34.700 align:middle line:90%


00:03:34.700 --> 00:03:38.100 align:middle line:84%
Now, as we discussed last time,
if you have a random

00:03:38.100 --> 00:03:41.220 align:middle line:84%
variable, you can take a
function of a random variable.

00:03:41.220 --> 00:03:43.860 align:middle line:84%
And that's going to be a
new random variable.

00:03:43.860 --> 00:03:47.870 align:middle line:84%
So if capital X is a random
variable and g is a function,

00:03:47.870 --> 00:03:51.920 align:middle line:84%
g of X is a new random
variable.

00:03:51.920 --> 00:03:53.090 align:middle line:90%
You do the experiment.

00:03:53.090 --> 00:03:54.310 align:middle line:90%
You get an outcome.

00:03:54.310 --> 00:03:56.950 align:middle line:84%
This determines the value of
X. And that determines the

00:03:56.950 --> 00:03:58.400 align:middle line:90%
value of g of X.

00:03:58.400 --> 00:04:01.790 align:middle line:84%
So the numerical value of g of
X is determined by whatever

00:04:01.790 --> 00:04:03.330 align:middle line:90%
happens in the experiment.

00:04:03.330 --> 00:04:04.150 align:middle line:90%
It's random.

00:04:04.150 --> 00:04:06.440 align:middle line:84%
And that makes it a
random variable.

00:04:06.440 --> 00:04:09.040 align:middle line:84%
Since it's a random variable,
it has an

00:04:09.040 --> 00:04:11.320 align:middle line:90%
expectation of its own.

00:04:11.320 --> 00:04:14.860 align:middle line:84%
So how would we calculate the
expectation of g of X?

00:04:14.860 --> 00:04:18.430 align:middle line:84%
You could proceed by just using
the definition, which

00:04:18.430 --> 00:04:23.170 align:middle line:84%
would require you to find the
PMF of the random variable g

00:04:23.170 --> 00:04:29.480 align:middle line:84%
of X. So find the PMF of g of X,
and then apply the formula

00:04:29.480 --> 00:04:31.260 align:middle line:84%
for the expected value
of a random

00:04:31.260 --> 00:04:33.360 align:middle line:90%
variable with known PMF.

00:04:33.360 --> 00:04:36.820 align:middle line:84%
But there is also a shortcut,
which is just a different way

00:04:36.820 --> 00:04:40.530 align:middle line:84%
of doing the counting and the
calculations, in which we do

00:04:40.530 --> 00:04:44.580 align:middle line:84%
not need to find the PMF of g
of X. We just work with the

00:04:44.580 --> 00:04:47.290 align:middle line:84%
PMF of the original
random variable.

00:04:47.290 --> 00:04:50.010 align:middle line:84%
And what this is saying is that
the average value of g of

00:04:50.010 --> 00:04:51.800 align:middle line:90%
X is obtained as follows.

00:04:51.800 --> 00:04:55.020 align:middle line:84%
You look at all the possible
results, the X's,

00:04:55.020 --> 00:04:56.460 align:middle line:90%
how likely they are.

00:04:56.460 --> 00:04:59.740 align:middle line:84%
And when that particular
X happens, this is

00:04:59.740 --> 00:05:01.470 align:middle line:90%
how much you get.

00:05:01.470 --> 00:05:05.690 align:middle line:84%
And so this way, you add
these things up.

00:05:05.690 --> 00:05:10.310 align:middle line:84%
And you get the average amount
that you're going to get, the

00:05:10.310 --> 00:05:13.120 align:middle line:84%
average value of g of X, where
you average over the

00:05:13.120 --> 00:05:16.130 align:middle line:84%
likelihoods of the
different X's.

00:05:16.130 --> 00:05:19.730 align:middle line:84%
Now expected values have some
properties that are always

00:05:19.730 --> 00:05:23.570 align:middle line:84%
true and some properties that
sometimes are not true.

00:05:23.570 --> 00:05:28.160 align:middle line:84%
So the property that is not
always true is that this would

00:05:28.160 --> 00:05:34.400 align:middle line:84%
be the same as g of the expected
value of X. So in

00:05:34.400 --> 00:05:36.730 align:middle line:90%
general, this is not true.

00:05:36.730 --> 00:05:40.260 align:middle line:84%
You cannot interchange function
and expectation,

00:05:40.260 --> 00:05:44.310 align:middle line:84%
which means you cannot reason
on the average, in general.

00:05:44.310 --> 00:05:45.780 align:middle line:90%
But there are some exceptions.

00:05:45.780 --> 00:05:49.780 align:middle line:84%
When g is a linear function,
then the expected value for a

00:05:49.780 --> 00:05:53.460 align:middle line:84%
linear function is the same as
that same linear function of

00:05:53.460 --> 00:05:54.470 align:middle line:90%
the expectation.

00:05:54.470 --> 00:05:57.470 align:middle line:84%
So for linear functions, so
for random variable, the

00:05:57.470 --> 00:06:00.010 align:middle line:90%
expectation behaves nicely.

00:06:00.010 --> 00:06:05.320 align:middle line:84%
So this is basically telling you
that, if X is degrees in

00:06:05.320 --> 00:06:09.440 align:middle line:84%
Celsius, alpha X plus b is
degrees in Fahrenheit, you can

00:06:09.440 --> 00:06:12.030 align:middle line:84%
first do the conversion
to Fahrenheit

00:06:12.030 --> 00:06:13.390 align:middle line:90%
and take the average.

00:06:13.390 --> 00:06:16.870 align:middle line:84%
Or you can find the average
temperature in Celsius, and

00:06:16.870 --> 00:06:21.270 align:middle line:84%
then do the conversion
to Fahrenheit.

00:06:21.270 --> 00:06:23.630 align:middle line:90%
Either is valid.

00:06:23.630 --> 00:06:27.370 align:middle line:84%
So the expected value tells us
something about where is the

00:06:27.370 --> 00:06:31.320 align:middle line:84%
center of the distribution, more
specifically, the center

00:06:31.320 --> 00:06:34.360 align:middle line:84%
of mass or the center of gravity
of the PMF, when you

00:06:34.360 --> 00:06:36.170 align:middle line:90%
plot it as a bar graph.

00:06:36.170 --> 00:06:39.880 align:middle line:84%
Besides the average value, you
may be interested in knowing

00:06:39.880 --> 00:06:45.810 align:middle line:84%
how far will you be from
the average, typically.

00:06:45.810 --> 00:06:48.880 align:middle line:84%
So let's look at this quantity,
X minus expected

00:06:48.880 --> 00:06:50.270 align:middle line:90%
value of X.

00:06:50.270 --> 00:06:53.620 align:middle line:84%
This is the distance from
the average value.

00:06:53.620 --> 00:06:58.230 align:middle line:84%
So for a random outcome of the
experiment, this quantity in

00:06:58.230 --> 00:07:01.260 align:middle line:84%
here measures how far
away from the mean

00:07:01.260 --> 00:07:03.380 align:middle line:90%
you happen to be.

00:07:03.380 --> 00:07:08.330 align:middle line:84%
This quantity inside the
brackets is a random variable.

00:07:08.330 --> 00:07:09.620 align:middle line:90%
Why?

00:07:09.620 --> 00:07:12.620 align:middle line:90%
Because capital X is random.

00:07:12.620 --> 00:07:15.910 align:middle line:84%
And what we have here is capital
X, which is random,

00:07:15.910 --> 00:07:17.620 align:middle line:90%
minus a number.

00:07:17.620 --> 00:07:20.130 align:middle line:84%
Remember, expected values
are numbers.

00:07:20.130 --> 00:07:22.340 align:middle line:84%
Now a random variable
minus a number is

00:07:22.340 --> 00:07:23.470 align:middle line:90%
a new random variable.

00:07:23.470 --> 00:07:26.340 align:middle line:84%
It has an expectation
of its own.

00:07:26.340 --> 00:07:30.600 align:middle line:84%
We can use the linearity rule,
expected value of something

00:07:30.600 --> 00:07:35.060 align:middle line:84%
minus something else is just
the difference of their

00:07:35.060 --> 00:07:36.120 align:middle line:90%
expected value.

00:07:36.120 --> 00:07:40.250 align:middle line:84%
So it's going to be expected
value of X minus the expected

00:07:40.250 --> 00:07:42.290 align:middle line:90%
value over this thing.

00:07:42.290 --> 00:07:44.180 align:middle line:90%
Now this thing is a number.

00:07:44.180 --> 00:07:46.700 align:middle line:84%
And the expected value
of a number is

00:07:46.700 --> 00:07:48.710 align:middle line:90%
just the number itself.

00:07:48.710 --> 00:07:51.660 align:middle line:84%
So we get from here that this
is expected value minus

00:07:51.660 --> 00:07:52.750 align:middle line:90%
expected value.

00:07:52.750 --> 00:07:55.510 align:middle line:90%
And we get zero.

00:07:55.510 --> 00:07:57.420 align:middle line:90%
What is this telling us?

00:07:57.420 --> 00:08:02.690 align:middle line:84%
That, on the average, the
assigned difference from the

00:08:02.690 --> 00:08:05.080 align:middle line:90%
mean is equal to zero.

00:08:05.080 --> 00:08:06.460 align:middle line:90%
That is, the mean is here.

00:08:06.460 --> 00:08:08.850 align:middle line:84%
Sometimes X will fall
to the right.

00:08:08.850 --> 00:08:11.770 align:middle line:84%
Sometimes X will fall
to the left.

00:08:11.770 --> 00:08:16.170 align:middle line:84%
On the average, the average
distance from the mean is

00:08:16.170 --> 00:08:19.390 align:middle line:84%
going to be zero, because
sometimes the realized

00:08:19.390 --> 00:08:21.950 align:middle line:84%
distance will be positive,
sometimes it will be negative.

00:08:21.950 --> 00:08:24.680 align:middle line:84%
Positives and negatives
cancel out.

00:08:24.680 --> 00:08:28.140 align:middle line:84%
So if we want to capture the
idea of how far are we from

00:08:28.140 --> 00:08:33.090 align:middle line:84%
the mean, just looking at the
assigned distance from the

00:08:33.090 --> 00:08:36.789 align:middle line:84%
mean is not going to give us
any useful information.

00:08:36.789 --> 00:08:40.200 align:middle line:84%
So if we want to say something
about how far we are,

00:08:40.200 --> 00:08:43.120 align:middle line:84%
typically, we should do
something different.

00:08:43.120 --> 00:08:47.810 align:middle line:84%
One possibility might be to take
the absolute values of

00:08:47.810 --> 00:08:49.520 align:middle line:90%
the differences.

00:08:49.520 --> 00:08:51.780 align:middle line:84%
And that's a quantity that
sometimes people are

00:08:51.780 --> 00:08:52.920 align:middle line:90%
interested in.

00:08:52.920 --> 00:08:58.540 align:middle line:84%
But it turns out that a more
useful quantity happens to be

00:08:58.540 --> 00:09:02.190 align:middle line:84%
the variance of a random
variable, which actually

00:09:02.190 --> 00:09:07.030 align:middle line:84%
measures the average squared
distance from the mean.

00:09:07.030 --> 00:09:11.230 align:middle line:84%
So you have a random outcome,
random results, random

00:09:11.230 --> 00:09:14.130 align:middle line:84%
numerical value of the
random variable.

00:09:14.130 --> 00:09:17.370 align:middle line:84%
It is a certain distance
away from the mean.

00:09:17.370 --> 00:09:19.390 align:middle line:84%
That certain distance
is random.

00:09:19.390 --> 00:09:21.210 align:middle line:90%
We take the square of that.

00:09:21.210 --> 00:09:23.200 align:middle line:84%
This is the squared distance
from the mean,

00:09:23.200 --> 00:09:24.750 align:middle line:90%
which is again random.

00:09:24.750 --> 00:09:27.990 align:middle line:84%
Since it's random, it has an
expected value of its own.

00:09:27.990 --> 00:09:32.000 align:middle line:84%
And that expected value, we call
it the variance of X. And

00:09:32.000 --> 00:09:35.350 align:middle line:84%
so we have this particular
definition.

00:09:35.350 --> 00:09:39.500 align:middle line:84%
Using the rule that we have up
here for how to calculate

00:09:39.500 --> 00:09:42.870 align:middle line:84%
expectations of functions
of a random variable,

00:09:42.870 --> 00:09:44.820 align:middle line:90%
why does that apply?

00:09:44.820 --> 00:09:48.980 align:middle line:84%
Well, what we have inside the
brackets here is a function of

00:09:48.980 --> 00:09:52.850 align:middle line:84%
the random variable, capital X.
So we can apply this rule

00:09:52.850 --> 00:09:55.750 align:middle line:84%
where g is this particular
function.

00:09:55.750 --> 00:09:58.540 align:middle line:84%
And we can use that to calculate
the variance,

00:09:58.540 --> 00:10:02.160 align:middle line:84%
starting with the PMF of the
random variable X. And then we

00:10:02.160 --> 00:10:05.200 align:middle line:84%
have a useful formula that's a
nice shortcut, sometimes, if

00:10:05.200 --> 00:10:07.620 align:middle line:84%
you want to do the
calculation.

00:10:07.620 --> 00:10:11.340 align:middle line:84%
Now one thing that's slightly
wrong with the variance is

00:10:11.340 --> 00:10:15.600 align:middle line:84%
that the units are not right, if
you want to talk about the

00:10:15.600 --> 00:10:16.930 align:middle line:90%
spread a of a distribution.

00:10:16.930 --> 00:10:20.830 align:middle line:84%
Suppose that X is a random
variable measured in meters.

00:10:20.830 --> 00:10:26.170 align:middle line:84%
The variance will have the
units of meters squared.

00:10:26.170 --> 00:10:28.390 align:middle line:84%
So it's a kind of a
different thing.

00:10:28.390 --> 00:10:31.120 align:middle line:84%
If you want to talk about the
spread of the distribution

00:10:31.120 --> 00:10:35.210 align:middle line:84%
using the same units as you have
for X, it's convenient to

00:10:35.210 --> 00:10:37.880 align:middle line:84%
take the square root
of the variance.

00:10:37.880 --> 00:10:39.580 align:middle line:84%
And that's something
that we define.

00:10:39.580 --> 00:10:42.980 align:middle line:84%
And we call it to the standard
deviation of X, or the

00:10:42.980 --> 00:10:46.140 align:middle line:84%
standard deviation of the
distribution of X. So it tells

00:10:46.140 --> 00:10:49.510 align:middle line:84%
you the amount of spread
in your distribution.

00:10:49.510 --> 00:10:52.980 align:middle line:84%
And it is in the same units as
the random variable itself

00:10:52.980 --> 00:10:54.230 align:middle line:90%
that you are dealing with.

00:10:54.230 --> 00:10:57.260 align:middle line:90%


00:10:57.260 --> 00:11:02.500 align:middle line:84%
And we can just illustrate
those quantities with an

00:11:02.500 --> 00:11:06.670 align:middle line:84%
example that's about as
simple as it can be.

00:11:06.670 --> 00:11:08.570 align:middle line:84%
So consider the following
experiment.

00:11:08.570 --> 00:11:11.060 align:middle line:84%
You're going to go from
here to New York,

00:11:11.060 --> 00:11:13.140 align:middle line:90%
let's say, 200 miles.

00:11:13.140 --> 00:11:15.500 align:middle line:90%
And you have two alternatives.

00:11:15.500 --> 00:11:20.640 align:middle line:84%
Either you'll get your private
plane and go at a speed of 200

00:11:20.640 --> 00:11:27.690 align:middle line:84%
miles per hour, constant speed
during your trip, or

00:11:27.690 --> 00:11:32.010 align:middle line:84%
otherwise, you'll decide to walk
really, really slowly, at

00:11:32.010 --> 00:11:35.120 align:middle line:84%
the leisurely pace of
one mile per hour.

00:11:35.120 --> 00:11:38.150 align:middle line:84%
So you pick the speed at
random by doing this

00:11:38.150 --> 00:11:39.820 align:middle line:84%
experiment, by flipping
a coin.

00:11:39.820 --> 00:11:41.510 align:middle line:84%
And with probability one-half,
you do one thing.

00:11:41.510 --> 00:11:44.230 align:middle line:84%
With probably one-half, you
do the other thing.

00:11:44.230 --> 00:11:47.890 align:middle line:84%
So your V is a random
variable.

00:11:47.890 --> 00:11:51.050 align:middle line:84%
In case you're interested in
how much time it's going to

00:11:51.050 --> 00:11:54.970 align:middle line:84%
take you to get there, well,
time is equal to distance

00:11:54.970 --> 00:11:56.920 align:middle line:90%
divided by speed.

00:11:56.920 --> 00:11:58.480 align:middle line:90%
So that's the formula.

00:11:58.480 --> 00:12:01.660 align:middle line:84%
The time itself is a random
variable, because it's a

00:12:01.660 --> 00:12:03.850 align:middle line:84%
function of V, which
is random.

00:12:03.850 --> 00:12:06.390 align:middle line:84%
How much time it's going to take
you depends on the coin

00:12:06.390 --> 00:12:09.270 align:middle line:84%
flip that you do in the
beginning to decide what speed

00:12:09.270 --> 00:12:11.920 align:middle line:90%
you are going to have.

00:12:11.920 --> 00:12:15.110 align:middle line:84%
OK, just as a warm up, the
trivial calculations.

00:12:15.110 --> 00:12:17.790 align:middle line:84%
To find the expected value of
V, you argue as follows.

00:12:17.790 --> 00:12:21.730 align:middle line:84%
With probability one-half,
V is going to be one.

00:12:21.730 --> 00:12:25.740 align:middle line:84%
And with probability one-half,
V is going to be 200.

00:12:25.740 --> 00:12:31.410 align:middle line:84%
And so the expected value
of your speed is 100.5.

00:12:31.410 --> 00:12:34.970 align:middle line:84%
If you wish to calculate the
variance of V, then you argue

00:12:34.970 --> 00:12:36.420 align:middle line:90%
as follows.

00:12:36.420 --> 00:12:40.650 align:middle line:84%
With probability one-half, I'm
going to travel at the speed

00:12:40.650 --> 00:12:45.920 align:middle line:84%
of one, whereas, the
mean is 100.5.

00:12:45.920 --> 00:12:49.090 align:middle line:84%
So this is the distance from
the mean, if I decide to

00:12:49.090 --> 00:12:50.380 align:middle line:90%
travel at the speed of one.

00:12:50.380 --> 00:12:52.920 align:middle line:84%
We take that distance from
the mean squared.

00:12:52.920 --> 00:12:55.330 align:middle line:84%
That's one contribution
to the variance.

00:12:55.330 --> 00:12:59.220 align:middle line:84%
And with probability one-half,
you're going to travel at the

00:12:59.220 --> 00:13:05.610 align:middle line:84%
speed of 200, which is this
much away from the mean.

00:13:05.610 --> 00:13:08.270 align:middle line:90%
You take the square of that.

00:13:08.270 --> 00:13:12.360 align:middle line:84%
OK, so approximately how
big is this number?

00:13:12.360 --> 00:13:14.370 align:middle line:84%
Well, this is roughly
100 squared.

00:13:14.370 --> 00:13:16.130 align:middle line:90%
That's also 100 squared.

00:13:16.130 --> 00:13:20.760 align:middle line:84%
So approximately, the variance
of this random

00:13:20.760 --> 00:13:25.060 align:middle line:90%
variable is 100 squared.

00:13:25.060 --> 00:13:28.090 align:middle line:84%
Now if I tell you that the
variance of this distribution

00:13:28.090 --> 00:13:31.980 align:middle line:84%
is 10,000, it doesn't
really help you to

00:13:31.980 --> 00:13:33.600 align:middle line:90%
relate it to this diagram.

00:13:33.600 --> 00:13:35.950 align:middle line:84%
Whereas, the standard deviation,
where you take the

00:13:35.950 --> 00:13:38.850 align:middle line:84%
square root, is more
interesting.

00:13:38.850 --> 00:13:43.770 align:middle line:84%
It's the square root of 100
squared, which is a 100.

00:13:43.770 --> 00:13:48.320 align:middle line:84%
And the standard deviation,
indeed, gives us a sense of

00:13:48.320 --> 00:13:53.280 align:middle line:84%
how spread out this distribution
is from the mean.

00:13:53.280 --> 00:13:57.440 align:middle line:84%
So the standard deviation
basically gives us some

00:13:57.440 --> 00:14:01.110 align:middle line:84%
indication about this spacing
that we have here.

00:14:01.110 --> 00:14:03.950 align:middle line:84%
It tells us the amount of spread
in our distribution.

00:14:03.950 --> 00:14:07.320 align:middle line:90%


00:14:07.320 --> 00:14:12.110 align:middle line:84%
OK, now let's look at what
happens to time.

00:14:12.110 --> 00:14:14.970 align:middle line:90%
V is a random variable.

00:14:14.970 --> 00:14:17.340 align:middle line:90%
T is a random variable.

00:14:17.340 --> 00:14:19.820 align:middle line:84%
So now let's look at the
expected values and all of

00:14:19.820 --> 00:14:23.730 align:middle line:90%
that for the time.

00:14:23.730 --> 00:14:29.250 align:middle line:84%
OK, so the time is a function
of a random variable.

00:14:29.250 --> 00:14:33.070 align:middle line:84%
We can find the expected time
by looking at all possible

00:14:33.070 --> 00:14:37.030 align:middle line:84%
outcomes of the experiment, the
V's, weigh them according

00:14:37.030 --> 00:14:39.820 align:middle line:84%
to their probabilities, and for
each particular V, keep

00:14:39.820 --> 00:14:42.880 align:middle line:84%
track of how much
time it took us.

00:14:42.880 --> 00:14:48.760 align:middle line:84%
So if V is one, which happens
with probability one-half, the

00:14:48.760 --> 00:14:52.440 align:middle line:84%
time it takes is going
to be 200.

00:14:52.440 --> 00:14:56.410 align:middle line:84%
If we travel at speed of one,
it takes us 200 time units.

00:14:56.410 --> 00:15:01.560 align:middle line:84%
And otherwise, if our speed is
equal to 200, the time is one.

00:15:01.560 --> 00:15:06.540 align:middle line:84%
So the expected value of T is
once more the same as before.

00:15:06.540 --> 00:15:09.740 align:middle line:90%
It's 100.5.

00:15:09.740 --> 00:15:13.840 align:middle line:84%
So the expected speed
is 100.5.

00:15:13.840 --> 00:15:19.320 align:middle line:84%
The expected time
is also 100.5.

00:15:19.320 --> 00:15:22.030 align:middle line:84%
So the product of these
expectations is

00:15:22.030 --> 00:15:24.350 align:middle line:90%
something like 10,000.

00:15:24.350 --> 00:15:29.230 align:middle line:84%
How about the expected value
of the product of T and V?

00:15:29.230 --> 00:15:32.460 align:middle line:90%
Well, T times V is 200.

00:15:32.460 --> 00:15:37.790 align:middle line:84%
No matter what outcome you have
in the experiment, in

00:15:37.790 --> 00:15:41.320 align:middle line:84%
that particular outcome, T
times V is total distance

00:15:41.320 --> 00:15:45.660 align:middle line:84%
traveled, which is
exactly 200.

00:15:45.660 --> 00:15:49.890 align:middle line:84%
And so what do we get in this
simple example is that the

00:15:49.890 --> 00:15:53.550 align:middle line:84%
expected value of the product of
these two random variables

00:15:53.550 --> 00:15:57.630 align:middle line:84%
is different than the product
of their expected values.

00:15:57.630 --> 00:16:01.120 align:middle line:84%
This is one more instance
of where we cannot

00:16:01.120 --> 00:16:03.500 align:middle line:90%
reason on the average.

00:16:03.500 --> 00:16:08.190 align:middle line:84%
So on the average, over a large
number of trips, your

00:16:08.190 --> 00:16:10.120 align:middle line:90%
average time would be 100.

00:16:10.120 --> 00:16:13.570 align:middle line:84%
On the average, over a large
number of trips, your average

00:16:13.570 --> 00:16:15.820 align:middle line:90%
speed would be 100.

00:16:15.820 --> 00:16:20.740 align:middle line:84%
But your average distance
traveled is not 100 times 100.

00:16:20.740 --> 00:16:22.580 align:middle line:90%
It's something else.

00:16:22.580 --> 00:16:26.850 align:middle line:84%
So you cannot reason on the
average, whenever you're

00:16:26.850 --> 00:16:28.850 align:middle line:84%
dealing with non-linear
things.

00:16:28.850 --> 00:16:31.410 align:middle line:84%
And the non-linear thing here
is that you have a function

00:16:31.410 --> 00:16:34.740 align:middle line:84%
which is a product of stuff,
as opposed to just

00:16:34.740 --> 00:16:37.330 align:middle line:90%
linear sums of stuff.

00:16:37.330 --> 00:16:42.530 align:middle line:84%
Another way to look at what's
happening here is the expected

00:16:42.530 --> 00:16:44.100 align:middle line:90%
value of the time.

00:16:44.100 --> 00:16:47.460 align:middle line:84%
Time, by definition, is
200 over the speed.

00:16:47.460 --> 00:16:52.100 align:middle line:84%
Expected value of the time, we
found it to be about a 100.

00:16:52.100 --> 00:17:02.800 align:middle line:84%
And so expected value of 200
over V is about a 100.

00:17:02.800 --> 00:17:07.280 align:middle line:84%
But it's different from this
quantity here, which is

00:17:07.280 --> 00:17:13.119 align:middle line:84%
roughly equal to
2, and so 200.

00:17:13.119 --> 00:17:15.430 align:middle line:84%
Expected value of
V is about 100.

00:17:15.430 --> 00:17:19.030 align:middle line:84%
So this quantity is about
equal to two.

00:17:19.030 --> 00:17:22.329 align:middle line:84%
Whereas, this quantity
up here is about 100.

00:17:22.329 --> 00:17:23.609 align:middle line:90%
So what do we have here?

00:17:23.609 --> 00:17:26.960 align:middle line:84%
We have a non-linear function
of V. And we find that the

00:17:26.960 --> 00:17:31.130 align:middle line:84%
expected value of this function
is not the same thing

00:17:31.130 --> 00:17:34.390 align:middle line:84%
as the function of the
expected value.

00:17:34.390 --> 00:17:38.560 align:middle line:84%
So again, that's an instance
where you cannot interchange

00:17:38.560 --> 00:17:40.820 align:middle line:90%
expected values and functions.

00:17:40.820 --> 00:17:42.770 align:middle line:84%
And that's because things
are non-linear.

00:17:42.770 --> 00:17:46.120 align:middle line:90%


00:17:46.120 --> 00:17:51.730 align:middle line:84%
OK, so now let us introduce
a new concept.

00:17:51.730 --> 00:17:56.030 align:middle line:84%
Or maybe it's not quite
a new concept.

00:17:56.030 --> 00:17:58.570 align:middle line:84%
So we discussed, in chapter
one, that we have

00:17:58.570 --> 00:17:59.470 align:middle line:90%
probabilities.

00:17:59.470 --> 00:18:03.170 align:middle line:84%
We also have conditional
probabilities.

00:18:03.170 --> 00:18:05.250 align:middle line:84%
What's the difference
between them?

00:18:05.250 --> 00:18:06.410 align:middle line:90%
Essentially, none.

00:18:06.410 --> 00:18:09.560 align:middle line:84%
Probabilities are just an
assignment of probability

00:18:09.560 --> 00:18:11.650 align:middle line:84%
values to give different
outcomes, given

00:18:11.650 --> 00:18:12.980 align:middle line:90%
a particular model.

00:18:12.980 --> 00:18:16.180 align:middle line:84%
Somebody comes and gives
you new information.

00:18:16.180 --> 00:18:18.370 align:middle line:84%
So you come up with
a new model.

00:18:18.370 --> 00:18:20.180 align:middle line:84%
And you have a new
probabilities.

00:18:20.180 --> 00:18:23.120 align:middle line:84%
We call these conditional
probabilities, but they taste

00:18:23.120 --> 00:18:27.230 align:middle line:84%
and behave exactly the same
as ordinary probabilities.

00:18:27.230 --> 00:18:30.110 align:middle line:84%
So since we can have conditional
probabilities, why

00:18:30.110 --> 00:18:35.020 align:middle line:84%
not have conditional PMFs as
well, since PMFs deal with

00:18:35.020 --> 00:18:37.120 align:middle line:90%
probabilities anyway.

00:18:37.120 --> 00:18:40.520 align:middle line:84%
So we have a random variable,
capital X. It has

00:18:40.520 --> 00:18:42.420 align:middle line:90%
a PMF of its own.

00:18:42.420 --> 00:18:46.810 align:middle line:84%
For example, it could be the PMF
in this picture, which is

00:18:46.810 --> 00:18:52.160 align:middle line:84%
a uniform PMF that takes for
possible different values.

00:18:52.160 --> 00:18:54.850 align:middle line:90%
And we also have an event.

00:18:54.850 --> 00:18:57.240 align:middle line:84%
And somebody comes
and tells us that

00:18:57.240 --> 00:18:59.920 align:middle line:90%
this event has occurred.

00:18:59.920 --> 00:19:02.700 align:middle line:84%
The PMF tells you the
probability that capital X

00:19:02.700 --> 00:19:04.730 align:middle line:90%
equals to some little x.

00:19:04.730 --> 00:19:08.930 align:middle line:84%
Somebody tells you that a
certain event has occurred

00:19:08.930 --> 00:19:12.320 align:middle line:84%
that's going to make you change
the probabilities that

00:19:12.320 --> 00:19:14.690 align:middle line:84%
you assign to the different
values.

00:19:14.690 --> 00:19:17.180 align:middle line:84%
You are going to use conditional
probabilities.

00:19:17.180 --> 00:19:20.880 align:middle line:84%
So this part, it's clear what
it means from chapter one.

00:19:20.880 --> 00:19:25.200 align:middle line:84%
And this part is just the new
notation we're using in this

00:19:25.200 --> 00:19:28.380 align:middle line:84%
chapter to talk about
conditional probabilities.

00:19:28.380 --> 00:19:31.200 align:middle line:90%
So this is just a definition.

00:19:31.200 --> 00:19:35.010 align:middle line:84%
So the conditional PMF
is an ordinary PMF.

00:19:35.010 --> 00:19:39.840 align:middle line:84%
But it's the PMF that applies
to a new model in which we

00:19:39.840 --> 00:19:42.480 align:middle line:84%
have been given some information
about the outcome

00:19:42.480 --> 00:19:43.840 align:middle line:90%
of the experiment.

00:19:43.840 --> 00:19:47.560 align:middle line:84%
So to make it concrete, consider
this event here.

00:19:47.560 --> 00:19:50.610 align:middle line:84%
Take the event that capital
X is bigger

00:19:50.610 --> 00:19:52.290 align:middle line:90%
than or equal to two.

00:19:52.290 --> 00:19:54.490 align:middle line:84%
In the picture, what
is the event A?

00:19:54.490 --> 00:19:57.165 align:middle line:84%
The event A consists of
these three outcomes.

00:19:57.165 --> 00:20:00.470 align:middle line:90%


00:20:00.470 --> 00:20:04.900 align:middle line:84%
OK, what is the conditional PMF,
given that we are told

00:20:04.900 --> 00:20:08.920 align:middle line:90%
that event A has occurred?

00:20:08.920 --> 00:20:11.900 align:middle line:84%
Given that the event A has
occurred, it basically tells

00:20:11.900 --> 00:20:15.390 align:middle line:84%
us that this outcome
has not occurred.

00:20:15.390 --> 00:20:18.670 align:middle line:84%
There's only three possible
outcomes now.

00:20:18.670 --> 00:20:21.350 align:middle line:84%
In the new universe, in the new
model where we condition

00:20:21.350 --> 00:20:23.840 align:middle line:84%
on A, there's only three
possible outcomes.

00:20:23.840 --> 00:20:25.900 align:middle line:84%
Those three possible outcomes
were equally

00:20:25.900 --> 00:20:27.660 align:middle line:90%
likely when we started.

00:20:27.660 --> 00:20:30.130 align:middle line:84%
So in the conditional universe,
they will remain

00:20:30.130 --> 00:20:31.270 align:middle line:90%
equally likely.

00:20:31.270 --> 00:20:34.230 align:middle line:84%
Remember, whenever you
condition, the relative

00:20:34.230 --> 00:20:36.330 align:middle line:90%
likelihoods remain the same.

00:20:36.330 --> 00:20:38.290 align:middle line:84%
They keep the same
proportions.

00:20:38.290 --> 00:20:40.860 align:middle line:84%
They just need to be
re-scaled, so that

00:20:40.860 --> 00:20:42.520 align:middle line:90%
they add up to one.

00:20:42.520 --> 00:20:46.380 align:middle line:84%
So each one of these will have
the same probability.

00:20:46.380 --> 00:20:48.310 align:middle line:84%
Now in the new world,
probabilities need

00:20:48.310 --> 00:20:49.330 align:middle line:90%
to add up to 1.

00:20:49.330 --> 00:20:54.130 align:middle line:84%
So each one of them is going to
get a probability of 1/3 in

00:20:54.130 --> 00:20:55.385 align:middle line:90%
the conditional universe.

00:20:55.385 --> 00:20:58.420 align:middle line:90%


00:20:58.420 --> 00:21:00.650 align:middle line:84%
So this is our conditional
model.

00:21:00.650 --> 00:21:08.276 align:middle line:84%
So our PMF is equal to 1/3 for
X equals to 2, 3 and 4.

00:21:08.276 --> 00:21:10.270 align:middle line:90%
All right.

00:21:10.270 --> 00:21:13.380 align:middle line:84%
Now whenever you have a
probabilistic model involving

00:21:13.380 --> 00:21:16.550 align:middle line:84%
a random variable and you have
a PMF for that random

00:21:16.550 --> 00:21:19.690 align:middle line:84%
variable, you can talk about
the expected value of that

00:21:19.690 --> 00:21:20.990 align:middle line:90%
random variable.

00:21:20.990 --> 00:21:25.340 align:middle line:84%
We defined expected values
just a few minutes ago.

00:21:25.340 --> 00:21:28.800 align:middle line:84%
Here, we're dealing with
a conditional model and

00:21:28.800 --> 00:21:30.680 align:middle line:90%
conditional probabilities.

00:21:30.680 --> 00:21:33.680 align:middle line:84%
And so we can also talk about
the expected value of the

00:21:33.680 --> 00:21:38.100 align:middle line:84%
random variable X in this new
universe, in this new

00:21:38.100 --> 00:21:40.830 align:middle line:84%
conditional model that
we're dealing with.

00:21:40.830 --> 00:21:43.680 align:middle line:84%
And this leads us to the
definition of the notion of a

00:21:43.680 --> 00:21:45.780 align:middle line:90%
conditional expectation.

00:21:45.780 --> 00:21:50.680 align:middle line:84%
The conditional expectation
is nothing but an ordinary

00:21:50.680 --> 00:21:56.720 align:middle line:84%
expectation, except that you
don't use the original PMF.

00:21:56.720 --> 00:21:58.600 align:middle line:90%
You use the conditional PMF.

00:21:58.600 --> 00:22:00.620 align:middle line:84%
You use the conditional
probabilities.

00:22:00.620 --> 00:22:05.550 align:middle line:84%
It's just an ordinary
expectation, but applied to

00:22:05.550 --> 00:22:09.400 align:middle line:84%
the new model that we have to
the conditional universe where

00:22:09.400 --> 00:22:13.270 align:middle line:84%
we are told that the certain
event has occurred.

00:22:13.270 --> 00:22:17.310 align:middle line:84%
So we can now calculate the
condition expectation, which,

00:22:17.310 --> 00:22:19.890 align:middle line:84%
in this particular example,
would be 1/3.

00:22:19.890 --> 00:22:24.150 align:middle line:84%
That's the probability of a
2, plus 1/3 which is the

00:22:24.150 --> 00:22:29.280 align:middle line:84%
probability of a 3 plus 1/3,
the probability of a 4.

00:22:29.280 --> 00:22:33.160 align:middle line:84%
And then you can use your
calculator to find the answer,

00:22:33.160 --> 00:22:35.780 align:middle line:84%
or you can just argue
by symmetry.

00:22:35.780 --> 00:22:39.360 align:middle line:84%
The expected value has to be the
center of gravity of the

00:22:39.360 --> 00:22:45.040 align:middle line:84%
PMF we're working with,
which is equal to 3.

00:22:45.040 --> 00:22:49.880 align:middle line:84%
So conditional expectations are
no different from ordinary

00:22:49.880 --> 00:22:51.230 align:middle line:90%
expectations.

00:22:51.230 --> 00:22:54.500 align:middle line:84%
They're just ordinary
expectations applied to a new

00:22:54.500 --> 00:22:57.600 align:middle line:84%
type of situation or a
new type of model.

00:22:57.600 --> 00:23:03.010 align:middle line:84%
Anything we might know about
expectations will remain valid

00:23:03.010 --> 00:23:04.930 align:middle line:84%
about conditional
expectations.

00:23:04.930 --> 00:23:07.880 align:middle line:84%
So for example, the conditional
expectation of a

00:23:07.880 --> 00:23:11.040 align:middle line:84%
linear function of a random
variable is going to be the

00:23:11.040 --> 00:23:14.250 align:middle line:84%
linear function of the
conditional expectations.

00:23:14.250 --> 00:23:18.200 align:middle line:84%
Or you can take any formula that
you might know, such as

00:23:18.200 --> 00:23:23.970 align:middle line:84%
the formula that expected value
of X is equal to the--

00:23:23.970 --> 00:23:24.310 align:middle line:90%
sorry--

00:23:24.310 --> 00:23:31.030 align:middle line:84%
expected value of g of X is the
sum over all X's of g of X

00:23:31.030 --> 00:23:37.700 align:middle line:84%
times the PMF of X. So this is
the formula that we already

00:23:37.700 --> 00:23:41.790 align:middle line:84%
know about how to calculate
expectations of a function of

00:23:41.790 --> 00:23:43.540 align:middle line:90%
a random variable.

00:23:43.540 --> 00:23:47.190 align:middle line:84%
If we move to the conditional
universe, what changes?

00:23:47.190 --> 00:23:51.170 align:middle line:84%
In the conditional universe,
we're talking about the

00:23:51.170 --> 00:23:55.150 align:middle line:84%
conditional expectation, given
that event A has occurred.

00:23:55.150 --> 00:23:59.390 align:middle line:84%
And we use the conditional
probabilities, given that A

00:23:59.390 --> 00:24:00.650 align:middle line:90%
has occurred.

00:24:00.650 --> 00:24:05.140 align:middle line:84%
So any formula has a conditional
counterpart.

00:24:05.140 --> 00:24:07.790 align:middle line:84%
In the conditional counterparts,
expectations get

00:24:07.790 --> 00:24:10.000 align:middle line:84%
replaced by conditional
expectations.

00:24:10.000 --> 00:24:13.940 align:middle line:84%
And probabilities get replaced
by conditional probabilities.

00:24:13.940 --> 00:24:17.190 align:middle line:84%
So once you know the first
formula and you know the

00:24:17.190 --> 00:24:21.210 align:middle line:84%
general idea, there's absolutely
no reason for you

00:24:21.210 --> 00:24:24.020 align:middle line:84%
to memorize a formula
like this one.

00:24:24.020 --> 00:24:27.070 align:middle line:84%
You shouldn't even have to write
it on your cheat sheet

00:24:27.070 --> 00:24:30.840 align:middle line:90%
for the exam, OK?

00:24:30.840 --> 00:24:40.980 align:middle line:84%
OK, all right, so now let's look
at an example of a random

00:24:40.980 --> 00:24:44.470 align:middle line:84%
variable that we've seen before,
the geometric random

00:24:44.470 --> 00:24:47.910 align:middle line:84%
variable, and this time do
something a little more

00:24:47.910 --> 00:24:51.660 align:middle line:90%
interesting with it.

00:24:51.660 --> 00:24:54.510 align:middle line:84%
Do you remember from last time
what the geometric random

00:24:54.510 --> 00:24:55.580 align:middle line:90%
variable is?

00:24:55.580 --> 00:24:56.560 align:middle line:90%
We do coin flips.

00:24:56.560 --> 00:24:59.580 align:middle line:84%
Each time there's a
probability of P

00:24:59.580 --> 00:25:01.250 align:middle line:90%
of obtaining heads.

00:25:01.250 --> 00:25:03.910 align:middle line:84%
And we're interested in the
number of tosses we're going

00:25:03.910 --> 00:25:07.580 align:middle line:84%
to need until we observe heads
for the first time.

00:25:07.580 --> 00:25:10.045 align:middle line:84%
The probability that the random
variable takes the

00:25:10.045 --> 00:25:13.290 align:middle line:84%
value K, this is the probability
that the first K

00:25:13.290 --> 00:25:15.620 align:middle line:90%
appeared at the K-th toss.

00:25:15.620 --> 00:25:20.900 align:middle line:84%
So this is the probability of
K minus 1 consecutive tails

00:25:20.900 --> 00:25:22.670 align:middle line:90%
followed by a head.

00:25:22.670 --> 00:25:28.360 align:middle line:84%
So this is the probability of
having to weight K tosses.

00:25:28.360 --> 00:25:32.280 align:middle line:84%
And when we plot this PMF, it
has this kind of shape, which

00:25:32.280 --> 00:25:36.020 align:middle line:84%
is the shape of a geometric
progression.

00:25:36.020 --> 00:25:40.550 align:middle line:84%
It starts at 1, and it goes
all the way to infinity.

00:25:40.550 --> 00:25:43.700 align:middle line:84%
So this is a discrete random
variable that takes values

00:25:43.700 --> 00:25:49.160 align:middle line:84%
over an infinite set, the set
of the positive integers.

00:25:49.160 --> 00:25:51.720 align:middle line:84%
So it's a random variable,
therefore, it has an

00:25:51.720 --> 00:25:53.140 align:middle line:90%
expectation.

00:25:53.140 --> 00:25:56.790 align:middle line:84%
And the expected value is, by
definition, we'll consider all

00:25:56.790 --> 00:25:59.180 align:middle line:84%
possible values of the
random variable.

00:25:59.180 --> 00:26:02.560 align:middle line:84%
And we weigh them according to
their probabilities, which

00:26:02.560 --> 00:26:05.860 align:middle line:90%
leads us to this expression.

00:26:05.860 --> 00:26:09.860 align:middle line:84%
You may have evaluated that
expression some time in your

00:26:09.860 --> 00:26:11.400 align:middle line:90%
previous life.

00:26:11.400 --> 00:26:15.730 align:middle line:84%
And there are tricks for how
to evaluate this and get a

00:26:15.730 --> 00:26:16.770 align:middle line:90%
closed-form answer.

00:26:16.770 --> 00:26:19.350 align:middle line:84%
But it's sort of an
algebraic trick.

00:26:19.350 --> 00:26:20.710 align:middle line:90%
You might not remember it.

00:26:20.710 --> 00:26:23.520 align:middle line:84%
How do we go about doing
this summation?

00:26:23.520 --> 00:26:26.830 align:middle line:84%
Well, we're going to use a
probabilistic trick and manage

00:26:26.830 --> 00:26:33.440 align:middle line:84%
to evaluate the expectation of
X, essentially, without doing

00:26:33.440 --> 00:26:34.870 align:middle line:90%
any algebra.

00:26:34.870 --> 00:26:38.600 align:middle line:84%
And in the process of doing so,
we're going to get some

00:26:38.600 --> 00:26:43.080 align:middle line:84%
intuition about what happens
in coin tosses and with

00:26:43.080 --> 00:26:45.750 align:middle line:90%
geometric random variables.

00:26:45.750 --> 00:26:48.930 align:middle line:84%
So we have two people who
are going to do the same

00:26:48.930 --> 00:26:53.870 align:middle line:84%
experiment, flip a coin until
they obtain heads for the

00:26:53.870 --> 00:26:55.550 align:middle line:90%
first time.

00:26:55.550 --> 00:27:00.170 align:middle line:84%
One of these people is going to
use the letter Y to count

00:27:00.170 --> 00:27:02.760 align:middle line:90%
how many heads it took.

00:27:02.760 --> 00:27:06.860 align:middle line:84%
So that person starts
flipping right now.

00:27:06.860 --> 00:27:08.710 align:middle line:90%
This is the current time.

00:27:08.710 --> 00:27:11.440 align:middle line:84%
And they are going to obtain
tails, tails, tails, until

00:27:11.440 --> 00:27:13.620 align:middle line:90%
eventually they obtain heads.

00:27:13.620 --> 00:27:20.750 align:middle line:84%
And this random variable Y is,
of course, geometric, so it

00:27:20.750 --> 00:27:22.560 align:middle line:90%
has a PMF of this form.

00:27:22.560 --> 00:27:25.510 align:middle line:90%


00:27:25.510 --> 00:27:32.090 align:middle line:84%
OK, now there is a second person
who is doing that same

00:27:32.090 --> 00:27:33.410 align:middle line:90%
experiment.

00:27:33.410 --> 00:27:37.400 align:middle line:84%
That second person is going to
take, again, a random number,

00:27:37.400 --> 00:27:40.160 align:middle line:84%
X, until they obtain heads
for the first time.

00:27:40.160 --> 00:27:44.560 align:middle line:84%
And of course, X is going to
have the same PMF as Y.

00:27:44.560 --> 00:27:47.140 align:middle line:90%
But that person was impatient.

00:27:47.140 --> 00:27:51.880 align:middle line:84%
And they actually started
flipping earlier, before the Y

00:27:51.880 --> 00:27:53.490 align:middle line:90%
person started flipping.

00:27:53.490 --> 00:27:55.400 align:middle line:90%
They flipped the coin twice.

00:27:55.400 --> 00:27:57.640 align:middle line:90%
And they were unlucky, and they

00:27:57.640 --> 00:27:59.655 align:middle line:90%
obtained tails both times.

00:27:59.655 --> 00:28:02.370 align:middle line:90%


00:28:02.370 --> 00:28:05.115 align:middle line:90%
And so they have to continue.

00:28:05.115 --> 00:28:09.100 align:middle line:90%


00:28:09.100 --> 00:28:13.300 align:middle line:84%
Looking at the situation at this
time, how do these two

00:28:13.300 --> 00:28:14.690 align:middle line:90%
people compare?

00:28:14.690 --> 00:28:20.260 align:middle line:84%
Who do you think is going
to obtain heads first?

00:28:20.260 --> 00:28:22.610 align:middle line:84%
Is one more likely
than the other?

00:28:22.610 --> 00:28:26.160 align:middle line:84%
So if you play at the casino a
lot, you'll say, oh, there

00:28:26.160 --> 00:28:29.810 align:middle line:84%
were two tails in a row, so
a head should be coming up

00:28:29.810 --> 00:28:31.350 align:middle line:90%
sometime soon.

00:28:31.350 --> 00:28:35.600 align:middle line:84%
But this is a wrong argument,
because coin flips, at least

00:28:35.600 --> 00:28:37.870 align:middle line:90%
in our model, are independent.

00:28:37.870 --> 00:28:41.750 align:middle line:84%
The fact that these two happened
to be tails doesn't

00:28:41.750 --> 00:28:45.230 align:middle line:84%
change anything about our
beliefs about what's going to

00:28:45.230 --> 00:28:46.900 align:middle line:90%
be happening here.

00:28:46.900 --> 00:28:49.900 align:middle line:84%
So what's going to be happening
to that person is

00:28:49.900 --> 00:28:53.140 align:middle line:84%
they will be flipping
independent coin flips.

00:28:53.140 --> 00:28:54.930 align:middle line:84%
That person will also
be flipping

00:28:54.930 --> 00:28:56.600 align:middle line:90%
independent coin flips.

00:28:56.600 --> 00:29:00.660 align:middle line:84%
And both of them wait until
the first head occurs.

00:29:00.660 --> 00:29:04.050 align:middle line:84%
They're facing an identical
situation,

00:29:04.050 --> 00:29:06.770 align:middle line:90%
starting from this time.

00:29:06.770 --> 00:29:11.850 align:middle line:84%
OK, now what's the probabilistic
model of what

00:29:11.850 --> 00:29:14.530 align:middle line:90%
this person is facing?

00:29:14.530 --> 00:29:18.940 align:middle line:84%
The time until that person
obtains heads for the first

00:29:18.940 --> 00:29:25.740 align:middle line:84%
time is X. So this number of
flips until they obtain heads

00:29:25.740 --> 00:29:30.080 align:middle line:84%
for the first time is going
to be X minus 2.

00:29:30.080 --> 00:29:35.810 align:middle line:84%
So X is the total number
until the first head.

00:29:35.810 --> 00:29:41.280 align:middle line:84%
X minus 2 is the number or
flips, starting from here.

00:29:41.280 --> 00:29:44.060 align:middle line:84%
Now what information do we
have about that person?

00:29:44.060 --> 00:29:45.910 align:middle line:84%
We have the information
that their first

00:29:45.910 --> 00:29:47.970 align:middle line:90%
two flips were tails.

00:29:47.970 --> 00:29:52.650 align:middle line:84%
So we're given the information
that X was bigger than 2.

00:29:52.650 --> 00:29:57.035 align:middle line:84%
So the probabilistic model that
describes this piece of

00:29:57.035 --> 00:30:01.790 align:middle line:84%
the experiment is that it's
going to take a random number

00:30:01.790 --> 00:30:04.270 align:middle line:90%
of flips until the first head.

00:30:04.270 --> 00:30:08.420 align:middle line:84%
That number of flips, starting
from here until the next head,

00:30:08.420 --> 00:30:10.980 align:middle line:90%
is that number X minus 2.

00:30:10.980 --> 00:30:13.210 align:middle line:84%
But we're given the information
that this person

00:30:13.210 --> 00:30:17.780 align:middle line:84%
has already wasted
2 coin flips.

00:30:17.780 --> 00:30:20.330 align:middle line:84%
Now we argued that
probabilistically, this

00:30:20.330 --> 00:30:24.780 align:middle line:84%
person, this part of the
experiment here is identical

00:30:24.780 --> 00:30:26.710 align:middle line:84%
with that part of
the experiment.

00:30:26.710 --> 00:30:31.650 align:middle line:84%
So the PMF of this random
variable, which is X minus 2,

00:30:31.650 --> 00:30:33.980 align:middle line:84%
conditioned on this information,
should be the

00:30:33.980 --> 00:30:39.150 align:middle line:84%
same as that PMF that
we have down there.

00:30:39.150 --> 00:30:46.290 align:middle line:84%
So the formal statement that
I'm making is that this PMF

00:30:46.290 --> 00:30:51.910 align:middle line:84%
here of X minus 2, given that
X is bigger than 2, is the

00:30:51.910 --> 00:30:58.060 align:middle line:90%
same as the PMF of X itself.

00:30:58.060 --> 00:31:00.280 align:middle line:90%
What is this saying?

00:31:00.280 --> 00:31:04.220 align:middle line:84%
Given that I tell you that you
already did a few flips and

00:31:04.220 --> 00:31:08.450 align:middle line:84%
they were failures, the
remaining number of flips

00:31:08.450 --> 00:31:13.260 align:middle line:84%
until the first head has the
same geometric distribution as

00:31:13.260 --> 00:31:16.130 align:middle line:84%
if you were starting
from scratch.

00:31:16.130 --> 00:31:19.590 align:middle line:84%
Whatever happened in the past,
it happened, but has no

00:31:19.590 --> 00:31:22.670 align:middle line:84%
bearing what's going to
happen in the future.

00:31:22.670 --> 00:31:27.660 align:middle line:84%
Remaining coin flips until
a head has the same

00:31:27.660 --> 00:31:32.220 align:middle line:84%
distribution, whether you're
starting right now, or whether

00:31:32.220 --> 00:31:35.590 align:middle line:84%
you had done some other
stuff in the past.

00:31:35.590 --> 00:31:38.860 align:middle line:84%
So this is a property that we
call the memorylessness

00:31:38.860 --> 00:31:42.550 align:middle line:84%
property of the geometric
distribution.

00:31:42.550 --> 00:31:45.560 align:middle line:84%
Essentially, it says that
whatever happens in the future

00:31:45.560 --> 00:31:48.920 align:middle line:84%
is independent from whatever
happened in the past.

00:31:48.920 --> 00:31:51.350 align:middle line:84%
And that's true almost by
definition, because we're

00:31:51.350 --> 00:31:53.750 align:middle line:84%
assuming independent
coin flips.

00:31:53.750 --> 00:31:56.750 align:middle line:84%
Really, independence means that
information about one

00:31:56.750 --> 00:32:00.390 align:middle line:84%
part of the experiment has no
bearing about what's going to

00:32:00.390 --> 00:32:04.280 align:middle line:84%
happen in the other parts
of the experiment.

00:32:04.280 --> 00:32:09.010 align:middle line:84%
The argument that I tried to
give using the intuition of

00:32:09.010 --> 00:32:14.240 align:middle line:84%
coin flips, you can make it
formal by just manipulating

00:32:14.240 --> 00:32:16.110 align:middle line:90%
PMFs formally.

00:32:16.110 --> 00:32:19.450 align:middle line:84%
So this is the original
PMF of X.

00:32:19.450 --> 00:32:22.090 align:middle line:84%
Suppose that you condition
on the event that X

00:32:22.090 --> 00:32:24.030 align:middle line:90%
is bigger than 3.

00:32:24.030 --> 00:32:27.570 align:middle line:84%
This conditioning information,
what it does is it tells you

00:32:27.570 --> 00:32:30.450 align:middle line:84%
that this piece did
not happen.

00:32:30.450 --> 00:32:33.760 align:middle line:84%
You're conditioning just
on this event.

00:32:33.760 --> 00:32:37.430 align:middle line:84%
When you condition on that
event, what's left is the

00:32:37.430 --> 00:32:42.130 align:middle line:84%
conditional PMF, which has the
same shape as this one, except

00:32:42.130 --> 00:32:45.010 align:middle line:84%
that it needs to be
re-normalized up, so that the

00:32:45.010 --> 00:32:46.820 align:middle line:90%
probabilities add up to one.

00:32:46.820 --> 00:32:52.460 align:middle line:84%
So you take that picture, but
you need to change the height

00:32:52.460 --> 00:32:56.210 align:middle line:84%
of it, so that these
terms add up to 1.

00:32:56.210 --> 00:32:59.730 align:middle line:84%
And this is the conditional
PMF of X, given that X is

00:32:59.730 --> 00:33:01.310 align:middle line:90%
bigger than 2.

00:33:01.310 --> 00:33:04.360 align:middle line:84%
But we're talking here not about
X. We're talking about

00:33:04.360 --> 00:33:07.930 align:middle line:90%
the remaining number of heads.

00:33:07.930 --> 00:33:12.030 align:middle line:84%
Remaining number of heads
is X minus 2.

00:33:12.030 --> 00:33:17.120 align:middle line:84%
If we have the PMF of X, can we
find the PMF of X minus 2?

00:33:17.120 --> 00:33:22.870 align:middle line:84%
Well, if X is equal to 3, that
corresponds to X minus 2 being

00:33:22.870 --> 00:33:24.170 align:middle line:90%
equal to 1.

00:33:24.170 --> 00:33:26.730 align:middle line:84%
So this probability here
should be equal to that

00:33:26.730 --> 00:33:27.950 align:middle line:90%
probability.

00:33:27.950 --> 00:33:31.400 align:middle line:84%
The probability that X is equal
to 4 should be the same

00:33:31.400 --> 00:33:34.710 align:middle line:84%
as the probability that X
minus 2 is equal to 2.

00:33:34.710 --> 00:33:38.980 align:middle line:84%
So basically, the PMF of X minus
2 is the same as the PMF

00:33:38.980 --> 00:33:43.460 align:middle line:84%
of X, except that it gets
shifted by these 2 units.

00:33:43.460 --> 00:33:47.340 align:middle line:84%
So this way, we have formally
derived the conditional PMF of

00:33:47.340 --> 00:33:51.490 align:middle line:84%
the remaining number of coin
tosses, given that the first

00:33:51.490 --> 00:33:55.230 align:middle line:90%
two flips were tails.

00:33:55.230 --> 00:33:58.880 align:middle line:84%
And we see that it's exactly
the same as the PMF that we

00:33:58.880 --> 00:34:00.230 align:middle line:90%
started with.

00:34:00.230 --> 00:34:05.130 align:middle line:84%
And so this is the formal proof
of this statement here.

00:34:05.130 --> 00:34:10.010 align:middle line:84%
So it's useful here to digest
both these formal statements

00:34:10.010 --> 00:34:13.290 align:middle line:84%
and understand it and understand
the notation that

00:34:13.290 --> 00:34:17.050 align:middle line:84%
is involved here, but also
to really appreciate the

00:34:17.050 --> 00:34:21.409 align:middle line:84%
intuitive argument what
this is really saying.

00:34:21.409 --> 00:34:28.980 align:middle line:84%
OK, all right, so now we want to
use this observation, this

00:34:28.980 --> 00:34:32.389 align:middle line:84%
memorylessness, to eventually
calculate the expected value

00:34:32.389 --> 00:34:34.679 align:middle line:84%
for a geometric random
variable.

00:34:34.679 --> 00:34:38.150 align:middle line:84%
And the way we're going to do
it is by using a divide and

00:34:38.150 --> 00:34:41.650 align:middle line:84%
conquer tool, which is an analog
of what we have already

00:34:41.650 --> 00:34:44.489 align:middle line:90%
seen sometime before.

00:34:44.489 --> 00:34:48.230 align:middle line:84%
Remember our story that there's
a number of possible

00:34:48.230 --> 00:34:49.840 align:middle line:90%
scenarios about the world?

00:34:49.840 --> 00:34:54.120 align:middle line:84%
And there's a certain event, B,
that can happen under any

00:34:54.120 --> 00:34:55.889 align:middle line:90%
of these possible scenarios.

00:34:55.889 --> 00:34:57.980 align:middle line:84%
And we have the total
probability theory.

00:34:57.980 --> 00:35:00.970 align:middle line:84%
And that tells us that, to find
the probability of this

00:35:00.970 --> 00:35:03.360 align:middle line:84%
event, B, you consider
the probabilities of

00:35:03.360 --> 00:35:06.000 align:middle line:90%
B under each scenario.

00:35:06.000 --> 00:35:09.180 align:middle line:84%
And you weigh those
probabilities according to the

00:35:09.180 --> 00:35:12.190 align:middle line:84%
probabilities of the different
scenarios that we have.

00:35:12.190 --> 00:35:14.520 align:middle line:84%
So that's a formula that
we already know

00:35:14.520 --> 00:35:16.760 align:middle line:90%
and have worked with.

00:35:16.760 --> 00:35:18.020 align:middle line:90%
What's the next step?

00:35:18.020 --> 00:35:19.910 align:middle line:90%
Is it something deep?

00:35:19.910 --> 00:35:24.280 align:middle line:84%
No, it's just translation
in different notation.

00:35:24.280 --> 00:35:29.140 align:middle line:84%
This is the exactly same
formula, but with PMFs.

00:35:29.140 --> 00:35:32.720 align:middle line:84%
The event that capital X is
equal to little x can happen

00:35:32.720 --> 00:35:34.420 align:middle line:90%
in many different ways.

00:35:34.420 --> 00:35:37.580 align:middle line:84%
It can happen under
either scenario.

00:35:37.580 --> 00:35:40.910 align:middle line:84%
And within each scenario, you
need to use the conditional

00:35:40.910 --> 00:35:44.140 align:middle line:84%
probabilities of that event,
given that this

00:35:44.140 --> 00:35:46.270 align:middle line:90%
scenario has occurred.

00:35:46.270 --> 00:35:49.640 align:middle line:84%
So this formula is identical to
that one, except that we're

00:35:49.640 --> 00:35:53.440 align:middle line:84%
using conditional PMFs,
instead of conditional

00:35:53.440 --> 00:35:54.290 align:middle line:90%
probabilities.

00:35:54.290 --> 00:35:56.860 align:middle line:84%
But conditional PMFs, of
course, are nothing but

00:35:56.860 --> 00:35:59.710 align:middle line:84%
conditional probabilities
anyway.

00:35:59.710 --> 00:36:02.500 align:middle line:90%
So nothing new so far.

00:36:02.500 --> 00:36:08.700 align:middle line:84%
Then what I do is to take this
formula here and multiply both

00:36:08.700 --> 00:36:15.320 align:middle line:84%
sides by X and take the
sum over all X's.

00:36:15.320 --> 00:36:17.270 align:middle line:90%
What do we get on this side?

00:36:17.270 --> 00:36:19.430 align:middle line:84%
We get the expected
value of X.

00:36:19.430 --> 00:36:22.830 align:middle line:90%
What do we get on that side?

00:36:22.830 --> 00:36:24.290 align:middle line:90%
Probability of A1.

00:36:24.290 --> 00:36:29.770 align:middle line:84%
And then here, sum over all
X's of X times P. That's,

00:36:29.770 --> 00:36:33.030 align:middle line:84%
again, the same calculation
we have when we deal with

00:36:33.030 --> 00:36:36.450 align:middle line:84%
expectations, except that, since
here, we're dealing with

00:36:36.450 --> 00:36:39.010 align:middle line:84%
conditional probabilities,
we're going to get the

00:36:39.010 --> 00:36:41.220 align:middle line:90%
conditional expectation.

00:36:41.220 --> 00:36:44.160 align:middle line:84%
And this is the total
expectation theorem.

00:36:44.160 --> 00:36:47.300 align:middle line:84%
It's a very useful way for
calculating expectations using

00:36:47.300 --> 00:36:49.440 align:middle line:90%
a divide and conquer method.

00:36:49.440 --> 00:36:53.730 align:middle line:84%
We figure out the average value
of X under each one of

00:36:53.730 --> 00:36:55.590 align:middle line:90%
the possible scenarios.

00:36:55.590 --> 00:37:01.330 align:middle line:84%
The overall average value of
X is a weighted linear

00:37:01.330 --> 00:37:04.500 align:middle line:84%
combination of the expected
values of X in the different

00:37:04.500 --> 00:37:07.960 align:middle line:84%
scenarios where the weights are
chosen according to the

00:37:07.960 --> 00:37:09.210 align:middle line:90%
different probabilities.

00:37:09.210 --> 00:37:15.356 align:middle line:90%


00:37:15.356 --> 00:37:21.040 align:middle line:84%
OK, and now we're going to apply
this to the case of a

00:37:21.040 --> 00:37:23.070 align:middle line:90%
geometric random variable.

00:37:23.070 --> 00:37:26.410 align:middle line:84%
And we're going to divide and
conquer by considering

00:37:26.410 --> 00:37:31.350 align:middle line:84%
separately the two cases where
the first toss was heads, and

00:37:31.350 --> 00:37:35.820 align:middle line:84%
the other case where the
first toss was tails.

00:37:35.820 --> 00:37:40.160 align:middle line:84%
So the expected value of X is
the probability that the first

00:37:40.160 --> 00:37:44.020 align:middle line:84%
toss was heads, so that X is
equal to 1, and the expected

00:37:44.020 --> 00:37:46.520 align:middle line:90%
value if that happened.

00:37:46.520 --> 00:37:51.530 align:middle line:84%
What is the expected value of X,
given that X is equal to 1?

00:37:51.530 --> 00:37:55.770 align:middle line:84%
If X is known to be
equal to 1, then X

00:37:55.770 --> 00:37:57.390 align:middle line:90%
becomes just a number.

00:37:57.390 --> 00:38:01.100 align:middle line:84%
And the expected value of a
number is the number itself.

00:38:01.100 --> 00:38:05.120 align:middle line:84%
So this first line here is the
probability of heads in the

00:38:05.120 --> 00:38:07.660 align:middle line:90%
first toss times the number 1.

00:38:07.660 --> 00:38:13.320 align:middle line:90%


00:38:13.320 --> 00:38:21.000 align:middle line:84%
So the probability that X is
bigger than 1 is 1 minus P.

00:38:21.000 --> 00:38:25.400 align:middle line:84%
And then we need to do
something about this

00:38:25.400 --> 00:38:26.650 align:middle line:90%
conditional expectation.

00:38:26.650 --> 00:38:29.830 align:middle line:90%


00:38:29.830 --> 00:38:33.400 align:middle line:90%
What is it?

00:38:33.400 --> 00:38:39.420 align:middle line:84%
I can write it in, perhaps,
a more suggested form, as

00:38:39.420 --> 00:38:51.360 align:middle line:84%
expected the value of X minus
1, given that X minus 1 is

00:38:51.360 --> 00:38:52.610 align:middle line:90%
bigger than 1.

00:38:52.610 --> 00:38:56.590 align:middle line:90%


00:38:56.590 --> 00:38:57.840 align:middle line:90%
Ah.

00:38:57.840 --> 00:39:02.420 align:middle line:90%


00:39:02.420 --> 00:39:07.453 align:middle line:84%
OK, X bigger than 1 is the
same as X minus 1 being

00:39:07.453 --> 00:39:10.770 align:middle line:90%
positive, this way.

00:39:10.770 --> 00:39:15.500 align:middle line:90%
X minus 1 is positive plus 1.

00:39:15.500 --> 00:39:16.680 align:middle line:90%
What did I do here?

00:39:16.680 --> 00:39:21.250 align:middle line:90%
I added and subtracted 1.

00:39:21.250 --> 00:39:24.110 align:middle line:90%
Now what is this?

00:39:24.110 --> 00:39:29.240 align:middle line:84%
This is the expected value of
the remaining coin flips,

00:39:29.240 --> 00:39:34.660 align:middle line:84%
until I obtain heads, given that
the first one was tails.

00:39:34.660 --> 00:39:38.690 align:middle line:84%
It's the same story that we were
going through down there.

00:39:38.690 --> 00:39:41.860 align:middle line:84%
Given that the first coin flip
was tails doesn't tell me

00:39:41.860 --> 00:39:43.790 align:middle line:84%
anything about the
future, about the

00:39:43.790 --> 00:39:45.750 align:middle line:90%
remaining coin flips.

00:39:45.750 --> 00:39:49.610 align:middle line:84%
So this expectation should be
the same as the expectation

00:39:49.610 --> 00:39:53.740 align:middle line:84%
faced by a person who was
starting just now.

00:39:53.740 --> 00:39:59.080 align:middle line:84%
So this should be equal to the
expected value of X itself.

00:39:59.080 --> 00:40:04.120 align:middle line:84%
And then we have the plus 1
that's come from there, OK?

00:40:04.120 --> 00:40:07.830 align:middle line:90%


00:40:07.830 --> 00:40:11.140 align:middle line:84%
Remaining coin flips until a
head, given that I had a tail

00:40:11.140 --> 00:40:15.710 align:middle line:84%
yesterday, is the same as
expected number of flips until

00:40:15.710 --> 00:40:20.160 align:middle line:84%
heads for a person just starting
now and wasn't doing

00:40:20.160 --> 00:40:21.280 align:middle line:90%
anything yesterday.

00:40:21.280 --> 00:40:24.640 align:middle line:84%
So the fact that they I had a
coin flip yesterday doesn't

00:40:24.640 --> 00:40:28.990 align:middle line:84%
change my beliefs about how
long it's going to take me

00:40:28.990 --> 00:40:31.170 align:middle line:90%
until the first head.

00:40:31.170 --> 00:40:34.700 align:middle line:84%
So once we believe that
relation, than

00:40:34.700 --> 00:40:37.750 align:middle line:90%
we plug this here.

00:40:37.750 --> 00:40:42.346 align:middle line:84%
And this red term becomes
expected value of X plus 1.

00:40:42.346 --> 00:40:46.000 align:middle line:90%


00:40:46.000 --> 00:40:50.850 align:middle line:84%
So now we didn't exactly get the
answer we wanted, but we

00:40:50.850 --> 00:40:56.110 align:middle line:84%
got an equation that involves
the expected value of X. And

00:40:56.110 --> 00:40:58.230 align:middle line:84%
it's the only unknown
in that equation.

00:40:58.230 --> 00:41:03.990 align:middle line:84%
Expected value of X equals to P
plus (1 minus P) times this

00:41:03.990 --> 00:41:05.190 align:middle line:90%
expression.

00:41:05.190 --> 00:41:08.480 align:middle line:84%
You solve this equation for
expected value of X, and you

00:41:08.480 --> 00:41:12.990 align:middle line:90%
get the value of 1/P.

00:41:12.990 --> 00:41:16.920 align:middle line:84%
The final answer does make
intuitive sense.

00:41:16.920 --> 00:41:21.030 align:middle line:84%
If P is small, heads are
difficult to obtain.

00:41:21.030 --> 00:41:24.050 align:middle line:84%
So you expect that it's going to
take you a long time until

00:41:24.050 --> 00:41:26.310 align:middle line:84%
you see heads for
the first time.

00:41:26.310 --> 00:41:29.243 align:middle line:84%
So it is definitely a
reasonable answer.

00:41:29.243 --> 00:41:32.960 align:middle line:84%
Now the trick that we used here,
the divide and conquer

00:41:32.960 --> 00:41:36.860 align:middle line:90%
trick, is a really nice one.

00:41:36.860 --> 00:41:39.760 align:middle line:84%
It gives us a very good shortcut
in this problem.

00:41:39.760 --> 00:41:44.230 align:middle line:84%
But you must definitely spend
some time making sure you

00:41:44.230 --> 00:41:48.670 align:middle line:84%
understand why this expression
here is the same as that

00:41:48.670 --> 00:41:50.020 align:middle line:90%
expression there.

00:41:50.020 --> 00:41:53.790 align:middle line:84%
Essentially, what it's saying is
that, if I tell you that X

00:41:53.790 --> 00:41:57.460 align:middle line:84%
is bigger than 1, that the first
coin flip was tails, all

00:41:57.460 --> 00:42:02.040 align:middle line:84%
I'm telling you is that that
person has wasted a coin flip,

00:42:02.040 --> 00:42:05.310 align:middle line:84%
and they are starting
all over again.

00:42:05.310 --> 00:42:08.510 align:middle line:90%
So they've wasted 1 coin flip.

00:42:08.510 --> 00:42:10.670 align:middle line:84%
And they're starting
all over again.

00:42:10.670 --> 00:42:13.220 align:middle line:84%
If I tell you that the first
flip was tails, that's the

00:42:13.220 --> 00:42:16.800 align:middle line:84%
only information that I'm
basically giving you, a wasted

00:42:16.800 --> 00:42:19.970 align:middle line:84%
flip, and then starts
all over again.

00:42:19.970 --> 00:42:23.180 align:middle line:84%
All right, so in the few
remaining minutes now, we're

00:42:23.180 --> 00:42:26.970 align:middle line:84%
going to quickly introduce a few
new concepts that we will

00:42:26.970 --> 00:42:31.050 align:middle line:84%
be playing with in the
next ten days or so.

00:42:31.050 --> 00:42:33.300 align:middle line:84%
And you will get plenty
of opportunities

00:42:33.300 --> 00:42:34.860 align:middle line:90%
to manipulate them.

00:42:34.860 --> 00:42:37.180 align:middle line:90%
So here's the idea.

00:42:37.180 --> 00:42:40.370 align:middle line:84%
A typical experiment may have
several random variables

00:42:40.370 --> 00:42:43.310 align:middle line:84%
associated with that
experiment.

00:42:43.310 --> 00:42:46.370 align:middle line:84%
So a typical student has
height and weight.

00:42:46.370 --> 00:42:48.800 align:middle line:84%
If I give you the PMF of
height, that tells me

00:42:48.800 --> 00:42:51.060 align:middle line:84%
something about distribution
of heights in the class.

00:42:51.060 --> 00:42:57.110 align:middle line:84%
I give you the PMF of weight,
it tells me something about

00:42:57.110 --> 00:42:58.990 align:middle line:84%
the different weights
in this class.

00:42:58.990 --> 00:43:01.690 align:middle line:84%
But if I want to ask a
question, is there an

00:43:01.690 --> 00:43:05.910 align:middle line:84%
association between height and
weight, then I need to know a

00:43:05.910 --> 00:43:09.730 align:middle line:84%
little more how height and
weight relate to each other.

00:43:09.730 --> 00:43:13.480 align:middle line:84%
And the PMF of height
individuality and PMF of

00:43:13.480 --> 00:43:16.130 align:middle line:84%
weight just by itself
do not tell me

00:43:16.130 --> 00:43:18.240 align:middle line:90%
anything about those relations.

00:43:18.240 --> 00:43:21.730 align:middle line:84%
To be able to say something
about those relations, I need

00:43:21.730 --> 00:43:27.230 align:middle line:84%
to know something about joint
probabilities, how likely is

00:43:27.230 --> 00:43:31.500 align:middle line:84%
it that certain X's go together
with certain Y's.

00:43:31.500 --> 00:43:34.180 align:middle line:84%
So these probabilities,
essentially, capture

00:43:34.180 --> 00:43:37.930 align:middle line:84%
associations between these
two random variables.

00:43:37.930 --> 00:43:40.910 align:middle line:84%
And it's the information I would
need to have to do any

00:43:40.910 --> 00:43:44.900 align:middle line:84%
kind of statistical study that
tries to relate the two random

00:43:44.900 --> 00:43:47.600 align:middle line:90%
variables with each other.

00:43:47.600 --> 00:43:49.440 align:middle line:84%
These are ordinary
probabilities.

00:43:49.440 --> 00:43:50.750 align:middle line:90%
This is an event.

00:43:50.750 --> 00:43:52.630 align:middle line:84%
It's the event that
this thing happens

00:43:52.630 --> 00:43:55.230 align:middle line:90%
and that thing happens.

00:43:55.230 --> 00:43:58.460 align:middle line:84%
This is just the notation
that we will be using.

00:43:58.460 --> 00:44:00.840 align:middle line:90%
It's called the joint PMF.

00:44:00.840 --> 00:44:04.560 align:middle line:84%
It's the joint Probability
Mass Function of the two

00:44:04.560 --> 00:44:09.170 align:middle line:84%
random variables X and Y looked
at together, jointly.

00:44:09.170 --> 00:44:11.740 align:middle line:84%
And it gives me the probability
that any

00:44:11.740 --> 00:44:17.100 align:middle line:84%
particular numerical outcome
pair does happen.

00:44:17.100 --> 00:44:20.580 align:middle line:84%
So in the finite case, you can
represent joint PMFs, for

00:44:20.580 --> 00:44:22.660 align:middle line:90%
example, by a table.

00:44:22.660 --> 00:44:25.940 align:middle line:84%
This particular table here would
give you information

00:44:25.940 --> 00:44:31.870 align:middle line:84%
such as, let's see, the joint
PMF evaluated at 2, 3.

00:44:31.870 --> 00:44:35.240 align:middle line:84%
This is the probability that
X is equal to 3 and,

00:44:35.240 --> 00:44:38.200 align:middle line:84%
simultaneously, Y
is equal to 3.

00:44:38.200 --> 00:44:40.370 align:middle line:84%
So it would be that
number here.

00:44:40.370 --> 00:44:41.620 align:middle line:90%
It's 4/20.

00:44:41.620 --> 00:44:44.290 align:middle line:90%


00:44:44.290 --> 00:44:47.330 align:middle line:84%
OK, what is a basic
property of PMFs?

00:44:47.330 --> 00:44:49.920 align:middle line:84%
First, these are probabilities,
so all of the

00:44:49.920 --> 00:44:52.240 align:middle line:84%
entries have to be
non-negative.

00:44:52.240 --> 00:44:57.470 align:middle line:84%
If you adopt the probabilities
over all possible numerical

00:44:57.470 --> 00:45:01.070 align:middle line:84%
pairs that you could get, of
course, the total probability

00:45:01.070 --> 00:45:03.050 align:middle line:90%
must be equal to 1.

00:45:03.050 --> 00:45:06.070 align:middle line:84%
So that's another thing
that we want.

00:45:06.070 --> 00:45:10.090 align:middle line:84%
Now suppose somebody gives
me this model, but I

00:45:10.090 --> 00:45:12.410 align:middle line:90%
don't care about Y's.

00:45:12.410 --> 00:45:15.760 align:middle line:84%
All I care is the distribution
of the X's.

00:45:15.760 --> 00:45:18.550 align:middle line:84%
So I'm going to find the
probability that X takes on a

00:45:18.550 --> 00:45:20.060 align:middle line:90%
particular value.

00:45:20.060 --> 00:45:22.230 align:middle line:90%
Can I find it from the table?

00:45:22.230 --> 00:45:23.190 align:middle line:90%
Of course, I can.

00:45:23.190 --> 00:45:27.930 align:middle line:84%
If you ask me what's the
probability that X is equal to

00:45:27.930 --> 00:45:31.890 align:middle line:84%
3, what I'm going to do is
to add up those three

00:45:31.890 --> 00:45:33.790 align:middle line:90%
probabilities together.

00:45:33.790 --> 00:45:37.680 align:middle line:84%
And those probabilities, taken
all together, give me the

00:45:37.680 --> 00:45:40.400 align:middle line:84%
probability that X
is equal to 3.

00:45:40.400 --> 00:45:43.950 align:middle line:84%
These are all the possible ways
that the event X equals

00:45:43.950 --> 00:45:45.310 align:middle line:90%
to 3 can happen.

00:45:45.310 --> 00:45:49.850 align:middle line:84%
So we add these, and
we get the 6/20.

00:45:49.850 --> 00:45:53.180 align:middle line:84%
What I just did, can we
translate it to a formula?

00:45:53.180 --> 00:45:55.510 align:middle line:90%
What did I do?

00:45:55.510 --> 00:45:59.790 align:middle line:84%
I fixed the particular value
of X. And I added up the

00:45:59.790 --> 00:46:05.100 align:middle line:84%
values of the joint PMF over all
the possible values of Y.

00:46:05.100 --> 00:46:07.710 align:middle line:90%
So that's how you do it.

00:46:07.710 --> 00:46:08.990 align:middle line:90%
You take the joint.

00:46:08.990 --> 00:46:13.020 align:middle line:84%
You take one slice of the joint,
keeping X fixed, and

00:46:13.020 --> 00:46:16.110 align:middle line:84%
adding up over the different
values of Y.

00:46:16.110 --> 00:46:18.930 align:middle line:84%
The moral of this example is
that, if you know the joint

00:46:18.930 --> 00:46:22.590 align:middle line:84%
PMFs, then you can find the
individual PMFs of every

00:46:22.590 --> 00:46:24.310 align:middle line:90%
individual random variable.

00:46:24.310 --> 00:46:25.980 align:middle line:90%
And we have a name for these.

00:46:25.980 --> 00:46:28.840 align:middle line:84%
We call them the
marginal PMFs.

00:46:28.840 --> 00:46:31.900 align:middle line:84%
We have the joint that talks
about both together, and the

00:46:31.900 --> 00:46:35.170 align:middle line:84%
marginal that talks about
them one at the time.

00:46:35.170 --> 00:46:38.590 align:middle line:84%
And finally, since we love
conditional probabilities, we

00:46:38.590 --> 00:46:41.150 align:middle line:84%
will certainly want to define
an object called the

00:46:41.150 --> 00:46:44.160 align:middle line:90%
conditional PMF.

00:46:44.160 --> 00:46:46.940 align:middle line:84%
So this quantity here
is a familiar one.

00:46:46.940 --> 00:46:49.310 align:middle line:84%
It's just a conditional
probability.

00:46:49.310 --> 00:46:54.940 align:middle line:84%
It's the probability that X
takes on a particular value,

00:46:54.940 --> 00:46:58.210 align:middle line:84%
given that Y takes
a certain value.

00:46:58.210 --> 00:47:02.060 align:middle line:90%


00:47:02.060 --> 00:47:07.160 align:middle line:84%
For our example, let's take
little y to be equal to 2,

00:47:07.160 --> 00:47:10.890 align:middle line:84%
which means that we're
conditioning to live inside

00:47:10.890 --> 00:47:12.490 align:middle line:90%
this universe.

00:47:12.490 --> 00:47:17.920 align:middle line:84%
This red universe here is the
y equal to 2 universe.

00:47:17.920 --> 00:47:20.860 align:middle line:84%
And these are the conditional
probabilities of the different

00:47:20.860 --> 00:47:22.935 align:middle line:90%
X's inside that universe.

00:47:22.935 --> 00:47:26.020 align:middle line:90%


00:47:26.020 --> 00:47:29.860 align:middle line:84%
OK, once more, just an
exercise in notation.

00:47:29.860 --> 00:47:34.850 align:middle line:84%
This is the chapter two version
of the notation of

00:47:34.850 --> 00:47:37.750 align:middle line:84%
what we were denoting this
way in chapter one.

00:47:37.750 --> 00:47:43.000 align:middle line:84%
The way to read this is that
it's a conditional PMF having

00:47:43.000 --> 00:47:46.450 align:middle line:84%
to do with two random variables,
the PMF of X

00:47:46.450 --> 00:47:51.150 align:middle line:84%
conditioned on information
about Y. We are fixing a

00:47:51.150 --> 00:47:54.830 align:middle line:84%
particular value of capital Y,
that's the value on which we

00:47:54.830 --> 00:47:56.610 align:middle line:90%
are conditioning.

00:47:56.610 --> 00:47:58.340 align:middle line:84%
And we're looking at the
probabilities of

00:47:58.340 --> 00:48:00.190 align:middle line:90%
the different X's.

00:48:00.190 --> 00:48:03.890 align:middle line:84%
So it's really a function
of two arguments, little

00:48:03.890 --> 00:48:05.340 align:middle line:90%
x and little y.

00:48:05.340 --> 00:48:10.490 align:middle line:84%
But the best way to think about
it is to fix little y

00:48:10.490 --> 00:48:15.030 align:middle line:84%
and think of it as a function
of X. So I'm fixing little y

00:48:15.030 --> 00:48:17.610 align:middle line:84%
here, let's say, to
y equal to 2.

00:48:17.610 --> 00:48:20.400 align:middle line:90%
So I'm considering only this.

00:48:20.400 --> 00:48:24.090 align:middle line:84%
And now, this quantity becomes
a function of little x.

00:48:24.090 --> 00:48:27.340 align:middle line:84%
For the different little x's,
we're going to have different

00:48:27.340 --> 00:48:29.290 align:middle line:90%
conditional probabilities.

00:48:29.290 --> 00:48:31.040 align:middle line:84%
What are those conditional
probabilities?

00:48:31.040 --> 00:48:36.230 align:middle line:90%


00:48:36.230 --> 00:48:40.760 align:middle line:84%
OK, conditional probabilities
are proportional to original

00:48:40.760 --> 00:48:41.940 align:middle line:90%
probabilities.

00:48:41.940 --> 00:48:45.200 align:middle line:84%
So it's going to be those
numbers, but scaled up.

00:48:45.200 --> 00:48:48.340 align:middle line:84%
And they need to be scaled
so that they add up to 1.

00:48:48.340 --> 00:48:50.280 align:middle line:90%
So we have 1, 3 and 1.

00:48:50.280 --> 00:48:51.970 align:middle line:90%
That's a total of 5.

00:48:51.970 --> 00:48:56.220 align:middle line:84%
So the conditional PMF would
have the shape zero,

00:48:56.220 --> 00:49:02.480 align:middle line:90%
1/5, 3/5, and 1/5.

00:49:02.480 --> 00:49:07.540 align:middle line:84%
This is the conditional PMF,
given a particular value of Y.

00:49:07.540 --> 00:49:13.180 align:middle line:84%
It has the same shape as those
numbers, where by shape, I

00:49:13.180 --> 00:49:15.850 align:middle line:84%
mean try to visualize
a bar graph.

00:49:15.850 --> 00:49:19.370 align:middle line:84%
The bar graph associated with
those numbers has exactly the

00:49:19.370 --> 00:49:23.630 align:middle line:84%
same shape as the bar graph
associated with those numbers.

00:49:23.630 --> 00:49:26.790 align:middle line:84%
The only thing that has changed
is the scaling.

00:49:26.790 --> 00:49:29.630 align:middle line:84%
Big moral, let me say in
different words, the

00:49:29.630 --> 00:49:34.250 align:middle line:84%
conditional PMF, given a
particular value of Y, is just

00:49:34.250 --> 00:49:39.790 align:middle line:84%
a slice of the joint PMF where
you maintain the same shape,

00:49:39.790 --> 00:49:44.320 align:middle line:84%
but you rescale the numbers
so that they add up to 1.

00:49:44.320 --> 00:49:48.410 align:middle line:84%
Now mathematically, of course,
what all of this is doing is

00:49:48.410 --> 00:49:54.750 align:middle line:84%
it's taking the original joint
PDF and it rescales it by a

00:49:54.750 --> 00:49:56.540 align:middle line:90%
certain factor.

00:49:56.540 --> 00:50:00.340 align:middle line:84%
This does not involve X, so the
shape, is a function of X,

00:50:00.340 --> 00:50:01.720 align:middle line:90%
has not changed.

00:50:01.720 --> 00:50:04.910 align:middle line:84%
We're keeping the same shape
as a function of X, but we

00:50:04.910 --> 00:50:06.420 align:middle line:90%
divide by a certain number.

00:50:06.420 --> 00:50:09.670 align:middle line:84%
And that's the number that we
need, so that the conditional

00:50:09.670 --> 00:50:12.810 align:middle line:90%
probabilities add up to 1.

00:50:12.810 --> 00:50:15.850 align:middle line:84%
Now where does this
formula come from?

00:50:15.850 --> 00:50:17.880 align:middle line:84%
Well, this is just the
definition of conditional

00:50:17.880 --> 00:50:19.000 align:middle line:90%
probabilities.

00:50:19.000 --> 00:50:21.870 align:middle line:84%
Probability of something
conditioned on something else

00:50:21.870 --> 00:50:24.620 align:middle line:84%
is the probability of both
things happening, the

00:50:24.620 --> 00:50:28.040 align:middle line:84%
intersection of the two divided
by the probability of

00:50:28.040 --> 00:50:29.890 align:middle line:90%
the conditioning event.

00:50:29.890 --> 00:50:33.100 align:middle line:84%
And last remark is that, as
I just said, conditional

00:50:33.100 --> 00:50:35.930 align:middle line:84%
probabilities are nothing
different than ordinary

00:50:35.930 --> 00:50:37.210 align:middle line:90%
probabilities.

00:50:37.210 --> 00:50:42.390 align:middle line:84%
So a conditional PMF must sum
to 1, no matter what you are

00:50:42.390 --> 00:50:44.360 align:middle line:90%
conditioning on.

00:50:44.360 --> 00:50:47.370 align:middle line:84%
All right, so this was sort of
quick introduction into our

00:50:47.370 --> 00:50:48.920 align:middle line:90%
new notation.

00:50:48.920 --> 00:50:53.030 align:middle line:84%
But you get a lot of practice
in the next days to come.