WEBVTT

00:00:00.000 --> 00:00:00.040 align:middle line:90%


00:00:00.040 --> 00:00:02.460 align:middle line:84%
The following content is
provided under a Creative

00:00:02.460 --> 00:00:03.870 align:middle line:90%
Commons license.

00:00:03.870 --> 00:00:06.910 align:middle line:84%
Your support will help MIT
OpenCourseWare continue to

00:00:06.910 --> 00:00:10.560 align:middle line:84%
offer high quality educational
resources for free.

00:00:10.560 --> 00:00:13.460 align:middle line:84%
To make a donation or view
additional materials from

00:00:13.460 --> 00:00:19.290 align:middle line:84%
hundreds of MIT courses, visit
MIT OpenCourseWare at

00:00:19.290 --> 00:00:20.540 align:middle line:90%
ocw.mit.edu.

00:00:20.540 --> 00:00:23.130 align:middle line:90%


00:00:23.130 --> 00:00:25.940 align:middle line:84%
PROFESSOR: So today's agenda
is to say a few more things

00:00:25.940 --> 00:00:28.050 align:middle line:84%
about continuous random
variables.

00:00:28.050 --> 00:00:32.049 align:middle line:84%
Mainly we're going to talk a
little bit about inference.

00:00:32.049 --> 00:00:35.080 align:middle line:84%
This is a topic that we're going
to revisit at the end of

00:00:35.080 --> 00:00:36.390 align:middle line:90%
the semester.

00:00:36.390 --> 00:00:38.070 align:middle line:84%
But there's a few things
that we can

00:00:38.070 --> 00:00:40.180 align:middle line:90%
already say at this point.

00:00:40.180 --> 00:00:44.060 align:middle line:84%
And then the new topic for
today is the subject of

00:00:44.060 --> 00:00:45.880 align:middle line:90%
derived distributions.

00:00:45.880 --> 00:00:48.140 align:middle line:84%
Basically if you know the
distribution of one random

00:00:48.140 --> 00:00:50.230 align:middle line:84%
variable, and you have a
function of that random

00:00:50.230 --> 00:00:52.010 align:middle line:90%
variable, how to find a

00:00:52.010 --> 00:00:54.840 align:middle line:90%
distribution for that function.

00:00:54.840 --> 00:00:58.180 align:middle line:84%
And it's a fairly mechanical
skill, but that's an important

00:00:58.180 --> 00:01:00.740 align:middle line:84%
one, so we're going
to go through it.

00:01:00.740 --> 00:01:02.200 align:middle line:90%
So let's see where we stand.

00:01:02.200 --> 00:01:03.540 align:middle line:90%
Here is the big picture.

00:01:03.540 --> 00:01:06.720 align:middle line:84%
That's all we have
done so far.

00:01:06.720 --> 00:01:09.460 align:middle line:84%
We have talked about discrete
random variables, which we

00:01:09.460 --> 00:01:11.970 align:middle line:84%
described by probability
mass function.

00:01:11.970 --> 00:01:14.900 align:middle line:84%
So if we have multiple random
variables, we describe them

00:01:14.900 --> 00:01:16.760 align:middle line:84%
with the a joint
mass function.

00:01:16.760 --> 00:01:19.810 align:middle line:84%
And then we define conditional
probabilities, or conditional

00:01:19.810 --> 00:01:24.310 align:middle line:84%
PMFs, and the three are related
according to this

00:01:24.310 --> 00:01:27.040 align:middle line:84%
formula, which is, you can
think of it either as the

00:01:27.040 --> 00:01:29.300 align:middle line:84%
definition of conditional
probability.

00:01:29.300 --> 00:01:32.170 align:middle line:84%
Or as the multiplication rule,
the probability of two things

00:01:32.170 --> 00:01:35.870 align:middle line:84%
happening is the product of the
probabilities of the first

00:01:35.870 --> 00:01:38.200 align:middle line:84%
thing happening, and then the
second happening, given that

00:01:38.200 --> 00:01:39.860 align:middle line:90%
the first has happened.

00:01:39.860 --> 00:01:42.830 align:middle line:84%
There's another relation between
this, which is the

00:01:42.830 --> 00:01:46.360 align:middle line:84%
probability of x occurring, is
the sum of the different

00:01:46.360 --> 00:01:50.560 align:middle line:84%
probabilities of the different
ways that x may occur, which

00:01:50.560 --> 00:01:53.700 align:middle line:84%
is in conjunction with different
values of y.

00:01:53.700 --> 00:01:57.730 align:middle line:84%
And there's an analog of all
that in the continuous world,

00:01:57.730 --> 00:02:02.430 align:middle line:84%
where all you do is to replace
p's by f's, and replace sums

00:02:02.430 --> 00:02:03.340 align:middle line:90%
by integrals.

00:02:03.340 --> 00:02:05.620 align:middle line:84%
So the formulas all
look the same.

00:02:05.620 --> 00:02:09.120 align:middle line:84%
The interpretations are a little
more subtle, so the f's

00:02:09.120 --> 00:02:11.720 align:middle line:84%
are not probabilities, they're
probability densities.

00:02:11.720 --> 00:02:16.010 align:middle line:84%
So they're probabilities per
unit length, or in the case of

00:02:16.010 --> 00:02:20.290 align:middle line:84%
joint PDf's, these are
probabilities per unit area.

00:02:20.290 --> 00:02:22.690 align:middle line:84%
So they're densities
of some sort.

00:02:22.690 --> 00:02:26.020 align:middle line:84%
Probably the more subtle concept
to understand what it

00:02:26.020 --> 00:02:29.250 align:middle line:84%
really is the conditional
density.

00:02:29.250 --> 00:02:30.590 align:middle line:90%
In some sense, it's simple.

00:02:30.590 --> 00:02:34.900 align:middle line:84%
It's just the density of X in
a world where you have been

00:02:34.900 --> 00:02:40.290 align:middle line:84%
told the value of the random
variable Y. It's a function

00:02:40.290 --> 00:02:44.510 align:middle line:84%
that has two arguments, but the
best way to think about it

00:02:44.510 --> 00:02:47.050 align:middle line:90%
is to say that we fixed y.

00:02:47.050 --> 00:02:50.980 align:middle line:84%
We're told the value of the
random variable Y, and we look

00:02:50.980 --> 00:02:52.930 align:middle line:90%
at it as a function of x.

00:02:52.930 --> 00:02:56.150 align:middle line:84%
So as a function of x, the
denominator is a constant, and

00:02:56.150 --> 00:02:59.650 align:middle line:84%
it just looks like the
joint density.

00:02:59.650 --> 00:03:01.620 align:middle line:90%
when we keep y fixed.

00:03:01.620 --> 00:03:05.570 align:middle line:84%
So it's really a function of
one argument, just the

00:03:05.570 --> 00:03:06.870 align:middle line:90%
argument x.

00:03:06.870 --> 00:03:10.080 align:middle line:84%
And it has the same shape as the
joint's density when you

00:03:10.080 --> 00:03:11.720 align:middle line:90%
take that slice of it.

00:03:11.720 --> 00:03:17.570 align:middle line:84%
So conditional PDFs are just
slices of joint PDFs.

00:03:17.570 --> 00:03:20.810 align:middle line:84%
There's a bunch of concepts,
expectations, variances,

00:03:20.810 --> 00:03:23.790 align:middle line:84%
cumulative distribution
functions that apply equally

00:03:23.790 --> 00:03:26.260 align:middle line:84%
well for to both universes
of discrete or

00:03:26.260 --> 00:03:28.800 align:middle line:90%
continuous random variables.

00:03:28.800 --> 00:03:31.330 align:middle line:90%
So why is probability useful?

00:03:31.330 --> 00:03:36.170 align:middle line:84%
Probability is useful because,
among other things, we use it

00:03:36.170 --> 00:03:38.420 align:middle line:84%
to make sense of the
world around us.

00:03:38.420 --> 00:03:41.870 align:middle line:84%
We use it to make inferences
about things that we do not

00:03:41.870 --> 00:03:43.280 align:middle line:90%
see directly.

00:03:43.280 --> 00:03:45.570 align:middle line:84%
And this is done in a
very simple manner

00:03:45.570 --> 00:03:46.840 align:middle line:90%
using the base rule.

00:03:46.840 --> 00:03:49.730 align:middle line:84%
We've already seen some of that,
and now we're going to

00:03:49.730 --> 00:03:55.070 align:middle line:84%
revisit it with a bunch of
different variations.

00:03:55.070 --> 00:03:58.240 align:middle line:84%
And the variations come because
sometimes our random

00:03:58.240 --> 00:04:01.040 align:middle line:84%
variable are discrete, sometimes
they're continuous,

00:04:01.040 --> 00:04:04.390 align:middle line:84%
or we can have a combination
of the two.

00:04:04.390 --> 00:04:08.170 align:middle line:84%
So the big picture is that
there's some unknown random

00:04:08.170 --> 00:04:11.660 align:middle line:84%
variable out of there, and we
know the distribution that's

00:04:11.660 --> 00:04:12.550 align:middle line:90%
random variable.

00:04:12.550 --> 00:04:16.360 align:middle line:84%
And in the discrete case, it's
going to be given by PMF.

00:04:16.360 --> 00:04:20.269 align:middle line:84%
In the continuous case,
it's given a PDF.

00:04:20.269 --> 00:04:24.060 align:middle line:84%
Then we have some phenomenon,
some noisy phenomenon or some

00:04:24.060 --> 00:04:28.380 align:middle line:84%
measuring device, and that
measuring device produces

00:04:28.380 --> 00:04:31.260 align:middle line:90%
observable random variables Y.

00:04:31.260 --> 00:04:34.930 align:middle line:84%
We don't know what x is, but we
have some beliefs about how

00:04:34.930 --> 00:04:36.310 align:middle line:90%
X is distributed.

00:04:36.310 --> 00:04:39.450 align:middle line:84%
We observe the random variable
Y. We need a

00:04:39.450 --> 00:04:41.300 align:middle line:90%
model of this box.

00:04:41.300 --> 00:04:46.170 align:middle line:84%
And the model of that box is
going to be either a PMF, for

00:04:46.170 --> 00:04:52.565 align:middle line:84%
the random variable Y. And that
model tells us, if the

00:04:52.565 --> 00:04:57.080 align:middle line:84%
true state of the world is X,
how do we expect to Y to be

00:04:57.080 --> 00:04:58.520 align:middle line:90%
distributed?

00:04:58.520 --> 00:05:01.610 align:middle line:84%
That's for the case where
Y is this discrete.

00:05:01.610 --> 00:05:06.350 align:middle line:84%
If Y is a continuous, you might
instead have a density

00:05:06.350 --> 00:05:10.820 align:middle line:84%
for Y, or something
of that form.

00:05:10.820 --> 00:05:13.980 align:middle line:84%
So in either case, this
should be a function

00:05:13.980 --> 00:05:15.520 align:middle line:90%
that's known to us.

00:05:15.520 --> 00:05:18.370 align:middle line:84%
This is our model of the
measuring device.

00:05:18.370 --> 00:05:20.950 align:middle line:84%
And now having observed
y, we want to make

00:05:20.950 --> 00:05:22.680 align:middle line:90%
inferences about x.

00:05:22.680 --> 00:05:25.140 align:middle line:84%
What does it mean to
make inferences?

00:05:25.140 --> 00:05:29.880 align:middle line:84%
Well the most complete answer in
the inference problem is to

00:05:29.880 --> 00:05:32.380 align:middle line:84%
tell me the probability
distribution

00:05:32.380 --> 00:05:34.830 align:middle line:90%
of the unknown quantity.

00:05:34.830 --> 00:05:36.900 align:middle line:84%
But when I say the probability
distribution, I

00:05:36.900 --> 00:05:38.540 align:middle line:90%
don't mean this one.

00:05:38.540 --> 00:05:41.280 align:middle line:84%
I mean the probability
distribution that takes into

00:05:41.280 --> 00:05:43.760 align:middle line:84%
account the measurements
that you got.

00:05:43.760 --> 00:05:48.270 align:middle line:84%
So the output of an inference
problem is to come up with the

00:05:48.270 --> 00:05:59.830 align:middle line:84%
distribution of X, the unknown
quantity, given what we have

00:05:59.830 --> 00:06:00.980 align:middle line:90%
already observed.

00:06:00.980 --> 00:06:04.110 align:middle line:84%
And in the discrete case, it
would be an object like that.

00:06:04.110 --> 00:06:08.920 align:middle line:84%
If X is continuous, it would
be an object of this kind.

00:06:08.920 --> 00:06:13.340 align:middle line:90%


00:06:13.340 --> 00:06:18.080 align:middle line:84%
OK, so we're given conditional
probabilities of this type,

00:06:18.080 --> 00:06:21.240 align:middle line:84%
and we want to get conditional
distributions of the opposite

00:06:21.240 --> 00:06:23.280 align:middle line:90%
type where the order of the

00:06:23.280 --> 00:06:25.580 align:middle line:90%
conditioning is being reversed.

00:06:25.580 --> 00:06:28.980 align:middle line:84%
So the starting point
is always a formula

00:06:28.980 --> 00:06:30.810 align:middle line:90%
such as this one.

00:06:30.810 --> 00:06:33.670 align:middle line:84%
The probability of x happening,
and then y

00:06:33.670 --> 00:06:36.280 align:middle line:84%
happening given that
x happens.

00:06:36.280 --> 00:06:40.910 align:middle line:84%
This is the probability that
a particular x and y happen

00:06:40.910 --> 00:06:42.370 align:middle line:90%
simultaneously.

00:06:42.370 --> 00:06:47.240 align:middle line:84%
But this is also equal to the
probability that y happens,

00:06:47.240 --> 00:06:50.377 align:middle line:84%
and then that x happens, given
that y has happened.

00:06:50.377 --> 00:06:53.060 align:middle line:90%


00:06:53.060 --> 00:06:57.140 align:middle line:84%
And you take this expression
and send one term to the

00:06:57.140 --> 00:07:00.950 align:middle line:84%
denominator of the other side,
and this gives us the base

00:07:00.950 --> 00:07:03.180 align:middle line:90%
rule for the discrete case.

00:07:03.180 --> 00:07:05.550 align:middle line:84%
Which is this one that you have
already seen, and you

00:07:05.550 --> 00:07:07.200 align:middle line:90%
have played with it.

00:07:07.200 --> 00:07:10.720 align:middle line:84%
So this is what the formula
looks like in

00:07:10.720 --> 00:07:12.030 align:middle line:90%
the discrete case.

00:07:12.030 --> 00:07:14.570 align:middle line:84%
And the typical example where
both random variables are

00:07:14.570 --> 00:07:18.000 align:middle line:84%
discrete is the one we discussed
some time ago.

00:07:18.000 --> 00:07:20.720 align:middle line:84%
X is, let's say, a binary
variable, or whether an

00:07:20.720 --> 00:07:22.960 align:middle line:84%
airplane is present
up there or not.

00:07:22.960 --> 00:07:27.790 align:middle line:84%
Y is a discrete measurement, for
example, whether our radar

00:07:27.790 --> 00:07:30.040 align:middle line:90%
beeped or it didn't beep.

00:07:30.040 --> 00:07:33.860 align:middle line:84%
And we make inferences and
calculate the probability that

00:07:33.860 --> 00:07:37.860 align:middle line:84%
the plane is there, or the
probability that the plane is

00:07:37.860 --> 00:07:41.000 align:middle line:84%
not there, given the measurement
that we have made.

00:07:41.000 --> 00:07:43.940 align:middle line:84%
And of course X and Y do not
need to be just binary.

00:07:43.940 --> 00:07:47.480 align:middle line:84%
They could be more general
discrete random variables.

00:07:47.480 --> 00:07:50.900 align:middle line:84%
So how does the story change
in the continuous case?

00:07:50.900 --> 00:07:53.290 align:middle line:84%
First, what's a possible
application of

00:07:53.290 --> 00:07:54.570 align:middle line:90%
the continuous case?

00:07:54.570 --> 00:07:59.620 align:middle line:84%
Well, think of X as being some
signal that takes values over

00:07:59.620 --> 00:08:00.630 align:middle line:90%
a continuous range.

00:08:00.630 --> 00:08:04.730 align:middle line:84%
Let's say X is the current
through a resistor.

00:08:04.730 --> 00:08:07.530 align:middle line:84%
And then you have some measuring
device that measures

00:08:07.530 --> 00:08:11.530 align:middle line:84%
currents, but that device is
noisy, it gets hit, let's say

00:08:11.530 --> 00:08:13.640 align:middle line:84%
for example, by Gaussian
noise.

00:08:13.640 --> 00:08:18.340 align:middle line:84%
And the Y that you observe is a
noisy version of X. But your

00:08:18.340 --> 00:08:22.410 align:middle line:84%
instruments are analog, so
you measure things on

00:08:22.410 --> 00:08:24.750 align:middle line:90%
a continuous scale.

00:08:24.750 --> 00:08:26.250 align:middle line:84%
What are you going to
do in that case?

00:08:26.250 --> 00:08:29.920 align:middle line:84%
Well the inference problem, the
output of the inference

00:08:29.920 --> 00:08:33.360 align:middle line:84%
problem, is going to be the
conditional distribution of X.

00:08:33.360 --> 00:08:38.950 align:middle line:84%
What do you think your current
is based on a particular value

00:08:38.950 --> 00:08:40.870 align:middle line:90%
of Y that you have observed?

00:08:40.870 --> 00:08:44.480 align:middle line:84%
So the output of our inference
problem is, given the specific

00:08:44.480 --> 00:08:48.560 align:middle line:84%
value of Y, to calculate this
entire function as a function

00:08:48.560 --> 00:08:51.050 align:middle line:90%
of x, and then go and plot it.

00:08:51.050 --> 00:08:53.570 align:middle line:90%
How do we calculate it?

00:08:53.570 --> 00:08:57.410 align:middle line:84%
You go through the same
calculation as in the discrete

00:08:57.410 --> 00:09:01.590 align:middle line:84%
case, except that all of the
x's gets replaced by p's.

00:09:01.590 --> 00:09:04.630 align:middle line:84%
In the continuous case, it's
equally true that the joint's

00:09:04.630 --> 00:09:07.790 align:middle line:84%
density is the product of the
marginal density with the

00:09:07.790 --> 00:09:09.220 align:middle line:90%
conditional density.

00:09:09.220 --> 00:09:11.400 align:middle line:84%
So the formula is still
valid with just a

00:09:11.400 --> 00:09:13.160 align:middle line:90%
little change of notation.

00:09:13.160 --> 00:09:16.480 align:middle line:84%
So we end up with the same
formula here, except that we

00:09:16.480 --> 00:09:18.990 align:middle line:90%
replace x's with p's.

00:09:18.990 --> 00:09:23.240 align:middle line:84%
So all of these functions
are known to us.

00:09:23.240 --> 00:09:25.500 align:middle line:90%
We have formulas for them.

00:09:25.500 --> 00:09:29.400 align:middle line:84%
We fix a specific value of y,
we plug it in, so we're left

00:09:29.400 --> 00:09:30.640 align:middle line:90%
with a function of x.

00:09:30.640 --> 00:09:33.420 align:middle line:84%
And that gives us the posterior
distribution.

00:09:33.420 --> 00:09:38.130 align:middle line:84%
Actually there's also a
denominator term that's not

00:09:38.130 --> 00:09:42.340 align:middle line:84%
necessarily given to us, but we
can always calculate it if

00:09:42.340 --> 00:09:45.650 align:middle line:84%
we have the marginal of X,
and we have the model for

00:09:45.650 --> 00:09:47.250 align:middle line:90%
measuring device.

00:09:47.250 --> 00:09:50.960 align:middle line:84%
Then we can always find the
marginal distribution of Y. So

00:09:50.960 --> 00:09:54.630 align:middle line:84%
this quantity, that number, is
in general a known one, as

00:09:54.630 --> 00:09:58.490 align:middle line:84%
well, and doesn't give
us any problems.

00:09:58.490 --> 00:10:03.140 align:middle line:84%
So to complicate things a little
bit, we can also look

00:10:03.140 --> 00:10:07.610 align:middle line:84%
into situations where our two
random variables are of

00:10:07.610 --> 00:10:09.080 align:middle line:90%
different kinds.

00:10:09.080 --> 00:10:12.290 align:middle line:84%
For example, one random variable
could be discrete,

00:10:12.290 --> 00:10:15.280 align:middle line:84%
and the other it might
be continuous.

00:10:15.280 --> 00:10:17.340 align:middle line:90%
And there's two versions.

00:10:17.340 --> 00:10:22.320 align:middle line:84%
Here one version is when X is
discrete, but Y is continuous.

00:10:22.320 --> 00:10:25.130 align:middle line:90%
What's an example of this?

00:10:25.130 --> 00:10:30.690 align:middle line:84%
Well suppose that I send a
single bit of information so

00:10:30.690 --> 00:10:34.620 align:middle line:90%
my X is 0 or 1.

00:10:34.620 --> 00:10:39.710 align:middle line:84%
And what I measure is Y,
which is X plus, let's

00:10:39.710 --> 00:10:42.360 align:middle line:90%
say, Gaussian noise.

00:10:42.360 --> 00:10:48.960 align:middle line:90%


00:10:48.960 --> 00:10:52.550 align:middle line:84%
This is the standard example
that shows up in any textbook

00:10:52.550 --> 00:10:55.220 align:middle line:84%
on communication, or
signal processing.

00:10:55.220 --> 00:10:58.530 align:middle line:84%
You send a single bit, but what
you observe is a noisy

00:10:58.530 --> 00:11:02.120 align:middle line:90%
version of that bit.

00:11:02.120 --> 00:11:05.150 align:middle line:84%
You start with a model
of your x's.

00:11:05.150 --> 00:11:07.610 align:middle line:84%
These would be your prior
probabilities.

00:11:07.610 --> 00:11:11.670 align:middle line:84%
For example, you might be
believe that either 0 or 1 are

00:11:11.670 --> 00:11:16.250 align:middle line:84%
equally likely, in which case
your PMF gives equal weight to

00:11:16.250 --> 00:11:18.320 align:middle line:90%
two possible values.

00:11:18.320 --> 00:11:21.840 align:middle line:84%
And then we need a model of
our measuring device.

00:11:21.840 --> 00:11:23.990 align:middle line:90%
This is one specific model.

00:11:23.990 --> 00:11:28.090 align:middle line:84%
The general model would have
a shape such as follows.

00:11:28.090 --> 00:11:37.560 align:middle line:84%
Y has a distribution,
its density.

00:11:37.560 --> 00:11:41.590 align:middle line:84%
And that density, however,
depends on the value of X.

00:11:41.590 --> 00:11:46.170 align:middle line:84%
So when x is 0, we might get
a density of this kind.

00:11:46.170 --> 00:11:50.250 align:middle line:84%
And when x is 1, we might
get the density

00:11:50.250 --> 00:11:52.210 align:middle line:90%
of a different kind.

00:11:52.210 --> 00:11:57.010 align:middle line:84%
So these are the conditional
densities of y in a universe

00:11:57.010 --> 00:11:59.730 align:middle line:84%
that's specified by a particular
value of x.

00:11:59.730 --> 00:12:04.660 align:middle line:90%


00:12:04.660 --> 00:12:09.040 align:middle line:84%
And then we go ahead and
do our inference.

00:12:09.040 --> 00:12:13.520 align:middle line:84%
OK, what's the right formula
for doing this inference?

00:12:13.520 --> 00:12:18.270 align:middle line:84%
We need a formula that's sort of
an analog of this one, but

00:12:18.270 --> 00:12:22.210 align:middle line:84%
applies to the case where we
have two random variables of

00:12:22.210 --> 00:12:23.670 align:middle line:90%
different kinds.

00:12:23.670 --> 00:12:29.370 align:middle line:84%
So let me just redo this
calculation here.

00:12:29.370 --> 00:12:33.250 align:middle line:84%
Except that I'm not going to
have a probability of taking

00:12:33.250 --> 00:12:34.340 align:middle line:90%
specific values.

00:12:34.340 --> 00:12:36.800 align:middle line:84%
It will have to be something
a little different.

00:12:36.800 --> 00:12:39.250 align:middle line:90%
So here's how it goes.

00:12:39.250 --> 00:12:44.340 align:middle line:84%
Let's look at the probability
that X takes a specific value

00:12:44.340 --> 00:12:47.510 align:middle line:84%
that makes sense in the discrete
case, but for the

00:12:47.510 --> 00:12:50.040 align:middle line:84%
continuous random variable,
let's look at the probability

00:12:50.040 --> 00:12:53.480 align:middle line:84%
that it takes values in
some little interval.

00:12:53.480 --> 00:12:55.940 align:middle line:84%
And now this probability of
two things happening, I'm

00:12:55.940 --> 00:12:57.520 align:middle line:84%
going to write it
as a product.

00:12:57.520 --> 00:12:59.450 align:middle line:84%
And I'm going to write
this as a product in

00:12:59.450 --> 00:13:01.350 align:middle line:90%
two different ways.

00:13:01.350 --> 00:13:09.360 align:middle line:84%
So one way is to say that this
is the probability that X

00:13:09.360 --> 00:13:13.670 align:middle line:84%
takes that value and then given
that X takes that value,

00:13:13.670 --> 00:13:19.310 align:middle line:84%
the probability that Y falls
inside that interval.

00:13:19.310 --> 00:13:21.810 align:middle line:84%
So this is our usual
multiplication rule for

00:13:21.810 --> 00:13:25.330 align:middle line:84%
multiplying probabilities, but
I can use the multiplication

00:13:25.330 --> 00:13:27.610 align:middle line:90%
rule also in a different way.

00:13:27.610 --> 00:13:30.210 align:middle line:84%
It's the probability
that Y falls in

00:13:30.210 --> 00:13:33.460 align:middle line:90%
the range of interest.

00:13:33.460 --> 00:13:36.990 align:middle line:84%
And then the probability that X
takes the value of interest

00:13:36.990 --> 00:13:41.145 align:middle line:84%
given that Y satisfies
the first condition.

00:13:41.145 --> 00:13:45.960 align:middle line:90%


00:13:45.960 --> 00:13:53.760 align:middle line:84%
So this is something that's
definitely true.

00:13:53.760 --> 00:13:57.410 align:middle line:84%
We're just using the
multiplication rule.

00:13:57.410 --> 00:14:02.240 align:middle line:84%
And now let's translate it
into PMF is PDF notation.

00:14:02.240 --> 00:14:07.130 align:middle line:84%
So the entry up there is the
PMF of X evaluated at x.

00:14:07.130 --> 00:14:10.030 align:middle line:90%
The second entry, what is it?

00:14:10.030 --> 00:14:12.230 align:middle line:84%
Well probabilities of
little intervals are

00:14:12.230 --> 00:14:13.480 align:middle line:90%
given to us by densities.

00:14:13.480 --> 00:14:16.010 align:middle line:90%


00:14:16.010 --> 00:14:19.160 align:middle line:84%
But we are in the conditional
universe where X takes on a

00:14:19.160 --> 00:14:20.430 align:middle line:90%
particular value.

00:14:20.430 --> 00:14:27.450 align:middle line:84%
So it's going to be the density
of Y given the value

00:14:27.450 --> 00:14:30.210 align:middle line:90%
of X times delta.

00:14:30.210 --> 00:14:32.790 align:middle line:84%
So probabilities of little
intervals are given by the

00:14:32.790 --> 00:14:36.430 align:middle line:84%
density times the length of
the little interval, but

00:14:36.430 --> 00:14:39.390 align:middle line:84%
because we're working in the
conditional universe, it has

00:14:39.390 --> 00:14:41.230 align:middle line:90%
to be the conditional density.

00:14:41.230 --> 00:14:43.860 align:middle line:84%
Now let's try the second
expression.

00:14:43.860 --> 00:14:46.690 align:middle line:84%
This is the probability
that the Y falls

00:14:46.690 --> 00:14:48.040 align:middle line:90%
into the little interval.

00:14:48.040 --> 00:14:51.160 align:middle line:84%
So that's the density
of Y times delta.

00:14:51.160 --> 00:14:53.950 align:middle line:84%
And then here we have an
object which is the

00:14:53.950 --> 00:14:59.690 align:middle line:84%
conditional probability X in a
universe where the value of Y

00:14:59.690 --> 00:15:00.940 align:middle line:90%
is given to us.

00:15:00.940 --> 00:15:04.900 align:middle line:90%


00:15:04.900 --> 00:15:08.830 align:middle line:84%
Now this relation is sort
of approximate.

00:15:08.830 --> 00:15:13.630 align:middle line:84%
This is true for very small
delta in the limit.

00:15:13.630 --> 00:15:17.880 align:middle line:84%
But we can cancel the deltas
from both sides, and we're

00:15:17.880 --> 00:15:21.800 align:middle line:84%
left with a formula that links
together PMFs and PDFs.

00:15:21.800 --> 00:15:25.340 align:middle line:84%
Now this may look terribly
confusing because there's both

00:15:25.340 --> 00:15:27.730 align:middle line:90%
p's and f's involved.

00:15:27.730 --> 00:15:29.850 align:middle line:90%
But the logic should be clear.

00:15:29.850 --> 00:15:32.590 align:middle line:84%
If a random variable
is discrete, it's

00:15:32.590 --> 00:15:34.480 align:middle line:90%
described by PMF.

00:15:34.480 --> 00:15:38.120 align:middle line:84%
So here we're talking about
the PMF of X in some

00:15:38.120 --> 00:15:39.130 align:middle line:90%
particular universe.

00:15:39.130 --> 00:15:41.210 align:middle line:84%
X is discrete, so
it has a PMF.

00:15:41.210 --> 00:15:42.320 align:middle line:90%
Similarly here.

00:15:42.320 --> 00:15:45.380 align:middle line:84%
Y is continuous so it's
described by a PDF.

00:15:45.380 --> 00:15:47.840 align:middle line:84%
And even in the conditional
universe where I tell you the

00:15:47.840 --> 00:15:50.900 align:middle line:84%
value of X, Y is still a
continuous random variable, so

00:15:50.900 --> 00:15:53.280 align:middle line:90%
it's been described by a PDF.

00:15:53.280 --> 00:15:55.430 align:middle line:84%
So this is the basic
relation that links

00:15:55.430 --> 00:15:57.360 align:middle line:90%
together PMF and PDFs.

00:15:57.360 --> 00:15:59.080 align:middle line:90%
In this mixed the world.

00:15:59.080 --> 00:16:04.270 align:middle line:84%
And now in this inequality,
you can take this term and

00:16:04.270 --> 00:16:07.830 align:middle line:84%
send it to the new denominator
to the other side.

00:16:07.830 --> 00:16:10.070 align:middle line:84%
And what you end up with
is the formula

00:16:10.070 --> 00:16:11.830 align:middle line:90%
that we have up here.

00:16:11.830 --> 00:16:15.640 align:middle line:84%
And this is a formula that we
can use to make inferences

00:16:15.640 --> 00:16:18.780 align:middle line:84%
about the discrete random
variable X when we're told the

00:16:18.780 --> 00:16:26.540 align:middle line:84%
value of the continuous random
variable Y. The probability

00:16:26.540 --> 00:16:29.690 align:middle line:84%
that X takes on a particular
value has something

00:16:29.690 --> 00:16:31.330 align:middle line:90%
to do with the prior.

00:16:31.330 --> 00:16:36.520 align:middle line:84%
And other than that, it's
proportional to this quantity,

00:16:36.520 --> 00:16:41.720 align:middle line:84%
the conditional of Y given X.
So these are the quantities

00:16:41.720 --> 00:16:43.190 align:middle line:90%
that we plotted here.

00:16:43.190 --> 00:16:47.550 align:middle line:84%
Suppose that the x's are equally
likely in your prior,

00:16:47.550 --> 00:16:50.210 align:middle line:84%
so we don't really care
about that term.

00:16:50.210 --> 00:16:55.530 align:middle line:84%
It tells us that the posterior
of X is proportional to that

00:16:55.530 --> 00:16:58.520 align:middle line:84%
particular density under
the given x's.

00:16:58.520 --> 00:17:03.350 align:middle line:84%
So in this picture, if I were to
get a particular y here, I

00:17:03.350 --> 00:17:07.200 align:middle line:84%
would say that x equals 1
has a probability that's

00:17:07.200 --> 00:17:09.220 align:middle line:90%
proportional to this quantity.

00:17:09.220 --> 00:17:11.470 align:middle line:84%
x equals 0 has a probability
that's

00:17:11.470 --> 00:17:13.599 align:middle line:90%
proportional to this quantity.

00:17:13.599 --> 00:17:16.910 align:middle line:84%
So the ratio of these two
quantities gives us the

00:17:16.910 --> 00:17:21.200 align:middle line:84%
relative odds of the different
x's given the y

00:17:21.200 --> 00:17:24.010 align:middle line:90%
that we have observed.

00:17:24.010 --> 00:17:28.099 align:middle line:84%
So we're going to come back to
this topic and redo plenty of

00:17:28.099 --> 00:17:31.350 align:middle line:84%
examples of these kinds towards
the end of the class,

00:17:31.350 --> 00:17:34.480 align:middle line:84%
when we spend some
time dedicated

00:17:34.480 --> 00:17:36.130 align:middle line:90%
to inference problems.

00:17:36.130 --> 00:17:39.890 align:middle line:84%
But already at this stage, we
sort of have the basic skills

00:17:39.890 --> 00:17:42.000 align:middle line:90%
to deal with a lot of that.

00:17:42.000 --> 00:17:43.840 align:middle line:84%
And it's useful at this
point to pull all

00:17:43.840 --> 00:17:45.610 align:middle line:90%
the formulas together.

00:17:45.610 --> 00:17:49.990 align:middle line:84%
So finally let's look at the
last case that's remaining.

00:17:49.990 --> 00:17:54.440 align:middle line:84%
Here we have a continuous
phenomenon that we're trying

00:17:54.440 --> 00:17:57.770 align:middle line:84%
to measure, but our measurements
are discrete.

00:17:57.770 --> 00:18:00.780 align:middle line:84%
What's an example where
this might happen?

00:18:00.780 --> 00:18:05.270 align:middle line:84%
So you have some device that
emits light, and you drive it

00:18:05.270 --> 00:18:07.500 align:middle line:84%
with a current that has
a certain intensity.

00:18:07.500 --> 00:18:09.910 align:middle line:84%
You don't know what that
current is, and it's a

00:18:09.910 --> 00:18:12.120 align:middle line:90%
continuous random variable.

00:18:12.120 --> 00:18:14.600 align:middle line:84%
But the device emits
light by sending

00:18:14.600 --> 00:18:16.580 align:middle line:90%
out individual photons.

00:18:16.580 --> 00:18:20.480 align:middle line:84%
And your measurement is some
other device that counts how

00:18:20.480 --> 00:18:23.250 align:middle line:84%
many photons did you get
in a single second.

00:18:23.250 --> 00:18:28.020 align:middle line:84%
So if we have devices that emit
a very low intensity you

00:18:28.020 --> 00:18:31.720 align:middle line:84%
can actually start counting
individual photons as they're

00:18:31.720 --> 00:18:32.980 align:middle line:90%
being observed.

00:18:32.980 --> 00:18:35.390 align:middle line:84%
So we have a discrete
measurement, which is the

00:18:35.390 --> 00:18:38.920 align:middle line:84%
number of problems, and we
have a continuous hidden

00:18:38.920 --> 00:18:43.060 align:middle line:84%
random variable that we're
trying to estimate.

00:18:43.060 --> 00:18:45.790 align:middle line:90%
What do we do in this case?

00:18:45.790 --> 00:18:52.600 align:middle line:84%
Well we start again with a
formula of this kind, and send

00:18:52.600 --> 00:18:55.560 align:middle line:90%
the p term to the denominator.

00:18:55.560 --> 00:18:58.180 align:middle line:84%
And that's the formula that we
use there, except that the

00:18:58.180 --> 00:19:01.100 align:middle line:84%
roles of x's and y's
are interchanged.

00:19:01.100 --> 00:19:06.810 align:middle line:84%
So since here we have Y being
discrete, we should change all

00:19:06.810 --> 00:19:07.590 align:middle line:90%
the subscripts.

00:19:07.590 --> 00:19:15.490 align:middle line:84%
It would be p_Y f_X given
y f_X, and P(Y given X).

00:19:15.490 --> 00:19:19.230 align:middle line:84%
So just change all
those subscripts.

00:19:19.230 --> 00:19:22.740 align:middle line:84%
Because now what we're used to
be continuous became discrete,

00:19:22.740 --> 00:19:25.310 align:middle line:90%
and vice versa.

00:19:25.310 --> 00:19:27.360 align:middle line:84%
Take that formula, send
the other terms to the

00:19:27.360 --> 00:19:32.140 align:middle line:84%
denominator, and we have a
formula for the density, or X,

00:19:32.140 --> 00:19:34.370 align:middle line:84%
given the particular
measurements for Y that we

00:19:34.370 --> 00:19:36.350 align:middle line:90%
have obtained.

00:19:36.350 --> 00:19:41.420 align:middle line:84%
In some sense that's all there
is in Bayesian inference.

00:19:41.420 --> 00:19:46.540 align:middle line:84%
It's using these very simple
one line formulas.

00:19:46.540 --> 00:19:51.210 align:middle line:84%
But why are there people then
who make their living solving

00:19:51.210 --> 00:19:52.550 align:middle line:90%
inference problems?

00:19:52.550 --> 00:19:54.990 align:middle line:84%
Well, the devil is
in the details.

00:19:54.990 --> 00:19:57.460 align:middle line:84%
As we're going to discuss,
there are some real world

00:19:57.460 --> 00:20:01.150 align:middle line:84%
issues of how exactly do you
design your f's, how do you

00:20:01.150 --> 00:20:04.680 align:middle line:84%
model your system, then how do
you do your calculations.

00:20:04.680 --> 00:20:06.940 align:middle line:90%
This might not be always easy.

00:20:06.940 --> 00:20:09.710 align:middle line:84%
For example, there's certain
integrals or sums that have to

00:20:09.710 --> 00:20:12.900 align:middle line:84%
be evaluated, which may be
hard to do and so on.

00:20:12.900 --> 00:20:14.900 align:middle line:84%
So this object is
a lot of richer

00:20:14.900 --> 00:20:16.730 align:middle line:90%
than just these formulas.

00:20:16.730 --> 00:20:21.270 align:middle line:84%
On the other hand, at the
conceptual level, that's the

00:20:21.270 --> 00:20:23.910 align:middle line:84%
basis for Bayesian inference,
that these

00:20:23.910 --> 00:20:25.160 align:middle line:90%
are the basic concepts.

00:20:25.160 --> 00:20:27.570 align:middle line:90%


00:20:27.570 --> 00:20:30.850 align:middle line:84%
All right, so now let's change
gear and move to the new

00:20:30.850 --> 00:20:36.180 align:middle line:84%
subject, which is the topic of
finding the distribution of a

00:20:36.180 --> 00:20:38.360 align:middle line:84%
functional for a random
variable.

00:20:38.360 --> 00:20:42.820 align:middle line:84%
We call those distributions
derived distributions, because

00:20:42.820 --> 00:20:45.480 align:middle line:84%
we're given the distribution
of X. We're interested in a

00:20:45.480 --> 00:20:48.980 align:middle line:84%
function of X. We want to derive
the distribution of

00:20:48.980 --> 00:20:51.020 align:middle line:84%
that function based on
the distribution

00:20:51.020 --> 00:20:53.060 align:middle line:90%
that we already know.

00:20:53.060 --> 00:20:56.610 align:middle line:84%
So it could be a function of
just one random variable.

00:20:56.610 --> 00:20:59.170 align:middle line:84%
It could be a function of
several random variables.

00:20:59.170 --> 00:21:02.880 align:middle line:84%
So one example that we are going
to solve at some point,

00:21:02.880 --> 00:21:05.830 align:middle line:84%
let's say you have to run the
variables X and Y. Somebody

00:21:05.830 --> 00:21:09.055 align:middle line:84%
tells you their distribution,
for example, is a uniform of

00:21:09.055 --> 00:21:10.000 align:middle line:90%
the square.

00:21:10.000 --> 00:21:12.120 align:middle line:84%
For some reason, you're
interested in the ratio of

00:21:12.120 --> 00:21:14.660 align:middle line:84%
these two random variables,
and you want to find the

00:21:14.660 --> 00:21:16.910 align:middle line:90%
distribution of that ratio.

00:21:16.910 --> 00:21:21.810 align:middle line:84%
You can think of lots of cases
where your random variable of

00:21:21.810 --> 00:21:25.950 align:middle line:84%
interest is created by taking
some other unknown variables

00:21:25.950 --> 00:21:27.570 align:middle line:90%
and taking a function of them.

00:21:27.570 --> 00:21:31.170 align:middle line:84%
And so it's legitimate to care
about the distribution of that

00:21:31.170 --> 00:21:33.310 align:middle line:90%
random variable.

00:21:33.310 --> 00:21:35.560 align:middle line:90%
A caveat, however.

00:21:35.560 --> 00:21:39.480 align:middle line:84%
There's an important case where
you don't need to find

00:21:39.480 --> 00:21:41.840 align:middle line:84%
the distribution of that
random variable.

00:21:41.840 --> 00:21:44.600 align:middle line:84%
And this is when you want to
calculate the expectations.

00:21:44.600 --> 00:21:47.750 align:middle line:84%
If all you care about is the
expected value of this

00:21:47.750 --> 00:21:50.580 align:middle line:84%
function of the random
variables, you can work

00:21:50.580 --> 00:21:53.800 align:middle line:84%
directly with the distribution
of the original random

00:21:53.800 --> 00:21:58.490 align:middle line:84%
variables without ever having
to find the PDF of g.

00:21:58.490 --> 00:22:03.790 align:middle line:84%
So you don't do unnecessary work
if it's not needed, but

00:22:03.790 --> 00:22:06.290 align:middle line:84%
if it's needed, or if you're
asked to do it,

00:22:06.290 --> 00:22:08.470 align:middle line:90%
then you just do it.

00:22:08.470 --> 00:22:13.040 align:middle line:84%
So how do we find the
distribution of the function?

00:22:13.040 --> 00:22:17.690 align:middle line:84%
As a warm-up, let's look
at the discrete case.

00:22:17.690 --> 00:22:21.120 align:middle line:84%
Suppose that X is a discrete
random variable and takes

00:22:21.120 --> 00:22:22.550 align:middle line:90%
certain values.

00:22:22.550 --> 00:22:27.070 align:middle line:84%
We have a function g that
maps x's into y's.

00:22:27.070 --> 00:22:30.430 align:middle line:84%
And we want to find the
probability mass function for

00:22:30.430 --> 00:22:31.930 align:middle line:90%
Y.

00:22:31.930 --> 00:22:36.780 align:middle line:84%
So for example, if I'm
interested in finding the

00:22:36.780 --> 00:22:41.020 align:middle line:84%
probability that Y takes on
this particular value, how

00:22:41.020 --> 00:22:42.910 align:middle line:90%
would they find it?

00:22:42.910 --> 00:22:46.890 align:middle line:84%
Well I ask, what are the
different ways that these

00:22:46.890 --> 00:22:49.390 align:middle line:90%
particular y value can happen?

00:22:49.390 --> 00:22:53.390 align:middle line:84%
And the different ways that it
can happen is either if x

00:22:53.390 --> 00:22:56.800 align:middle line:84%
takes this value, or if
X takes that value.

00:22:56.800 --> 00:23:02.650 align:middle line:84%
So we identify this event in the
y space with that event in

00:23:02.650 --> 00:23:04.220 align:middle line:90%
the x space.

00:23:04.220 --> 00:23:06.790 align:middle line:84%
These two events
are identical.

00:23:06.790 --> 00:23:12.350 align:middle line:84%
X falls in this set if and only
if Y falls in that set.

00:23:12.350 --> 00:23:15.060 align:middle line:84%
Therefore, the probability of
Y falling in that set is the

00:23:15.060 --> 00:23:17.540 align:middle line:84%
probability of X falling
in that set.

00:23:17.540 --> 00:23:20.890 align:middle line:84%
The probability of X falling in
that set is just the sum of

00:23:20.890 --> 00:23:24.650 align:middle line:84%
the individual probabilities
of the x's in this set.

00:23:24.650 --> 00:23:27.360 align:middle line:84%
So we just add the probabilities
of the different

00:23:27.360 --> 00:23:31.300 align:middle line:84%
x's where the summation is taken
over all x's that leads

00:23:31.300 --> 00:23:35.070 align:middle line:90%
to that particular value of y.

00:23:35.070 --> 00:23:35.860 align:middle line:90%
Very good.

00:23:35.860 --> 00:23:39.090 align:middle line:84%
So that's all there is
in the discrete case.

00:23:39.090 --> 00:23:41.070 align:middle line:90%
It's a very nice and simple.

00:23:41.070 --> 00:23:43.460 align:middle line:84%
So let's transfer
these methods to

00:23:43.460 --> 00:23:45.810 align:middle line:90%
the continuous case.

00:23:45.810 --> 00:23:47.890 align:middle line:84%
Suppose we are in the
continuous case.

00:23:47.890 --> 00:23:52.140 align:middle line:84%
Suppose that X and Y now can
take values anywhere.

00:23:52.140 --> 00:23:55.440 align:middle line:84%
And I try to use same methods
and I ask, what is the

00:23:55.440 --> 00:24:00.340 align:middle line:84%
probability that Y is going
to take this value?

00:24:00.340 --> 00:24:03.100 align:middle line:84%
At least if the diagram is this
way, you would say this

00:24:03.100 --> 00:24:06.990 align:middle line:84%
is the same as the probability
that X takes this value.

00:24:06.990 --> 00:24:10.220 align:middle line:84%
So I can find the probability
of Y being this in terms of

00:24:10.220 --> 00:24:12.600 align:middle line:84%
the probability of
X being that.

00:24:12.600 --> 00:24:14.610 align:middle line:90%
Is this useful?

00:24:14.610 --> 00:24:16.480 align:middle line:84%
In the continuous
case, it's not.

00:24:16.480 --> 00:24:19.830 align:middle line:84%
Because in the continuous case,
any single value has 0

00:24:19.830 --> 00:24:21.020 align:middle line:90%
probability.

00:24:21.020 --> 00:24:25.450 align:middle line:84%
So what you're going to get out
of this argument is that

00:24:25.450 --> 00:24:29.530 align:middle line:84%
the probability Y takes this
value is 0, is equal to the

00:24:29.530 --> 00:24:32.800 align:middle line:84%
probability that X takes that
value which also 0.

00:24:32.800 --> 00:24:34.650 align:middle line:90%
That doesn't help us.

00:24:34.650 --> 00:24:36.060 align:middle line:90%
We want to do something more.

00:24:36.060 --> 00:24:40.650 align:middle line:84%
We want to actually find,
perhaps, the density of Y, as

00:24:40.650 --> 00:24:43.550 align:middle line:84%
opposed to the probabilities
of individual y's.

00:24:43.550 --> 00:24:47.620 align:middle line:84%
So to find the density of Y,
you might argue as follows.

00:24:47.620 --> 00:24:51.100 align:middle line:84%
I'm looking at an interval for
y, and I ask what's the

00:24:51.100 --> 00:24:53.510 align:middle line:84%
probability of falling
in this interval.

00:24:53.510 --> 00:24:57.890 align:middle line:84%
And you go back and find the
corresponding set of x's that

00:24:57.890 --> 00:25:02.090 align:middle line:84%
leads to those y's, and equate
those two probabilities.

00:25:02.090 --> 00:25:04.960 align:middle line:84%
The probability of all of those
y's collectively should

00:25:04.960 --> 00:25:09.710 align:middle line:84%
be equal to the probability of
all of the x's that map into

00:25:09.710 --> 00:25:11.930 align:middle line:90%
that interval collectively.

00:25:11.930 --> 00:25:16.010 align:middle line:84%
And this way you can
relate the two.

00:25:16.010 --> 00:25:22.870 align:middle line:84%
As far as the mechanics go, in
many cases it's easier to not

00:25:22.870 --> 00:25:26.670 align:middle line:84%
to work with little intervals,
but instead to work with

00:25:26.670 --> 00:25:30.110 align:middle line:84%
cumulative distribution
functions that used to work

00:25:30.110 --> 00:25:32.600 align:middle line:90%
with sort of big intervals.

00:25:32.600 --> 00:25:35.460 align:middle line:84%
So you can instead do
a different picture.

00:25:35.460 --> 00:25:38.250 align:middle line:90%
Look at this set of y's.

00:25:38.250 --> 00:25:41.690 align:middle line:84%
This is the set of y's
that are smaller

00:25:41.690 --> 00:25:43.200 align:middle line:90%
than a certain value.

00:25:43.200 --> 00:25:46.990 align:middle line:84%
The probability of this set
is given by the cumulative

00:25:46.990 --> 00:25:49.740 align:middle line:84%
distribution of the
random variable Y.

00:25:49.740 --> 00:25:54.450 align:middle line:84%
Now this set of y's gets
produced by some corresponding

00:25:54.450 --> 00:25:56.850 align:middle line:90%
set of x's.

00:25:56.850 --> 00:26:04.120 align:middle line:84%
Maybe these are the x's that
map into y's in that set.

00:26:04.120 --> 00:26:06.040 align:middle line:90%
And then we argue as follows.

00:26:06.040 --> 00:26:08.870 align:middle line:84%
The probability that the Y falls
in this interval is the

00:26:08.870 --> 00:26:12.600 align:middle line:84%
same as the probability that
X falls in that interval.

00:26:12.600 --> 00:26:15.810 align:middle line:84%
So the event of Y falling here
and the event of X falling

00:26:15.810 --> 00:26:19.330 align:middle line:84%
there are the same, so their
probabilities must be equal.

00:26:19.330 --> 00:26:22.010 align:middle line:84%
And then I do the calculations
here.

00:26:22.010 --> 00:26:25.050 align:middle line:84%
And I end up getting the
cumulative distribution

00:26:25.050 --> 00:26:28.760 align:middle line:84%
function of Y. Once I have the
cumulative, I can get the

00:26:28.760 --> 00:26:31.670 align:middle line:90%
density by just differentiating.

00:26:31.670 --> 00:26:34.900 align:middle line:84%
So this is the general cookbook
procedure that we

00:26:34.900 --> 00:26:37.886 align:middle line:84%
will be using to calculate
it derived distributions.

00:26:37.886 --> 00:26:40.450 align:middle line:90%


00:26:40.450 --> 00:26:43.500 align:middle line:84%
We're interested in a random
variable Y, which is a

00:26:43.500 --> 00:26:45.320 align:middle line:90%
function of the x's.

00:26:45.320 --> 00:26:50.070 align:middle line:84%
We will aim at obtaining the
cumulative distribution of Y.

00:26:50.070 --> 00:26:54.040 align:middle line:84%
Somehow, manage to calculate the
probability of this event.

00:26:54.040 --> 00:26:58.120 align:middle line:84%
Once we get it, and what I mean
by get it, I don't mean

00:26:58.120 --> 00:27:00.980 align:middle line:84%
getting it for a single
value of little y.

00:27:00.980 --> 00:27:04.640 align:middle line:84%
You need to get this
for all little y's.

00:27:04.640 --> 00:27:07.930 align:middle line:84%
So you need to get the
function itself, the

00:27:07.930 --> 00:27:09.480 align:middle line:90%
cumulative distribution.

00:27:09.480 --> 00:27:12.750 align:middle line:84%
Once you get it in that form,
then you can calculate the

00:27:12.750 --> 00:27:15.260 align:middle line:84%
derivative at any particular
point.

00:27:15.260 --> 00:27:18.000 align:middle line:84%
And this is going to give
you the density of Y.

00:27:18.000 --> 00:27:19.690 align:middle line:84%
So a simple two-step
procedure.

00:27:19.690 --> 00:27:24.050 align:middle line:84%
The devil is in the details of
how you carry the mechanics.

00:27:24.050 --> 00:27:27.580 align:middle line:90%
So let's do one first example.

00:27:27.580 --> 00:27:31.020 align:middle line:84%
Suppose that X is a uniform
random variable, takes values

00:27:31.020 --> 00:27:32.660 align:middle line:90%
between 0 and 2.

00:27:32.660 --> 00:27:35.605 align:middle line:84%
We're interested in the random
variable Y, which is the cube

00:27:35.605 --> 00:27:37.500 align:middle line:90%
of X. What kind of distribution

00:27:37.500 --> 00:27:38.840 align:middle line:90%
is it going to have?

00:27:38.840 --> 00:27:44.960 align:middle line:84%
Now first notice that Y takes
values between 0 and 8.

00:27:44.960 --> 00:27:48.810 align:middle line:84%
So X is uniform, so all the
x's are equally likely.

00:27:48.810 --> 00:27:51.680 align:middle line:90%


00:27:51.680 --> 00:27:55.340 align:middle line:84%
You might then say, well, in
that case, all the y's should

00:27:55.340 --> 00:27:56.740 align:middle line:90%
be equally likely.

00:27:56.740 --> 00:28:00.630 align:middle line:84%
So Y might also have a
uniform distribution.

00:28:00.630 --> 00:28:02.210 align:middle line:90%
Is this true?

00:28:02.210 --> 00:28:04.040 align:middle line:90%
We'll find out.

00:28:04.040 --> 00:28:06.990 align:middle line:84%
So let's start applying the
cookbook procedure.

00:28:06.990 --> 00:28:10.410 align:middle line:84%
We want to find first the
cumulative distribution of the

00:28:10.410 --> 00:28:14.890 align:middle line:84%
random variable Y, which by
definition is the probability

00:28:14.890 --> 00:28:17.370 align:middle line:84%
that the random variable is
less than or equal to a

00:28:17.370 --> 00:28:18.850 align:middle line:90%
certain number.

00:28:18.850 --> 00:28:20.680 align:middle line:90%
That's what we want to find.

00:28:20.680 --> 00:28:24.440 align:middle line:84%
What we have in our hands is the
distribution of X. That's

00:28:24.440 --> 00:28:26.320 align:middle line:90%
what we need to work with.

00:28:26.320 --> 00:28:30.090 align:middle line:84%
So the first step that you need
to do is to look at this

00:28:30.090 --> 00:28:33.680 align:middle line:84%
events and translate it, and
write it in terms of the

00:28:33.680 --> 00:28:39.040 align:middle line:84%
random variable about which you
know you have information.

00:28:39.040 --> 00:28:44.320 align:middle line:84%
So Y is X cubed, so this event
is the same as that event.

00:28:44.320 --> 00:28:46.760 align:middle line:84%
So now we can forget
about the y's.

00:28:46.760 --> 00:28:49.860 align:middle line:84%
It's just an exercise involving
a single random

00:28:49.860 --> 00:28:52.750 align:middle line:84%
variable with a known
distribution and we want to

00:28:52.750 --> 00:28:56.610 align:middle line:84%
calculate the probability
of some event.

00:28:56.610 --> 00:28:58.780 align:middle line:84%
So we're looking
at this event.

00:28:58.780 --> 00:29:02.230 align:middle line:84%
X cubed being less than or equal
to Y. We massage that

00:29:02.230 --> 00:29:06.130 align:middle line:84%
expression so that's it involves
X directly, so let's

00:29:06.130 --> 00:29:08.960 align:middle line:84%
take cubic roots of both sides
of this inequality.

00:29:08.960 --> 00:29:12.130 align:middle line:84%
This event is the same as the
event that X is less than or

00:29:12.130 --> 00:29:14.820 align:middle line:90%
equal to Y to the 1/3.

00:29:14.820 --> 00:29:19.300 align:middle line:84%
Now with a uniform distribution
on [0,2], what is

00:29:19.300 --> 00:29:22.070 align:middle line:90%
that probability going to be?

00:29:22.070 --> 00:29:27.710 align:middle line:84%
It's the probability of being in
the interval from 0 to y to

00:29:27.710 --> 00:29:34.680 align:middle line:84%
the 1/3, so it's going to be in
the area under the uniform

00:29:34.680 --> 00:29:37.010 align:middle line:90%
going up to that point.

00:29:37.010 --> 00:29:39.315 align:middle line:84%
And what's the area under
that uniform?

00:29:39.315 --> 00:29:42.650 align:middle line:90%


00:29:42.650 --> 00:29:44.290 align:middle line:90%
So here's x.

00:29:44.290 --> 00:29:50.810 align:middle line:84%
Here is the distribution
of X. It goes up to 2.

00:29:50.810 --> 00:29:53.330 align:middle line:84%
The distribution of
X is this one.

00:29:53.330 --> 00:29:56.860 align:middle line:84%
We want to go up to
y to the 1/3.

00:29:56.860 --> 00:30:02.390 align:middle line:84%
So the probability for this
event happening is this area.

00:30:02.390 --> 00:30:06.590 align:middle line:84%
And the area is equal to the
base, which is y to the 1/3

00:30:06.590 --> 00:30:08.250 align:middle line:90%
times the height.

00:30:08.250 --> 00:30:09.720 align:middle line:90%
What is the height?

00:30:09.720 --> 00:30:13.480 align:middle line:84%
Well since the density must
integrate to 1, the total area

00:30:13.480 --> 00:30:15.340 align:middle line:90%
under the curve has to be 1.

00:30:15.340 --> 00:30:19.660 align:middle line:84%
So the height here is 1/2, and
that explains why we get the

00:30:19.660 --> 00:30:22.530 align:middle line:90%
1/2 factor down there.

00:30:22.530 --> 00:30:24.900 align:middle line:84%
So that's the formula for the
cumulative distribution.

00:30:24.900 --> 00:30:26.070 align:middle line:90%
And then the rest is easy.

00:30:26.070 --> 00:30:28.340 align:middle line:90%
You just take derivatives.

00:30:28.340 --> 00:30:32.650 align:middle line:84%
You differentiate this
expression with respect to y

00:30:32.650 --> 00:30:36.240 align:middle line:84%
1/2 times 1/3, and y
drops by one power.

00:30:36.240 --> 00:30:39.670 align:middle line:84%
So you get y to 2/3 in
the denominator.

00:30:39.670 --> 00:30:55.490 align:middle line:84%
So if you wish to plot this,
it's 1/y to the 2/3.

00:30:55.490 --> 00:31:00.480 align:middle line:84%
So when y goes to 0, it sort
of blows up and it

00:31:00.480 --> 00:31:03.090 align:middle line:90%
goes on this way.

00:31:03.090 --> 00:31:06.090 align:middle line:84%
Is this picture correct
the way I've drawn it?

00:31:06.090 --> 00:31:08.900 align:middle line:90%


00:31:08.900 --> 00:31:11.256 align:middle line:90%
What's wrong with it?

00:31:11.256 --> 00:31:12.630 align:middle line:90%
[? AUDIENCE:  Something. ?]

00:31:12.630 --> 00:31:13.420 align:middle line:90%
PROFESSOR: Yes.

00:31:13.420 --> 00:31:17.610 align:middle line:84%
y only takes values
from 0 to 8.

00:31:17.610 --> 00:31:21.890 align:middle line:84%
This formula that I wrote here
is only correct when the

00:31:21.890 --> 00:31:25.000 align:middle line:90%
preview picture applies.

00:31:25.000 --> 00:31:31.070 align:middle line:84%
I took my y to the 1/3 to
be between 0 and 2.

00:31:31.070 --> 00:31:40.650 align:middle line:84%
So this formula here is only
correct for y between 0 and 8.

00:31:40.650 --> 00:31:43.770 align:middle line:90%


00:31:43.770 --> 00:31:46.610 align:middle line:84%
And for that reason, the formula
for the derivative is

00:31:46.610 --> 00:31:50.700 align:middle line:84%
also true only for a
y between 0 and 8.

00:31:50.700 --> 00:31:55.630 align:middle line:84%
And any other values of why are
impossible, so they get

00:31:55.630 --> 00:31:57.880 align:middle line:90%
zero density.

00:31:57.880 --> 00:32:04.070 align:middle line:84%
So to complete the picture
here, the PDF of y has a

00:32:04.070 --> 00:32:09.290 align:middle line:84%
cut-off of 8, and it's also
0 everywhere else.

00:32:09.290 --> 00:32:13.330 align:middle line:90%


00:32:13.330 --> 00:32:16.640 align:middle line:84%
And one thing that we see is
that the distribution of Y is

00:32:16.640 --> 00:32:17.980 align:middle line:90%
not uniform.

00:32:17.980 --> 00:32:24.240 align:middle line:84%
Certain y's are more likely than
others, even though we

00:32:24.240 --> 00:32:26.130 align:middle line:90%
started with a uniform random

00:32:26.130 --> 00:32:32.240 align:middle line:90%
variable X. All right.

00:32:32.240 --> 00:32:36.530 align:middle line:84%
So we will keep doing examples
of this kind, a sequence of

00:32:36.530 --> 00:32:40.350 align:middle line:84%
progressively more interesting
or more complicated.

00:32:40.350 --> 00:32:42.530 align:middle line:84%
So that's going to continue
in the next lecture.

00:32:42.530 --> 00:32:45.930 align:middle line:84%
You're going to see plenty of
examples in your recitations

00:32:45.930 --> 00:32:48.060 align:middle line:90%
and tutorials and so on.

00:32:48.060 --> 00:32:52.420 align:middle line:84%
So let's do one that's pretty
similar to the one that we

00:32:52.420 --> 00:32:57.730 align:middle line:84%
did, but it's going to add to
just a small twist in how we

00:32:57.730 --> 00:33:00.470 align:middle line:90%
do the mechanics.

00:33:00.470 --> 00:33:02.780 align:middle line:84%
OK so you set your
cruise control

00:33:02.780 --> 00:33:04.010 align:middle line:90%
when you start driving.

00:33:04.010 --> 00:33:06.310 align:middle line:84%
And you keep driving at the
constants based at the

00:33:06.310 --> 00:33:07.870 align:middle line:90%
constant speed.

00:33:07.870 --> 00:33:09.980 align:middle line:84%
Where you set your cruise
control is somewhere

00:33:09.980 --> 00:33:11.660 align:middle line:90%
between 30 and 60.

00:33:11.660 --> 00:33:14.520 align:middle line:84%
You're going to drive
a distance of 200.

00:33:14.520 --> 00:33:18.660 align:middle line:84%
And so the time it's going to
take for your trip is 200 over

00:33:18.660 --> 00:33:20.530 align:middle line:84%
the setting of your
cruise control.

00:33:20.530 --> 00:33:22.610 align:middle line:90%
So it's 200/V.

00:33:22.610 --> 00:33:26.210 align:middle line:84%
Somebody gives you the
distribution of V, and they

00:33:26.210 --> 00:33:29.490 align:middle line:84%
tell you not only it's between
30 and 60, it's roughly

00:33:29.490 --> 00:33:33.530 align:middle line:84%
equally likely to be anything
between 30 and 60, so we have

00:33:33.530 --> 00:33:36.280 align:middle line:84%
a uniform distribution
over that range.

00:33:36.280 --> 00:33:40.060 align:middle line:84%
So we have a distribution of
V. We want to find the

00:33:40.060 --> 00:33:43.460 align:middle line:84%
distribution of the random
variable T, which is the time

00:33:43.460 --> 00:33:46.540 align:middle line:90%
it takes till your trip ends.

00:33:46.540 --> 00:33:49.200 align:middle line:90%


00:33:49.200 --> 00:33:51.790 align:middle line:84%
So how are we going
to proceed?

00:33:51.790 --> 00:33:55.170 align:middle line:84%
We'll use the exact same
cookbook procedure.

00:33:55.170 --> 00:33:57.360 align:middle line:84%
We're going to start by
finding the cumulative

00:33:57.360 --> 00:34:02.920 align:middle line:84%
distribution of T.
What is this?

00:34:02.920 --> 00:34:05.730 align:middle line:84%
By definition, the cumulative
distribution is the

00:34:05.730 --> 00:34:10.230 align:middle line:84%
probability that T is less
than a certain number.

00:34:10.230 --> 00:34:12.070 align:middle line:90%
OK.

00:34:12.070 --> 00:34:15.340 align:middle line:84%
Now we don't know the
distribution of T, so we

00:34:15.340 --> 00:34:17.989 align:middle line:84%
cannot to work with these
event directly.

00:34:17.989 --> 00:34:21.960 align:middle line:84%
But we take that event and
translate it into T-space.

00:34:21.960 --> 00:34:28.205 align:middle line:84%
So we replace the t's by what we
know T to be in terms of V

00:34:28.205 --> 00:34:28.271 align:middle line:90%
or

00:34:28.271 --> 00:34:33.565 align:middle line:90%
the v's All right.

00:34:33.565 --> 00:34:36.230 align:middle line:90%


00:34:36.230 --> 00:34:39.659 align:middle line:84%
So we have the distribution
of V. So now let's

00:34:39.659 --> 00:34:41.739 align:middle line:90%
calculate this quantity.

00:34:41.739 --> 00:34:42.179 align:middle line:90%
OK.

00:34:42.179 --> 00:34:46.210 align:middle line:84%
Let's massage this event and
rewrite it as the probability

00:34:46.210 --> 00:35:06.880 align:middle line:84%
that V is larger or
equal to 200/T.

00:35:06.880 --> 00:35:10.870 align:middle line:90%
So what is this going to be?

00:35:10.870 --> 00:35:14.400 align:middle line:84%
So let's say that 200/T
is some number that

00:35:14.400 --> 00:35:16.015 align:middle line:90%
falls inside the range.

00:35:16.015 --> 00:35:19.150 align:middle line:90%


00:35:19.150 --> 00:35:24.630 align:middle line:84%
So that's going to be true if
200/T is bigger than 30, and

00:35:24.630 --> 00:35:26.610 align:middle line:90%
less than 60.

00:35:26.610 --> 00:35:37.110 align:middle line:84%
Which means that t is
less than 30/200.

00:35:37.110 --> 00:35:38.360 align:middle line:90%
No, 200/30.

00:35:38.360 --> 00:35:41.300 align:middle line:90%


00:35:41.300 --> 00:35:44.570 align:middle line:90%
And bigger than 200/60.

00:35:44.570 --> 00:35:51.360 align:middle line:84%
So for t's inside that range,
this number 200/t falls inside

00:35:51.360 --> 00:35:52.230 align:middle line:90%
that range.

00:35:52.230 --> 00:35:55.960 align:middle line:84%
This is the range of t's that
are possible, given the

00:35:55.960 --> 00:35:59.240 align:middle line:84%
description of the problem
the we have set up.

00:35:59.240 --> 00:36:04.940 align:middle line:84%
So for t's in that range, what
is the probability that V is

00:36:04.940 --> 00:36:07.900 align:middle line:90%
bigger than this number?

00:36:07.900 --> 00:36:11.550 align:middle line:84%
So V being bigger than that
number is the probability of

00:36:11.550 --> 00:36:17.000 align:middle line:84%
this event, so it's going to be
the area under this curve.

00:36:17.000 --> 00:36:22.880 align:middle line:84%
So the area under that curve
is the height of the curve,

00:36:22.880 --> 00:36:27.300 align:middle line:84%
which is 1/3 over 30
times the base.

00:36:27.300 --> 00:36:28.910 align:middle line:90%
How big is the base?

00:36:28.910 --> 00:36:33.060 align:middle line:84%
Well it's from that point to 60,
so the base has a length

00:36:33.060 --> 00:36:36.500 align:middle line:90%
of 60 minus 200/t.

00:36:36.500 --> 00:36:45.470 align:middle line:90%


00:36:45.470 --> 00:36:50.580 align:middle line:84%
And this is a formula which is
valid for those t's for which

00:36:50.580 --> 00:36:52.420 align:middle line:90%
this picture is correct.

00:36:52.420 --> 00:36:57.410 align:middle line:84%
And this picture is correct if
200/T happens to fall in this

00:36:57.410 --> 00:37:01.540 align:middle line:84%
interval, which is the same as
T falling in that interval,

00:37:01.540 --> 00:37:03.980 align:middle line:84%
which are the t's that
are possible.

00:37:03.980 --> 00:37:07.390 align:middle line:84%
So finally let's find the
density of T, which is what

00:37:07.390 --> 00:37:09.430 align:middle line:90%
we're looking for.

00:37:09.430 --> 00:37:12.450 align:middle line:84%
We find this by taking the
derivative in this expression

00:37:12.450 --> 00:37:14.370 align:middle line:90%
with respect to t.

00:37:14.370 --> 00:37:18.150 align:middle line:84%
We only get one term
from here.

00:37:18.150 --> 00:37:26.045 align:middle line:84%
And this is going to be 200/30,
1 over t squared.

00:37:26.045 --> 00:37:30.820 align:middle line:90%


00:37:30.820 --> 00:37:34.020 align:middle line:84%
And this is the formula for
the density for t's in the

00:37:34.020 --> 00:37:35.270 align:middle line:90%
allowed to range.

00:37:35.270 --> 00:37:46.890 align:middle line:90%


00:37:46.890 --> 00:37:51.130 align:middle line:84%
OK, so that's the end of the
solution to this particular

00:37:51.130 --> 00:37:52.880 align:middle line:90%
problem as well.

00:37:52.880 --> 00:37:55.640 align:middle line:84%
I said that there was a little
twist compared to

00:37:55.640 --> 00:37:57.130 align:middle line:90%
the previous one.

00:37:57.130 --> 00:37:58.410 align:middle line:90%
What was the twist?

00:37:58.410 --> 00:38:01.380 align:middle line:84%
Well the twist was that in the
previous problem we dealt with

00:38:01.380 --> 00:38:05.580 align:middle line:84%
the X cubed function, which was
monotonically increasing.

00:38:05.580 --> 00:38:07.760 align:middle line:84%
Here we dealt with the
function that was

00:38:07.760 --> 00:38:09.850 align:middle line:90%
monotonically decreasing.

00:38:09.850 --> 00:38:13.850 align:middle line:84%
So when we had to find the
probability that T is less

00:38:13.850 --> 00:38:17.220 align:middle line:84%
than something, that translated
into an event that

00:38:17.220 --> 00:38:19.640 align:middle line:90%
V was bigger than something.

00:38:19.640 --> 00:38:22.410 align:middle line:84%
Your time is less than something
if and only if your

00:38:22.410 --> 00:38:25.090 align:middle line:84%
velocity is bigger
than something.

00:38:25.090 --> 00:38:27.510 align:middle line:84%
So for when you're dealing
with the monotonically

00:38:27.510 --> 00:38:31.950 align:middle line:84%
decreasing function, at some
point some inequalities will

00:38:31.950 --> 00:38:33.200 align:middle line:90%
have to get reversed.

00:38:33.200 --> 00:38:38.540 align:middle line:90%


00:38:38.540 --> 00:38:43.700 align:middle line:84%
Finally let's look at
a very useful one.

00:38:43.700 --> 00:38:47.990 align:middle line:84%
Which is the case where we take
a linear function of a

00:38:47.990 --> 00:38:49.700 align:middle line:90%
random variable.

00:38:49.700 --> 00:38:55.810 align:middle line:84%
So X is a random variable with
given distribution, and we can

00:38:55.810 --> 00:38:57.110 align:middle line:84%
see there is a linear
function.

00:38:57.110 --> 00:38:59.920 align:middle line:84%
So in this particular instance,
we take a to be

00:38:59.920 --> 00:39:03.590 align:middle line:90%
equal to 2 and b equal to 5.

00:39:03.590 --> 00:39:08.680 align:middle line:84%
And let us first argue
just by picture.

00:39:08.680 --> 00:39:13.920 align:middle line:84%
So X is a random variable that
has a given distribution.

00:39:13.920 --> 00:39:16.150 align:middle line:84%
Let's say it's this
weird shape here.

00:39:16.150 --> 00:39:20.170 align:middle line:90%
And x ranges from -1 to +2.

00:39:20.170 --> 00:39:22.140 align:middle line:84%
Let's do things one
step at the time.

00:39:22.140 --> 00:39:26.190 align:middle line:84%
Let's first find the
distribution of 2X.

00:39:26.190 --> 00:39:28.960 align:middle line:84%
Why do you think you
know about 2X?

00:39:28.960 --> 00:39:35.330 align:middle line:84%
Well if x ranges from -1 to 2,
then the random variable X is

00:39:35.330 --> 00:39:36.580 align:middle line:90%
going to range from -2 to +4.

00:39:36.580 --> 00:39:39.560 align:middle line:90%


00:39:39.560 --> 00:39:42.360 align:middle line:84%
So that's what the range
is going to be.

00:39:42.360 --> 00:39:48.840 align:middle line:84%
Now dealing with the random
variable 2X, as opposed to the

00:39:48.840 --> 00:39:52.520 align:middle line:84%
random variable X, in some sense
it's just changing the

00:39:52.520 --> 00:39:55.270 align:middle line:84%
units in which we measure
that random variable.

00:39:55.270 --> 00:39:58.130 align:middle line:84%
It's just changing the
scale on which we

00:39:58.130 --> 00:39:59.730 align:middle line:90%
draw and plot things.

00:39:59.730 --> 00:40:03.180 align:middle line:84%
So if it's just a scale change,
then intuition should

00:40:03.180 --> 00:40:08.120 align:middle line:84%
tell you that the random
variable X should have a PDF

00:40:08.120 --> 00:40:12.850 align:middle line:84%
of the same shape, except that
it's scaled out by a factor of

00:40:12.850 --> 00:40:16.540 align:middle line:84%
2, because our random variable
of 2X now has a range that's

00:40:16.540 --> 00:40:18.570 align:middle line:90%
twice as large.

00:40:18.570 --> 00:40:23.720 align:middle line:84%
So we take the same PDF and
scale it up by stretching the

00:40:23.720 --> 00:40:26.790 align:middle line:90%
x-axis by a factor of 2.

00:40:26.790 --> 00:40:30.330 align:middle line:84%
So what does scaling
correspond to

00:40:30.330 --> 00:40:33.870 align:middle line:90%
in terms of a formula?

00:40:33.870 --> 00:40:39.500 align:middle line:84%
So the distribution of 2X as a
function, let's say, a generic

00:40:39.500 --> 00:40:45.760 align:middle line:84%
argument z, is going to be the
distribution of X, but scaled

00:40:45.760 --> 00:40:47.010 align:middle line:90%
by a factor of 2.

00:40:47.010 --> 00:40:50.060 align:middle line:90%


00:40:50.060 --> 00:40:54.100 align:middle line:84%
So taking a function and
replacing its arguments by the

00:40:54.100 --> 00:40:58.740 align:middle line:84%
argument over 2, what it
does is it stretches it

00:40:58.740 --> 00:41:00.430 align:middle line:90%
by a factor of 2.

00:41:00.430 --> 00:41:04.410 align:middle line:84%
You have probably been tortured
ever since middle

00:41:04.410 --> 00:41:08.150 align:middle line:84%
school to figure out when need
to stretch a function, whether

00:41:08.150 --> 00:41:12.470 align:middle line:90%
you need to put 2z or z/2.

00:41:12.470 --> 00:41:15.450 align:middle line:84%
And the one that actually does
the stretching is to put the

00:41:15.450 --> 00:41:18.000 align:middle line:90%
z/2 in that place.

00:41:18.000 --> 00:41:21.180 align:middle line:84%
So that's what the
stretching does.

00:41:21.180 --> 00:41:23.670 align:middle line:84%
Could that to be the
full answer?

00:41:23.670 --> 00:41:24.930 align:middle line:90%
Well there's a catch.

00:41:24.930 --> 00:41:29.730 align:middle line:84%
If you stretch this function by
a factor of 2, what happens

00:41:29.730 --> 00:41:32.100 align:middle line:84%
to the area under
the function?

00:41:32.100 --> 00:41:34.120 align:middle line:90%
It's going to get doubled.

00:41:34.120 --> 00:41:38.670 align:middle line:84%
But the total probability must
add up to 1, so we need to do

00:41:38.670 --> 00:41:41.840 align:middle line:84%
something else to make sure that
the area under the curve

00:41:41.840 --> 00:41:44.300 align:middle line:90%
stays to 1.

00:41:44.300 --> 00:41:47.980 align:middle line:84%
So we need to take that function
and scale it down by

00:41:47.980 --> 00:41:51.720 align:middle line:90%
this factor of 2.

00:41:51.720 --> 00:41:55.580 align:middle line:84%
So when you're dealing with a
multiple of a random variable,

00:41:55.580 --> 00:42:00.580 align:middle line:84%
what happens to the PDF is you
stretch it according to the

00:42:00.580 --> 00:42:04.320 align:middle line:84%
multiple, and then scale it
down by the same number so

00:42:04.320 --> 00:42:07.460 align:middle line:84%
that you preserve the area
under that curve.

00:42:07.460 --> 00:42:10.800 align:middle line:84%
So now we found the distribution
of 2X.

00:42:10.800 --> 00:42:14.910 align:middle line:84%
How about the distribution
of 2X + 5?

00:42:14.910 --> 00:42:18.560 align:middle line:84%
Well what does adding 5
to random variable do?

00:42:18.560 --> 00:42:20.940 align:middle line:84%
You're going to get essentially
the same values

00:42:20.940 --> 00:42:23.720 align:middle line:84%
with the same probability,
except that those values all

00:42:23.720 --> 00:42:26.260 align:middle line:90%
get shifted by 5.

00:42:26.260 --> 00:42:30.650 align:middle line:84%
So all that you need to do is
to take this PDF here, and

00:42:30.650 --> 00:42:32.690 align:middle line:90%
shift it by 5 units.

00:42:32.690 --> 00:42:35.530 align:middle line:84%
So the range used to
be from -2 to 4.

00:42:35.530 --> 00:42:38.750 align:middle line:84%
The new range is going
to be from 3 to 9.

00:42:38.750 --> 00:42:40.390 align:middle line:90%
And that's the final answer.

00:42:40.390 --> 00:42:44.900 align:middle line:84%
This is the distribution of
2X + 5, starting with this

00:42:44.900 --> 00:42:48.240 align:middle line:90%
particular distribution of X.

00:42:48.240 --> 00:42:53.600 align:middle line:84%
Now shifting to the
right by b, what

00:42:53.600 --> 00:42:55.700 align:middle line:90%
does it do to a function?

00:42:55.700 --> 00:42:58.620 align:middle line:84%
Shifting to the right to
by a certain amount,

00:42:58.620 --> 00:43:04.960 align:middle line:84%
mathematically, it corresponds
to putting -b in the argument

00:43:04.960 --> 00:43:06.000 align:middle line:90%
of the function.

00:43:06.000 --> 00:43:09.750 align:middle line:84%
So I'm taking the formula that
I had here, which is the

00:43:09.750 --> 00:43:12.220 align:middle line:90%
scaling by a factor of a.

00:43:12.220 --> 00:43:17.200 align:middle line:84%
The scaling down to keep the
total area equal to 1.

00:43:17.200 --> 00:43:19.740 align:middle line:84%
And then I need to introduce
this extra

00:43:19.740 --> 00:43:20.990 align:middle line:90%
term to do the shifting.

00:43:20.990 --> 00:43:23.300 align:middle line:90%


00:43:23.300 --> 00:43:26.200 align:middle line:84%
So this is a plausible
argument.

00:43:26.200 --> 00:43:31.080 align:middle line:84%
The proof by picture that this
should be the right answer.

00:43:31.080 --> 00:43:38.295 align:middle line:84%
But just in order to keep our
skills tuned and refined, let

00:43:38.295 --> 00:43:42.950 align:middle line:84%
us do this derivation in a
more formal way using our

00:43:42.950 --> 00:43:45.135 align:middle line:90%
two-step cookbook procedure.

00:43:45.135 --> 00:43:48.000 align:middle line:90%


00:43:48.000 --> 00:43:51.010 align:middle line:84%
And I'm going to do it under
the assumption that a is

00:43:51.010 --> 00:43:54.910 align:middle line:84%
positive, as in the example
that's we just did.

00:43:54.910 --> 00:43:59.090 align:middle line:84%
So what's the two-step
procedure?

00:43:59.090 --> 00:44:03.700 align:middle line:84%
We want to find the cumulative
of Y, and after that we're

00:44:03.700 --> 00:44:05.720 align:middle line:90%
going to differentiate.

00:44:05.720 --> 00:44:09.220 align:middle line:84%
By definition the cumulative
is the probability that the

00:44:09.220 --> 00:44:13.280 align:middle line:84%
random variable takes values
less than a certain number.

00:44:13.280 --> 00:44:17.190 align:middle line:84%
And now we need to take this
event and translate it, and

00:44:17.190 --> 00:44:21.110 align:middle line:84%
express it in terms of the
original random variables.

00:44:21.110 --> 00:44:24.970 align:middle line:84%
So Y is, by definition,
aX + b, so we're

00:44:24.970 --> 00:44:28.970 align:middle line:90%
looking at this event.

00:44:28.970 --> 00:44:33.580 align:middle line:84%
And now we want to express this
event in a clean form

00:44:33.580 --> 00:44:39.730 align:middle line:84%
where X shows up in
a straight way.

00:44:39.730 --> 00:44:42.740 align:middle line:84%
Let's say I'm going to massage
this event and

00:44:42.740 --> 00:44:44.640 align:middle line:90%
write it in this form.

00:44:44.640 --> 00:44:48.070 align:middle line:84%
For this inequality to be true,
x should be less than or

00:44:48.070 --> 00:44:53.820 align:middle line:84%
equal to (y minus
b) divided by a.

00:44:53.820 --> 00:44:56.820 align:middle line:90%
OK, now what is this?

00:44:56.820 --> 00:45:01.330 align:middle line:84%
This is the cumulative
distribution of X evaluated at

00:45:01.330 --> 00:45:02.580 align:middle line:90%
the particular point.

00:45:02.580 --> 00:45:07.850 align:middle line:90%


00:45:07.850 --> 00:45:14.760 align:middle line:84%
So we got a formula for the
cumulative Y based on the

00:45:14.760 --> 00:45:17.880 align:middle line:84%
cumulative of X. What's
the next step?

00:45:17.880 --> 00:45:21.550 align:middle line:84%
Next step is to take derivatives
of both sides.

00:45:21.550 --> 00:45:28.810 align:middle line:84%
So the density of Y is going to
be the derivative of this

00:45:28.810 --> 00:45:31.270 align:middle line:90%
expression with respect to y.

00:45:31.270 --> 00:45:36.830 align:middle line:84%
OK, so now here we need
to use the chain rule.

00:45:36.830 --> 00:45:40.670 align:middle line:84%
It's going to be the derivative
of the F function

00:45:40.670 --> 00:45:43.080 align:middle line:90%
with respect to its argument.

00:45:43.080 --> 00:45:46.930 align:middle line:84%
And then we need to take the
derivative of the argument

00:45:46.930 --> 00:45:48.780 align:middle line:90%
with respect to y.

00:45:48.780 --> 00:45:51.530 align:middle line:84%
What is the derivative
of the cumulative?

00:45:51.530 --> 00:45:53.190 align:middle line:84%
The derivative of
the cumulative

00:45:53.190 --> 00:45:56.290 align:middle line:90%
is the density itself.

00:45:56.290 --> 00:45:59.578 align:middle line:84%
And we evaluate it at the
point of interest.

00:45:59.578 --> 00:46:02.180 align:middle line:90%


00:46:02.180 --> 00:46:05.340 align:middle line:84%
And then the chain rule tells
us that we need to take the

00:46:05.340 --> 00:46:08.800 align:middle line:84%
derivative of this with
respect to y, and the

00:46:08.800 --> 00:46:11.370 align:middle line:84%
derivative of this with
respect to y is 1/a.

00:46:11.370 --> 00:46:14.290 align:middle line:90%


00:46:14.290 --> 00:46:18.330 align:middle line:84%
And this gives us the formula
which is consistent with what

00:46:18.330 --> 00:46:21.810 align:middle line:84%
I had written down here,
for the case where a

00:46:21.810 --> 00:46:25.030 align:middle line:90%
is a positive number.

00:46:25.030 --> 00:46:27.915 align:middle line:84%
What if a was a negative
number?

00:46:27.915 --> 00:46:30.570 align:middle line:90%


00:46:30.570 --> 00:46:31.910 align:middle line:90%
Could this formula be true?

00:46:31.910 --> 00:46:35.120 align:middle line:90%


00:46:35.120 --> 00:46:36.140 align:middle line:90%
Of course not.

00:46:36.140 --> 00:46:39.000 align:middle line:84%
Densities cannot be
negative, right?

00:46:39.000 --> 00:46:41.180 align:middle line:84%
So that formula cannot
be true.

00:46:41.180 --> 00:46:43.750 align:middle line:90%
Something needs to change.

00:46:43.750 --> 00:46:45.140 align:middle line:90%
What should change?

00:46:45.140 --> 00:46:50.970 align:middle line:84%
Where does this argument break
down when a is negative?

00:46:50.970 --> 00:46:56.470 align:middle line:90%


00:46:56.470 --> 00:47:01.570 align:middle line:84%
So when I write this inequality
in this form, I

00:47:01.570 --> 00:47:03.940 align:middle line:90%
divide by a.

00:47:03.940 --> 00:47:07.730 align:middle line:84%
But when you divide by a
negative number, the direction

00:47:07.730 --> 00:47:10.390 align:middle line:84%
of an inequality is
going to change.

00:47:10.390 --> 00:47:14.520 align:middle line:84%
So when a is negative, this
inequality becomes larger than

00:47:14.520 --> 00:47:16.190 align:middle line:90%
or equal to.

00:47:16.190 --> 00:47:18.770 align:middle line:84%
And in that case, the expression
that I have up

00:47:18.770 --> 00:47:24.360 align:middle line:84%
there would change when this
is larger than here.

00:47:24.360 --> 00:47:27.900 align:middle line:84%
Instead of getting the
cumulative, I would get 1

00:47:27.900 --> 00:47:32.350 align:middle line:84%
minus the cumulative of (y
minus b) divided by a.

00:47:32.350 --> 00:47:35.240 align:middle line:90%


00:47:35.240 --> 00:47:39.890 align:middle line:84%
So this is the probability that
X is bigger than this

00:47:39.890 --> 00:47:41.170 align:middle line:90%
particular number.

00:47:41.170 --> 00:47:44.000 align:middle line:84%
And now when you take the
derivatives, there's going to

00:47:44.000 --> 00:47:46.570 align:middle line:90%
be a minus sign that shows up.

00:47:46.570 --> 00:47:49.810 align:middle line:84%
And that minus sign will
end up being here.

00:47:49.810 --> 00:47:53.730 align:middle line:84%
And so we're taking the negative
of a negative number,

00:47:53.730 --> 00:47:56.420 align:middle line:84%
and that basically is equivalent
to taking the

00:47:56.420 --> 00:47:58.660 align:middle line:90%
absolute value of that number.

00:47:58.660 --> 00:48:03.830 align:middle line:84%
So all that happens when we have
a negative a is that we

00:48:03.830 --> 00:48:07.010 align:middle line:84%
have to take the absolute value
of the scaling factor

00:48:07.010 --> 00:48:10.250 align:middle line:90%
instead of the factor itself.

00:48:10.250 --> 00:48:14.020 align:middle line:84%
All right, so this general
formula is quite useful for

00:48:14.020 --> 00:48:16.690 align:middle line:84%
dealing with linear functions
of random variables.

00:48:16.690 --> 00:48:21.330 align:middle line:84%
And one nice application of it
is to take the formula for a

00:48:21.330 --> 00:48:25.460 align:middle line:84%
normal random variable, consider
a linear function of

00:48:25.460 --> 00:48:29.600 align:middle line:84%
a normal random variable, plug
into this formula, and what

00:48:29.600 --> 00:48:34.000 align:middle line:84%
you will find is that Y also
has a normal distribution.

00:48:34.000 --> 00:48:37.310 align:middle line:84%
So using this formula, now we
can prove a statement that I

00:48:37.310 --> 00:48:40.565 align:middle line:84%
had made a couple of lectures
ago, that a linear function of

00:48:40.565 --> 00:48:43.900 align:middle line:84%
a normal random variable
is also linear.

00:48:43.900 --> 00:48:47.600 align:middle line:90%
That's how you would prove it.

00:48:47.600 --> 00:48:51.190 align:middle line:84%
I think this is it
for today so.

00:48:51.190 --> 00:48:52.440 align:middle line:90%