WEBVTT

00:00:00.000 --> 00:00:00.040 align:middle line:90%


00:00:00.040 --> 00:00:02.460 align:middle line:84%
The following content is
provided under a Creative

00:00:02.460 --> 00:00:03.870 align:middle line:90%
Commons license.

00:00:03.870 --> 00:00:06.910 align:middle line:84%
Your support will help MIT
OpenCourseWare continue to

00:00:06.910 --> 00:00:10.560 align:middle line:84%
offer high-quality educational
resources for free.

00:00:10.560 --> 00:00:13.460 align:middle line:84%
To make a donation or view
additional materials from

00:00:13.460 --> 00:00:19.290 align:middle line:84%
hundreds of MIT courses, visit
MIT OpenCourseWare at

00:00:19.290 --> 00:00:20.540 align:middle line:90%
ocw.mit.edu.

00:00:20.540 --> 00:00:23.050 align:middle line:90%


00:00:23.050 --> 00:00:25.080 align:middle line:84%
JOHN TSITSIKLIS:
OK let's start.

00:00:25.080 --> 00:00:26.560 align:middle line:90%
So we've had the quiz.

00:00:26.560 --> 00:00:29.760 align:middle line:84%
And I guess there's both good
and bad news in it.

00:00:29.760 --> 00:00:31.590 align:middle line:84%
Yesterday, as you know,
the bad news.

00:00:31.590 --> 00:00:33.910 align:middle line:84%
The average was a little
lower than what

00:00:33.910 --> 00:00:36.260 align:middle line:90%
we would have wanted.

00:00:36.260 --> 00:00:39.580 align:middle line:84%
On the other hand, the good news
is that the distribution

00:00:39.580 --> 00:00:41.770 align:middle line:90%
was nicely spread.

00:00:41.770 --> 00:00:44.890 align:middle line:84%
And that's the main purpose of
this quiz is basically for you

00:00:44.890 --> 00:00:48.260 align:middle line:84%
to calibrate and see roughly
where you are standing.

00:00:48.260 --> 00:00:50.650 align:middle line:84%
The other piece of the good
news is that, as you know,

00:00:50.650 --> 00:00:53.590 align:middle line:84%
this quiz doesn't count for very
much in your final grade.

00:00:53.590 --> 00:00:58.230 align:middle line:84%
So it's really a matter of
calibration and to get your

00:00:58.230 --> 00:01:02.810 align:middle line:84%
mind set appropriately to
prepare for the second quiz,

00:01:02.810 --> 00:01:04.470 align:middle line:90%
which counts a lot more.

00:01:04.470 --> 00:01:06.370 align:middle line:90%
And it's more substantial.

00:01:06.370 --> 00:01:08.810 align:middle line:84%
And we'll make sure that
the second quiz will

00:01:08.810 --> 00:01:12.110 align:middle line:90%
have a higher average.

00:01:12.110 --> 00:01:12.520 align:middle line:90%
All right.

00:01:12.520 --> 00:01:15.410 align:middle line:90%
So let's go to our material.

00:01:15.410 --> 00:01:18.190 align:middle line:84%
We're talking now
these days about

00:01:18.190 --> 00:01:20.440 align:middle line:90%
continuous random variables.

00:01:20.440 --> 00:01:23.240 align:middle line:84%
And I'll remind you what
we discussed last time.

00:01:23.240 --> 00:01:25.970 align:middle line:84%
I'll remind you of the concept
of the probability density

00:01:25.970 --> 00:01:28.230 align:middle line:84%
function of a single
random variable.

00:01:28.230 --> 00:01:31.090 align:middle line:84%
And then we're going to rush
through all the concepts that

00:01:31.090 --> 00:01:34.230 align:middle line:84%
we covered for the case of
discrete random variables and

00:01:34.230 --> 00:01:37.770 align:middle line:84%
discuss their analogs for
the continuous case.

00:01:37.770 --> 00:01:40.410 align:middle line:84%
And talk about notions
such as conditioning

00:01:40.410 --> 00:01:42.170 align:middle line:90%
independence and so on.

00:01:42.170 --> 00:01:46.420 align:middle line:90%
So the big picture is here.

00:01:46.420 --> 00:01:49.590 align:middle line:84%
We have all those concepts that
we developed for the case

00:01:49.590 --> 00:01:52.350 align:middle line:90%
of discrete random variables.

00:01:52.350 --> 00:01:55.560 align:middle line:84%
And now we will just talk about
their analogs in the

00:01:55.560 --> 00:01:56.840 align:middle line:90%
continuous case.

00:01:56.840 --> 00:02:00.800 align:middle line:84%
We already discussed this analog
last week, the density

00:02:00.800 --> 00:02:04.520 align:middle line:90%
of a single random variable.

00:02:04.520 --> 00:02:08.570 align:middle line:84%
Then there are certain concepts
that show up both in

00:02:08.570 --> 00:02:10.780 align:middle line:84%
the discrete and the
continuous case.

00:02:10.780 --> 00:02:14.560 align:middle line:84%
So we have the cumulative
distribution function, which

00:02:14.560 --> 00:02:18.070 align:middle line:84%
is a description of the
probability distribution of a

00:02:18.070 --> 00:02:21.470 align:middle line:84%
random variable and which
applies whether you have a

00:02:21.470 --> 00:02:23.780 align:middle line:84%
discrete or continuous
random variable.

00:02:23.780 --> 00:02:26.500 align:middle line:84%
Then there's the notion
of the expected value.

00:02:26.500 --> 00:02:29.990 align:middle line:84%
And in the two cases, the
expected value is calculated

00:02:29.990 --> 00:02:32.990 align:middle line:84%
in a slightly different way,
but not very different.

00:02:32.990 --> 00:02:36.080 align:middle line:84%
We have sums in one case,
integrals in the other.

00:02:36.080 --> 00:02:37.720 align:middle line:84%
And this is the general
pattern that

00:02:37.720 --> 00:02:39.030 align:middle line:90%
we're going to have.

00:02:39.030 --> 00:02:42.120 align:middle line:84%
Formulas for the discrete case
translate to corresponding

00:02:42.120 --> 00:02:44.920 align:middle line:84%
formulas or expressions in
the continuous case.

00:02:44.920 --> 00:02:50.010 align:middle line:84%
We generically replace sums by
integrals, and we replace must

00:02:50.010 --> 00:02:54.230 align:middle line:84%
functions with density
functions.

00:02:54.230 --> 00:02:58.330 align:middle line:84%
Then the new pieces for today
are going to be mostly the

00:02:58.330 --> 00:03:01.570 align:middle line:84%
notion of a joint density
function, which is how we

00:03:01.570 --> 00:03:04.330 align:middle line:84%
describe the probability
distribution of two random

00:03:04.330 --> 00:03:08.370 align:middle line:84%
variables that are somehow
related, in general, and then

00:03:08.370 --> 00:03:11.780 align:middle line:84%
the notion of a conditional
density function that tells us

00:03:11.780 --> 00:03:15.160 align:middle line:84%
the distribution of one random
variable X when you're told

00:03:15.160 --> 00:03:19.200 align:middle line:84%
the value of another random
variable Y. There's another

00:03:19.200 --> 00:03:22.680 align:middle line:84%
concept, which is the
conditional PDF given that the

00:03:22.680 --> 00:03:24.860 align:middle line:90%
certain event has happened.

00:03:24.860 --> 00:03:27.420 align:middle line:84%
This is a concept that's
in some ways simpler.

00:03:27.420 --> 00:03:31.360 align:middle line:84%
You've already seen a little
bit of that in last week's

00:03:31.360 --> 00:03:33.140 align:middle line:90%
recitation and tutorial.

00:03:33.140 --> 00:03:35.640 align:middle line:84%
The idea is that we have a
single random variable.

00:03:35.640 --> 00:03:37.710 align:middle line:90%
It's described by a density.

00:03:37.710 --> 00:03:41.110 align:middle line:84%
Then you're told that the
certain event has occurred.

00:03:41.110 --> 00:03:42.880 align:middle line:84%
Your model changes
the universe that

00:03:42.880 --> 00:03:43.910 align:middle line:90%
you are dealing with.

00:03:43.910 --> 00:03:46.640 align:middle line:84%
In the new universe, you are
dealing with a new density

00:03:46.640 --> 00:03:51.310 align:middle line:84%
function, the one that applies
given the knowledge that we

00:03:51.310 --> 00:03:55.700 align:middle line:84%
have that the certain
event has occurred.

00:03:55.700 --> 00:03:56.160 align:middle line:90%
All right.

00:03:56.160 --> 00:03:59.870 align:middle line:84%
So what exactly did
we say about

00:03:59.870 --> 00:04:02.140 align:middle line:90%
continuous random variables?

00:04:02.140 --> 00:04:05.020 align:middle line:84%
The first thing is the
definition, that a random

00:04:05.020 --> 00:04:09.370 align:middle line:84%
variable is said to be
continuous if we are given a

00:04:09.370 --> 00:04:12.220 align:middle line:84%
certain object that we call
the probability density

00:04:12.220 --> 00:04:17.050 align:middle line:84%
function and we can calculate
interval probabilities given

00:04:17.050 --> 00:04:18.709 align:middle line:90%
this density function.

00:04:18.709 --> 00:04:21.589 align:middle line:84%
So the definition is that the
random variable is continuous

00:04:21.589 --> 00:04:24.490 align:middle line:84%
if you can calculate
probabilities associated with

00:04:24.490 --> 00:04:27.380 align:middle line:84%
that random variable
given that formula.

00:04:27.380 --> 00:04:29.770 align:middle line:84%
So this formula tells you that
the probability that your

00:04:29.770 --> 00:04:33.340 align:middle line:84%
random variable falls inside
this interval is the area

00:04:33.340 --> 00:04:34.880 align:middle line:90%
under the density curve.

00:04:34.880 --> 00:04:37.390 align:middle line:90%


00:04:37.390 --> 00:04:37.700 align:middle line:90%
OK.

00:04:37.700 --> 00:04:39.720 align:middle line:84%
There's a few properties
that a density

00:04:39.720 --> 00:04:41.020 align:middle line:90%
function must satisfy.

00:04:41.020 --> 00:04:42.900 align:middle line:84%
Since we're talking about
probabilities, and

00:04:42.900 --> 00:04:45.890 align:middle line:84%
probabilities are non-negative,
we have that the

00:04:45.890 --> 00:04:49.530 align:middle line:84%
density function is always
a non-negative function.

00:04:49.530 --> 00:04:52.790 align:middle line:84%
The total probability over
the entire real line

00:04:52.790 --> 00:04:54.690 align:middle line:90%
must be equal to 1.

00:04:54.690 --> 00:04:58.070 align:middle line:84%
So the integral when you
integrate over the entire real

00:04:58.070 --> 00:04:59.590 align:middle line:90%
line has to be equal to 1.

00:04:59.590 --> 00:05:01.800 align:middle line:90%
That's the second property.

00:05:01.800 --> 00:05:05.200 align:middle line:84%
Another property that you get is
that if you let a equal to

00:05:05.200 --> 00:05:07.720 align:middle line:90%
b, this integral becomes 0.

00:05:07.720 --> 00:05:11.390 align:middle line:84%
And that tells you that the
probability of a single point

00:05:11.390 --> 00:05:15.990 align:middle line:84%
in the continuous case
is always equal to 0.

00:05:15.990 --> 00:05:17.780 align:middle line:84%
So these are formal
properties.

00:05:17.780 --> 00:05:21.290 align:middle line:84%
When you want to think
intuitively, the best way to

00:05:21.290 --> 00:05:25.540 align:middle line:84%
think about what the density
function is to think in terms

00:05:25.540 --> 00:05:28.320 align:middle line:84%
of little intervals, the
probability that my random

00:05:28.320 --> 00:05:31.540 align:middle line:84%
variable falls inside
the little interval.

00:05:31.540 --> 00:05:35.170 align:middle line:84%
Well, inside that little
interval, the density function

00:05:35.170 --> 00:05:36.940 align:middle line:90%
here is roughly constant.

00:05:36.940 --> 00:05:42.430 align:middle line:84%
So that integral becomes the
value of the density times the

00:05:42.430 --> 00:05:45.340 align:middle line:84%
length of the interval over
which you are integrating,

00:05:45.340 --> 00:05:47.070 align:middle line:90%
which is delta.

00:05:47.070 --> 00:05:50.240 align:middle line:84%
And so the density function
basically gives us

00:05:50.240 --> 00:05:54.990 align:middle line:84%
probabilities of little events,
of small events.

00:05:54.990 --> 00:05:59.200 align:middle line:84%
And the density is to be
interpreted as probability per

00:05:59.200 --> 00:06:02.290 align:middle line:84%
unit length at a certain
place in the diagram.

00:06:02.290 --> 00:06:04.800 align:middle line:84%
So in that place in the diagram,
the probability per

00:06:04.800 --> 00:06:07.870 align:middle line:84%
unit length around this
neighborhood would be the

00:06:07.870 --> 00:06:12.320 align:middle line:84%
height of the density function
at that point.

00:06:12.320 --> 00:06:13.270 align:middle line:90%
What else?

00:06:13.270 --> 00:06:16.440 align:middle line:84%
We have a formula for
calculating expected values of

00:06:16.440 --> 00:06:17.980 align:middle line:90%
functions of random variables.

00:06:17.980 --> 00:06:21.310 align:middle line:84%
In the discrete case, we had the
formula where here we had

00:06:21.310 --> 00:06:25.430 align:middle line:84%
the sum, and instead of the
density, we had the PMF.

00:06:25.430 --> 00:06:29.188 align:middle line:84%
The same formula is also valid
in the continuous case.

00:06:29.188 --> 00:06:35.120 align:middle line:84%
And it's not too hard to derive,
but we will not do it.

00:06:35.120 --> 00:06:36.910 align:middle line:84%
But let's think of the
intuition of what

00:06:36.910 --> 00:06:38.420 align:middle line:90%
this formula says.

00:06:38.420 --> 00:06:41.670 align:middle line:84%
You're trying to figure out on
the average how much g(X) is

00:06:41.670 --> 00:06:42.780 align:middle line:90%
going to be.

00:06:42.780 --> 00:06:47.130 align:middle line:84%
And then you reason, and you
say, well, X may turn out to

00:06:47.130 --> 00:06:52.560 align:middle line:84%
take a particular value or a
small interval of values.

00:06:52.560 --> 00:06:54.780 align:middle line:84%
This is the probability
that X falls

00:06:54.780 --> 00:06:56.640 align:middle line:90%
inside the small interval.

00:06:56.640 --> 00:07:00.310 align:middle line:84%
And when that happens, g(X)
takes that value.

00:07:00.310 --> 00:07:03.930 align:middle line:84%
So this fraction of the time,
you fall in the little

00:07:03.930 --> 00:07:07.350 align:middle line:84%
neighborhood of x, and
you get so much.

00:07:07.350 --> 00:07:10.860 align:middle line:84%
Then you average over all the
possible x's that can happen.

00:07:10.860 --> 00:07:13.930 align:middle line:84%
And that gives you the average
value of the function g(X).

00:07:13.930 --> 00:07:17.730 align:middle line:90%


00:07:17.730 --> 00:07:18.045 align:middle line:90%
OK.

00:07:18.045 --> 00:07:20.650 align:middle line:90%
So this is the easy stuff.

00:07:20.650 --> 00:07:23.690 align:middle line:84%
Now let's get to the
new material.

00:07:23.690 --> 00:07:26.330 align:middle line:84%
We want to talk about multiple
random variables

00:07:26.330 --> 00:07:27.320 align:middle line:90%
simultaneously.

00:07:27.320 --> 00:07:31.530 align:middle line:84%
So we want to talk now about two
random variables that are

00:07:31.530 --> 00:07:35.020 align:middle line:84%
continuous, and in some sense
that they are jointly

00:07:35.020 --> 00:07:35.840 align:middle line:90%
continuous.

00:07:35.840 --> 00:07:38.080 align:middle line:90%
And let's see what this means.

00:07:38.080 --> 00:07:40.840 align:middle line:84%
The definition is similar to
the definition we had for a

00:07:40.840 --> 00:07:44.850 align:middle line:84%
single random variable, where
I take this formula here as

00:07:44.850 --> 00:07:49.510 align:middle line:84%
the definition of continuous
random variables.

00:07:49.510 --> 00:07:53.830 align:middle line:84%
Two random variables are said to
be jointly continuous if we

00:07:53.830 --> 00:07:58.190 align:middle line:84%
can calculate probabilities by
integrating a certain function

00:07:58.190 --> 00:08:01.070 align:middle line:84%
that we call the joint
density function

00:08:01.070 --> 00:08:03.310 align:middle line:90%
over the set of interest.

00:08:03.310 --> 00:08:08.690 align:middle line:84%
So we have our two-dimensional
plane.

00:08:08.690 --> 00:08:10.900 align:middle line:90%
This is the x-y plane.

00:08:10.900 --> 00:08:13.810 align:middle line:84%
There's a certain event S that
we're interested in.

00:08:13.810 --> 00:08:15.860 align:middle line:84%
We want to calculate
the probability.

00:08:15.860 --> 00:08:17.370 align:middle line:90%
How do we do that?

00:08:17.370 --> 00:08:22.660 align:middle line:84%
We are given this function
f_(X,Y), the joint density.

00:08:22.660 --> 00:08:25.910 align:middle line:84%
It's a function of the two
arguments x and y.

00:08:25.910 --> 00:08:29.530 align:middle line:84%
So think of that function as
being some kind of surface

00:08:29.530 --> 00:08:34.809 align:middle line:84%
that sits on top of the
two-dimensional plane.

00:08:34.809 --> 00:08:39.140 align:middle line:84%
The probability of falling
inside the set S, we calculate

00:08:39.140 --> 00:08:45.350 align:middle line:84%
it by looking at the volume
under the surface, that volume

00:08:45.350 --> 00:08:50.470 align:middle line:84%
that sits on top of S. So the
surface underneath it has a

00:08:50.470 --> 00:08:52.010 align:middle line:90%
certain total volume.

00:08:52.010 --> 00:08:54.650 align:middle line:84%
What should that total
volume be?

00:08:54.650 --> 00:08:57.050 align:middle line:84%
Well, we think of these volumes
as probabilities.

00:08:57.050 --> 00:09:00.180 align:middle line:84%
So the total probability
should be equal to 1.

00:09:00.180 --> 00:09:05.430 align:middle line:84%
The total volume under this
surface, should be equal to 1.

00:09:05.430 --> 00:09:08.220 align:middle line:84%
So that's one property
that we want our

00:09:08.220 --> 00:09:10.138 align:middle line:90%
density function to have.

00:09:10.138 --> 00:09:16.080 align:middle line:90%


00:09:16.080 --> 00:09:20.500 align:middle line:84%
So when you integrate over the
entire space, this is of the

00:09:20.500 --> 00:09:22.400 align:middle line:90%
volume under your surface.

00:09:22.400 --> 00:09:24.090 align:middle line:90%
That should be equal to 1.

00:09:24.090 --> 00:09:27.280 align:middle line:84%
Of course, since we're talking
about probabilities, the joint

00:09:27.280 --> 00:09:29.560 align:middle line:84%
density should be a non-negative
function.

00:09:29.560 --> 00:09:34.140 align:middle line:84%
So think of the situation
as having one pound of

00:09:34.140 --> 00:09:38.230 align:middle line:84%
probability that's spread
all over your space.

00:09:38.230 --> 00:09:41.430 align:middle line:84%
And the height of this joint
density function basically

00:09:41.430 --> 00:09:45.470 align:middle line:84%
tells you how much probability
tends to be accumulated in

00:09:45.470 --> 00:09:48.400 align:middle line:84%
certain regions of space
as opposed to other

00:09:48.400 --> 00:09:49.870 align:middle line:90%
parts of the space.

00:09:49.870 --> 00:09:53.130 align:middle line:84%
So wherever the density is big,
that means that this is

00:09:53.130 --> 00:09:54.920 align:middle line:84%
an area of the two-dimensional
plane that's

00:09:54.920 --> 00:09:56.340 align:middle line:90%
more likely to occur.

00:09:56.340 --> 00:09:59.160 align:middle line:84%
Where the density is small, that
means that those x-y's

00:09:59.160 --> 00:10:01.100 align:middle line:90%
are less likely to occur.

00:10:01.100 --> 00:10:03.070 align:middle line:84%
You have already seen
one example

00:10:03.070 --> 00:10:06.050 align:middle line:90%
of continuous densities.

00:10:06.050 --> 00:10:08.730 align:middle line:84%
That was the example we had in
the very beginning of the

00:10:08.730 --> 00:10:10.700 align:middle line:90%
class with a uniform

00:10:10.700 --> 00:10:13.380 align:middle line:90%
distribution on the unit square.

00:10:13.380 --> 00:10:15.510 align:middle line:84%
That was a special
case of a density

00:10:15.510 --> 00:10:17.250 align:middle line:90%
function that was constant.

00:10:17.250 --> 00:10:20.090 align:middle line:84%
So all places in the unit square
were roughly equally

00:10:20.090 --> 00:10:22.010 align:middle line:90%
likely as any other places.

00:10:22.010 --> 00:10:25.580 align:middle line:84%
But in other models, some parts
of the space may be more

00:10:25.580 --> 00:10:27.000 align:middle line:90%
likely than others.

00:10:27.000 --> 00:10:29.470 align:middle line:84%
And we describe those relative
likelihoods using

00:10:29.470 --> 00:10:31.120 align:middle line:90%
this density function.

00:10:31.120 --> 00:10:33.420 align:middle line:84%
So if somebody gives us the
density function, this

00:10:33.420 --> 00:10:38.480 align:middle line:84%
determines for us probabilities
of all the

00:10:38.480 --> 00:10:41.520 align:middle line:84%
subsets of the two-dimensional
plane.

00:10:41.520 --> 00:10:45.710 align:middle line:84%
Now for an intuitive
interpretation, it's good to

00:10:45.710 --> 00:10:47.460 align:middle line:90%
think about small events.

00:10:47.460 --> 00:10:51.220 align:middle line:84%
So let's take a particular x
here and then x plus delta.

00:10:51.220 --> 00:10:53.020 align:middle line:90%
So this is a small interval.

00:10:53.020 --> 00:10:56.190 align:middle line:84%
Take another small interval
here that goes from y to y

00:10:56.190 --> 00:10:57.560 align:middle line:90%
plus delta.

00:10:57.560 --> 00:11:03.270 align:middle line:84%
And let's look at the event that
x falls here and y falls

00:11:03.270 --> 00:11:04.780 align:middle line:90%
right there.

00:11:04.780 --> 00:11:05.780 align:middle line:90%
What is this event?

00:11:05.780 --> 00:11:07.760 align:middle line:84%
Well, this is the event
that will fall

00:11:07.760 --> 00:11:11.030 align:middle line:90%
inside this little rectangle.

00:11:11.030 --> 00:11:15.820 align:middle line:84%
Using this rule for calculating
probabilities,

00:11:15.820 --> 00:11:19.040 align:middle line:84%
what is the probability of that
rectangle going to be?

00:11:19.040 --> 00:11:23.130 align:middle line:84%
Well, it should be the integral
of the density over

00:11:23.130 --> 00:11:24.300 align:middle line:90%
this rectangle.

00:11:24.300 --> 00:11:29.720 align:middle line:84%
Or it's the volume under the
surface that sits on top of

00:11:29.720 --> 00:11:31.010 align:middle line:90%
that rectangle.

00:11:31.010 --> 00:11:34.300 align:middle line:84%
Now, if the rectangle is very
small, the joint density is

00:11:34.300 --> 00:11:36.760 align:middle line:84%
not going to change very much
in that neighborhood.

00:11:36.760 --> 00:11:38.770 align:middle line:84%
So we can treat it
as a constant.

00:11:38.770 --> 00:11:42.350 align:middle line:84%
So the volume is going to
be the height times

00:11:42.350 --> 00:11:44.030 align:middle line:90%
the area of the base.

00:11:44.030 --> 00:11:47.150 align:middle line:84%
The height at that point is
whatever the function happens

00:11:47.150 --> 00:11:49.460 align:middle line:90%
to be around that point.

00:11:49.460 --> 00:11:52.590 align:middle line:84%
And the area of the base
is delta squared.

00:11:52.590 --> 00:11:58.750 align:middle line:84%
So this is the intuitive way
to understand what a joint

00:11:58.750 --> 00:12:01.070 align:middle line:84%
density function really
tells you.

00:12:01.070 --> 00:12:04.200 align:middle line:84%
It specifies for you
probabilities of little

00:12:04.200 --> 00:12:08.500 align:middle line:90%
squares, of little rectangles.

00:12:08.500 --> 00:12:11.880 align:middle line:84%
And it allows you to think of
the joint density function as

00:12:11.880 --> 00:12:15.310 align:middle line:90%
probability per unit area.

00:12:15.310 --> 00:12:18.790 align:middle line:84%
So these are the units of the
density, its probability per

00:12:18.790 --> 00:12:23.800 align:middle line:84%
unit area in the neighborhood
of a certain point.

00:12:23.800 --> 00:12:26.970 align:middle line:84%
So what do we do with this
density function once we have

00:12:26.970 --> 00:12:28.410 align:middle line:90%
it in our hands?

00:12:28.410 --> 00:12:32.640 align:middle line:84%
Well, we can use it to calculate
expected values.

00:12:32.640 --> 00:12:34.880 align:middle line:84%
Suppose that you have a
function of two random

00:12:34.880 --> 00:12:38.040 align:middle line:84%
variables described by
a joint density.

00:12:38.040 --> 00:12:41.580 align:middle line:84%
You can find, perhaps, the
distribution of this random

00:12:41.580 --> 00:12:45.330 align:middle line:84%
variable and then use the
basic definition of the

00:12:45.330 --> 00:12:46.150 align:middle line:90%
expectation.

00:12:46.150 --> 00:12:49.260 align:middle line:84%
Or you can calculate
expectations directly, using

00:12:49.260 --> 00:12:52.010 align:middle line:84%
the distribution of the original
random variables.

00:12:52.010 --> 00:12:55.280 align:middle line:84%
This is a formula that's again
identical to the formula that

00:12:55.280 --> 00:12:57.290 align:middle line:90%
we had for the discrete case.

00:12:57.290 --> 00:12:59.500 align:middle line:84%
In the discrete case,
we had a double sum

00:12:59.500 --> 00:13:02.590 align:middle line:90%
here, and we had PMFs.

00:13:02.590 --> 00:13:06.290 align:middle line:84%
So the intuition behind this
formula is the same that one

00:13:06.290 --> 00:13:08.220 align:middle line:90%
had for the discrete case.

00:13:08.220 --> 00:13:12.550 align:middle line:84%
It's just that the mechanics
are different.

00:13:12.550 --> 00:13:16.220 align:middle line:84%
Then something that we did in
the discrete case was to find

00:13:16.220 --> 00:13:21.510 align:middle line:84%
a way to go from the joint
density of the two random

00:13:21.510 --> 00:13:25.750 align:middle line:84%
variables taken together to the
density of just one of the

00:13:25.750 --> 00:13:28.190 align:middle line:90%
random variables.

00:13:28.190 --> 00:13:30.570 align:middle line:84%
So we had a formula for
the discrete case.

00:13:30.570 --> 00:13:33.450 align:middle line:84%
Let's see how things are
going to work out in

00:13:33.450 --> 00:13:35.800 align:middle line:90%
the continuous case.

00:13:35.800 --> 00:13:40.560 align:middle line:84%
So in the continuous
case, we have here

00:13:40.560 --> 00:13:42.330 align:middle line:90%
our two random variables.

00:13:42.330 --> 00:13:45.030 align:middle line:84%
And we have a density
for them.

00:13:45.030 --> 00:13:48.340 align:middle line:84%
And let's say that we want to
calculate the probability that

00:13:48.340 --> 00:13:51.570 align:middle line:90%
x falls inside this interval.

00:13:51.570 --> 00:13:53.510 align:middle line:84%
So we're looking at the
probability that our random

00:13:53.510 --> 00:13:58.380 align:middle line:84%
variable X falls in the interval
from little x to x

00:13:58.380 --> 00:13:59.630 align:middle line:90%
plus delta.

00:13:59.630 --> 00:14:02.130 align:middle line:90%


00:14:02.130 --> 00:14:08.260 align:middle line:84%
Now, by the properties that we
already have for interpreting

00:14:08.260 --> 00:14:11.460 align:middle line:84%
the density function of a single
random variable, the

00:14:11.460 --> 00:14:14.100 align:middle line:84%
probability of a little interval
is approximately the

00:14:14.100 --> 00:14:18.750 align:middle line:84%
density of that single random
variable times delta.

00:14:18.750 --> 00:14:22.120 align:middle line:84%
And now we want to find a
formula for this marginal

00:14:22.120 --> 00:14:26.540 align:middle line:84%
density in terms of
the joint density.

00:14:26.540 --> 00:14:26.890 align:middle line:90%
OK.

00:14:26.890 --> 00:14:28.930 align:middle line:84%
So this is the probability
that x

00:14:28.930 --> 00:14:30.970 align:middle line:90%
falls inside this interval.

00:14:30.970 --> 00:14:34.070 align:middle line:84%
In terms of the two-dimensional
plane, this is

00:14:34.070 --> 00:14:40.030 align:middle line:84%
the probability that (x,y)
falls inside this strip.

00:14:40.030 --> 00:14:44.520 align:middle line:84%
So to find that probability,
we need to calculate the

00:14:44.520 --> 00:14:48.530 align:middle line:84%
probability that (x,y) falls in
here, which is going to be

00:14:48.530 --> 00:14:55.780 align:middle line:84%
the double integral over the
interval over this strip, of

00:14:55.780 --> 00:14:57.030 align:middle line:90%
the joint density.

00:14:57.030 --> 00:15:05.080 align:middle line:90%


00:15:05.080 --> 00:15:07.920 align:middle line:84%
And what are we integrating
over?

00:15:07.920 --> 00:15:11.185 align:middle line:84%
y goes from minus infinity
to plus infinity.

00:15:11.185 --> 00:15:15.680 align:middle line:90%


00:15:15.680 --> 00:15:22.755 align:middle line:84%
And the dummy variable x goes
from little x to x plus delta.

00:15:22.755 --> 00:15:27.240 align:middle line:90%


00:15:27.240 --> 00:15:31.580 align:middle line:84%
So to integrate over this strip,
what we do is for any

00:15:31.580 --> 00:15:34.810 align:middle line:84%
given y, we integrate
in this dimension.

00:15:34.810 --> 00:15:36.770 align:middle line:90%
This is the x integral.

00:15:36.770 --> 00:15:40.220 align:middle line:84%
And then we integrate over
the y dimension.

00:15:40.220 --> 00:15:42.920 align:middle line:84%
Now what is this
inner integral?

00:15:42.920 --> 00:15:50.250 align:middle line:84%
Because x only varies very
little, this is approximately

00:15:50.250 --> 00:15:53.040 align:middle line:90%
constant in that range.

00:15:53.040 --> 00:15:56.210 align:middle line:84%
So the integral with
respect to x just

00:15:56.210 --> 00:15:58.840 align:middle line:90%
becomes delta times f(x,y).

00:15:58.840 --> 00:16:02.010 align:middle line:90%


00:16:02.010 --> 00:16:03.490 align:middle line:90%
And then we've got our dy.

00:16:03.490 --> 00:16:06.930 align:middle line:90%


00:16:06.930 --> 00:16:11.760 align:middle line:84%
So this is what the inner
integral will evaluate to.

00:16:11.760 --> 00:16:15.280 align:middle line:84%
We are integrating over
the little interval.

00:16:15.280 --> 00:16:17.450 align:middle line:90%
So we're keeping y fixed.

00:16:17.450 --> 00:16:22.020 align:middle line:84%
Integrating over here, we take
the value of the density times

00:16:22.020 --> 00:16:24.940 align:middle line:84%
how much we're integrating
over.

00:16:24.940 --> 00:16:27.890 align:middle line:90%
And we get this formula.

00:16:27.890 --> 00:16:28.410 align:middle line:90%
OK.

00:16:28.410 --> 00:16:33.170 align:middle line:84%
Now, this expression must be
equal to that expression.

00:16:33.170 --> 00:16:40.060 align:middle line:84%
So if we cancel the deltas, we
see that the marginal density

00:16:40.060 --> 00:16:44.000 align:middle line:84%
must be equal to the integral of
the joint density, where we

00:16:44.000 --> 00:16:48.200 align:middle line:84%
have integrated out
the value of y.

00:16:48.200 --> 00:16:54.060 align:middle line:90%


00:16:54.060 --> 00:16:59.000 align:middle line:84%
So this formula should come as
no surprise at this point.

00:16:59.000 --> 00:17:01.380 align:middle line:84%
It's exactly the same as the
formula that we had for

00:17:01.380 --> 00:17:03.270 align:middle line:90%
discrete random variables.

00:17:03.270 --> 00:17:06.800 align:middle line:84%
But now we are replacing the
sum with an integral.

00:17:06.800 --> 00:17:14.690 align:middle line:84%
And instead of using the
joint PMF, we are

00:17:14.690 --> 00:17:18.480 align:middle line:90%
using the joint PDF.

00:17:18.480 --> 00:17:21.810 align:middle line:84%
Then, continuing going down the
list of things we did for

00:17:21.810 --> 00:17:24.839 align:middle line:84%
discrete random variables, we
can now introduce a definition

00:17:24.839 --> 00:17:28.310 align:middle line:84%
of the notion of independence
of two random variables.

00:17:28.310 --> 00:17:31.050 align:middle line:84%
And by analogy with the discrete
case, we define

00:17:31.050 --> 00:17:33.940 align:middle line:84%
independence to be the
following condition.

00:17:33.940 --> 00:17:37.210 align:middle line:84%
Two random variables are
independent if and only if

00:17:37.210 --> 00:17:42.220 align:middle line:84%
their joint density function
factors out as a product of

00:17:42.220 --> 00:17:44.390 align:middle line:90%
their marginal densities.

00:17:44.390 --> 00:17:48.000 align:middle line:84%
And this property needs to
be true for all x and y.

00:17:48.000 --> 00:17:49.890 align:middle line:84%
So this is the formal
definition.

00:17:49.890 --> 00:17:53.020 align:middle line:84%
Operationally and intuitively,
what does it mean?

00:17:53.020 --> 00:17:55.110 align:middle line:84%
Well, intuitively it means
the same thing as in

00:17:55.110 --> 00:17:56.600 align:middle line:90%
the discrete case.

00:17:56.600 --> 00:18:00.610 align:middle line:84%
Knowing anything about X
shouldn't tell you anything

00:18:00.610 --> 00:18:05.320 align:middle line:84%
about Y. That is, information
about X is not going to change

00:18:05.320 --> 00:18:10.120 align:middle line:84%
your beliefs about Y. We are
going to come back to this

00:18:10.120 --> 00:18:11.370 align:middle line:90%
statement in a second.

00:18:11.370 --> 00:18:14.320 align:middle line:90%


00:18:14.320 --> 00:18:16.920 align:middle line:84%
The other thing that it
allows you to do--

00:18:16.920 --> 00:18:20.750 align:middle line:84%
I'm not going to derive this--
is it allows you to calculate

00:18:20.750 --> 00:18:25.650 align:middle line:84%
probabilities by multiplying
individual probabilities.

00:18:25.650 --> 00:18:28.110 align:middle line:84%
So if you ask for the
probability that x falls in a

00:18:28.110 --> 00:18:34.220 align:middle line:84%
certain set A and y falls in a
certain set B, then you can

00:18:34.220 --> 00:18:37.670 align:middle line:84%
calculate that probability
by multiplying individual

00:18:37.670 --> 00:18:38.920 align:middle line:90%
probabilities.

00:18:38.920 --> 00:18:41.860 align:middle line:90%


00:18:41.860 --> 00:18:46.090 align:middle line:84%
This takes just two lines of
derivation, which I'm not

00:18:46.090 --> 00:18:47.710 align:middle line:90%
going to do.

00:18:47.710 --> 00:18:51.240 align:middle line:84%
But it comes back to
the usual notion of

00:18:51.240 --> 00:18:53.370 align:middle line:90%
independence of events.

00:18:53.370 --> 00:18:56.340 align:middle line:84%
Basically, operationally
independence means that you

00:18:56.340 --> 00:18:57.660 align:middle line:90%
can multiply probabilities.

00:18:57.660 --> 00:19:00.190 align:middle line:90%


00:19:00.190 --> 00:19:04.380 align:middle line:84%
So now let's look
at an example.

00:19:04.380 --> 00:19:08.150 align:middle line:84%
There's a sort of pretty famous
and classical one.

00:19:08.150 --> 00:19:12.540 align:middle line:84%
It goes back a lot more
than a 100 years.

00:19:12.540 --> 00:19:16.290 align:middle line:84%
And it's the famous
Needle of Buffon.

00:19:16.290 --> 00:19:19.860 align:middle line:84%
Buffon was a French naturalist
who, for some reason, also

00:19:19.860 --> 00:19:22.150 align:middle line:84%
decided to play with
probability.

00:19:22.150 --> 00:19:24.590 align:middle line:84%
And look at the following
problem.

00:19:24.590 --> 00:19:28.400 align:middle line:84%
So you have the two-dimensional
plane.

00:19:28.400 --> 00:19:33.870 align:middle line:84%
And on the plane we draw a
bunch of parallel lines.

00:19:33.870 --> 00:19:37.575 align:middle line:84%
And those parallel lines are
separated by a length.

00:19:37.575 --> 00:19:46.830 align:middle line:90%


00:19:46.830 --> 00:19:52.270 align:middle line:84%
And the lines are apart
at distance d.

00:19:52.270 --> 00:19:58.780 align:middle line:84%
And we throw a needle at random,
completely at random.

00:19:58.780 --> 00:20:01.510 align:middle line:84%
And we'll have to give a meaning
to what "completely at

00:20:01.510 --> 00:20:03.180 align:middle line:90%
random" means.

00:20:03.180 --> 00:20:06.490 align:middle line:84%
And when we throw a needle,
there's two possibilities.

00:20:06.490 --> 00:20:09.640 align:middle line:84%
Either the needle is going to
fall in a way that does not

00:20:09.640 --> 00:20:13.120 align:middle line:84%
intersect any of the lines, or
it's going to fall in a way

00:20:13.120 --> 00:20:15.700 align:middle line:84%
that it intersects
one of the lines.

00:20:15.700 --> 00:20:19.470 align:middle line:84%
We're taking the needle to be
shorter than this distance, so

00:20:19.470 --> 00:20:22.185 align:middle line:84%
the needle cannot intersect
two lines simultaneously.

00:20:22.185 --> 00:20:26.230 align:middle line:84%
It either intersects 0, or it
intersects one of the lines.

00:20:26.230 --> 00:20:29.610 align:middle line:84%
The question is to find the
probability that the needle is

00:20:29.610 --> 00:20:32.100 align:middle line:90%
going to intersect a line.

00:20:32.100 --> 00:20:34.650 align:middle line:84%
What's the probability
of this?

00:20:34.650 --> 00:20:35.010 align:middle line:90%
OK.

00:20:35.010 --> 00:20:40.020 align:middle line:84%
We are going to approach this
problem by using our standard

00:20:40.020 --> 00:20:42.110 align:middle line:90%
four-step procedure.

00:20:42.110 --> 00:20:46.560 align:middle line:84%
Set up your sample space,
describe a probability law on

00:20:46.560 --> 00:20:51.460 align:middle line:84%
that sample space, identify
the event of interest, and

00:20:51.460 --> 00:20:53.370 align:middle line:90%
then calculate.

00:20:53.370 --> 00:20:58.470 align:middle line:84%
These four steps basically
correspond to these three

00:20:58.470 --> 00:21:04.110 align:middle line:84%
bullets and then the last
equation down here.

00:21:04.110 --> 00:21:06.510 align:middle line:84%
So first thing is to set
up a sample space.

00:21:06.510 --> 00:21:09.470 align:middle line:84%
We need some variables to
describe what happened in the

00:21:09.470 --> 00:21:10.780 align:middle line:90%
experiment.

00:21:10.780 --> 00:21:14.300 align:middle line:84%
So what happens in the
experiment is that the needle

00:21:14.300 --> 00:21:16.500 align:middle line:90%
lands somewhere.

00:21:16.500 --> 00:21:20.450 align:middle line:84%
And where it lands, we can
describe this by specifying

00:21:20.450 --> 00:21:24.160 align:middle line:84%
the location of the center
of the needle.

00:21:24.160 --> 00:21:27.020 align:middle line:84%
And what do we mean by the
location of the center?

00:21:27.020 --> 00:21:30.310 align:middle line:84%
Well, we can take as our
variable to be the distance

00:21:30.310 --> 00:21:33.035 align:middle line:84%
from the center of the needle
to the nearest line.

00:21:33.035 --> 00:21:36.280 align:middle line:90%


00:21:36.280 --> 00:21:42.520 align:middle line:84%
So it tells us the vertical
distance of the center of the

00:21:42.520 --> 00:21:45.930 align:middle line:90%
needle from the nearest line.

00:21:45.930 --> 00:21:47.500 align:middle line:84%
The other thing that
matters is the

00:21:47.500 --> 00:21:49.400 align:middle line:90%
orientation of the needle.

00:21:49.400 --> 00:21:53.820 align:middle line:84%
So we need one more variable,
which we take to be the angle

00:21:53.820 --> 00:21:56.940 align:middle line:84%
that the needle is forming
with the lines.

00:21:56.940 --> 00:22:00.260 align:middle line:84%
We can put the angle here,
or you can put in there.

00:22:00.260 --> 00:22:02.620 align:middle line:84%
Yes, it's still the
same angle.

00:22:02.620 --> 00:22:06.850 align:middle line:84%
So we have these two variables
that described what happened

00:22:06.850 --> 00:22:08.190 align:middle line:90%
in the experiment.

00:22:08.190 --> 00:22:11.280 align:middle line:84%
And we can take our sample space
to be the set of all

00:22:11.280 --> 00:22:14.390 align:middle line:90%
possible x's and theta's.

00:22:14.390 --> 00:22:16.770 align:middle line:90%
What are the possible x's?

00:22:16.770 --> 00:22:20.800 align:middle line:84%
The lines are d apart, so the
nearest line is going to be

00:22:20.800 --> 00:22:24.400 align:middle line:84%
anywhere between
0 and d/2 away.

00:22:24.400 --> 00:22:28.630 align:middle line:84%
So that tells us what the
possible x's will be.

00:22:28.630 --> 00:22:31.420 align:middle line:84%
As for theta, it really
depends how

00:22:31.420 --> 00:22:33.230 align:middle line:90%
you define your angle.

00:22:33.230 --> 00:22:37.510 align:middle line:84%
We are going to define our theta
to be the acute angle

00:22:37.510 --> 00:22:44.020 align:middle line:84%
that's formed between the needle
and a line, if you were

00:22:44.020 --> 00:22:45.130 align:middle line:90%
to extend it.

00:22:45.130 --> 00:22:50.180 align:middle line:84%
So theta is going to be
something between 0 and pi/2.

00:22:50.180 --> 00:22:54.140 align:middle line:84%
So I guess these red pieces
really correspond to the part

00:22:54.140 --> 00:22:58.490 align:middle line:84%
of setting up the
sample space.

00:22:58.490 --> 00:22:58.810 align:middle line:90%
OK.

00:22:58.810 --> 00:23:00.270 align:middle line:90%
So that's part one.

00:23:00.270 --> 00:23:03.390 align:middle line:84%
Second part is we
need a model.

00:23:03.390 --> 00:23:03.690 align:middle line:90%
OK.

00:23:03.690 --> 00:23:08.140 align:middle line:84%
Let's take our model to be that
we basically know nothing

00:23:08.140 --> 00:23:10.600 align:middle line:90%
about how the needle falls.

00:23:10.600 --> 00:23:13.890 align:middle line:84%
It can fall in any possible way,
and all possible ways are

00:23:13.890 --> 00:23:15.230 align:middle line:90%
equally likely.

00:23:15.230 --> 00:23:18.910 align:middle line:84%
Now, if you have those parallel
lines, and you close

00:23:18.910 --> 00:23:22.330 align:middle line:84%
your eyes completely and throw a
needle completely at random,

00:23:22.330 --> 00:23:25.260 align:middle line:84%
any x should be equally
likely.

00:23:25.260 --> 00:23:29.490 align:middle line:84%
So we describe that situation by
saying that X should have a

00:23:29.490 --> 00:23:31.360 align:middle line:90%
uniform distribution.

00:23:31.360 --> 00:23:33.880 align:middle line:84%
That is, it should have a
constant density over the

00:23:33.880 --> 00:23:35.410 align:middle line:90%
range of interest.

00:23:35.410 --> 00:23:39.160 align:middle line:84%
Similarly, if you kind of spin
your needle completely at

00:23:39.160 --> 00:23:43.580 align:middle line:84%
random, any angle should be as
likely as any other angle.

00:23:43.580 --> 00:23:47.160 align:middle line:84%
And we decide to model this
situation by saying that theta

00:23:47.160 --> 00:23:49.680 align:middle line:84%
also has a uniform
distribution over

00:23:49.680 --> 00:23:50.995 align:middle line:90%
the range of interest.

00:23:50.995 --> 00:23:54.220 align:middle line:90%


00:23:54.220 --> 00:23:58.500 align:middle line:84%
And finally, where we put it
should have nothing to do with

00:23:58.500 --> 00:24:00.370 align:middle line:90%
how much we rotate it.

00:24:00.370 --> 00:24:04.320 align:middle line:84%
And we capture this
mathematically by saying that

00:24:04.320 --> 00:24:07.480 align:middle line:84%
X is going to be independent
of theta.

00:24:07.480 --> 00:24:09.220 align:middle line:84%
Now, this is going
to be our model.

00:24:09.220 --> 00:24:11.920 align:middle line:84%
I'm not deriving the model
from anything.

00:24:11.920 --> 00:24:15.480 align:middle line:84%
I'm only saying that this sounds
like a model that does

00:24:15.480 --> 00:24:19.800 align:middle line:84%
not assume any knowledge or
preference for certain values

00:24:19.800 --> 00:24:22.360 align:middle line:84%
of x rather than other
values of theta.

00:24:22.360 --> 00:24:25.660 align:middle line:84%
In the absence of any other
particular information you

00:24:25.660 --> 00:24:28.420 align:middle line:84%
might have in your hands, that's
the most reasonable

00:24:28.420 --> 00:24:30.520 align:middle line:90%
model to come up with.

00:24:30.520 --> 00:24:32.150 align:middle line:84%
So you model the problem
that way.

00:24:32.150 --> 00:24:35.490 align:middle line:84%
So what's the formula for
the joint density?

00:24:35.490 --> 00:24:37.590 align:middle line:84%
It's going to be the
product of the

00:24:37.590 --> 00:24:41.200 align:middle line:90%
densities of X and Theta.

00:24:41.200 --> 00:24:42.410 align:middle line:90%
Why is it the product?

00:24:42.410 --> 00:24:45.530 align:middle line:84%
This is because we assumed
independence.

00:24:45.530 --> 00:24:48.910 align:middle line:84%
And the density of X, since
it's uniform, and since it

00:24:48.910 --> 00:24:54.630 align:middle line:84%
needs to integrate to 1, that
density needs to be 2/d.

00:24:54.630 --> 00:24:57.580 align:middle line:84%
That's the density of X.
And the density of

00:24:57.580 --> 00:25:00.740 align:middle line:90%
Theta needs to be 2/pi.

00:25:00.740 --> 00:25:03.660 align:middle line:84%
That's the value for the density
of Theta so that the

00:25:03.660 --> 00:25:07.920 align:middle line:84%
overall probability over this
interval ends up being 1.

00:25:07.920 --> 00:25:12.390 align:middle line:84%
So now we do have our joint
density in our hands.

00:25:12.390 --> 00:25:14.690 align:middle line:84%
The next thing to do
is to identify

00:25:14.690 --> 00:25:17.920 align:middle line:90%
the event of interest.

00:25:17.920 --> 00:25:20.720 align:middle line:84%
And this is best done
in a picture.

00:25:20.720 --> 00:25:23.380 align:middle line:84%
And there's two possible
situations

00:25:23.380 --> 00:25:25.450 align:middle line:90%
that one could have.

00:25:25.450 --> 00:25:33.450 align:middle line:84%
Either the needle falls this
way, or it falls this way.

00:25:33.450 --> 00:25:38.300 align:middle line:84%
So how can we tell if one or the
other is going to happen?

00:25:38.300 --> 00:25:45.470 align:middle line:84%
It has to do with whether this
interval here is smaller than

00:25:45.470 --> 00:25:50.130 align:middle line:90%
that or bigger than that.

00:25:50.130 --> 00:25:52.260 align:middle line:84%
So we are comparing
the height of this

00:25:52.260 --> 00:25:55.460 align:middle line:90%
interval to that interval.

00:25:55.460 --> 00:25:58.220 align:middle line:84%
This interval here
is capital X.

00:25:58.220 --> 00:26:02.350 align:middle line:84%
This interval here,
what is it?

00:26:02.350 --> 00:26:07.040 align:middle line:84%
This is half of the length of
the needle, which is l/2.

00:26:07.040 --> 00:26:10.590 align:middle line:84%
To find this height, we take l/2
and multiply it with the

00:26:10.590 --> 00:26:13.700 align:middle line:84%
sine of the angle
that we have.

00:26:13.700 --> 00:26:18.330 align:middle line:84%
So the length of this
interval up here is

00:26:18.330 --> 00:26:23.500 align:middle line:90%
l/2 times sine theta.

00:26:23.500 --> 00:26:28.520 align:middle line:84%
If this is smaller than
x, the needle does not

00:26:28.520 --> 00:26:30.010 align:middle line:90%
intersect the line.

00:26:30.010 --> 00:26:33.130 align:middle line:84%
If this is bigger than
x, then the needle

00:26:33.130 --> 00:26:34.920 align:middle line:90%
intersects the line.

00:26:34.920 --> 00:26:37.870 align:middle line:84%
So the event of interest, that
the needle intersects the

00:26:37.870 --> 00:26:42.740 align:middle line:84%
line, is described this way
in terms of x and theta.

00:26:42.740 --> 00:26:46.170 align:middle line:84%
And now that we have the event
of interest described

00:26:46.170 --> 00:26:50.100 align:middle line:84%
mathematically, all that we
need to do is to find the

00:26:50.100 --> 00:26:54.800 align:middle line:84%
probability of this event, we
integrate the joint density

00:26:54.800 --> 00:26:59.560 align:middle line:84%
over the part of (x, theta)
space in which this

00:26:59.560 --> 00:27:01.320 align:middle line:90%
inequality is true.

00:27:01.320 --> 00:27:04.670 align:middle line:84%
So it's a double integral over
the set of all x's and theta's

00:27:04.670 --> 00:27:06.450 align:middle line:90%
where this is true.

00:27:06.450 --> 00:27:11.430 align:middle line:84%
The way to do this integral is
we fix theta, and we integrate

00:27:11.430 --> 00:27:15.150 align:middle line:84%
for x's that go from 0
up to that number.

00:27:15.150 --> 00:27:19.030 align:middle line:84%
And theta can be anything
between 0 and pi/2.

00:27:19.030 --> 00:27:23.620 align:middle line:84%
So the integral over this set
is basically this double

00:27:23.620 --> 00:27:24.980 align:middle line:90%
integral here.

00:27:24.980 --> 00:27:27.475 align:middle line:84%
We already have a formula
for the joint density.

00:27:27.475 --> 00:27:30.930 align:middle line:84%
It's 4 over pi d, so
we put it here.

00:27:30.930 --> 00:27:32.640 align:middle line:84%
And now, fortunately,
this is a pretty

00:27:32.640 --> 00:27:34.645 align:middle line:90%
easy integral to evaluate.

00:27:34.645 --> 00:27:37.650 align:middle line:84%
The integral with respect to x
-- there's nothing in here.

00:27:37.650 --> 00:27:40.950 align:middle line:84%
So the integral is just the
length of the interval over

00:27:40.950 --> 00:27:42.370 align:middle line:90%
which we're integrating.

00:27:42.370 --> 00:27:44.950 align:middle line:90%
It's l/2 sine theta.

00:27:44.950 --> 00:27:47.870 align:middle line:84%
And then we need to integrate
this with respect to theta.

00:27:47.870 --> 00:27:53.990 align:middle line:84%
We know that the integral of a
sine is a negative cosine.

00:27:53.990 --> 00:27:56.990 align:middle line:84%
You plug in the values for
the negative cosine

00:27:56.990 --> 00:27:58.390 align:middle line:90%
at the two end points.

00:27:58.390 --> 00:28:00.260 align:middle line:84%
I'm sure you can do
this integral .

00:28:00.260 --> 00:28:04.540 align:middle line:84%
And we finally obtain the
answer, which is amazingly

00:28:04.540 --> 00:28:08.210 align:middle line:84%
simple for such a pretty
complicated-looking problem.

00:28:08.210 --> 00:28:09.910 align:middle line:90%
It's 2l over pi d.

00:28:09.910 --> 00:28:12.420 align:middle line:90%


00:28:12.420 --> 00:28:15.360 align:middle line:84%
So some people a long, long time
ago, after they looked at

00:28:15.360 --> 00:28:19.290 align:middle line:84%
this answer, they said that
maybe that gives us an

00:28:19.290 --> 00:28:22.910 align:middle line:84%
interesting way where one could
estimate the value by

00:28:22.910 --> 00:28:26.130 align:middle line:84%
pi, for example,
experimentally.

00:28:26.130 --> 00:28:27.690 align:middle line:90%
How do you do that?

00:28:27.690 --> 00:28:32.360 align:middle line:84%
Fix l and d, the dimensions
of the problem.

00:28:32.360 --> 00:28:36.680 align:middle line:84%
Throw a million needles on
your piece of paper.

00:28:36.680 --> 00:28:40.690 align:middle line:84%
See how often your needless
do intersect the line.

00:28:40.690 --> 00:28:43.540 align:middle line:84%
That gives you a number
for this quantity.

00:28:43.540 --> 00:28:48.540 align:middle line:84%
You know l and d, so you can
use that to infer pi.

00:28:48.540 --> 00:28:52.330 align:middle line:84%
And there's an apocryphal story
about a wounded soldier

00:28:52.330 --> 00:28:55.300 align:middle line:84%
in a hospital after the
American Civil War who

00:28:55.300 --> 00:28:58.490 align:middle line:84%
actually had heard about this
and was spending his time in

00:28:58.490 --> 00:29:02.680 align:middle line:84%
the hospital throwing needles
on pieces of paper.

00:29:02.680 --> 00:29:04.350 align:middle line:84%
I don't know if it's
true or not.

00:29:04.350 --> 00:29:07.330 align:middle line:84%
But let's do something
similar here.

00:29:07.330 --> 00:29:11.720 align:middle line:90%
So let's look at this diagram.

00:29:11.720 --> 00:29:14.110 align:middle line:90%
We fix the dimensions.

00:29:14.110 --> 00:29:15.920 align:middle line:84%
This is supposed to
be our little d.

00:29:15.920 --> 00:29:18.330 align:middle line:84%
That's supposed to
be our little l.

00:29:18.330 --> 00:29:22.430 align:middle line:84%
We have the formula from the
previous slide that p

00:29:22.430 --> 00:29:25.230 align:middle line:90%
is 2l over pi d.

00:29:25.230 --> 00:29:29.230 align:middle line:84%
In this instance, we choose
d to be twice l.

00:29:29.230 --> 00:29:32.170 align:middle line:90%
So this number is 1/pi.

00:29:32.170 --> 00:29:37.770 align:middle line:84%
So the probability that the
needle hits the line is 1/pi.

00:29:37.770 --> 00:29:41.150 align:middle line:84%
So I need needles that are
3.1 centimeters long.

00:29:41.150 --> 00:29:42.730 align:middle line:90%
I couldn't find such needles.

00:29:42.730 --> 00:29:47.360 align:middle line:84%
But I could find paper clips
that are 3.1 centimeters long.

00:29:47.360 --> 00:29:51.510 align:middle line:84%
So let's start throwing paper
clips at random and see how

00:29:51.510 --> 00:29:55.285 align:middle line:84%
many of them will end up
intersecting the lines.

00:29:55.285 --> 00:30:00.501 align:middle line:90%


00:30:00.501 --> 00:30:01.920 align:middle line:90%
Good.

00:30:01.920 --> 00:30:02.400 align:middle line:90%
OK.

00:30:02.400 --> 00:30:09.350 align:middle line:84%
So out of eight paper clips,
we have exactly four that

00:30:09.350 --> 00:30:11.510 align:middle line:90%
intersected the line.

00:30:11.510 --> 00:30:13.620 align:middle line:84%
So our estimate for the
probability of intersecting

00:30:13.620 --> 00:30:18.970 align:middle line:84%
the line is 1/2, which gives us
an estimate for the value

00:30:18.970 --> 00:30:22.010 align:middle line:90%
of pi, which is two.

00:30:22.010 --> 00:30:24.960 align:middle line:84%
Well, I mean, within an
engineering approximation,

00:30:24.960 --> 00:30:29.090 align:middle line:84%
we're in the right
ballpark, right?

00:30:29.090 --> 00:30:32.890 align:middle line:84%
So this might look like a
silly way of trying to

00:30:32.890 --> 00:30:33.920 align:middle line:90%
estimate pi.

00:30:33.920 --> 00:30:36.420 align:middle line:90%
And it probably is.

00:30:36.420 --> 00:30:41.200 align:middle line:84%
On the other hand, this kind of
methodology is being used

00:30:41.200 --> 00:30:44.930 align:middle line:84%
especially by physicists and
also by statisticians.

00:30:44.930 --> 00:30:46.550 align:middle line:90%
It's used a lot.

00:30:46.550 --> 00:30:48.260 align:middle line:90%
When is it used?

00:30:48.260 --> 00:30:52.300 align:middle line:84%
If you have an integral to
calculate, such as this

00:30:52.300 --> 00:30:55.980 align:middle line:84%
integral, but you're not lucky,
and your functions are

00:30:55.980 --> 00:30:59.980 align:middle line:84%
not so simple where you can do
your calculations by hand, and

00:30:59.980 --> 00:31:02.590 align:middle line:84%
maybe the dimensions are
larger-- instead of two random

00:31:02.590 --> 00:31:04.590 align:middle line:84%
variables you have 100
random variables, so

00:31:04.590 --> 00:31:08.210 align:middle line:90%
it's a 100-fold integral--

00:31:08.210 --> 00:31:10.830 align:middle line:84%
then there's no way to do
that in the computer.

00:31:10.830 --> 00:31:14.230 align:middle line:84%
But the way that you can
actually do it is by

00:31:14.230 --> 00:31:18.290 align:middle line:84%
generating random samples of
your random variables, doing

00:31:18.290 --> 00:31:21.220 align:middle line:84%
that simulation over and
over many times.

00:31:21.220 --> 00:31:25.010 align:middle line:84%
That is, by interpreting an
integral as a probability, you

00:31:25.010 --> 00:31:29.060 align:middle line:84%
can use simulation to estimate
that probability.

00:31:29.060 --> 00:31:32.470 align:middle line:84%
And that gives you a way of
calculating integrals.

00:31:32.470 --> 00:31:36.850 align:middle line:84%
And physicists do actually use
that a lot, as well as

00:31:36.850 --> 00:31:39.630 align:middle line:84%
statisticians, computer
scientists, and so on.

00:31:39.630 --> 00:31:41.760 align:middle line:84%
It's a so-called Monte
Carlo method

00:31:41.760 --> 00:31:43.990 align:middle line:90%
for evaluating integrals.

00:31:43.990 --> 00:31:50.250 align:middle line:84%
And it's a basic piece of the
toolbox in science these days.

00:31:50.250 --> 00:31:54.610 align:middle line:84%
Finally, the harder concept
of the day is the idea of

00:31:54.610 --> 00:31:55.770 align:middle line:90%
conditioning.

00:31:55.770 --> 00:31:58.740 align:middle line:84%
And here things become a little
subtle when you deal

00:31:58.740 --> 00:32:00.970 align:middle line:84%
with continuous random
variables.

00:32:00.970 --> 00:32:02.290 align:middle line:90%
OK.

00:32:02.290 --> 00:32:05.810 align:middle line:84%
First, remember again our basic
interpretation of what a

00:32:05.810 --> 00:32:06.860 align:middle line:90%
density is.

00:32:06.860 --> 00:32:08.200 align:middle line:90%
A density gives us

00:32:08.200 --> 00:32:10.500 align:middle line:90%
probabilities of little intervals.

00:32:10.500 --> 00:32:13.560 align:middle line:84%
So how should we define
conditional densities?

00:32:13.560 --> 00:32:16.600 align:middle line:84%
Conditional densities should
again give us probabilities of

00:32:16.600 --> 00:32:21.290 align:middle line:84%
little intervals, but inside a
conditional world where we

00:32:21.290 --> 00:32:24.530 align:middle line:84%
have been told something about
the other random variable.

00:32:24.530 --> 00:32:28.090 align:middle line:84%
So what we would like to be
true is the following.

00:32:28.090 --> 00:32:31.340 align:middle line:84%
We would like to define a
concept of a conditional

00:32:31.340 --> 00:32:34.530 align:middle line:84%
density of a random variable X
given the value of another

00:32:34.530 --> 00:32:37.860 align:middle line:84%
random variable Y. And it should
behave the following

00:32:37.860 --> 00:32:40.570 align:middle line:84%
way, that the conditional
density gives us the

00:32:40.570 --> 00:32:42.690 align:middle line:84%
probability of little
intervals--

00:32:42.690 --> 00:32:44.260 align:middle line:90%
same as here--

00:32:44.260 --> 00:32:48.440 align:middle line:84%
given that we are told
the value of y.

00:32:48.440 --> 00:32:50.930 align:middle line:84%
And here's where the
subtleties come.

00:32:50.930 --> 00:32:54.420 align:middle line:84%
The main thing to notice is
that here I didn't write

00:32:54.420 --> 00:32:59.000 align:middle line:84%
"equal," I wrote "approximately
equal." Why do

00:32:59.000 --> 00:33:01.250 align:middle line:90%
we need that?

00:33:01.250 --> 00:33:04.460 align:middle line:84%
Well, the thing is that
conditional probabilities are

00:33:04.460 --> 00:33:08.840 align:middle line:84%
not defined when you condition
on an event that has 0

00:33:08.840 --> 00:33:10.180 align:middle line:90%
probability.

00:33:10.180 --> 00:33:13.400 align:middle line:84%
So we need the conditioning
event here to have posed this

00:33:13.400 --> 00:33:14.430 align:middle line:90%
probability.

00:33:14.430 --> 00:33:18.840 align:middle line:84%
So instead of saying that Y is
exactly equal to little y, we

00:33:18.840 --> 00:33:22.900 align:middle line:84%
want to instead say we're in a
new universe where capital Y

00:33:22.900 --> 00:33:27.070 align:middle line:90%
is very close to little y.

00:33:27.070 --> 00:33:31.410 align:middle line:84%
And then this notion of "very
close" kind of takes the limit

00:33:31.410 --> 00:33:34.910 align:middle line:84%
and takes it to be
infinitesimally close.

00:33:34.910 --> 00:33:38.610 align:middle line:84%
So this is the way to interpret
conditional

00:33:38.610 --> 00:33:40.120 align:middle line:90%
probabilities.

00:33:40.120 --> 00:33:42.550 align:middle line:90%
That's what they should mean.

00:33:42.550 --> 00:33:45.330 align:middle line:84%
Now, in practice, when you
actually use probability, you

00:33:45.330 --> 00:33:46.780 align:middle line:90%
forget about that subtlety.

00:33:46.780 --> 00:33:50.940 align:middle line:84%
And you say, well, I've been
told that Y is equal to 1.3.

00:33:50.940 --> 00:33:53.780 align:middle line:84%
Give me the conditional
distribution of X. But

00:33:53.780 --> 00:33:58.080 align:middle line:84%
formally or rigorously, you
should say I'm being told that

00:33:58.080 --> 00:34:01.400 align:middle line:84%
Y is infinitesimally
close to 1.3.

00:34:01.400 --> 00:34:03.620 align:middle line:90%
Tell me the distribution of X.

00:34:03.620 --> 00:34:08.580 align:middle line:84%
Now, if this is what we want,
what should this quantity be?

00:34:08.580 --> 00:34:10.489 align:middle line:84%
It's a conditional probability,
so it should be

00:34:10.489 --> 00:34:12.800 align:middle line:84%
the probability of two
things happening--

00:34:12.800 --> 00:34:16.550 align:middle line:84%
X being close to little x, Y
being close to little y.

00:34:16.550 --> 00:34:20.010 align:middle line:84%
And that's basically given to
us by the joint density

00:34:20.010 --> 00:34:23.920 align:middle line:84%
divided by the probability of
the conditioning event, which

00:34:23.920 --> 00:34:27.449 align:middle line:84%
has something to do with the
density of Y itself.

00:34:27.449 --> 00:34:30.840 align:middle line:84%
And if you do things carefully,
you see that the

00:34:30.840 --> 00:34:34.350 align:middle line:84%
only way to satisfy this
relation is to define the

00:34:34.350 --> 00:34:38.065 align:middle line:84%
conditional density by this
particular formula.

00:34:38.065 --> 00:34:38.590 align:middle line:90%
OK.

00:34:38.590 --> 00:34:44.159 align:middle line:84%
Big discussion to come down in
the end to what you should

00:34:44.159 --> 00:34:46.120 align:middle line:90%
have probably guessed by now.

00:34:46.120 --> 00:34:49.170 align:middle line:84%
We just take any formulas and
expressions from the discrete

00:34:49.170 --> 00:34:53.570 align:middle line:90%
case and replace PMFs by PDFs.

00:34:53.570 --> 00:34:58.030 align:middle line:84%
So the conditional PDF is
defined by this formula where

00:34:58.030 --> 00:35:02.450 align:middle line:84%
here we have joint PDF and
marginal PDF, as opposed to

00:35:02.450 --> 00:35:05.450 align:middle line:84%
the discrete case where we
had the joint PMF and

00:35:05.450 --> 00:35:07.540 align:middle line:90%
the marginal PMF.

00:35:07.540 --> 00:35:11.850 align:middle line:84%
So in some sense, it's just
a syntactic change.

00:35:11.850 --> 00:35:14.510 align:middle line:84%
In another sense, it's a little
subtler on how you

00:35:14.510 --> 00:35:17.130 align:middle line:90%
actually interpret it.

00:35:17.130 --> 00:35:20.230 align:middle line:84%
Speaking about interpretation,
what are some ways of thinking

00:35:20.230 --> 00:35:22.170 align:middle line:90%
about the joint density?

00:35:22.170 --> 00:35:24.740 align:middle line:84%
Well, the best way to think
about it is that somebody has

00:35:24.740 --> 00:35:27.720 align:middle line:90%
fixed little y for you.

00:35:27.720 --> 00:35:31.980 align:middle line:84%
So little y is being
fixed here.

00:35:31.980 --> 00:35:35.350 align:middle line:84%
And we look at this density
as a function of X.

00:35:35.350 --> 00:35:37.020 align:middle line:90%
I've told you what Y is.

00:35:37.020 --> 00:35:39.870 align:middle line:84%
Tell me what you know about X.
And you tell me that X has a

00:35:39.870 --> 00:35:42.070 align:middle line:90%
certain distribution.

00:35:42.070 --> 00:35:44.840 align:middle line:84%
What does that distribution
look like?

00:35:44.840 --> 00:35:50.070 align:middle line:84%
It has exactly the same shape
as the joint density.

00:35:50.070 --> 00:35:53.390 align:middle line:84%
Remember, we fixed Y. So
this is a constant.

00:35:53.390 --> 00:35:57.200 align:middle line:84%
So the only thing that varies
is X. So we get the function

00:35:57.200 --> 00:36:01.320 align:middle line:84%
that behaves like the joint
density when you fix y, which

00:36:01.320 --> 00:36:04.100 align:middle line:84%
is really you take the joint
density, and you

00:36:04.100 --> 00:36:05.650 align:middle line:90%
take a slice of it.

00:36:05.650 --> 00:36:09.200 align:middle line:84%
You fix a y, and you see
how it varies with x.

00:36:09.200 --> 00:36:11.810 align:middle line:84%
So in that sense, the
conditional PDF is just a

00:36:11.810 --> 00:36:14.150 align:middle line:90%
slice of the joint PDF.

00:36:14.150 --> 00:36:17.230 align:middle line:84%
But we need to divide by a
certain number, which just

00:36:17.230 --> 00:36:19.480 align:middle line:84%
scales it and changes
its shape.

00:36:19.480 --> 00:36:21.950 align:middle line:84%
We're coming back to a
picture in a second.

00:36:21.950 --> 00:36:25.410 align:middle line:84%
But before going to the picture,
lets go back to the

00:36:25.410 --> 00:36:27.840 align:middle line:84%
interpretation of
independence.

00:36:27.840 --> 00:36:30.230 align:middle line:84%
If the two random the variables
are independent,

00:36:30.230 --> 00:36:33.550 align:middle line:84%
according to our definition in
the previous slide, the joint

00:36:33.550 --> 00:36:36.130 align:middle line:84%
density is going to factor
as the product of

00:36:36.130 --> 00:36:37.820 align:middle line:90%
the marginal densities.

00:36:37.820 --> 00:36:40.850 align:middle line:84%
The density of Y in the
numerator cancels the density

00:36:40.850 --> 00:36:42.010 align:middle line:90%
in the denominator.

00:36:42.010 --> 00:36:44.410 align:middle line:84%
And we're just left with
the density of X.

00:36:44.410 --> 00:36:46.940 align:middle line:84%
So in the case of independence,
what we get is

00:36:46.940 --> 00:36:49.870 align:middle line:84%
that the conditional is the
same as the marginal.

00:36:49.870 --> 00:36:52.980 align:middle line:84%
And that solidifies our
intuition that in the case of

00:36:52.980 --> 00:36:58.080 align:middle line:84%
independence, being told
something about the value of Y

00:36:58.080 --> 00:37:02.540 align:middle line:84%
does not change our beliefs
about how X is distributed.

00:37:02.540 --> 00:37:06.110 align:middle line:84%
So whatever we expected about X
is going to remain true even

00:37:06.110 --> 00:37:09.180 align:middle line:84%
after we are told something
about Y.

00:37:09.180 --> 00:37:12.680 align:middle line:84%
So let's look at
some pictures.

00:37:12.680 --> 00:37:16.110 align:middle line:84%
Here is what the joint
PDF might look like.

00:37:16.110 --> 00:37:19.480 align:middle line:84%
Here we've got our
x and y-axis.

00:37:19.480 --> 00:37:23.100 align:middle line:84%
And if you want to calculate the
probability of a certain

00:37:23.100 --> 00:37:27.240 align:middle line:84%
event, what you do is you look
at that event and you see how

00:37:27.240 --> 00:37:31.740 align:middle line:84%
much of that mass is sitting
on top of that event.

00:37:31.740 --> 00:37:35.180 align:middle line:90%
Now let's start slicing.

00:37:35.180 --> 00:37:43.360 align:middle line:84%
Let's fix a value of x and look
along that slice where we

00:37:43.360 --> 00:37:48.610 align:middle line:90%
obtain this function.

00:37:48.610 --> 00:37:52.280 align:middle line:90%
Now what does that slice do?

00:37:52.280 --> 00:37:56.100 align:middle line:84%
That slice tells us for that
particular x what the possible

00:37:56.100 --> 00:38:00.330 align:middle line:84%
values of y are going to be
and how likely they are.

00:38:00.330 --> 00:38:05.440 align:middle line:84%
If we integrate over all
y's, what do we get?

00:38:05.440 --> 00:38:10.400 align:middle line:84%
Integrating over all y's just
gives us the marginal density

00:38:10.400 --> 00:38:15.270 align:middle line:84%
of X. It's the calculation
that we did here.

00:38:15.270 --> 00:38:19.820 align:middle line:84%
By integrating over all y's, we
find the marginal density

00:38:19.820 --> 00:38:27.850 align:middle line:84%
of X. So the total area under
that slice gives us the

00:38:27.850 --> 00:38:31.340 align:middle line:84%
marginal density of X. And by
looking at the different

00:38:31.340 --> 00:38:35.430 align:middle line:84%
slices, we find how likely the
different values of x are

00:38:35.430 --> 00:38:36.660 align:middle line:90%
going to be.

00:38:36.660 --> 00:38:39.410 align:middle line:90%
How about the conditional?

00:38:39.410 --> 00:38:48.790 align:middle line:84%
If we're interested in the
conditional of Y given X, how

00:38:48.790 --> 00:38:51.200 align:middle line:90%
would you think about it?

00:38:51.200 --> 00:38:54.620 align:middle line:84%
This refers to a universe where
we are told that capital

00:38:54.620 --> 00:38:57.550 align:middle line:90%
X takes on a specific value.

00:38:57.550 --> 00:39:00.010 align:middle line:84%
So we put ourselves in
the universe where

00:39:00.010 --> 00:39:01.810 align:middle line:90%
this line has happened.

00:39:01.810 --> 00:39:05.940 align:middle line:84%
There's still possible values
of y that can happen.

00:39:05.940 --> 00:39:09.270 align:middle line:84%
And this shape kind of tells us
the relative likelihoods of

00:39:09.270 --> 00:39:10.760 align:middle line:90%
the different y's.

00:39:10.760 --> 00:39:14.060 align:middle line:84%
And this is indeed going to be
the shape of the conditional

00:39:14.060 --> 00:39:17.850 align:middle line:84%
distribution of Y given
that X has occurred.

00:39:17.850 --> 00:39:21.090 align:middle line:84%
On the other hand, the
conditional distribution must

00:39:21.090 --> 00:39:22.630 align:middle line:90%
add up to 1.

00:39:22.630 --> 00:39:25.920 align:middle line:84%
So the total probability over
all of the different y's in

00:39:25.920 --> 00:39:27.730 align:middle line:84%
this universe, that
total probability

00:39:27.730 --> 00:39:29.540 align:middle line:90%
should be equal to 1.

00:39:29.540 --> 00:39:31.450 align:middle line:90%
Here it's not equal to 1.

00:39:31.450 --> 00:39:34.290 align:middle line:84%
The total area is the
marginal density.

00:39:34.290 --> 00:39:38.590 align:middle line:84%
To make it equal to 1, we need
to divide by the marginal

00:39:38.590 --> 00:39:44.160 align:middle line:84%
density, which is basically to
renormalize this shape so that

00:39:44.160 --> 00:39:48.500 align:middle line:84%
the total area under that slice,
under that shape, is

00:39:48.500 --> 00:39:50.400 align:middle line:90%
equal to 1.

00:39:50.400 --> 00:39:53.430 align:middle line:90%
So we start with the joint.

00:39:53.430 --> 00:39:55.730 align:middle line:90%
We take the slices.

00:39:55.730 --> 00:40:00.280 align:middle line:84%
And then we adjust the slices
so that every slice has an

00:40:00.280 --> 00:40:03.610 align:middle line:90%
area underneath equal to 1.

00:40:03.610 --> 00:40:05.650 align:middle line:84%
And this gives us
the conditional.

00:40:05.650 --> 00:40:09.160 align:middle line:90%
So for example, down here--

00:40:09.160 --> 00:40:11.840 align:middle line:84%
you can not even see it
in this diagram--

00:40:11.840 --> 00:40:15.410 align:middle line:84%
but after you renormalize it
so that its total area is

00:40:15.410 --> 00:40:20.160 align:middle line:84%
equal to 1, you get this sort of
narrow spike that goes up.

00:40:20.160 --> 00:40:22.980 align:middle line:84%
And so this is a plot of the
conditional distributions that

00:40:22.980 --> 00:40:26.060 align:middle line:84%
you get for the different
values of x.

00:40:26.060 --> 00:40:29.050 align:middle line:84%
Given a particular value of x,
you're going to get this

00:40:29.050 --> 00:40:31.460 align:middle line:84%
certain conditional
distribution.

00:40:31.460 --> 00:40:36.460 align:middle line:84%
So this picture is worth about
as much as anything else in

00:40:36.460 --> 00:40:38.840 align:middle line:90%
this particular chapter.

00:40:38.840 --> 00:40:42.990 align:middle line:84%
Make sure you kind of understand
exactly all these

00:40:42.990 --> 00:40:44.240 align:middle line:90%
pieces of the picture.

00:40:44.240 --> 00:40:47.130 align:middle line:90%


00:40:47.130 --> 00:40:49.870 align:middle line:84%
And finally, let's go, in the
remaining time, through an

00:40:49.870 --> 00:40:55.240 align:middle line:84%
example where we're going to
throw in the bucket all the

00:40:55.240 --> 00:40:58.320 align:middle line:84%
concepts and notations that
we have introduced so far.

00:40:58.320 --> 00:40:59.960 align:middle line:90%
So the example is as follows.

00:40:59.960 --> 00:41:04.210 align:middle line:84%
We start with a stick that
has a certain length.

00:41:04.210 --> 00:41:07.790 align:middle line:84%
And we break it a completely
random location.

00:41:07.790 --> 00:41:09.390 align:middle line:90%
And--

00:41:09.390 --> 00:41:13.686 align:middle line:90%
yes, this 1 should be l.

00:41:13.686 --> 00:41:14.130 align:middle line:90%
OK.

00:41:14.130 --> 00:41:15.770 align:middle line:90%
So it has length l.

00:41:15.770 --> 00:41:19.210 align:middle line:84%
And we're going to break
it at the random place.

00:41:19.210 --> 00:41:21.970 align:middle line:84%
And we call that random place
where we break it, we call it

00:41:21.970 --> 00:41:24.210 align:middle line:90%
X.

00:41:24.210 --> 00:41:26.670 align:middle line:84%
X can be anywhere, uniform
distribution.

00:41:26.670 --> 00:41:31.800 align:middle line:84%
So this means that X has a
density that goes from 0 to l.

00:41:31.800 --> 00:41:34.760 align:middle line:84%
I guess this capital L is
supposed to be the same as the

00:41:34.760 --> 00:41:36.190 align:middle line:90%
lower-case l.

00:41:36.190 --> 00:41:39.430 align:middle line:84%
So that's the density of X. And
since the density needs to

00:41:39.430 --> 00:41:43.160 align:middle line:84%
integrate to 1, the height of
that density has to be 1/l.

00:41:43.160 --> 00:41:46.330 align:middle line:90%


00:41:46.330 --> 00:41:49.660 align:middle line:84%
Now, having broken the stick
and given that we are left

00:41:49.660 --> 00:41:53.080 align:middle line:84%
with this piece of the stick,
I'm now going to break it

00:41:53.080 --> 00:41:56.900 align:middle line:84%
again at a completely random
place, meaning I'm going to

00:41:56.900 --> 00:41:59.940 align:middle line:84%
choose a point where I break it
uniformly over the length

00:41:59.940 --> 00:42:00.940 align:middle line:90%
of the stick.

00:42:00.940 --> 00:42:02.750 align:middle line:90%
What does this mean?

00:42:02.750 --> 00:42:05.720 align:middle line:84%
And let's call Y the location
where I break it.

00:42:05.720 --> 00:42:10.290 align:middle line:84%
So Y is going to range
between 0 and x.

00:42:10.290 --> 00:42:11.850 align:middle line:84%
x is the stick that
I'm left with.

00:42:11.850 --> 00:42:14.190 align:middle line:84%
So I'm going to break it
somewhere in between.

00:42:14.190 --> 00:42:21.140 align:middle line:90%
So I pick a y between 0 and x.

00:42:21.140 --> 00:42:24.480 align:middle line:84%
And of course, x
is less than l.

00:42:24.480 --> 00:42:26.150 align:middle line:84%
And I'm going to
break it there.

00:42:26.150 --> 00:42:30.640 align:middle line:84%
So y is uniform between
0 and x.

00:42:30.640 --> 00:42:36.460 align:middle line:84%
What does that mean, that the
density of y, given that you

00:42:36.460 --> 00:42:42.940 align:middle line:84%
have already told me x, ranges
from 0 to little x?

00:42:42.940 --> 00:42:46.170 align:middle line:84%
If I told you that the first
break happened at a particular

00:42:46.170 --> 00:42:50.850 align:middle line:84%
x, then y can only range
over this interval.

00:42:50.850 --> 00:42:52.830 align:middle line:90%
And I'm assuming a uniform

00:42:52.830 --> 00:42:54.330 align:middle line:90%
distribution over that interval.

00:42:54.330 --> 00:42:56.420 align:middle line:90%
So we have this kind of shape.

00:42:56.420 --> 00:43:00.700 align:middle line:84%
And that fixes for
us the height of

00:43:00.700 --> 00:43:01.950 align:middle line:90%
the conditional density.

00:43:01.950 --> 00:43:05.380 align:middle line:90%


00:43:05.380 --> 00:43:11.690 align:middle line:84%
So what's the joint density of
those two random variables?

00:43:11.690 --> 00:43:14.440 align:middle line:84%
By the definition of conditional
densities, the

00:43:14.440 --> 00:43:18.290 align:middle line:84%
conditional was defined as the
ratio of this divided by that.

00:43:18.290 --> 00:43:21.500 align:middle line:84%
So we can find the joint density
by taking the marginal

00:43:21.500 --> 00:43:23.630 align:middle line:84%
and then multiplying
by the conditional.

00:43:23.630 --> 00:43:26.120 align:middle line:84%
This is the same formula as
in the discrete case.

00:43:26.120 --> 00:43:29.770 align:middle line:84%
This is our very familiar
multiplication rule, but

00:43:29.770 --> 00:43:32.150 align:middle line:84%
adjusted to the case of
continuous random variables.

00:43:32.150 --> 00:43:34.871 align:middle line:90%
So Ps become Fs.

00:43:34.871 --> 00:43:35.290 align:middle line:90%
OK.

00:43:35.290 --> 00:43:37.560 align:middle line:84%
So we do have a formula
for this.

00:43:37.560 --> 00:43:38.540 align:middle line:90%
What is it?

00:43:38.540 --> 00:43:40.190 align:middle line:90%
It's 1/l--

00:43:40.190 --> 00:43:42.140 align:middle line:90%
that's the density of X --

00:43:42.140 --> 00:43:46.460 align:middle line:84%
times 1/x, which is the
conditional density of Y. This

00:43:46.460 --> 00:43:48.630 align:middle line:84%
is the formula for the
joint density.

00:43:48.630 --> 00:43:50.140 align:middle line:90%
But we must be careful.

00:43:50.140 --> 00:43:53.230 align:middle line:84%
This is a formula that's
not valid anywhere.

00:43:53.230 --> 00:43:57.150 align:middle line:84%
It's only valid for the x's
and y's that are possible.

00:43:57.150 --> 00:44:00.840 align:middle line:84%
And the x's and y's that are
possible are given by these

00:44:00.840 --> 00:44:01.900 align:middle line:90%
inequalities.

00:44:01.900 --> 00:44:05.940 align:middle line:84%
So x can range from 0 to
l, and y can only be

00:44:05.940 --> 00:44:07.270 align:middle line:90%
smaller than x.

00:44:07.270 --> 00:44:09.780 align:middle line:84%
So this is the formula
for the density on

00:44:09.780 --> 00:44:12.310 align:middle line:90%
this part of our space.

00:44:12.310 --> 00:44:16.270 align:middle line:84%
The density is 0
anywhere else.

00:44:16.270 --> 00:44:18.430 align:middle line:90%
So what does it look like?

00:44:18.430 --> 00:44:20.950 align:middle line:90%
It's basically a 1/x function.

00:44:20.950 --> 00:44:23.460 align:middle line:84%
So it's sort of constant
along that dimension.

00:44:23.460 --> 00:44:27.600 align:middle line:84%
But as x goes to 0, your
density goes up and

00:44:27.600 --> 00:44:29.280 align:middle line:90%
can even blow up.

00:44:29.280 --> 00:44:33.400 align:middle line:84%
It sort of looks like a sail
that's raised and somewhat

00:44:33.400 --> 00:44:37.640 align:middle line:84%
curved and has a point up
there going to infinity.

00:44:37.640 --> 00:44:39.680 align:middle line:90%
So this is the joint density.

00:44:39.680 --> 00:44:43.480 align:middle line:84%
Now once you have in your hands
a joint density, then

00:44:43.480 --> 00:44:46.010 align:middle line:84%
you can answer in principle
any problem.

00:44:46.010 --> 00:44:50.550 align:middle line:84%
It's just a matter of plugging
in and doing computations.

00:44:50.550 --> 00:44:53.650 align:middle line:84%
How about calculating something
like a conditional

00:44:53.650 --> 00:44:59.040 align:middle line:84%
expectation of Y given
a value of x?

00:44:59.040 --> 00:44:59.430 align:middle line:90%
OK.

00:44:59.430 --> 00:45:02.530 align:middle line:84%
That's a concept we have
not defined so far.

00:45:02.530 --> 00:45:04.860 align:middle line:90%
But how should we define it?

00:45:04.860 --> 00:45:06.080 align:middle line:90%
Means the reasonable thing.

00:45:06.080 --> 00:45:09.930 align:middle line:84%
We'll define it the same way
as ordinary expectations

00:45:09.930 --> 00:45:14.160 align:middle line:84%
except that since we're given
some conditioning information,

00:45:14.160 --> 00:45:17.130 align:middle line:84%
we should use the probability
distribution that applies to

00:45:17.130 --> 00:45:18.840 align:middle line:90%
that particular situation.

00:45:18.840 --> 00:45:22.570 align:middle line:84%
So in a situation where we are
told the value of x, the

00:45:22.570 --> 00:45:25.760 align:middle line:84%
distribution that applies is the
conditional distribution

00:45:25.760 --> 00:45:29.950 align:middle line:84%
of Y. So it's going to be the
conditional density of Y given

00:45:29.950 --> 00:45:31.470 align:middle line:90%
the value of x.

00:45:31.470 --> 00:45:34.120 align:middle line:90%
Now, we know what this is.

00:45:34.120 --> 00:45:37.860 align:middle line:90%
It's given by 1/x.

00:45:37.860 --> 00:45:46.160 align:middle line:84%
So we need to integrate
y times 1/x dy.

00:45:46.160 --> 00:45:48.920 align:middle line:84%
And what should we
integrate over?

00:45:48.920 --> 00:45:53.930 align:middle line:84%
Well, given the value of x, y
can only range from 0 to x.

00:45:53.930 --> 00:45:56.150 align:middle line:90%
So this is what we get.

00:45:56.150 --> 00:46:01.690 align:middle line:84%
And you do your integral, and
you get that this is x/2.

00:46:01.690 --> 00:46:03.060 align:middle line:90%
Is it a surprise?

00:46:03.060 --> 00:46:04.450 align:middle line:90%
It shouldn't be.

00:46:04.450 --> 00:46:10.890 align:middle line:84%
This is just the expected value
of Y in a universe where

00:46:10.890 --> 00:46:14.560 align:middle line:84%
X has been realized and Y is
given by this distribution.

00:46:14.560 --> 00:46:17.390 align:middle line:90%
Y is uniform between 0 and x.

00:46:17.390 --> 00:46:20.820 align:middle line:84%
The expected value of Y should
be the midpoint of this

00:46:20.820 --> 00:46:22.100 align:middle line:90%
interval, which is x/2.

00:46:22.100 --> 00:46:25.090 align:middle line:90%


00:46:25.090 --> 00:46:28.580 align:middle line:90%
Now let's do fancier stuff.

00:46:28.580 --> 00:46:31.850 align:middle line:84%
Since we have the joint
distribution, we should be

00:46:31.850 --> 00:46:34.250 align:middle line:84%
able to calculate
the marginal.

00:46:34.250 --> 00:46:36.500 align:middle line:90%
What is the distribution of Y?

00:46:36.500 --> 00:46:40.510 align:middle line:84%
After breaking the stick twice,
how big is the little

00:46:40.510 --> 00:46:42.890 align:middle line:90%
piece that I'm left with?

00:46:42.890 --> 00:46:44.630 align:middle line:90%
How do we find this?

00:46:44.630 --> 00:46:48.850 align:middle line:84%
To find the marginal, we just
take the joint and integrate

00:46:48.850 --> 00:46:52.670 align:middle line:84%
out the variable that
we don't want.

00:46:52.670 --> 00:46:55.220 align:middle line:84%
A particular y can happen
in many ways.

00:46:55.220 --> 00:46:57.800 align:middle line:84%
It can happen together
with any x.

00:46:57.800 --> 00:47:00.700 align:middle line:84%
So we consider all the possible
x's that can go

00:47:00.700 --> 00:47:05.940 align:middle line:84%
together with this y and average
over all those x's.

00:47:05.940 --> 00:47:09.330 align:middle line:84%
So we plug in the formula for
the joint density from the

00:47:09.330 --> 00:47:10.140 align:middle line:90%
previous slide.

00:47:10.140 --> 00:47:13.070 align:middle line:90%
We know that it's 1/lx.

00:47:13.070 --> 00:47:16.880 align:middle line:84%
And what's the range
of the x's?

00:47:16.880 --> 00:47:22.880 align:middle line:84%
So to find the density of Y for
a particular y up here,

00:47:22.880 --> 00:47:26.480 align:middle line:84%
I'm going to integrate
over x's.

00:47:26.480 --> 00:47:29.040 align:middle line:84%
The density is 0
here and there.

00:47:29.040 --> 00:47:32.160 align:middle line:84%
The density is nonzero
only in this part.

00:47:32.160 --> 00:47:37.260 align:middle line:84%
So I need to integrate over x's
going from here to there.

00:47:37.260 --> 00:47:39.120 align:middle line:90%
So what's the "here"?

00:47:39.120 --> 00:47:42.200 align:middle line:84%
This line goes up at
the slope of 1.

00:47:42.200 --> 00:47:45.420 align:middle line:84%
So this is the line
x equals y.

00:47:45.420 --> 00:47:49.835 align:middle line:84%
So if I fix y, it means that
my integral starts from a

00:47:49.835 --> 00:47:53.670 align:middle line:84%
value of x that is
also equal to y.

00:47:53.670 --> 00:47:58.330 align:middle line:84%
So where the integral starts
from is at x equals y.

00:47:58.330 --> 00:48:01.770 align:middle line:84%
And it goes all the way until
the end of the length of our

00:48:01.770 --> 00:48:03.660 align:middle line:90%
stick, which is l.

00:48:03.660 --> 00:48:08.760 align:middle line:84%
So we need to integrate
from little y up to l.

00:48:08.760 --> 00:48:12.520 align:middle line:84%
So that's something that
almost always comes up.

00:48:12.520 --> 00:48:15.690 align:middle line:84%
It's not enough to have just
this formula for integrating

00:48:15.690 --> 00:48:16.640 align:middle line:90%
the joint density.

00:48:16.640 --> 00:48:19.160 align:middle line:84%
You need to keep track
of different regions.

00:48:19.160 --> 00:48:23.920 align:middle line:84%
And if the joint density is 0
in some regions, then you

00:48:23.920 --> 00:48:28.250 align:middle line:84%
exclude those regions from
the range of integration.

00:48:28.250 --> 00:48:32.380 align:middle line:84%
So the range of integration is
only over those values where

00:48:32.380 --> 00:48:35.600 align:middle line:84%
the particular formula is valid,
the places where the

00:48:35.600 --> 00:48:37.990 align:middle line:90%
joint density is nonzero.

00:48:37.990 --> 00:48:38.360 align:middle line:90%
All right.

00:48:38.360 --> 00:48:41.760 align:middle line:84%
The integral of 1/x dx, that
gives you a logarithm.

00:48:41.760 --> 00:48:45.460 align:middle line:84%
So we evaluate this integral,
and we get an

00:48:45.460 --> 00:48:47.410 align:middle line:90%
expression of this kind.

00:48:47.410 --> 00:48:53.660 align:middle line:84%
So the density of Y has a
somewhat unexpected shape.

00:48:53.660 --> 00:48:55.470 align:middle line:84%
So it's a logarithmic
function.

00:48:55.470 --> 00:48:59.860 align:middle line:90%
And it goes this way.

00:48:59.860 --> 00:49:02.980 align:middle line:84%
It's for y going all
the way to l.

00:49:02.980 --> 00:49:07.860 align:middle line:84%
When y is equal to l, the
logarithm of 1 is equal to 0.

00:49:07.860 --> 00:49:12.660 align:middle line:84%
But when y approaches 0,
logarithm of something big

00:49:12.660 --> 00:49:15.740 align:middle line:84%
blows up, and we get a
shape of this form.

00:49:15.740 --> 00:49:21.900 align:middle line:90%


00:49:21.900 --> 00:49:22.330 align:middle line:90%
OK.

00:49:22.330 --> 00:49:25.960 align:middle line:84%
Finally, we can calculate the
expected value of Y. And we

00:49:25.960 --> 00:49:29.430 align:middle line:84%
can do this by using the
definition of the expectation.

00:49:29.430 --> 00:49:33.300 align:middle line:84%
So integral of y times
the density of y.

00:49:33.300 --> 00:49:36.290 align:middle line:84%
We already found what that
density is, so we

00:49:36.290 --> 00:49:38.030 align:middle line:90%
can plug it in here.

00:49:38.030 --> 00:49:40.470 align:middle line:84%
And we're integrating over
the range of possible

00:49:40.470 --> 00:49:42.470 align:middle line:90%
y's, from 0 to l.

00:49:42.470 --> 00:49:46.930 align:middle line:84%
Now this involves the integral
for y log y, which I'm sure

00:49:46.930 --> 00:49:49.500 align:middle line:84%
you have encountered in your
calculus classes but maybe do

00:49:49.500 --> 00:49:51.350 align:middle line:90%
not remember how to do it.

00:49:51.350 --> 00:49:53.650 align:middle line:84%
In any case, you look it
up in some integral

00:49:53.650 --> 00:49:55.300 align:middle line:90%
tables or do it by parts.

00:49:55.300 --> 00:49:59.360 align:middle line:84%
And you get the final
answer of l/4.

00:49:59.360 --> 00:50:02.400 align:middle line:84%
And at this point, you say,
that's a really simple answer.

00:50:02.400 --> 00:50:06.200 align:middle line:84%
Shouldn't I have expected
it to be l/4?

00:50:06.200 --> 00:50:07.680 align:middle line:90%
I guess, yes.

00:50:07.680 --> 00:50:11.070 align:middle line:84%
I mean, when you break it once,
the expected value of

00:50:11.070 --> 00:50:14.220 align:middle line:84%
what you are left with is going
to be 1/2 of what you

00:50:14.220 --> 00:50:15.860 align:middle line:90%
started with.

00:50:15.860 --> 00:50:19.320 align:middle line:84%
When you break it the next time,
the expected length of

00:50:19.320 --> 00:50:23.380 align:middle line:84%
what you're left with should be
1/2 of the piece that you

00:50:23.380 --> 00:50:24.550 align:middle line:90%
are now breaking.

00:50:24.550 --> 00:50:27.350 align:middle line:84%
So each time that you break it
at random, you expected it to

00:50:27.350 --> 00:50:29.840 align:middle line:84%
become smaller by
a factor of 1/2.

00:50:29.840 --> 00:50:31.960 align:middle line:84%
So if you break it twice, you
are left something that's

00:50:31.960 --> 00:50:33.940 align:middle line:90%
expected to be 1/4.

00:50:33.940 --> 00:50:37.350 align:middle line:84%
This is reasoning on the
average, which happens to give

00:50:37.350 --> 00:50:39.010 align:middle line:84%
you the right answer
in this case.

00:50:39.010 --> 00:50:41.800 align:middle line:84%
But again, there's the warning
that reasoning on the average

00:50:41.800 --> 00:50:44.230 align:middle line:84%
doesn't always give you
the right answer.

00:50:44.230 --> 00:50:48.100 align:middle line:84%
So be careful about doing
arguments of this type.

00:50:48.100 --> 00:50:48.620 align:middle line:90%
Very good.

00:50:48.620 --> 00:50:49.870 align:middle line:90%
See you on Wednesday.

00:50:49.870 --> 00:50:50.870 align:middle line:90%