WEBVTT

00:00:00.000 --> 00:00:00.040 align:middle line:90%


00:00:00.040 --> 00:00:02.460 align:middle line:84%
The following content is
provided under a Creative

00:00:02.460 --> 00:00:03.870 align:middle line:90%
Commons license.

00:00:03.870 --> 00:00:06.910 align:middle line:84%
Your support will help MIT
OpenCourseWare continue to

00:00:06.910 --> 00:00:10.560 align:middle line:84%
offer high quality educational
resources for free.

00:00:10.560 --> 00:00:13.460 align:middle line:84%
To make a donation, or view
additional materials from

00:00:13.460 --> 00:00:17.390 align:middle line:84%
hundreds of MIT courses, visit
MIT OpenCourseWare at

00:00:17.390 --> 00:00:22.620 align:middle line:90%
ocw.mit.edu,

00:00:22.620 --> 00:00:23.380 align:middle line:90%
OK.

00:00:23.380 --> 00:00:24.630 align:middle line:90%
So let us start.

00:00:24.630 --> 00:00:55.510 align:middle line:90%


00:00:55.510 --> 00:00:55.890 align:middle line:90%
All right.

00:00:55.890 --> 00:01:00.230 align:middle line:84%
So today we're starting a
new unit in this class.

00:01:00.230 --> 00:01:03.390 align:middle line:84%
We have covered, so far, the
basics of probability theory--

00:01:03.390 --> 00:01:06.890 align:middle line:84%
the main concepts and tools, as
far as just probabilities

00:01:06.890 --> 00:01:08.150 align:middle line:90%
are concerned.

00:01:08.150 --> 00:01:11.300 align:middle line:84%
But if that was all that there
is in this subject, the

00:01:11.300 --> 00:01:13.060 align:middle line:84%
subject would not
be rich enough.

00:01:13.060 --> 00:01:16.070 align:middle line:84%
What makes probability theory
a lot more interesting and

00:01:16.070 --> 00:01:19.590 align:middle line:84%
richer is that we can also talk
about random variables,

00:01:19.590 --> 00:01:25.230 align:middle line:84%
which are ways of assigning
numerical results to the

00:01:25.230 --> 00:01:27.430 align:middle line:90%
outcomes of an experiment.

00:01:27.430 --> 00:01:32.500 align:middle line:84%
So we're going to define what
random variables are, and then

00:01:32.500 --> 00:01:35.560 align:middle line:84%
we're going to describe them
using so-called probability

00:01:35.560 --> 00:01:36.410 align:middle line:90%
mass functions.

00:01:36.410 --> 00:01:39.770 align:middle line:84%
Basically some numerical values
are more likely to

00:01:39.770 --> 00:01:43.260 align:middle line:84%
occur than other numerical
values, and we capture this by

00:01:43.260 --> 00:01:46.260 align:middle line:84%
assigning probabilities
to them the usual way.

00:01:46.260 --> 00:01:49.870 align:middle line:84%
And we represent these in
a compact way using the

00:01:49.870 --> 00:01:51.950 align:middle line:84%
so-called probability
mass functions.

00:01:51.950 --> 00:01:55.340 align:middle line:84%
We're going to see a couple of
examples of random variables,

00:01:55.340 --> 00:01:58.260 align:middle line:84%
some of which we have already
seen but with different

00:01:58.260 --> 00:01:59.930 align:middle line:90%
terminology.

00:01:59.930 --> 00:02:04.950 align:middle line:84%
And so far, it's going to be
just a couple of definitions

00:02:04.950 --> 00:02:06.790 align:middle line:84%
and calculations of
the type that you

00:02:06.790 --> 00:02:08.810 align:middle line:90%
already know how to do.

00:02:08.810 --> 00:02:11.370 align:middle line:84%
But then we're going to
introduce the one new, big

00:02:11.370 --> 00:02:12.980 align:middle line:90%
concept of the day.

00:02:12.980 --> 00:02:17.040 align:middle line:84%
So up to here it's going to be
mostly an exercise in notation

00:02:17.040 --> 00:02:18.190 align:middle line:90%
and definitions.

00:02:18.190 --> 00:02:20.850 align:middle line:84%
But then we got our big concept
which is the concept

00:02:20.850 --> 00:02:24.190 align:middle line:84%
of the expected value of a
random variable, which is some

00:02:24.190 --> 00:02:27.290 align:middle line:84%
kind of average value of
the random variable.

00:02:27.290 --> 00:02:30.260 align:middle line:84%
And then we're going to also
talk, very briefly, about

00:02:30.260 --> 00:02:33.560 align:middle line:84%
close distance of the
expectation, which is the

00:02:33.560 --> 00:02:37.010 align:middle line:84%
concept of the variance
of a random variable.

00:02:37.010 --> 00:02:37.910 align:middle line:90%
OK.

00:02:37.910 --> 00:02:40.455 align:middle line:90%
So what is a random variable?

00:02:40.455 --> 00:02:43.860 align:middle line:90%


00:02:43.860 --> 00:02:47.710 align:middle line:84%
It's an assignment of a
numerical value to every

00:02:47.710 --> 00:02:49.950 align:middle line:84%
possible outcome of
the experiment.

00:02:49.950 --> 00:02:51.430 align:middle line:90%
So here's the picture.

00:02:51.430 --> 00:02:54.800 align:middle line:84%
The sample space is this class,
and we've got lots of

00:02:54.800 --> 00:02:56.900 align:middle line:90%
students in here.

00:02:56.900 --> 00:03:00.480 align:middle line:84%
This is our sample
space, omega.

00:03:00.480 --> 00:03:04.220 align:middle line:84%
I'm interested in the height
of a random student.

00:03:04.220 --> 00:03:08.990 align:middle line:84%
So I'm going to use a real line
where I record height,

00:03:08.990 --> 00:03:12.490 align:middle line:84%
and let's say this is
height in inches.

00:03:12.490 --> 00:03:16.380 align:middle line:84%
And the experiment happens,
I pick a random student.

00:03:16.380 --> 00:03:19.510 align:middle line:84%
And I go and measure the height
of that random student,

00:03:19.510 --> 00:03:22.200 align:middle line:84%
and that gives me a
specific number.

00:03:22.200 --> 00:03:25.310 align:middle line:84%
So what's a good number
in inches?

00:03:25.310 --> 00:03:28.430 align:middle line:90%
Let's say 60.

00:03:28.430 --> 00:03:28.840 align:middle line:90%
OK.

00:03:28.840 --> 00:03:33.480 align:middle line:84%
Or I pick another student, and
that student has a height of

00:03:33.480 --> 00:03:36.030 align:middle line:90%
71 inches, and so on.

00:03:36.030 --> 00:03:37.420 align:middle line:90%
So this is the experiment.

00:03:37.420 --> 00:03:39.020 align:middle line:90%
These are the outcomes.

00:03:39.020 --> 00:03:43.070 align:middle line:84%
These are the numerical values
of the random variable that we

00:03:43.070 --> 00:03:46.340 align:middle line:90%
call height.

00:03:46.340 --> 00:03:46.690 align:middle line:90%
OK.

00:03:46.690 --> 00:03:49.770 align:middle line:84%
So mathematically, what are
we dealing with here?

00:03:49.770 --> 00:03:54.100 align:middle line:84%
We're basically dealing with a
function from the sample space

00:03:54.100 --> 00:03:56.230 align:middle line:90%
into the real numbers.

00:03:56.230 --> 00:04:01.280 align:middle line:84%
That function takes as argument,
an outcome of the

00:04:01.280 --> 00:04:04.720 align:middle line:84%
experiment, that is a typical
student, and produces the

00:04:04.720 --> 00:04:07.520 align:middle line:84%
value of that function, which
is the height of that

00:04:07.520 --> 00:04:09.100 align:middle line:90%
particular student.

00:04:09.100 --> 00:04:12.600 align:middle line:84%
So we think of an abstract
object that we denote by a

00:04:12.600 --> 00:04:17.480 align:middle line:84%
capital H, which is the random
variable called height.

00:04:17.480 --> 00:04:21.279 align:middle line:84%
And that random variable is
essentially this particular

00:04:21.279 --> 00:04:25.870 align:middle line:84%
function that we talked
about here.

00:04:25.870 --> 00:04:26.490 align:middle line:90%
OK.

00:04:26.490 --> 00:04:29.190 align:middle line:84%
So there's a distinction that
we're making here--

00:04:29.190 --> 00:04:32.170 align:middle line:90%
H is height in the abstract.

00:04:32.170 --> 00:04:33.570 align:middle line:90%
It's the function.

00:04:33.570 --> 00:04:36.440 align:middle line:84%
These numbers here are
particular numerical values

00:04:36.440 --> 00:04:39.580 align:middle line:84%
that this function takes when
you choose one particular

00:04:39.580 --> 00:04:41.400 align:middle line:90%
outcome of the experiment.

00:04:41.400 --> 00:04:44.910 align:middle line:84%
Now, when you have a single
probability experiment, you

00:04:44.910 --> 00:04:47.690 align:middle line:84%
can have multiple random
variables.

00:04:47.690 --> 00:04:52.690 align:middle line:84%
So perhaps, instead of just
height, I'm also interested in

00:04:52.690 --> 00:04:55.890 align:middle line:84%
the weight of a typical
student.

00:04:55.890 --> 00:04:58.680 align:middle line:84%
And so when the experiment
happens, I

00:04:58.680 --> 00:05:00.690 align:middle line:90%
pick that random student--

00:05:00.690 --> 00:05:02.250 align:middle line:84%
this is the height
of the student.

00:05:02.250 --> 00:05:05.450 align:middle line:84%
But that student would also
have a weight, and I could

00:05:05.450 --> 00:05:06.880 align:middle line:90%
record it here.

00:05:06.880 --> 00:05:09.570 align:middle line:84%
And similarly, every student
is going to have their own

00:05:09.570 --> 00:05:10.910 align:middle line:90%
particular weight.

00:05:10.910 --> 00:05:13.850 align:middle line:84%
So the weight function is a
different function from the

00:05:13.850 --> 00:05:17.330 align:middle line:84%
sample space to the real
numbers, and it's a different

00:05:17.330 --> 00:05:18.680 align:middle line:90%
random variable.

00:05:18.680 --> 00:05:21.840 align:middle line:84%
So the point I'm making here is
that a single probabilistic

00:05:21.840 --> 00:05:26.825 align:middle line:84%
experiment may involve several
interesting random variables.

00:05:26.825 --> 00:05:30.190 align:middle line:84%
I may be interested in the
height of a random student or

00:05:30.190 --> 00:05:31.760 align:middle line:84%
the weight of the
random student.

00:05:31.760 --> 00:05:33.300 align:middle line:84%
These are different
random variables

00:05:33.300 --> 00:05:35.140 align:middle line:90%
that could be of interest.

00:05:35.140 --> 00:05:37.580 align:middle line:90%
I can also do other things.

00:05:37.580 --> 00:05:44.000 align:middle line:84%
Suppose I define an object such
as H bar, which is 2.58.

00:05:44.000 --> 00:05:46.420 align:middle line:90%
What does that correspond to?

00:05:46.420 --> 00:05:50.540 align:middle line:84%
Well, this is the height
in centimeters.

00:05:50.540 --> 00:05:55.160 align:middle line:84%
Now, H bar is a function of H
itself, but if you were to

00:05:55.160 --> 00:05:57.740 align:middle line:84%
draw the picture, the picture
would go this way.

00:05:57.740 --> 00:06:03.100 align:middle line:84%
60 gets mapped to 150, 71 gets
mapped to, oh, that's

00:06:03.100 --> 00:06:04.720 align:middle line:90%
too hard for me.

00:06:04.720 --> 00:06:10.040 align:middle line:84%
OK, gets mapped to something,
and so on.

00:06:10.040 --> 00:06:14.220 align:middle line:84%
So H bar is also a
random variable.

00:06:14.220 --> 00:06:15.100 align:middle line:90%
Why?

00:06:15.100 --> 00:06:19.140 align:middle line:84%
Once I pick a particular
student, that particular

00:06:19.140 --> 00:06:24.120 align:middle line:84%
outcome determines completely
the numerical value of H bar,

00:06:24.120 --> 00:06:29.070 align:middle line:84%
which is the height of that
student but measured in

00:06:29.070 --> 00:06:31.080 align:middle line:90%
centimeters.

00:06:31.080 --> 00:06:33.820 align:middle line:84%
What we have here is actually
a random variable, which is

00:06:33.820 --> 00:06:37.860 align:middle line:84%
defined as a function of another
random variable.

00:06:37.860 --> 00:06:41.010 align:middle line:84%
And the point that this example
is trying to make is

00:06:41.010 --> 00:06:42.640 align:middle line:84%
that functions of
random variables

00:06:42.640 --> 00:06:44.630 align:middle line:90%
are also random variables.

00:06:44.630 --> 00:06:47.390 align:middle line:84%
The experiment happens, the
experiment determines a

00:06:47.390 --> 00:06:49.410 align:middle line:84%
numerical value for
this object.

00:06:49.410 --> 00:06:51.630 align:middle line:84%
And once you have the numerical
value for this

00:06:51.630 --> 00:06:54.180 align:middle line:84%
object, that determines
also the numerical

00:06:54.180 --> 00:06:55.900 align:middle line:90%
value for that object.

00:06:55.900 --> 00:06:59.080 align:middle line:84%
So given an outcome, the
numerical value of this

00:06:59.080 --> 00:07:00.840 align:middle line:84%
particular object
is determined.

00:07:00.840 --> 00:07:05.730 align:middle line:84%
So H bar is itself a function
from the sample space, from

00:07:05.730 --> 00:07:08.340 align:middle line:90%
outcomes to numerical values.

00:07:08.340 --> 00:07:11.350 align:middle line:84%
And that makes it a random
variable according to the

00:07:11.350 --> 00:07:13.920 align:middle line:84%
formal definition that
we have here.

00:07:13.920 --> 00:07:18.610 align:middle line:84%
So the formal definition is that
the random variable is

00:07:18.610 --> 00:07:22.570 align:middle line:84%
not random, it's not a variable,
it's just a function

00:07:22.570 --> 00:07:25.000 align:middle line:84%
from the sample space
to the real numbers.

00:07:25.000 --> 00:07:29.220 align:middle line:84%
That's the abstract, right way
of thinking about them.

00:07:29.220 --> 00:07:32.000 align:middle line:84%
Now, random variables can
be of different types.

00:07:32.000 --> 00:07:34.440 align:middle line:84%
They can be discrete
or continuous.

00:07:34.440 --> 00:07:38.490 align:middle line:84%
Suppose that I measure the
heights in inches, but I round

00:07:38.490 --> 00:07:40.270 align:middle line:90%
to the nearest inch.

00:07:40.270 --> 00:07:43.490 align:middle line:84%
Then the numerical values I'm
going to get here would be

00:07:43.490 --> 00:07:45.000 align:middle line:90%
just integers.

00:07:45.000 --> 00:07:47.080 align:middle line:84%
So that would make
it an integer

00:07:47.080 --> 00:07:48.520 align:middle line:90%
valued random variable.

00:07:48.520 --> 00:07:50.870 align:middle line:84%
And this is a discrete
random variable.

00:07:50.870 --> 00:07:54.950 align:middle line:84%
Or maybe I have a scale for
measuring height which is

00:07:54.950 --> 00:07:58.640 align:middle line:84%
infinitely precise and records
your height to an infinite

00:07:58.640 --> 00:08:00.830 align:middle line:90%
number of digits of precision.

00:08:00.830 --> 00:08:03.390 align:middle line:84%
In that case, your height
would be just a

00:08:03.390 --> 00:08:05.430 align:middle line:90%
general real number.

00:08:05.430 --> 00:08:08.510 align:middle line:84%
So we would have a random
variable that takes values in

00:08:08.510 --> 00:08:10.720 align:middle line:84%
the entire set of
real numbers.

00:08:10.720 --> 00:08:14.490 align:middle line:84%
Well, I guess not really
negative numbers, but the set

00:08:14.490 --> 00:08:16.350 align:middle line:90%
of non-negative numbers.

00:08:16.350 --> 00:08:19.880 align:middle line:84%
And that would be a continuous
random variable.

00:08:19.880 --> 00:08:22.250 align:middle line:84%
It takes values in
a continuous set.

00:08:22.250 --> 00:08:25.020 align:middle line:84%
So we will be talking about both
discrete and continuous

00:08:25.020 --> 00:08:26.330 align:middle line:90%
random variables.

00:08:26.330 --> 00:08:28.790 align:middle line:84%
The first thing we will do
will be to devote a few

00:08:28.790 --> 00:08:32.030 align:middle line:84%
lectures on discrete random
variables, because discrete is

00:08:32.030 --> 00:08:33.380 align:middle line:90%
always easier.

00:08:33.380 --> 00:08:35.640 align:middle line:84%
And then we're going to
repeat everything in

00:08:35.640 --> 00:08:37.710 align:middle line:90%
the continuous setting.

00:08:37.710 --> 00:08:41.760 align:middle line:84%
So discrete is easier, and
it's the right place to

00:08:41.760 --> 00:08:45.140 align:middle line:84%
understand all the concepts,
even those who may appear to

00:08:45.140 --> 00:08:47.090 align:middle line:90%
be elementary.

00:08:47.090 --> 00:08:50.110 align:middle line:84%
And then you will be set to
understand what's going on

00:08:50.110 --> 00:08:51.650 align:middle line:84%
when we go to the
continuous case.

00:08:51.650 --> 00:08:54.260 align:middle line:84%
So in the continuous case, you
get all the complications of

00:08:54.260 --> 00:08:57.840 align:middle line:84%
calculus and some extra math
that comes in there.

00:08:57.840 --> 00:09:00.910 align:middle line:84%
So it's important to have been
down all the concepts very

00:09:00.910 --> 00:09:04.320 align:middle line:84%
well in the easy, discrete case
so that you don't have

00:09:04.320 --> 00:09:06.770 align:middle line:84%
conceptual hurdles when
you move on to

00:09:06.770 --> 00:09:08.730 align:middle line:90%
the continuous case.

00:09:08.730 --> 00:09:13.610 align:middle line:84%
Now, one important remark that
may seem trivial but it's

00:09:13.610 --> 00:09:17.920 align:middle line:84%
actually very important so that
you don't get tangled up

00:09:17.920 --> 00:09:20.550 align:middle line:84%
between different types
of concepts--

00:09:20.550 --> 00:09:23.440 align:middle line:84%
there's a fundamental
distinction between the random

00:09:23.440 --> 00:09:26.220 align:middle line:84%
variable itself, and
the numerical

00:09:26.220 --> 00:09:29.080 align:middle line:90%
values that it takes.

00:09:29.080 --> 00:09:31.670 align:middle line:84%
Abstractly speaking, or
mathematically speaking, a

00:09:31.670 --> 00:09:38.350 align:middle line:84%
random variable, x, or H in this
example, is a function.

00:09:38.350 --> 00:09:39.140 align:middle line:90%
OK.

00:09:39.140 --> 00:09:43.750 align:middle line:84%
Maybe if you like programming
the words "procedure" or

00:09:43.750 --> 00:09:45.820 align:middle line:90%
"sub-routine" might be better.

00:09:45.820 --> 00:09:48.290 align:middle line:84%
So what's the sub-routine
height?

00:09:48.290 --> 00:09:51.470 align:middle line:84%
Given a student, I take that
student, force them on the

00:09:51.470 --> 00:09:53.110 align:middle line:90%
scale and measure them.

00:09:53.110 --> 00:09:56.610 align:middle line:84%
That's the sub-routine that
measures heights.

00:09:56.610 --> 00:10:00.270 align:middle line:84%
It's really a function that
takes students as input and

00:10:00.270 --> 00:10:02.670 align:middle line:90%
produces numbers as output.

00:10:02.670 --> 00:10:05.790 align:middle line:84%
The sub-routine we denoted
by capital H.

00:10:05.790 --> 00:10:07.450 align:middle line:90%
That's the random variable.

00:10:07.450 --> 00:10:10.500 align:middle line:84%
But once you plug in a
particular student into that

00:10:10.500 --> 00:10:14.460 align:middle line:84%
sub-routine, you end up getting
a particular number.

00:10:14.460 --> 00:10:17.610 align:middle line:84%
This is the numerical output
of that sub-routine or the

00:10:17.610 --> 00:10:19.900 align:middle line:84%
numerical value of
that function.

00:10:19.900 --> 00:10:24.390 align:middle line:84%
And that numerical value is an
element of the real numbers.

00:10:24.390 --> 00:10:29.040 align:middle line:84%
So the numerical value is a
real number, where this

00:10:29.040 --> 00:10:35.670 align:middle line:84%
capital X is a function from
omega to the real numbers.

00:10:35.670 --> 00:10:38.400 align:middle line:84%
So they are very different
types of objects.

00:10:38.400 --> 00:10:41.510 align:middle line:84%
And the way that we keep track
of what we're talking about at

00:10:41.510 --> 00:10:45.020 align:middle line:84%
any given time is by using
capital letters for random

00:10:45.020 --> 00:10:49.150 align:middle line:84%
variables and lower case
letters for numbers.

00:10:49.150 --> 00:10:52.260 align:middle line:90%


00:10:52.260 --> 00:10:52.700 align:middle line:90%
OK.

00:10:52.700 --> 00:11:00.520 align:middle line:84%
So now once we have a random
variable at hand, that random

00:11:00.520 --> 00:11:04.410 align:middle line:84%
variable takes on different
numerical values.

00:11:04.410 --> 00:11:08.980 align:middle line:84%
And we want to describe to say
something about the relative

00:11:08.980 --> 00:11:12.760 align:middle line:84%
likelihoods of the different
numerical values that the

00:11:12.760 --> 00:11:15.350 align:middle line:90%
random variable can take.

00:11:15.350 --> 00:11:21.300 align:middle line:84%
So here's our sample space,
and here's the real line.

00:11:21.300 --> 00:11:23.850 align:middle line:90%


00:11:23.850 --> 00:11:28.900 align:middle line:84%
And there's a bunch of outcomes
that gave rise to one

00:11:28.900 --> 00:11:30.940 align:middle line:90%
particular numerical value.

00:11:30.940 --> 00:11:33.660 align:middle line:84%
There's another numerical value
that arises if we have

00:11:33.660 --> 00:11:34.270 align:middle line:90%
this outcome.

00:11:34.270 --> 00:11:37.620 align:middle line:84%
There's another numerical value
that arises if we have

00:11:37.620 --> 00:11:38.390 align:middle line:90%
this outcome.

00:11:38.390 --> 00:11:40.430 align:middle line:90%
So our sample space is here.

00:11:40.430 --> 00:11:42.530 align:middle line:90%
The real numbers are here.

00:11:42.530 --> 00:11:46.600 align:middle line:84%
And what we want to do is to ask
the question, how likely

00:11:46.600 --> 00:11:50.000 align:middle line:84%
is that particular numerical
value to occur?

00:11:50.000 --> 00:11:53.620 align:middle line:84%
So what we're essentially asking
is, how likely is it

00:11:53.620 --> 00:11:57.450 align:middle line:84%
that we obtain an outcome that
leads to that particular

00:11:57.450 --> 00:11:59.000 align:middle line:90%
numerical value?

00:11:59.000 --> 00:12:02.680 align:middle line:84%
We calculate that overall
probability of that numerical

00:12:02.680 --> 00:12:07.810 align:middle line:84%
value and we represent that
probability using a bar so

00:12:07.810 --> 00:12:13.300 align:middle line:84%
that we end up generating
a bar graph.

00:12:13.300 --> 00:12:16.550 align:middle line:84%
So that could be a possible
bar graph

00:12:16.550 --> 00:12:19.210 align:middle line:90%
associated with this picture.

00:12:19.210 --> 00:12:22.860 align:middle line:84%
The size of this bar is the
total probability that our

00:12:22.860 --> 00:12:27.590 align:middle line:84%
random variable took on this
numerical value, which is just

00:12:27.590 --> 00:12:32.240 align:middle line:84%
the sum of the probabilities of
the different outcomes that

00:12:32.240 --> 00:12:34.310 align:middle line:90%
led to that numerical value.

00:12:34.310 --> 00:12:37.100 align:middle line:84%
So the thing that we're plotting
here, the bar graph--

00:12:37.100 --> 00:12:39.370 align:middle line:90%
we give a name to it.

00:12:39.370 --> 00:12:43.910 align:middle line:84%
It's a function, which we denote
by lowercase b, capital

00:12:43.910 --> 00:12:47.690 align:middle line:84%
X. The capital X indicates which
random variable we're

00:12:47.690 --> 00:12:48.920 align:middle line:90%
talking about.

00:12:48.920 --> 00:12:54.670 align:middle line:84%
And it's a function of little
x, which is the range of

00:12:54.670 --> 00:12:58.530 align:middle line:84%
values that our random
variable is taking.

00:12:58.530 --> 00:13:04.830 align:middle line:84%
So in mathematical notation, the
value of the PMF at some

00:13:04.830 --> 00:13:09.220 align:middle line:84%
particular number, little x,
is the probability that our

00:13:09.220 --> 00:13:14.060 align:middle line:84%
random variable takes on the
numerical value, little x.

00:13:14.060 --> 00:13:17.510 align:middle line:84%
And if you want to be precise
about what this means, it's

00:13:17.510 --> 00:13:23.020 align:middle line:84%
the overall probability of all
outcomes for which the random

00:13:23.020 --> 00:13:26.770 align:middle line:84%
variable ends up taking
that value, little x.

00:13:26.770 --> 00:13:34.110 align:middle line:84%
So this is the overall
probability of all omegas that

00:13:34.110 --> 00:13:36.950 align:middle line:84%
lead to that particular
numerical

00:13:36.950 --> 00:13:39.260 align:middle line:90%
value, x, of interest.

00:13:39.260 --> 00:13:44.630 align:middle line:90%
So what do we know about PMFs?

00:13:44.630 --> 00:13:47.600 align:middle line:84%
Since there are probabilities,
all these entries in the bar

00:13:47.600 --> 00:13:49.880 align:middle line:90%
graph have to be non-negative.

00:13:49.880 --> 00:13:54.610 align:middle line:84%
Also, if you exhaust all the
possible values of little x's,

00:13:54.610 --> 00:13:57.840 align:middle line:84%
you will have exhausted all the
possible outcomes here.

00:13:57.840 --> 00:14:01.030 align:middle line:84%
Because every outcome leads
to some particular x.

00:14:01.030 --> 00:14:03.160 align:middle line:84%
So the sum of these
probabilities

00:14:03.160 --> 00:14:04.760 align:middle line:90%
should be equal to one.

00:14:04.760 --> 00:14:06.890 align:middle line:84%
This is the second
relation here.

00:14:06.890 --> 00:14:10.970 align:middle line:84%
So this relation tell
us that some little

00:14:10.970 --> 00:14:13.150 align:middle line:90%
x is going to happen.

00:14:13.150 --> 00:14:15.500 align:middle line:84%
They happen with different
probabilities, but when you

00:14:15.500 --> 00:14:19.370 align:middle line:84%
consider all the possible little
x's together, one of

00:14:19.370 --> 00:14:21.750 align:middle line:84%
those little x's is going
to be realized.

00:14:21.750 --> 00:14:25.640 align:middle line:84%
Probabilities need
to add to one.

00:14:25.640 --> 00:14:26.090 align:middle line:90%
OK.

00:14:26.090 --> 00:14:31.200 align:middle line:84%
So let's get our first example
of a non-trivial bar graph.

00:14:31.200 --> 00:14:35.780 align:middle line:84%
Consider the experiment where
I start with a coin and I

00:14:35.780 --> 00:14:38.270 align:middle line:84%
start flipping it
over and over.

00:14:38.270 --> 00:14:42.670 align:middle line:84%
And I do this until I obtain
heads for the first time.

00:14:42.670 --> 00:14:45.610 align:middle line:84%
So what are possible outcomes
of this experiment?

00:14:45.610 --> 00:14:48.850 align:middle line:84%
One possible outcome is that
I obtain heads at the first

00:14:48.850 --> 00:14:50.930 align:middle line:90%
toss, and then I stop.

00:14:50.930 --> 00:14:55.040 align:middle line:84%
In this case, my random variable
takes the value 1.

00:14:55.040 --> 00:14:59.360 align:middle line:84%
Or it's possible that I obtain
tails and then heads.

00:14:59.360 --> 00:15:02.500 align:middle line:84%
How many tosses did it take
until heads appeared?

00:15:02.500 --> 00:15:04.390 align:middle line:90%
This would be x equals to 2.

00:15:04.390 --> 00:15:09.930 align:middle line:84%
Or more generally, I might
obtain tails for k minus 1

00:15:09.930 --> 00:15:15.000 align:middle line:84%
times, and then obtain heads
at the k-th time, in which

00:15:15.000 --> 00:15:19.930 align:middle line:84%
case, our random variable takes
the value, little k.

00:15:19.930 --> 00:15:21.210 align:middle line:90%
So that's the experiment.

00:15:21.210 --> 00:15:25.350 align:middle line:84%
So capital X is a well defined
random variable.

00:15:25.350 --> 00:15:29.070 align:middle line:84%
It's the number of tosses it
takes until I see heads for

00:15:29.070 --> 00:15:30.710 align:middle line:90%
the first time.

00:15:30.710 --> 00:15:32.330 align:middle line:84%
These are the possible
outcomes.

00:15:32.330 --> 00:15:34.710 align:middle line:84%
These are elements of
our sample space.

00:15:34.710 --> 00:15:38.750 align:middle line:84%
And these are the values of X
depending on the outcome.

00:15:38.750 --> 00:15:43.950 align:middle line:84%
Clearly X is a function
of the outcome.

00:15:43.950 --> 00:15:47.570 align:middle line:84%
You tell me the outcome, I'm
going to tell you what X is.

00:15:47.570 --> 00:15:54.520 align:middle line:84%
So what we want to do now is to
calculate the PMF of X. So

00:15:54.520 --> 00:15:59.210 align:middle line:84%
Px of k is, by definition, the
probability that our random

00:15:59.210 --> 00:16:02.810 align:middle line:90%
variable takes the value k.

00:16:02.810 --> 00:16:07.250 align:middle line:84%
For the random variable to take
the value of k, the first

00:16:07.250 --> 00:16:09.680 align:middle line:90%
head appears at toss number k.

00:16:09.680 --> 00:16:13.550 align:middle line:84%
The only way that this event
can happen is if we obtain

00:16:13.550 --> 00:16:15.650 align:middle line:90%
this sequence of events.

00:16:15.650 --> 00:16:19.290 align:middle line:84%
T's the first k minus
1 times, tails, and

00:16:19.290 --> 00:16:21.820 align:middle line:90%
heads at the k-th flip.

00:16:21.820 --> 00:16:25.980 align:middle line:84%
So this event, that the random
variable is equal to k, is the

00:16:25.980 --> 00:16:30.590 align:middle line:84%
same as this event, k minus 1
tails followed by 1 head.

00:16:30.590 --> 00:16:32.780 align:middle line:84%
What's the probability
of that event?

00:16:32.780 --> 00:16:36.450 align:middle line:84%
We're assuming that the coin
tosses are independent.

00:16:36.450 --> 00:16:39.190 align:middle line:84%
So to find the probability
of this event, we need to

00:16:39.190 --> 00:16:41.860 align:middle line:84%
multiply the probability of
tails, times the probability

00:16:41.860 --> 00:16:43.680 align:middle line:84%
of tails, times the probability
of tails.

00:16:43.680 --> 00:16:47.370 align:middle line:84%
We multiply k minus one times,
times the probability of

00:16:47.370 --> 00:16:50.660 align:middle line:84%
heads, which puts an
extra p at the end.

00:16:50.660 --> 00:16:56.470 align:middle line:84%
And this is the formula for the
so-called geometric PMF.

00:16:56.470 --> 00:16:58.650 align:middle line:84%
And why do we call
it geometric?

00:16:58.650 --> 00:17:04.859 align:middle line:84%
Because if you go and plot the
bar graph of this random

00:17:04.859 --> 00:17:10.510 align:middle line:84%
variable, X, we start
at 1 with a certain

00:17:10.510 --> 00:17:14.300 align:middle line:90%
number, which is p.

00:17:14.300 --> 00:17:20.550 align:middle line:90%
And then at 2 we get p(1-p).

00:17:20.550 --> 00:17:23.640 align:middle line:84%
At 3 we're going to get
something smaller, it's p

00:17:23.640 --> 00:17:25.900 align:middle line:90%
times (1-p)-squared.

00:17:25.900 --> 00:17:29.740 align:middle line:84%
And the bars keep going down
at the rate of geometric

00:17:29.740 --> 00:17:30.730 align:middle line:90%
progression.

00:17:30.730 --> 00:17:34.490 align:middle line:84%
Each bar is smaller than the
previous bar, because each

00:17:34.490 --> 00:17:38.380 align:middle line:84%
time we get an extra factor
of 1-p involved.

00:17:38.380 --> 00:17:42.480 align:middle line:84%
So the shape of this
PMF is the graph

00:17:42.480 --> 00:17:44.300 align:middle line:90%
of a geometric sequence.

00:17:44.300 --> 00:17:48.330 align:middle line:84%
For that reason, we say that
it's the geometric PMF, and we

00:17:48.330 --> 00:17:51.860 align:middle line:84%
call X also a geometric
random variable.

00:17:51.860 --> 00:17:55.730 align:middle line:84%
So the number of coin tosses
until the first head is a

00:17:55.730 --> 00:17:58.290 align:middle line:90%
geometric random variable.

00:17:58.290 --> 00:18:00.730 align:middle line:84%
So this was an example
of how to compute the

00:18:00.730 --> 00:18:02.630 align:middle line:90%
PMF of a random variable.

00:18:02.630 --> 00:18:06.510 align:middle line:84%
This was an easy example,
because this event could be

00:18:06.510 --> 00:18:09.520 align:middle line:84%
realized in one and
only one way.

00:18:09.520 --> 00:18:12.650 align:middle line:84%
So to find the probability of
this, we just needed to find

00:18:12.650 --> 00:18:15.510 align:middle line:84%
the probability of this
particular outcome.

00:18:15.510 --> 00:18:18.680 align:middle line:84%
More generally, there's going
to be many outcomes that can

00:18:18.680 --> 00:18:22.120 align:middle line:84%
lead to the same numerical
value.

00:18:22.120 --> 00:18:25.010 align:middle line:84%
And we need to keep track
of all of them.

00:18:25.010 --> 00:18:28.030 align:middle line:84%
For example, in this picture,
if I want to find this value

00:18:28.030 --> 00:18:31.610 align:middle line:84%
of the PMF, I need to add up the
probabilities of all the

00:18:31.610 --> 00:18:34.550 align:middle line:84%
outcomes that leads
to that value.

00:18:34.550 --> 00:18:37.070 align:middle line:84%
So the general procedure
is exactly what

00:18:37.070 --> 00:18:38.240 align:middle line:90%
this picture suggests.

00:18:38.240 --> 00:18:43.050 align:middle line:84%
To find this probability, you go
and identify which outcomes

00:18:43.050 --> 00:18:47.770 align:middle line:84%
lead to this numerical value,
and add their probabilities.

00:18:47.770 --> 00:18:49.590 align:middle line:90%
So let's do a simple example.

00:18:49.590 --> 00:18:51.820 align:middle line:90%
I take a tetrahedral die.

00:18:51.820 --> 00:18:53.820 align:middle line:90%
I toss it twice.

00:18:53.820 --> 00:18:55.850 align:middle line:84%
And there's lots of random
variables that you can

00:18:55.850 --> 00:18:57.850 align:middle line:84%
associate with the
same experiment.

00:18:57.850 --> 00:19:01.450 align:middle line:84%
So the outcome of the first
throw, we can call it F.

00:19:01.450 --> 00:19:05.300 align:middle line:84%
That's a random variable because
it's determined once

00:19:05.300 --> 00:19:09.840 align:middle line:84%
you tell me what happens
in the experiment.

00:19:09.840 --> 00:19:11.615 align:middle line:84%
The outcome of the
second throw is

00:19:11.615 --> 00:19:13.470 align:middle line:90%
another random variable.

00:19:13.470 --> 00:19:16.890 align:middle line:84%
The minimum of the two throws
is also a random variable.

00:19:16.890 --> 00:19:20.580 align:middle line:84%
Once I do the experiment, this
random variable takes on a

00:19:20.580 --> 00:19:22.580 align:middle line:90%
specific numerical value.

00:19:22.580 --> 00:19:26.850 align:middle line:84%
So suppose I do the experiment
and I get a 2 and a 3.

00:19:26.850 --> 00:19:29.530 align:middle line:84%
So this random variable is going
to take the numerical

00:19:29.530 --> 00:19:30.420 align:middle line:90%
value of 2.

00:19:30.420 --> 00:19:32.440 align:middle line:84%
This is going to take the
numerical value of 3.

00:19:32.440 --> 00:19:35.500 align:middle line:84%
This is going to take the
numerical value of 2.

00:19:35.500 --> 00:19:38.830 align:middle line:84%
And now suppose that I want to
calculate the PMF of this

00:19:38.830 --> 00:19:40.490 align:middle line:90%
random variable.

00:19:40.490 --> 00:19:43.311 align:middle line:84%
What I will need to do is to
calculate Px(0), Px(1), Px(2),

00:19:43.311 --> 00:19:47.980 align:middle line:90%
Px(3), and so on.

00:19:47.980 --> 00:19:50.680 align:middle line:84%
Let's not do the entire
calculation then, let's just

00:19:50.680 --> 00:19:54.770 align:middle line:84%
calculate one of the
entries of the PMF.

00:19:54.770 --> 00:19:56.010 align:middle line:90%
So Px(2)--

00:19:56.010 --> 00:19:58.870 align:middle line:84%
that's the probability that the
minimum of the two throws

00:19:58.870 --> 00:20:00.280 align:middle line:90%
gives us a 2.

00:20:00.280 --> 00:20:04.080 align:middle line:84%
And this can happen
in many ways.

00:20:04.080 --> 00:20:06.390 align:middle line:84%
There are five ways that
it can happen.

00:20:06.390 --> 00:20:11.010 align:middle line:84%
Those are all of the outcomes
for which the smallest of the

00:20:11.010 --> 00:20:13.780 align:middle line:90%
two is equal to 2.

00:20:13.780 --> 00:20:18.090 align:middle line:84%
That's five outcomes assuming
that the tetrahedral die is

00:20:18.090 --> 00:20:20.920 align:middle line:84%
fair and the tosses
are independent.

00:20:20.920 --> 00:20:24.450 align:middle line:84%
Each one of these outcomes
has probability of 1/16.

00:20:24.450 --> 00:20:27.185 align:middle line:84%
There's five of them, so
we get an answer, 5/16.

00:20:27.185 --> 00:20:30.490 align:middle line:90%


00:20:30.490 --> 00:20:33.770 align:middle line:84%
Conceptually, this is just the
procedure that you use to

00:20:33.770 --> 00:20:37.280 align:middle line:84%
calculate PMFs the way that
you construct this

00:20:37.280 --> 00:20:38.730 align:middle line:90%
particular bar graph.

00:20:38.730 --> 00:20:41.340 align:middle line:84%
You consider all the possible
values of your random

00:20:41.340 --> 00:20:43.860 align:middle line:84%
variable, and for each one of
those random variables you

00:20:43.860 --> 00:20:47.090 align:middle line:84%
find the probability that the
random variable takes on that

00:20:47.090 --> 00:20:49.710 align:middle line:84%
value by adding the
probabilities of all the

00:20:49.710 --> 00:20:51.750 align:middle line:84%
possible outcomes that
leads to that

00:20:51.750 --> 00:20:54.100 align:middle line:90%
particular numerical value.

00:20:54.100 --> 00:20:57.620 align:middle line:84%
So let's do another, more
interesting one.

00:20:57.620 --> 00:21:00.270 align:middle line:84%
So let's revisit the
coin tossing

00:21:00.270 --> 00:21:02.490 align:middle line:90%
problem from last time.

00:21:02.490 --> 00:21:11.600 align:middle line:84%
Let us fix a number n, and we
decide to flip a coin n

00:21:11.600 --> 00:21:13.080 align:middle line:90%
consecutive times.

00:21:13.080 --> 00:21:16.100 align:middle line:84%
Each time the coin tosses
are independent.

00:21:16.100 --> 00:21:19.300 align:middle line:84%
And each one of the tosses will
have a probability, p, of

00:21:19.300 --> 00:21:20.960 align:middle line:90%
obtaining heads.

00:21:20.960 --> 00:21:23.590 align:middle line:84%
Let's consider the random
variable, which is the total

00:21:23.590 --> 00:21:26.325 align:middle line:84%
number of heads that
have been obtained.

00:21:26.325 --> 00:21:29.690 align:middle line:84%
Well, that's something that
we dealt with last time.

00:21:29.690 --> 00:21:33.380 align:middle line:84%
We know the probabilities for
different numbers of heads,

00:21:33.380 --> 00:21:35.530 align:middle line:84%
but we're just going
to do the same now

00:21:35.530 --> 00:21:37.960 align:middle line:90%
using today's notation.

00:21:37.960 --> 00:21:41.610 align:middle line:84%
So let's, for concreteness,
n equal to 4.

00:21:41.610 --> 00:21:48.410 align:middle line:84%
Px is the PMF of that random
variable, X. Px(2) is meant to

00:21:48.410 --> 00:21:52.130 align:middle line:84%
be, by definition, it's the
probability that a random

00:21:52.130 --> 00:21:54.420 align:middle line:90%
variable takes the value of 2.

00:21:54.420 --> 00:21:57.410 align:middle line:84%
So this is the probability
that we have, exactly two

00:21:57.410 --> 00:22:00.080 align:middle line:90%
heads in our four tosses.

00:22:00.080 --> 00:22:03.910 align:middle line:84%
The event of exactly two heads
can happen in multiple ways.

00:22:03.910 --> 00:22:05.740 align:middle line:84%
And here I've written
down the different

00:22:05.740 --> 00:22:06.920 align:middle line:90%
ways that it can happen.

00:22:06.920 --> 00:22:09.230 align:middle line:84%
It turns out that there's
exactly six

00:22:09.230 --> 00:22:10.920 align:middle line:90%
ways that it can happen.

00:22:10.920 --> 00:22:15.010 align:middle line:84%
And each one of these ways,
luckily enough, has the same

00:22:15.010 --> 00:22:16.180 align:middle line:90%
probability--

00:22:16.180 --> 00:22:19.460 align:middle line:90%
p-squared times (1-p)-squared.

00:22:19.460 --> 00:22:24.690 align:middle line:84%
So that gives us the value for
the PMF evaluated at 2.

00:22:24.690 --> 00:22:28.370 align:middle line:84%
So here we just counted
explicitly that we have six

00:22:28.370 --> 00:22:31.170 align:middle line:84%
possible ways that this can
happen, and this gave rise to

00:22:31.170 --> 00:22:32.900 align:middle line:90%
this factor of 6.

00:22:32.900 --> 00:22:37.360 align:middle line:84%
But this factor of 6 turns
out to be the same as

00:22:37.360 --> 00:22:39.350 align:middle line:90%
this 4 choose 2.

00:22:39.350 --> 00:22:42.490 align:middle line:84%
If you remember definition from
last time, 4 choose 2 is

00:22:42.490 --> 00:22:45.650 align:middle line:84%
4 factorial divided by 2
factorial, divided by 2

00:22:45.650 --> 00:22:49.940 align:middle line:84%
factorial, which is
indeed equal to 6.

00:22:49.940 --> 00:22:52.370 align:middle line:84%
And this is the more general
formula that

00:22:52.370 --> 00:22:53.830 align:middle line:90%
you would be using.

00:22:53.830 --> 00:22:59.190 align:middle line:84%
In general, if you have n tosses
and you're interested

00:22:59.190 --> 00:23:02.540 align:middle line:84%
in the probability of obtaining
k heads, the

00:23:02.540 --> 00:23:05.560 align:middle line:84%
probability of that event is
given by this formula.

00:23:05.560 --> 00:23:08.710 align:middle line:84%
So that's the formula that
we derived last time.

00:23:08.710 --> 00:23:11.230 align:middle line:84%
Except that last time we didn't
use this notation.

00:23:11.230 --> 00:23:15.300 align:middle line:84%
We just said the probability of
k heads is equal to this.

00:23:15.300 --> 00:23:18.020 align:middle line:84%
Today we introduce the
extra notation.

00:23:18.020 --> 00:23:22.470 align:middle line:84%
And also having that notation,
we may be tempted to also plot

00:23:22.470 --> 00:23:26.310 align:middle line:90%
a bar graph for the Px.

00:23:26.310 --> 00:23:29.390 align:middle line:84%
In this case, for the coin
tossing problem.

00:23:29.390 --> 00:23:35.090 align:middle line:84%
And if you plot that bar graph
as a function of k when n is a

00:23:35.090 --> 00:23:40.850 align:middle line:84%
fairly large number, what you
will end up obtaining is a bar

00:23:40.850 --> 00:23:47.525 align:middle line:84%
graph that has a shape of
something like this.

00:23:47.525 --> 00:23:53.840 align:middle line:90%


00:23:53.840 --> 00:23:58.800 align:middle line:84%
So certain values of k are more
likely than others, and

00:23:58.800 --> 00:24:00.790 align:middle line:84%
the more likely values
are somewhere in the

00:24:00.790 --> 00:24:02.230 align:middle line:90%
middle of the range.

00:24:02.230 --> 00:24:03.490 align:middle line:90%
And extreme values--

00:24:03.490 --> 00:24:07.110 align:middle line:84%
too few heads or too many
heads, are unlikely.

00:24:07.110 --> 00:24:09.870 align:middle line:84%
Now, the miraculous thing is
that it turns out that this

00:24:09.870 --> 00:24:15.550 align:middle line:84%
curve gets a pretty definite
shape, like a so-called bell

00:24:15.550 --> 00:24:18.210 align:middle line:90%
curve, when n is big.

00:24:18.210 --> 00:24:20.770 align:middle line:90%


00:24:20.770 --> 00:24:24.920 align:middle line:84%
This is a very deep and central
fact from probability

00:24:24.920 --> 00:24:30.210 align:middle line:84%
theory that we will get to
in a couple of months.

00:24:30.210 --> 00:24:33.900 align:middle line:84%
For now, it just could be
a curious observation.

00:24:33.900 --> 00:24:38.390 align:middle line:84%
If you go into MATLAB and put
this formula in and ask MATLAB

00:24:38.390 --> 00:24:41.540 align:middle line:84%
to plot it for you, you're going
to get an interesting

00:24:41.540 --> 00:24:43.140 align:middle line:90%
shape of this form.

00:24:43.140 --> 00:24:46.700 align:middle line:84%
And later on we will have to
sort of understand where this

00:24:46.700 --> 00:24:50.760 align:middle line:84%
is coming from and whether
there's a nice, simple formula

00:24:50.760 --> 00:24:54.920 align:middle line:84%
for the asymptotic
form that we get.

00:24:54.920 --> 00:24:55.370 align:middle line:90%
All right.

00:24:55.370 --> 00:25:00.580 align:middle line:84%
So, so far I've said essentially
nothing new, just

00:25:00.580 --> 00:25:05.240 align:middle line:84%
a little bit of notation and
this little conceptual thing

00:25:05.240 --> 00:25:07.900 align:middle line:84%
that you have to think of random
variables as functions

00:25:07.900 --> 00:25:09.060 align:middle line:90%
in the sample space.

00:25:09.060 --> 00:25:11.620 align:middle line:84%
So now it's time to introduce
something new.

00:25:11.620 --> 00:25:14.250 align:middle line:84%
This is the big concept
of the day.

00:25:14.250 --> 00:25:17.180 align:middle line:84%
In some sense it's
an easy concept.

00:25:17.180 --> 00:25:23.420 align:middle line:84%
But it's the most central, most
important concept that we

00:25:23.420 --> 00:25:26.970 align:middle line:84%
have to deal with random
variables.

00:25:26.970 --> 00:25:28.790 align:middle line:84%
It's the concept
of the expected

00:25:28.790 --> 00:25:30.860 align:middle line:90%
value of a random variable.

00:25:30.860 --> 00:25:34.570 align:middle line:84%
So the expected value is meant
to be, let's speak loosely,

00:25:34.570 --> 00:25:38.100 align:middle line:84%
something like an average,
where you interpret

00:25:38.100 --> 00:25:41.520 align:middle line:84%
probabilities as something
like frequencies.

00:25:41.520 --> 00:25:46.490 align:middle line:84%
So you play a certain game and
your rewards are going to be--

00:25:46.490 --> 00:25:49.660 align:middle line:90%


00:25:49.660 --> 00:25:52.010 align:middle line:90%
use my standard numbers--

00:25:52.010 --> 00:25:54.530 align:middle line:84%
your rewards are going
to be one dollar

00:25:54.530 --> 00:25:58.040 align:middle line:90%
with probability 1/6.

00:25:58.040 --> 00:26:04.711 align:middle line:84%
It's going to be 2 dollars with
probability 1/2, and four

00:26:04.711 --> 00:26:08.670 align:middle line:90%
dollars with probability 1/3.

00:26:08.670 --> 00:26:11.920 align:middle line:84%
So this is a plot of the PMF
of some random variable.

00:26:11.920 --> 00:26:15.270 align:middle line:84%
If you play that game and you
get so many dollars with this

00:26:15.270 --> 00:26:18.520 align:middle line:84%
probability, and so on, how much
do you expect to get on

00:26:18.520 --> 00:26:21.670 align:middle line:84%
the average if you play the
game a zillion times?

00:26:21.670 --> 00:26:23.420 align:middle line:84%
Well, you can think
as follows--

00:26:23.420 --> 00:26:27.990 align:middle line:84%
one sixth of the time I'm
going to get one dollar.

00:26:27.990 --> 00:26:31.620 align:middle line:84%
One half of the time that
outcome is going to happen and

00:26:31.620 --> 00:26:34.140 align:middle line:90%
I'm going to get two dollars.

00:26:34.140 --> 00:26:37.920 align:middle line:84%
And one third of the time the
other outcome happens, and I'm

00:26:37.920 --> 00:26:40.690 align:middle line:90%
going to get four dollars.

00:26:40.690 --> 00:26:45.230 align:middle line:84%
And you evaluate that number
and it turns out to be 2.5.

00:26:45.230 --> 00:26:45.490 align:middle line:90%
OK.

00:26:45.490 --> 00:26:50.410 align:middle line:84%
So that's a reasonable way of
calculating the average payoff

00:26:50.410 --> 00:26:52.550 align:middle line:84%
if you think of these
probabilities as the

00:26:52.550 --> 00:26:56.440 align:middle line:84%
frequencies with which you
obtain the different payoffs.

00:26:56.440 --> 00:26:59.430 align:middle line:84%
And loosely speaking, it doesn't
hurt to think of

00:26:59.430 --> 00:27:02.430 align:middle line:84%
probabilities as frequencies
when you try to make sense of

00:27:02.430 --> 00:27:04.990 align:middle line:90%
various things.

00:27:04.990 --> 00:27:06.480 align:middle line:90%
So what did we do here?

00:27:06.480 --> 00:27:11.710 align:middle line:84%
We took the probabilities of the
different outcomes, of the

00:27:11.710 --> 00:27:15.370 align:middle line:84%
different numerical values, and
multiplied them with the

00:27:15.370 --> 00:27:17.610 align:middle line:90%
corresponding numerical value.

00:27:17.610 --> 00:27:19.910 align:middle line:84%
Similarly here, we have
a probability and the

00:27:19.910 --> 00:27:24.890 align:middle line:84%
corresponding numerical value
and we added up over all x's.

00:27:24.890 --> 00:27:26.430 align:middle line:90%
So that's what we did.

00:27:26.430 --> 00:27:29.750 align:middle line:84%
It looks like an interesting
quantity to deal with.

00:27:29.750 --> 00:27:32.800 align:middle line:84%
So we're going to give a name to
it, and we're going to call

00:27:32.800 --> 00:27:35.740 align:middle line:84%
it the expected value of
a random variable.

00:27:35.740 --> 00:27:39.610 align:middle line:84%
So this formula just captures
the calculation that we did.

00:27:39.610 --> 00:27:43.520 align:middle line:84%
How do we interpret the
expected value?

00:27:43.520 --> 00:27:46.490 align:middle line:84%
So the one interpretation
is the one that I

00:27:46.490 --> 00:27:48.110 align:middle line:90%
used in this example.

00:27:48.110 --> 00:27:52.290 align:middle line:84%
You can think of it as the
average that you get over a

00:27:52.290 --> 00:27:56.200 align:middle line:84%
large number of repetitions
of an experiment where you

00:27:56.200 --> 00:27:59.330 align:middle line:84%
interpret the probabilities as
the frequencies with which the

00:27:59.330 --> 00:28:02.090 align:middle line:84%
different numerical
values can happen.

00:28:02.090 --> 00:28:04.870 align:middle line:84%
There's another interpretation
that's a little more visual

00:28:04.870 --> 00:28:07.550 align:middle line:84%
and that's kind of insightful,
if you remember your freshman

00:28:07.550 --> 00:28:10.860 align:middle line:84%
physics, this kind of formula
gives you the center of

00:28:10.860 --> 00:28:14.390 align:middle line:84%
gravity of an object
of this kind.

00:28:14.390 --> 00:28:17.290 align:middle line:84%
If you take that picture
literally and think of this as

00:28:17.290 --> 00:28:20.700 align:middle line:84%
a mass of one sixth sitting
here, and the mass of one half

00:28:20.700 --> 00:28:24.000 align:middle line:84%
sitting here, and one third
sitting there, and you ask me

00:28:24.000 --> 00:28:26.920 align:middle line:84%
what's the center of gravity
of that structure.

00:28:26.920 --> 00:28:29.320 align:middle line:84%
This is the formula that gives
you the center of gravity.

00:28:29.320 --> 00:28:30.900 align:middle line:84%
Now what's the center
of gravity?

00:28:30.900 --> 00:28:34.960 align:middle line:84%
It's the place where if you put
your pen right underneath,

00:28:34.960 --> 00:28:38.050 align:middle line:84%
that diagram will stay in place
and will not fall on one

00:28:38.050 --> 00:28:40.440 align:middle line:84%
side and will not fall
on the other side.

00:28:40.440 --> 00:28:44.950 align:middle line:84%
So in this thing, by picture,
since the 4 is a little more

00:28:44.950 --> 00:28:47.880 align:middle line:84%
to the right and a little
heavier, the center of gravity

00:28:47.880 --> 00:28:50.200 align:middle line:84%
should be somewhere
around here.

00:28:50.200 --> 00:28:52.290 align:middle line:84%
And that's what for
the math gave us.

00:28:52.290 --> 00:28:54.740 align:middle line:84%
It turns out to be
two and a half.

00:28:54.740 --> 00:28:56.920 align:middle line:84%
Once you have this
interpretation about centers

00:28:56.920 --> 00:28:58.890 align:middle line:84%
of gravity, sometimes
you can calculate

00:28:58.890 --> 00:29:01.090 align:middle line:90%
expectations pretty fast.

00:29:01.090 --> 00:29:04.410 align:middle line:84%
So here's our new
random variable.

00:29:04.410 --> 00:29:07.840 align:middle line:84%
It's the uniform random variable
in which each one of

00:29:07.840 --> 00:29:10.420 align:middle line:84%
the numerical values
is equally likely.

00:29:10.420 --> 00:29:13.980 align:middle line:84%
Here there's a total of n plus
1 possible numerical values.

00:29:13.980 --> 00:29:17.600 align:middle line:84%
So each one of them has
probability 1 over (n + 1).

00:29:17.600 --> 00:29:20.650 align:middle line:84%
Let's calculate the expected
value of this random variable.

00:29:20.650 --> 00:29:24.620 align:middle line:84%
We can take the formula
literally and consider all

00:29:24.620 --> 00:29:28.920 align:middle line:84%
possible outcomes, or all
possible numerical values, and

00:29:28.920 --> 00:29:32.330 align:middle line:84%
weigh them by their
corresponding probability, and

00:29:32.330 --> 00:29:35.170 align:middle line:84%
do this calculation and
obtain an answer.

00:29:35.170 --> 00:29:38.520 align:middle line:84%
But I gave you the intuition
of centers of gravity.

00:29:38.520 --> 00:29:41.990 align:middle line:84%
Can you use that intuition
to guess the answer?

00:29:41.990 --> 00:29:46.680 align:middle line:84%
What's the center of gravity
infrastructure of this kind?

00:29:46.680 --> 00:29:47.860 align:middle line:90%
We have symmetry.

00:29:47.860 --> 00:29:50.710 align:middle line:90%
So it should be in the middle.

00:29:50.710 --> 00:29:51.970 align:middle line:90%
And what's the middle?

00:29:51.970 --> 00:29:54.850 align:middle line:84%
It's the average of the
two end points.

00:29:54.850 --> 00:29:57.850 align:middle line:84%
So without having to do the
algebra, you know that's the

00:29:57.850 --> 00:30:01.200 align:middle line:84%
answer is going to
be n over 2.

00:30:01.200 --> 00:30:05.850 align:middle line:84%
So this is a moral that you
should keep whenever you have

00:30:05.850 --> 00:30:11.460 align:middle line:84%
PMF, which is symmetric around
a certain point.

00:30:11.460 --> 00:30:15.350 align:middle line:84%
That certain point is going
to be the expected value

00:30:15.350 --> 00:30:17.460 align:middle line:84%
associated with this
particular PMF.

00:30:17.460 --> 00:30:21.920 align:middle line:90%


00:30:21.920 --> 00:30:22.380 align:middle line:90%
OK.

00:30:22.380 --> 00:30:29.610 align:middle line:84%
So having defined the expected
value, what is there that's

00:30:29.610 --> 00:30:31.810 align:middle line:90%
left for us to do?

00:30:31.810 --> 00:30:37.290 align:middle line:84%
Well, we want to investigate how
it behaves, what kind of

00:30:37.290 --> 00:30:43.390 align:middle line:84%
properties does it have, and
also how do you calculate

00:30:43.390 --> 00:30:48.040 align:middle line:84%
expected values of complicated
random variables.

00:30:48.040 --> 00:30:52.130 align:middle line:84%
So the first complication that
we're going to start with is

00:30:52.130 --> 00:30:54.985 align:middle line:84%
the case where we deal with a
function of a random variable.

00:30:54.985 --> 00:30:59.002 align:middle line:90%


00:30:59.002 --> 00:30:59.680 align:middle line:90%
OK.

00:30:59.680 --> 00:31:05.890 align:middle line:84%
So let me redraw this same
picture as before.

00:31:05.890 --> 00:31:07.090 align:middle line:90%
We have omega.

00:31:07.090 --> 00:31:09.580 align:middle line:90%
This is our sample space.

00:31:09.580 --> 00:31:12.310 align:middle line:90%
This is the real line.

00:31:12.310 --> 00:31:17.370 align:middle line:84%
And we have a random variable
that gives rise to various

00:31:17.370 --> 00:31:24.400 align:middle line:84%
values for X. So the random
variable is capital X, and

00:31:24.400 --> 00:31:28.690 align:middle line:84%
every outcome leads to a
particular numerical value x

00:31:28.690 --> 00:31:33.030 align:middle line:84%
for our random variable X. So
capital X is really the

00:31:33.030 --> 00:31:37.930 align:middle line:84%
function that maps these points
into the real line.

00:31:37.930 --> 00:31:42.710 align:middle line:84%
And then I consider a function
of this random variable, call

00:31:42.710 --> 00:31:47.080 align:middle line:84%
it capital Y, and it's
a function of my

00:31:47.080 --> 00:31:49.980 align:middle line:90%
previous random variable.

00:31:49.980 --> 00:31:54.190 align:middle line:84%
And this new random variable Y
takes numerical values that

00:31:54.190 --> 00:31:58.140 align:middle line:84%
are completely determined once
I know the numerical value of

00:31:58.140 --> 00:32:03.090 align:middle line:84%
capital X. And perhaps you get
a diagram of this kind.

00:32:03.090 --> 00:32:08.520 align:middle line:90%


00:32:08.520 --> 00:32:10.970 align:middle line:90%
So X is a random variable.

00:32:10.970 --> 00:32:14.506 align:middle line:84%
Once you have an outcome, this
determines the value of x.

00:32:14.506 --> 00:32:16.760 align:middle line:90%
Y is also a random variable.

00:32:16.760 --> 00:32:19.200 align:middle line:84%
Once you have the outcome,
that determines

00:32:19.200 --> 00:32:21.230 align:middle line:90%
the value of y.

00:32:21.230 --> 00:32:26.630 align:middle line:84%
Y is completely determined once
you know X. We have a

00:32:26.630 --> 00:32:31.560 align:middle line:84%
formula for how to calculate
the expected value of X.

00:32:31.560 --> 00:32:34.380 align:middle line:84%
Suppose that you're interested
in calculating the expected

00:32:34.380 --> 00:32:39.910 align:middle line:84%
value of Y. How would
you go about it?

00:32:39.910 --> 00:32:40.710 align:middle line:90%
OK.

00:32:40.710 --> 00:32:43.580 align:middle line:84%
The only thing you have in your
hands is the definition,

00:32:43.580 --> 00:32:47.750 align:middle line:84%
so you could start by just
using the definition.

00:32:47.750 --> 00:32:50.150 align:middle line:90%
And what does this entail?

00:32:50.150 --> 00:32:55.330 align:middle line:84%
It entails for every particular
value of y, collect

00:32:55.330 --> 00:32:59.160 align:middle line:84%
all the outcomes that leads
to that value of y.

00:32:59.160 --> 00:33:01.010 align:middle line:90%
Find their probability.

00:33:01.010 --> 00:33:02.280 align:middle line:90%
Do the same here.

00:33:02.280 --> 00:33:04.550 align:middle line:84%
For that value, collect
those outcomes.

00:33:04.550 --> 00:33:07.700 align:middle line:84%
Find their probability
and weight by y.

00:33:07.700 --> 00:33:13.050 align:middle line:84%
So this formula does the
addition over this line.

00:33:13.050 --> 00:33:17.060 align:middle line:84%
We consider the different
outcomes and add things up.

00:33:17.060 --> 00:33:20.290 align:middle line:84%
There's an alternative way of
doing the same accounting

00:33:20.290 --> 00:33:23.930 align:middle line:84%
where instead of doing the
addition over those numbers,

00:33:23.930 --> 00:33:26.540 align:middle line:90%
we do the addition up here.

00:33:26.540 --> 00:33:30.250 align:middle line:84%
We consider the different
possible values of x, and we

00:33:30.250 --> 00:33:31.500 align:middle line:90%
think as follows--

00:33:31.500 --> 00:33:34.500 align:middle line:90%


00:33:34.500 --> 00:33:38.900 align:middle line:84%
for each possible value of x,
that value is going to occur

00:33:38.900 --> 00:33:41.310 align:middle line:90%
with this probability.

00:33:41.310 --> 00:33:45.890 align:middle line:84%
And if that value has occurred,
this is how much I'm

00:33:45.890 --> 00:33:47.840 align:middle line:90%
getting, the g of x.

00:33:47.840 --> 00:33:52.990 align:middle line:84%
So I'm considering the
probability of this outcome.

00:33:52.990 --> 00:33:56.050 align:middle line:84%
And in that case, y
takes this value.

00:33:56.050 --> 00:34:00.240 align:middle line:84%
Then I'm considering the
probabilities of this outcome.

00:34:00.240 --> 00:34:04.650 align:middle line:84%
And in that case, g of x
takes again that value.

00:34:04.650 --> 00:34:08.100 align:middle line:84%
Then I consider this particular
x, it happens with

00:34:08.100 --> 00:34:11.280 align:middle line:84%
this much probability, and in
that case, g of x takes that

00:34:11.280 --> 00:34:14.300 align:middle line:90%
value, and similarly here.

00:34:14.300 --> 00:34:18.170 align:middle line:84%
We end up doing exactly the same
arithmetic, it's only a

00:34:18.170 --> 00:34:21.760 align:middle line:84%
question whether we bundle
things together.

00:34:21.760 --> 00:34:25.790 align:middle line:84%
That is, if we calculate the
probability of this, then

00:34:25.790 --> 00:34:28.239 align:middle line:84%
we're bundling these
two cases together.

00:34:28.239 --> 00:34:32.110 align:middle line:84%
Whereas if we do the addition
up here, we do a separate

00:34:32.110 --> 00:34:32.949 align:middle line:90%
calculation--

00:34:32.949 --> 00:34:35.639 align:middle line:84%
this probability times this
number, and then this

00:34:35.639 --> 00:34:37.989 align:middle line:90%
probability times that number.

00:34:37.989 --> 00:34:41.420 align:middle line:84%
So it's just a simple
rearrangement of the way that

00:34:41.420 --> 00:34:45.330 align:middle line:84%
we do the calculations, but it
does make a big difference in

00:34:45.330 --> 00:34:49.010 align:middle line:84%
practice if you actually want
to calculate expectations.

00:34:49.010 --> 00:34:52.389 align:middle line:84%
So the second procedure that I
mentioned, where you do the

00:34:52.389 --> 00:34:56.790 align:middle line:84%
addition by running
over the x-axis

00:34:56.790 --> 00:34:59.710 align:middle line:90%
corresponds to this formula.

00:34:59.710 --> 00:35:05.830 align:middle line:84%
Consider all possibilities for x
and when that x happens, how

00:35:05.830 --> 00:35:07.530 align:middle line:90%
much money are you getting?

00:35:07.530 --> 00:35:10.850 align:middle line:84%
That gives you the average money
that you are getting.

00:35:10.850 --> 00:35:11.270 align:middle line:90%
All right.

00:35:11.270 --> 00:35:14.840 align:middle line:84%
So I kind of hand waved and
argued that it's just a

00:35:14.840 --> 00:35:17.690 align:middle line:84%
different way of accounting, of
course one needs to prove

00:35:17.690 --> 00:35:19.060 align:middle line:90%
this formula.

00:35:19.060 --> 00:35:20.950 align:middle line:84%
And fortunately it
can be proved.

00:35:20.950 --> 00:35:23.470 align:middle line:84%
You're going to see that
in recitation.

00:35:23.470 --> 00:35:25.710 align:middle line:84%
Most people, once they're a
little comfortable with the

00:35:25.710 --> 00:35:28.570 align:middle line:84%
concepts of probability,
actually believe that this is

00:35:28.570 --> 00:35:30.130 align:middle line:90%
true by definition.

00:35:30.130 --> 00:35:31.860 align:middle line:84%
In fact it's not true
by definition.

00:35:31.860 --> 00:35:34.610 align:middle line:84%
It's called the law of the
unconscious statistician.

00:35:34.610 --> 00:35:37.930 align:middle line:84%
It's something that you always
do, but it's something that

00:35:37.930 --> 00:35:40.750 align:middle line:90%
does require justification.

00:35:40.750 --> 00:35:41.100 align:middle line:90%
All right.

00:35:41.100 --> 00:35:44.160 align:middle line:84%
So this gives us basically a
shortcut for calculating

00:35:44.160 --> 00:35:47.770 align:middle line:84%
expected values of functions of
a random variable without

00:35:47.770 --> 00:35:51.990 align:middle line:84%
having to find the PMF
of that function.

00:35:51.990 --> 00:35:54.470 align:middle line:84%
We can work with the PMF of
the original function.

00:35:54.470 --> 00:35:57.140 align:middle line:90%


00:35:57.140 --> 00:35:57.430 align:middle line:90%
All right.

00:35:57.430 --> 00:36:00.940 align:middle line:84%
So we're going to use this
property over and over.

00:36:00.940 --> 00:36:06.570 align:middle line:84%
Before we start using it, one
general word of caution--

00:36:06.570 --> 00:36:10.640 align:middle line:84%
the average of a function of a
random variable, in general,

00:36:10.640 --> 00:36:16.400 align:middle line:84%
is not the same as the function
of the average.

00:36:16.400 --> 00:36:20.820 align:middle line:84%
So these two operations of
taking averages and taking

00:36:20.820 --> 00:36:23.180 align:middle line:90%
functions do not commute.

00:36:23.180 --> 00:36:28.130 align:middle line:84%
What this inequality tells you
is that, in general, you can

00:36:28.130 --> 00:36:30.280 align:middle line:90%
not reason on the average.

00:36:30.280 --> 00:36:34.420 align:middle line:90%


00:36:34.420 --> 00:36:38.600 align:middle line:84%
So we're going to see instances
where this property

00:36:38.600 --> 00:36:39.610 align:middle line:90%
is not true.

00:36:39.610 --> 00:36:41.080 align:middle line:84%
You're going to see
lots of them.

00:36:41.080 --> 00:36:43.920 align:middle line:84%
Let me just throw it here that
it's something that's not true

00:36:43.920 --> 00:36:47.710 align:middle line:84%
in general, but we will be
interested in the exceptions

00:36:47.710 --> 00:36:51.480 align:middle line:84%
where a relation like
this is true.

00:36:51.480 --> 00:36:53.360 align:middle line:84%
But these will be
the exceptions.

00:36:53.360 --> 00:36:56.960 align:middle line:84%
So in general, expectations
are average,

00:36:56.960 --> 00:36:58.850 align:middle line:90%
something like averages.

00:36:58.850 --> 00:37:02.400 align:middle line:84%
But the function of an average
is not the same as the average

00:37:02.400 --> 00:37:05.070 align:middle line:90%
of the function.

00:37:05.070 --> 00:37:05.440 align:middle line:90%
OK.

00:37:05.440 --> 00:37:09.530 align:middle line:84%
So now let's go to properties
of expectations.

00:37:09.530 --> 00:37:15.170 align:middle line:84%
Suppose that alpha is a real
number, and I ask you, what's

00:37:15.170 --> 00:37:17.740 align:middle line:84%
the expected value of
that real number?

00:37:17.740 --> 00:37:21.010 align:middle line:84%
So for example, if I write
down this expression--

00:37:21.010 --> 00:37:23.070 align:middle line:90%
expected value of 2.

00:37:23.070 --> 00:37:25.930 align:middle line:90%
What is this?

00:37:25.930 --> 00:37:29.470 align:middle line:84%
Well, we defined random
variables and we defined

00:37:29.470 --> 00:37:31.860 align:middle line:84%
expectations of random
variables.

00:37:31.860 --> 00:37:35.870 align:middle line:84%
So for this to make syntactic
sense, this thing inside here

00:37:35.870 --> 00:37:37.670 align:middle line:90%
should be a random variable.

00:37:37.670 --> 00:37:39.260 align:middle line:90%
Is 2 --

00:37:39.260 --> 00:37:41.140 align:middle line:84%
the number 2 --- is it
a random variable?

00:37:41.140 --> 00:37:44.740 align:middle line:90%


00:37:44.740 --> 00:37:48.420 align:middle line:90%
In some sense, yes.

00:37:48.420 --> 00:37:55.750 align:middle line:84%
It's the random variable that
takes, always, the value of 2.

00:37:55.750 --> 00:37:59.220 align:middle line:84%
So suppose that you have some
experiment and that experiment

00:37:59.220 --> 00:38:02.580 align:middle line:84%
always outputs 2 whenever
it happens.

00:38:02.580 --> 00:38:05.880 align:middle line:84%
Then you can say, yes, it's
a random experiment but it

00:38:05.880 --> 00:38:06.960 align:middle line:90%
always gives me 2.

00:38:06.960 --> 00:38:08.600 align:middle line:84%
The value of the random
variable is

00:38:08.600 --> 00:38:10.460 align:middle line:90%
always 2 no matter what.

00:38:10.460 --> 00:38:13.200 align:middle line:84%
It's kind of a degenerate random
variable that doesn't

00:38:13.200 --> 00:38:17.230 align:middle line:84%
have any real randomness in it,
but it's still useful to

00:38:17.230 --> 00:38:20.130 align:middle line:90%
think of it as a special case.

00:38:20.130 --> 00:38:23.000 align:middle line:84%
So it corresponds to a function
from the sample space

00:38:23.000 --> 00:38:26.750 align:middle line:84%
to the real line that takes
only one value.

00:38:26.750 --> 00:38:30.390 align:middle line:84%
No matter what the outcome is,
it always gives me a 2.

00:38:30.390 --> 00:38:30.770 align:middle line:90%
OK.

00:38:30.770 --> 00:38:34.390 align:middle line:84%
If you have a random variable
that always gives you a 2,

00:38:34.390 --> 00:38:37.980 align:middle line:84%
what is the expected
value going to be?

00:38:37.980 --> 00:38:40.530 align:middle line:84%
The only entry that shows
up in this summation

00:38:40.530 --> 00:38:43.000 align:middle line:90%
is that number 2.

00:38:43.000 --> 00:38:46.270 align:middle line:84%
The probability of a 2 is equal
to 1, and the value of

00:38:46.270 --> 00:38:48.330 align:middle line:84%
that random variable
is equal to 2.

00:38:48.330 --> 00:38:51.030 align:middle line:90%
So it's the number itself.

00:38:51.030 --> 00:38:53.910 align:middle line:84%
So the average value in an
experiment that always gives

00:38:53.910 --> 00:38:56.580 align:middle line:90%
you 2's is 2.

00:38:56.580 --> 00:38:57.100 align:middle line:90%
All right.

00:38:57.100 --> 00:38:59.450 align:middle line:90%
So that's nice and simple.

00:38:59.450 --> 00:39:04.890 align:middle line:84%
Now let's go to our experiment
where age was

00:39:04.890 --> 00:39:07.310 align:middle line:90%
your height in inches.

00:39:07.310 --> 00:39:11.160 align:middle line:84%
And I know your height in
inches, but I'm interested in

00:39:11.160 --> 00:39:15.880 align:middle line:84%
your height measured
in centimeters.

00:39:15.880 --> 00:39:19.040 align:middle line:84%
How is that going
to be related to

00:39:19.040 --> 00:39:22.675 align:middle line:90%
your height in inches?

00:39:22.675 --> 00:39:27.440 align:middle line:84%
Well, if you take your height
in inches and convert it to

00:39:27.440 --> 00:39:30.690 align:middle line:84%
centimeters, I have another
random variable, which is

00:39:30.690 --> 00:39:34.280 align:middle line:84%
always, no matter what, two and
a half times bigger than

00:39:34.280 --> 00:39:36.570 align:middle line:84%
the random variable
I started with.

00:39:36.570 --> 00:39:40.470 align:middle line:84%
If you take some quantity and
always multiplied by two and a

00:39:40.470 --> 00:39:43.610 align:middle line:84%
half what happens to the average
of that quantity?

00:39:43.610 --> 00:39:46.990 align:middle line:84%
It also gets multiplied
by two and a half.

00:39:46.990 --> 00:39:52.030 align:middle line:84%
So you get a relation like
this, which says that the

00:39:52.030 --> 00:39:56.480 align:middle line:84%
average height of a student
measured in centimeters is two

00:39:56.480 --> 00:39:58.660 align:middle line:84%
and a half times the
average height of a

00:39:58.660 --> 00:40:01.660 align:middle line:90%
student measured in inches.

00:40:01.660 --> 00:40:03.730 align:middle line:84%
So that makes perfect
intuitive sense.

00:40:03.730 --> 00:40:07.490 align:middle line:84%
If you generalize it, it gives
us this relation, that if you

00:40:07.490 --> 00:40:13.790 align:middle line:84%
have a number, you can pull it
outside the expectation and

00:40:13.790 --> 00:40:16.210 align:middle line:90%
you get the right result.

00:40:16.210 --> 00:40:20.440 align:middle line:84%
So this is a case where you
can reason on the average.

00:40:20.440 --> 00:40:23.150 align:middle line:84%
If you take a number, such as
height, and multiply it by a

00:40:23.150 --> 00:40:25.500 align:middle line:84%
certain number, you can
reason on the average.

00:40:25.500 --> 00:40:27.650 align:middle line:84%
I multiply the numbers
by two, the averages

00:40:27.650 --> 00:40:29.630 align:middle line:90%
will go up by two.

00:40:29.630 --> 00:40:33.750 align:middle line:84%
So this is an exception to this
cautionary statement that

00:40:33.750 --> 00:40:35.460 align:middle line:90%
I had up there.

00:40:35.460 --> 00:40:39.860 align:middle line:84%
How do we prove that
this fact is true?

00:40:39.860 --> 00:40:44.360 align:middle line:84%
Well, we can use the expected
value rule here, which tells

00:40:44.360 --> 00:40:52.690 align:middle line:84%
us that the expected value of
alpha X, this is our g of X,

00:40:52.690 --> 00:40:59.720 align:middle line:84%
essentially, is going to be
the sum over all x's of my

00:40:59.720 --> 00:41:04.900 align:middle line:84%
function, g of X, times the
probability of the x's.

00:41:04.900 --> 00:41:11.270 align:middle line:84%
In our particular case, g of X
is alpha times X. And we have

00:41:11.270 --> 00:41:12.450 align:middle line:90%
those probabilities.

00:41:12.450 --> 00:41:15.600 align:middle line:84%
And the alpha goes outside
the summation.

00:41:15.600 --> 00:41:23.100 align:middle line:84%
So we get alpha, sum over x's,
x Px of x, which is alpha

00:41:23.100 --> 00:41:26.740 align:middle line:90%
times the expected value of X.

00:41:26.740 --> 00:41:30.580 align:middle line:84%
So that's how you prove this
relation formally using this

00:41:30.580 --> 00:41:32.490 align:middle line:90%
rule up here.

00:41:32.490 --> 00:41:35.810 align:middle line:84%
And the next formula that
I have here also gets

00:41:35.810 --> 00:41:37.310 align:middle line:90%
proved the same way.

00:41:37.310 --> 00:41:41.110 align:middle line:84%
What does this formula
tell you?

00:41:41.110 --> 00:41:46.560 align:middle line:84%
If I take everybody's height
in centimeters--

00:41:46.560 --> 00:41:49.030 align:middle line:84%
we already multiplied
by alpha--

00:41:49.030 --> 00:41:52.800 align:middle line:84%
and the gods give everyone
a bonus of ten extra

00:41:52.800 --> 00:41:54.670 align:middle line:90%
centimeters.

00:41:54.670 --> 00:41:57.720 align:middle line:84%
What's going to happen to the
average height of the class?

00:41:57.720 --> 00:42:02.800 align:middle line:84%
Well, it will just go up by
an extra ten centimeters.

00:42:02.800 --> 00:42:08.040 align:middle line:84%
So this expectation is going to
be giving you the bonus of

00:42:08.040 --> 00:42:15.710 align:middle line:84%
beta just adds a beta to the
average height in centimeters,

00:42:15.710 --> 00:42:20.740 align:middle line:84%
which we also know to be alpha
times the expected

00:42:20.740 --> 00:42:24.430 align:middle line:90%
value of X, plus beta.

00:42:24.430 --> 00:42:29.390 align:middle line:84%
So this is a linearity property
of expectations.

00:42:29.390 --> 00:42:34.140 align:middle line:84%
If you take a linear function
of a single random variable,

00:42:34.140 --> 00:42:38.390 align:middle line:84%
the expected value of that
linear function is the linear

00:42:38.390 --> 00:42:41.140 align:middle line:84%
function of the expected
value.

00:42:41.140 --> 00:42:44.100 align:middle line:84%
So this is our big exception to
this cautionary note, that

00:42:44.100 --> 00:42:48.710 align:middle line:90%
we have equal if g is linear.

00:42:48.710 --> 00:42:55.840 align:middle line:90%


00:42:55.840 --> 00:42:57.090 align:middle line:90%
OK.

00:42:57.090 --> 00:42:59.790 align:middle line:90%


00:42:59.790 --> 00:43:00.790 align:middle line:90%
All right.

00:43:00.790 --> 00:43:05.850 align:middle line:84%
So let's get to the last
concept of the day.

00:43:05.850 --> 00:43:07.470 align:middle line:84%
What kind of functions
of random

00:43:07.470 --> 00:43:11.010 align:middle line:90%
variables may be of interest?

00:43:11.010 --> 00:43:15.660 align:middle line:84%
One possibility might be the
average value of X-squared.

00:43:15.660 --> 00:43:18.780 align:middle line:90%


00:43:18.780 --> 00:43:20.150 align:middle line:90%
Why is it interesting?

00:43:20.150 --> 00:43:21.760 align:middle line:90%
Well, why not.

00:43:21.760 --> 00:43:24.290 align:middle line:84%
It's the simplest function
that you can think of.

00:43:24.290 --> 00:43:27.560 align:middle line:90%


00:43:27.560 --> 00:43:30.800 align:middle line:84%
So if you want to calculate
the expected value of

00:43:30.800 --> 00:43:35.260 align:middle line:84%
X-squared, you would use this
general rule for how you can

00:43:35.260 --> 00:43:39.470 align:middle line:84%
calculate expected values of
functions of random variables.

00:43:39.470 --> 00:43:41.340 align:middle line:84%
You consider all the
possible x's.

00:43:41.340 --> 00:43:45.550 align:middle line:84%
For each x, you see what's the
probability that it occurs.

00:43:45.550 --> 00:43:49.790 align:middle line:84%
And if that x occurs, you
consider and see how big

00:43:49.790 --> 00:43:52.090 align:middle line:90%
x-squared is.

00:43:52.090 --> 00:43:54.810 align:middle line:84%
Now, the more interesting
quantity, a more interesting

00:43:54.810 --> 00:43:58.580 align:middle line:84%
expectation that you can
calculate has to do not with

00:43:58.580 --> 00:44:03.570 align:middle line:84%
x-squared, but with the distance
of x from the mean

00:44:03.570 --> 00:44:05.710 align:middle line:90%
and then squared.

00:44:05.710 --> 00:44:10.800 align:middle line:84%
So let's try to parse what
we've got up here.

00:44:10.800 --> 00:44:14.970 align:middle line:84%
Let's look just at the
quantity inside here.

00:44:14.970 --> 00:44:16.610 align:middle line:90%
What kind of quantity is it?

00:44:16.610 --> 00:44:19.190 align:middle line:90%


00:44:19.190 --> 00:44:21.030 align:middle line:90%
It's a random variable.

00:44:21.030 --> 00:44:22.370 align:middle line:90%
Why?

00:44:22.370 --> 00:44:26.540 align:middle line:84%
X is random, the random
variable, expected value of X

00:44:26.540 --> 00:44:28.090 align:middle line:90%
is a number.

00:44:28.090 --> 00:44:30.800 align:middle line:84%
Subtract a number from a random
variable, you get

00:44:30.800 --> 00:44:32.460 align:middle line:90%
another random variable.

00:44:32.460 --> 00:44:35.630 align:middle line:84%
Take a random variable and
square it, you get another

00:44:35.630 --> 00:44:36.810 align:middle line:90%
random variable.

00:44:36.810 --> 00:44:40.590 align:middle line:84%
So the thing inside here is a
legitimate random variable.

00:44:40.590 --> 00:44:44.950 align:middle line:84%
What kind of random
variable is it?

00:44:44.950 --> 00:44:47.720 align:middle line:84%
So suppose that we have our
experiment and we have

00:44:47.720 --> 00:44:49.452 align:middle line:90%
different x's that can happen.

00:44:49.452 --> 00:44:52.310 align:middle line:90%


00:44:52.310 --> 00:44:56.090 align:middle line:84%
And the mean of X in this
picture might be somewhere

00:44:56.090 --> 00:44:57.340 align:middle line:90%
around here.

00:44:57.340 --> 00:45:00.500 align:middle line:90%


00:45:00.500 --> 00:45:02.570 align:middle line:90%
I do the experiment.

00:45:02.570 --> 00:45:05.350 align:middle line:84%
I obtain some numerical
value of x.

00:45:05.350 --> 00:45:09.610 align:middle line:84%
Let's say I obtain this
numerical value.

00:45:09.610 --> 00:45:13.810 align:middle line:84%
I look at the distance from
the mean, which is this

00:45:13.810 --> 00:45:18.460 align:middle line:84%
length, and I take the
square of that.

00:45:18.460 --> 00:45:22.730 align:middle line:84%
Each time that I do the
experiment, I go and record my

00:45:22.730 --> 00:45:25.780 align:middle line:84%
distance from the mean
and square it.

00:45:25.780 --> 00:45:29.490 align:middle line:84%
So I give more emphasis
to big distances.

00:45:29.490 --> 00:45:33.370 align:middle line:84%
And then I take the average over
all possible outcomes,

00:45:33.370 --> 00:45:35.520 align:middle line:90%
all possible numerical values.

00:45:35.520 --> 00:45:39.510 align:middle line:84%
So I'm trying to compute
the average squared

00:45:39.510 --> 00:45:42.980 align:middle line:90%
distance from the mean.

00:45:42.980 --> 00:45:47.770 align:middle line:84%
This corresponds to
this formula here.

00:45:47.770 --> 00:45:51.110 align:middle line:84%
So the picture that I drew
corresponds to that.

00:45:51.110 --> 00:45:55.920 align:middle line:84%
For every possible numerical
value of x, that numerical

00:45:55.920 --> 00:45:59.010 align:middle line:84%
value corresponds to a certain
distance from the mean

00:45:59.010 --> 00:46:03.580 align:middle line:84%
squared, and I weight according
to how likely is

00:46:03.580 --> 00:46:07.360 align:middle line:84%
that particular value
of x to arise.

00:46:07.360 --> 00:46:10.840 align:middle line:84%
So this measures the
average squared

00:46:10.840 --> 00:46:13.280 align:middle line:90%
distance from the mean.

00:46:13.280 --> 00:46:17.180 align:middle line:84%
Now, because of that expected
value rule, of course, this

00:46:17.180 --> 00:46:20.010 align:middle line:84%
thing is the same as
that expectation.

00:46:20.010 --> 00:46:23.880 align:middle line:84%
It's the average value of the
random variable, which is the

00:46:23.880 --> 00:46:26.300 align:middle line:84%
squared distance
from the mean.

00:46:26.300 --> 00:46:29.820 align:middle line:84%
With this probability, the
random variable takes on this

00:46:29.820 --> 00:46:33.050 align:middle line:84%
numerical value, and the squared
distance from the mean

00:46:33.050 --> 00:46:37.200 align:middle line:84%
ends up taking that particular
numerical value.

00:46:37.200 --> 00:46:37.680 align:middle line:90%
OK.

00:46:37.680 --> 00:46:40.560 align:middle line:84%
So why is the variance
interesting?

00:46:40.560 --> 00:46:45.380 align:middle line:84%
It tells us how far away from
the mean we expect to be on

00:46:45.380 --> 00:46:46.900 align:middle line:90%
the average.

00:46:46.900 --> 00:46:49.550 align:middle line:84%
Well, actually we're not
counting distances from the

00:46:49.550 --> 00:46:51.630 align:middle line:90%
mean, it's distances squared.

00:46:51.630 --> 00:46:56.500 align:middle line:84%
So it gives more emphasis to the
kind of outliers in here.

00:46:56.500 --> 00:46:59.090 align:middle line:84%
But it's a measure of
how spread out the

00:46:59.090 --> 00:47:01.180 align:middle line:90%
distribution is.

00:47:01.180 --> 00:47:05.240 align:middle line:84%
A big variance means that those
bars go far to the left

00:47:05.240 --> 00:47:07.010 align:middle line:90%
and to the right, typically.

00:47:07.010 --> 00:47:10.230 align:middle line:84%
Where as a small variance would
mean that all those bars

00:47:10.230 --> 00:47:13.850 align:middle line:84%
are tightly concentrated
around the mean value.

00:47:13.850 --> 00:47:16.190 align:middle line:84%
It's the average squared
deviation.

00:47:16.190 --> 00:47:18.970 align:middle line:84%
Small variance means that
we generally have small

00:47:18.970 --> 00:47:19.580 align:middle line:90%
deviations.

00:47:19.580 --> 00:47:22.500 align:middle line:84%
Large variances mean that
we generally have large

00:47:22.500 --> 00:47:24.210 align:middle line:90%
deviations.

00:47:24.210 --> 00:47:27.310 align:middle line:84%
Now as a practical matter, when
you want to calculate the

00:47:27.310 --> 00:47:31.140 align:middle line:84%
variance, there's a handy
formula which I'm not proving

00:47:31.140 --> 00:47:33.110 align:middle line:84%
but you will see it
in recitation.

00:47:33.110 --> 00:47:36.270 align:middle line:84%
It's just two lines
of algebra.

00:47:36.270 --> 00:47:40.680 align:middle line:84%
And it allows us to calculate it
in a somewhat simpler way.

00:47:40.680 --> 00:47:43.110 align:middle line:84%
We need to calculate the
expected value of the random

00:47:43.110 --> 00:47:45.210 align:middle line:84%
variable and the expected value
of the squares of the

00:47:45.210 --> 00:47:47.580 align:middle line:84%
random variable, and
these two are going

00:47:47.580 --> 00:47:49.710 align:middle line:90%
to give us the variance.

00:47:49.710 --> 00:47:53.970 align:middle line:84%
So to summarize what we did
up here, the variance, by

00:47:53.970 --> 00:47:57.370 align:middle line:84%
definition, is given
by this formula.

00:47:57.370 --> 00:48:01.470 align:middle line:84%
It's the expected value of
the squared deviation.

00:48:01.470 --> 00:48:06.380 align:middle line:84%
But we have the equivalent
formula, which comes from

00:48:06.380 --> 00:48:13.960 align:middle line:84%
application of the expected
value rule, to the function g

00:48:13.960 --> 00:48:18.690 align:middle line:84%
of X, equals to x minus the
(expected value of X)-squared.

00:48:18.690 --> 00:48:25.640 align:middle line:90%


00:48:25.640 --> 00:48:26.330 align:middle line:90%
OK.

00:48:26.330 --> 00:48:27.460 align:middle line:90%
So this is the definition.

00:48:27.460 --> 00:48:31.170 align:middle line:84%
This comes from the expected
value rule.

00:48:31.170 --> 00:48:35.010 align:middle line:84%
What are some properties
of the variance?

00:48:35.010 --> 00:48:38.650 align:middle line:84%
Of course variances are
always non-negative.

00:48:38.650 --> 00:48:40.880 align:middle line:90%
Why is it always non-negative?

00:48:40.880 --> 00:48:43.650 align:middle line:84%
Well, you look at the definition
and your just

00:48:43.650 --> 00:48:45.660 align:middle line:90%
adding up non-negative things.

00:48:45.660 --> 00:48:47.630 align:middle line:84%
We're adding squared
deviations.

00:48:47.630 --> 00:48:50.100 align:middle line:84%
So when you add non-negative
things, you get something

00:48:50.100 --> 00:48:51.400 align:middle line:90%
non-negative.

00:48:51.400 --> 00:48:55.800 align:middle line:84%
The next question is, how do
things scale if you take a

00:48:55.800 --> 00:48:59.880 align:middle line:84%
linear function of a
random variable?

00:48:59.880 --> 00:49:02.350 align:middle line:84%
Let's think about the
effects of beta.

00:49:02.350 --> 00:49:06.200 align:middle line:84%
If I take a random variable and
add the constant to it,

00:49:06.200 --> 00:49:09.820 align:middle line:84%
how does this affect the amount
of spread that we have?

00:49:09.820 --> 00:49:10.950 align:middle line:90%
It doesn't affect--

00:49:10.950 --> 00:49:14.610 align:middle line:84%
whatever the spread of this
thing is, if I add the

00:49:14.610 --> 00:49:18.840 align:middle line:84%
constant beta, it just moves
this diagram here, but the

00:49:18.840 --> 00:49:21.930 align:middle line:84%
spread doesn't grow
or get reduced.

00:49:21.930 --> 00:49:24.470 align:middle line:84%
The thing is that when I'm
adding a constant to a random

00:49:24.470 --> 00:49:28.160 align:middle line:84%
variable, all the x's that are
going to appear are further to

00:49:28.160 --> 00:49:32.890 align:middle line:84%
the right, but the expected
value also moves to the right.

00:49:32.890 --> 00:49:35.960 align:middle line:84%
And since we're only interested
in distances from

00:49:35.960 --> 00:49:39.500 align:middle line:84%
the mean, these distances
do not get affected.

00:49:39.500 --> 00:49:42.180 align:middle line:90%
x gets increased by something.

00:49:42.180 --> 00:49:44.390 align:middle line:84%
The mean gets increased by
that same something.

00:49:44.390 --> 00:49:46.180 align:middle line:90%
The difference stays the same.

00:49:46.180 --> 00:49:49.350 align:middle line:84%
So adding a constant to a random
variable doesn't do

00:49:49.350 --> 00:49:51.050 align:middle line:90%
anything to it's variance.

00:49:51.050 --> 00:49:54.940 align:middle line:84%
But if I multiply a random
variable by a constant alpha,

00:49:54.940 --> 00:49:58.730 align:middle line:84%
what is that going to
do to its variance?

00:49:58.730 --> 00:50:04.720 align:middle line:84%
Because we have a square here,
when I multiply my random

00:50:04.720 --> 00:50:08.430 align:middle line:84%
variable by a constant, this x
gets multiplied by a constant,

00:50:08.430 --> 00:50:12.310 align:middle line:84%
the mean gets multiplied by a
constant, the square gets

00:50:12.310 --> 00:50:15.650 align:middle line:84%
multiplied by the square
of that constant.

00:50:15.650 --> 00:50:18.960 align:middle line:84%
And because of that reason, we
get this square of alpha

00:50:18.960 --> 00:50:20.210 align:middle line:90%
showing up here.

00:50:20.210 --> 00:50:22.870 align:middle line:84%
So that's how variances
transform under linear

00:50:22.870 --> 00:50:23.650 align:middle line:90%
transformations.

00:50:23.650 --> 00:50:26.180 align:middle line:84%
You multiply your random
variable by constant, the

00:50:26.180 --> 00:50:30.540 align:middle line:84%
variance goes up by the square
of that same constant.

00:50:30.540 --> 00:50:31.290 align:middle line:90%
OK.

00:50:31.290 --> 00:50:32.950 align:middle line:90%
That's it for today.

00:50:32.950 --> 00:50:34.200 align:middle line:90%
See you on Wednesday.

00:50:34.200 --> 00:50:34.750 align:middle line:90%