WEBVTT

00:00:00.000 --> 00:00:00.040 align:middle line:90%


00:00:00.040 --> 00:00:02.460 align:middle line:84%
The following content is
provided under a Creative

00:00:02.460 --> 00:00:03.870 align:middle line:90%
Commons license.

00:00:03.870 --> 00:00:06.910 align:middle line:84%
Your support will help MIT
OpenCourseWare continue to

00:00:06.910 --> 00:00:08.700 align:middle line:90%
offer high quality, educational

00:00:08.700 --> 00:00:10.560 align:middle line:90%
resources for free.

00:00:10.560 --> 00:00:13.460 align:middle line:84%
To make a donation or view
additional materials from

00:00:13.460 --> 00:00:19.290 align:middle line:84%
hundreds of MIT courses, visit
MIT OpenCourseWare at

00:00:19.290 --> 00:00:20.540 align:middle line:90%
ocw.mit.edu.

00:00:20.540 --> 00:00:22.200 align:middle line:90%


00:00:22.200 --> 00:00:24.920 align:middle line:84%
PROFESSOR: So for the last three
lectures we're going to

00:00:24.920 --> 00:00:28.200 align:middle line:84%
talk about classical statistics,
the way statistics

00:00:28.200 --> 00:00:32.340 align:middle line:84%
can be done if you don't want to
assume a prior distribution

00:00:32.340 --> 00:00:34.800 align:middle line:90%
on the unknown parameters.

00:00:34.800 --> 00:00:38.290 align:middle line:84%
Today we're going to focus,
mostly, on the estimation side

00:00:38.290 --> 00:00:41.910 align:middle line:84%
and leave hypothesis testing
for the next two lectures.

00:00:41.910 --> 00:00:46.700 align:middle line:84%
So where there is one generic
method that one can use to

00:00:46.700 --> 00:00:50.850 align:middle line:84%
carry out parameter estimation,
that's the maximum

00:00:50.850 --> 00:00:51.850 align:middle line:90%
likelihood method.

00:00:51.850 --> 00:00:53.990 align:middle line:84%
We're going to define
what it is.

00:00:53.990 --> 00:00:58.200 align:middle line:84%
Then we will look at the most
common estimation problem

00:00:58.200 --> 00:01:00.620 align:middle line:84%
there is, which is to estimate
the mean of a given

00:01:00.620 --> 00:01:02.110 align:middle line:90%
distribution.

00:01:02.110 --> 00:01:05.540 align:middle line:84%
And we're going to talk about
confidence intervals, which

00:01:05.540 --> 00:01:09.130 align:middle line:84%
refers to providing an
interval around your

00:01:09.130 --> 00:01:13.330 align:middle line:84%
estimates, which has some
properties of the kind that

00:01:13.330 --> 00:01:17.640 align:middle line:84%
the parameter is highly likely
to be inside that interval,

00:01:17.640 --> 00:01:20.040 align:middle line:84%
but we will be careful about
how to interpret that

00:01:20.040 --> 00:01:22.220 align:middle line:90%
particular statement.

00:01:22.220 --> 00:01:22.345 align:middle line:90%
Ok.

00:01:22.345 --> 00:01:25.920 align:middle line:90%
So the big framework first.

00:01:25.920 --> 00:01:29.120 align:middle line:84%
The picture is almost the same
as the one that we had in the

00:01:29.120 --> 00:01:31.130 align:middle line:90%
case of Bayesian statistics.

00:01:31.130 --> 00:01:33.570 align:middle line:84%
We have some unknown
parameter.

00:01:33.570 --> 00:01:35.510 align:middle line:84%
And we have a measuring
device.

00:01:35.510 --> 00:01:38.150 align:middle line:84%
There is some noise,
some randomness.

00:01:38.150 --> 00:01:42.560 align:middle line:84%
And we get an observation, X,
whose distribution depends on

00:01:42.560 --> 00:01:44.560 align:middle line:90%
the value of the parameter.

00:01:44.560 --> 00:01:47.850 align:middle line:84%
However, the big change from the
Bayesian setting is that

00:01:47.850 --> 00:01:50.840 align:middle line:84%
here, this parameter
is just a number.

00:01:50.840 --> 00:01:53.200 align:middle line:84%
It's not modeled as
a random variable.

00:01:53.200 --> 00:01:55.900 align:middle line:84%
It does not have a probability
distribution.

00:01:55.900 --> 00:01:57.460 align:middle line:84%
There's nothing random
about it.

00:01:57.460 --> 00:01:58.720 align:middle line:90%
It's a constant.

00:01:58.720 --> 00:02:02.360 align:middle line:84%
It just happens that we don't
know what that constant is.

00:02:02.360 --> 00:02:05.970 align:middle line:84%
And in particular, this
probability distribution here,

00:02:05.970 --> 00:02:10.350 align:middle line:84%
the distribution of X,
depends on Theta.

00:02:10.350 --> 00:02:13.900 align:middle line:84%
But this is not a conditional
distribution in the usual

00:02:13.900 --> 00:02:15.450 align:middle line:90%
sense of the word.

00:02:15.450 --> 00:02:18.480 align:middle line:84%
Conditional distributions were
defined when we had two random

00:02:18.480 --> 00:02:21.800 align:middle line:84%
variables and we condition one
random variable on the other.

00:02:21.800 --> 00:02:25.890 align:middle line:84%
And we used the bar to separate
the X from the Theta.

00:02:25.890 --> 00:02:27.870 align:middle line:84%
To make the point that this
is not a conditioned

00:02:27.870 --> 00:02:29.840 align:middle line:84%
distribution, we use a
different notation.

00:02:29.840 --> 00:02:31.730 align:middle line:90%
We put a semicolon here.

00:02:31.730 --> 00:02:35.760 align:middle line:84%
And what this is meant to say is
that X has a distribution.

00:02:35.760 --> 00:02:39.640 align:middle line:84%
That distribution has
a certain parameter.

00:02:39.640 --> 00:02:42.240 align:middle line:84%
And we don't know what
that parameter is.

00:02:42.240 --> 00:02:46.270 align:middle line:84%
So for example, this might be
a normal distribution, with

00:02:46.270 --> 00:02:49.070 align:middle line:90%
variance 1 but a mean Theta.

00:02:49.070 --> 00:02:50.560 align:middle line:90%
We don't know what Theta is.

00:02:50.560 --> 00:02:52.980 align:middle line:90%
And we want to estimate it.

00:02:52.980 --> 00:02:55.970 align:middle line:84%
Now once we have this setting,
then your job is to design

00:02:55.970 --> 00:02:57.560 align:middle line:90%
this box, the estimator.

00:02:57.560 --> 00:03:00.620 align:middle line:84%
The estimator is some data
processing box that takes the

00:03:00.620 --> 00:03:03.950 align:middle line:84%
measurements and produces
an estimate

00:03:03.950 --> 00:03:06.300 align:middle line:90%
of the unknown parameter.

00:03:06.300 --> 00:03:11.950 align:middle line:84%
Now the notation that's used
here is as if X and Theta were

00:03:11.950 --> 00:03:13.640 align:middle line:90%
one-dimensional quantities.

00:03:13.640 --> 00:03:16.610 align:middle line:84%
But actually, everything we
say remains valid if you

00:03:16.610 --> 00:03:20.090 align:middle line:84%
interpret X and Theta as
vectors of parameters.

00:03:20.090 --> 00:03:22.180 align:middle line:84%
So for example, you
may obtain several

00:03:22.180 --> 00:03:25.050 align:middle line:90%
measurements, X1 up to 2Xn.

00:03:25.050 --> 00:03:27.980 align:middle line:84%
And there may be several unknown
parameters in the

00:03:27.980 --> 00:03:30.260 align:middle line:90%
background.

00:03:30.260 --> 00:03:34.200 align:middle line:84%
Once more, we do not have, and
we do not want to assume, a

00:03:34.200 --> 00:03:35.780 align:middle line:90%
prior distribution on Theta.

00:03:35.780 --> 00:03:37.070 align:middle line:90%
It's a constant.

00:03:37.070 --> 00:03:39.040 align:middle line:84%
And if you want to think
mathematically about this

00:03:39.040 --> 00:03:41.510 align:middle line:84%
situation, it's as if you
have many different

00:03:41.510 --> 00:03:43.340 align:middle line:90%
probabilistic models.

00:03:43.340 --> 00:03:46.360 align:middle line:84%
So a normal with this mean or
a normal with that mean or a

00:03:46.360 --> 00:03:49.020 align:middle line:84%
normal with that mean, these
are alternative candidate

00:03:49.020 --> 00:03:50.700 align:middle line:90%
probabilistic models.

00:03:50.700 --> 00:03:55.080 align:middle line:84%
And we want to try to make a
decision about which one is

00:03:55.080 --> 00:03:56.420 align:middle line:90%
the correct model.

00:03:56.420 --> 00:03:59.480 align:middle line:84%
In some cases, we have to choose
just between a small

00:03:59.480 --> 00:04:00.390 align:middle line:90%
number of models.

00:04:00.390 --> 00:04:03.400 align:middle line:84%
For example, you have a coin
with an unknown bias.

00:04:03.400 --> 00:04:06.410 align:middle line:90%
The bias is either 1/2 or 3/4.

00:04:06.410 --> 00:04:08.650 align:middle line:84%
You're going to flip the
coin a few times.

00:04:08.650 --> 00:04:13.150 align:middle line:84%
And you try to decide whether
the true bias is this one or

00:04:13.150 --> 00:04:14.150 align:middle line:90%
is that one.

00:04:14.150 --> 00:04:17.610 align:middle line:84%
So in this case, we have two
specific, alternative

00:04:17.610 --> 00:04:20.800 align:middle line:84%
probabilistic models from which
we want to distinguish.

00:04:20.800 --> 00:04:25.000 align:middle line:84%
But sometimes things are a
little more complicated.

00:04:25.000 --> 00:04:26.940 align:middle line:90%
For example, you have a coin.

00:04:26.940 --> 00:04:30.940 align:middle line:84%
And you have one hypothesis
that my coin is unbiased.

00:04:30.940 --> 00:04:34.650 align:middle line:84%
And the other hypothesis is
that my coin is biased.

00:04:34.650 --> 00:04:36.040 align:middle line:90%
And you do your experiments.

00:04:36.040 --> 00:04:40.840 align:middle line:84%
And you want to come up with a
decision that decides whether

00:04:40.840 --> 00:04:43.970 align:middle line:84%
this is true or this
one is true.

00:04:43.970 --> 00:04:46.630 align:middle line:84%
In this case, we're not
dealing with just two

00:04:46.630 --> 00:04:48.710 align:middle line:84%
alternative probabilistic
models.

00:04:48.710 --> 00:04:51.540 align:middle line:84%
This one is a specific
model for the coin.

00:04:51.540 --> 00:04:54.230 align:middle line:84%
But this one actually
corresponds to lots of

00:04:54.230 --> 00:04:56.890 align:middle line:84%
possible, alternative
coin models.

00:04:56.890 --> 00:05:00.420 align:middle line:84%
So this includes the model where
Theta is 0.6, the model

00:05:00.420 --> 00:05:03.860 align:middle line:84%
where Theta is 0.7, Theta
is 0.8, and so on.

00:05:03.860 --> 00:05:07.350 align:middle line:84%
So we're trying to discriminate
between one model

00:05:07.350 --> 00:05:09.510 align:middle line:84%
and lots of alternative
models.

00:05:09.510 --> 00:05:11.560 align:middle line:90%
How does one go about this?

00:05:11.560 --> 00:05:14.750 align:middle line:84%
Well, there's some systematic
ways that one can approach

00:05:14.750 --> 00:05:16.120 align:middle line:90%
problems of this kind.

00:05:16.120 --> 00:05:19.850 align:middle line:84%
And we will start talking
about these next time.

00:05:19.850 --> 00:05:22.380 align:middle line:84%
So today, we're going to focus
on estimation problems.

00:05:22.380 --> 00:05:27.080 align:middle line:84%
In estimation problems, theta is
a quantity, which is a real

00:05:27.080 --> 00:05:29.070 align:middle line:84%
number, a continuous
parameter.

00:05:29.070 --> 00:05:33.730 align:middle line:84%
We're to design this box, so
what we get out of this box is

00:05:33.730 --> 00:05:34.280 align:middle line:90%
an estimate.

00:05:34.280 --> 00:05:37.900 align:middle line:84%
Now notice that this estimate
here is a random variable.

00:05:37.900 --> 00:05:42.000 align:middle line:84%
Even though theta is
deterministic, this is random,

00:05:42.000 --> 00:05:45.110 align:middle line:84%
because it's a function of
the data that we observe.

00:05:45.110 --> 00:05:46.360 align:middle line:90%
The data are random.

00:05:46.360 --> 00:05:49.155 align:middle line:84%
We're applying a function
to the data to

00:05:49.155 --> 00:05:50.270 align:middle line:90%
construct our estimate.

00:05:50.270 --> 00:05:52.850 align:middle line:84%
So, since it's a function of
random variables, it's a

00:05:52.850 --> 00:05:54.630 align:middle line:90%
random variable itself.

00:05:54.630 --> 00:05:57.940 align:middle line:84%
The distribution of Theta hat
depends on the distribution of

00:05:57.940 --> 00:06:01.280 align:middle line:84%
X. The distribution of X
is affected by Theta.

00:06:01.280 --> 00:06:03.650 align:middle line:84%
So in the end, the distribution
of your estimate

00:06:03.650 --> 00:06:08.290 align:middle line:84%
Theta hat will also be affected
by whatever Theta

00:06:08.290 --> 00:06:09.920 align:middle line:90%
happens to be.

00:06:09.920 --> 00:06:12.950 align:middle line:84%
Our general objective, when
designing estimators, is that

00:06:12.950 --> 00:06:17.390 align:middle line:84%
we want to get, in the end, an
error, an estimation error,

00:06:17.390 --> 00:06:19.070 align:middle line:90%
which is not too large.

00:06:19.070 --> 00:06:21.500 align:middle line:84%
But we'll have to make
that specific.

00:06:21.500 --> 00:06:24.720 align:middle line:84%
Again, what exactly do
we mean by that?

00:06:24.720 --> 00:06:27.170 align:middle line:84%
So how do we go about
this problem?

00:06:27.170 --> 00:06:29.670 align:middle line:90%


00:06:29.670 --> 00:06:40.150 align:middle line:84%
One general approach is to pick
a Theta, under which the

00:06:40.150 --> 00:06:44.590 align:middle line:84%
data that we observe, that
this is the X's, our most

00:06:44.590 --> 00:06:47.180 align:middle line:90%
likely to have occurred.

00:06:47.180 --> 00:06:52.700 align:middle line:84%
So I observe X. For any given
Theta, I can calculate this

00:06:52.700 --> 00:06:56.630 align:middle line:84%
quantity, which tells me, under
this particular Theta,

00:06:56.630 --> 00:07:00.670 align:middle line:84%
the X that you observed had this
probability of occurring.

00:07:00.670 --> 00:07:03.270 align:middle line:84%
Under that Theta, the X that
you observe had that

00:07:03.270 --> 00:07:04.770 align:middle line:90%
probability of occurring.

00:07:04.770 --> 00:07:08.580 align:middle line:84%
You just choose that Theta,
which makes the data that you

00:07:08.580 --> 00:07:12.700 align:middle line:90%
observed most likely.

00:07:12.700 --> 00:07:15.810 align:middle line:84%
It's interesting to compare
this maximum likelihood

00:07:15.810 --> 00:07:19.120 align:middle line:84%
estimate with the estimates that
you would have, if you

00:07:19.120 --> 00:07:22.050 align:middle line:84%
were in a Bayesian setting,
and you were using maximum

00:07:22.050 --> 00:07:25.010 align:middle line:84%
approach theory probability
estimation.

00:07:25.010 --> 00:07:31.650 align:middle line:84%
In the Bayesian setting, what
we do is, given the data, we

00:07:31.650 --> 00:07:34.350 align:middle line:84%
use the prior distribution
on Theta.

00:07:34.350 --> 00:07:41.660 align:middle line:84%
And we calculate the posterior
distribution of Theta given X.

00:07:41.660 --> 00:07:44.350 align:middle line:84%
Notice that this is sort
of the opposite from

00:07:44.350 --> 00:07:46.040 align:middle line:90%
what we have here.

00:07:46.040 --> 00:07:49.180 align:middle line:84%
This is the probability of X
for a particular value of

00:07:49.180 --> 00:07:51.780 align:middle line:84%
Theta, whereas this is the
probability of Theta for a

00:07:51.780 --> 00:07:55.380 align:middle line:84%
particular X. So it's the
opposite type of conditioning.

00:07:55.380 --> 00:07:58.240 align:middle line:84%
In the Bayesian setting, Theta
is a random variable.

00:07:58.240 --> 00:07:59.890 align:middle line:84%
So we can talk about
the probability

00:07:59.890 --> 00:08:01.570 align:middle line:90%
distribution of Theta.

00:08:01.570 --> 00:08:04.740 align:middle line:84%
So how do these two compare,
except for this syntactic

00:08:04.740 --> 00:08:08.160 align:middle line:84%
difference that the order X's
and Theta's are reversed?

00:08:08.160 --> 00:08:11.410 align:middle line:84%
Let's write down, in full
detail, what this posterior

00:08:11.410 --> 00:08:13.280 align:middle line:90%
distribution of Theta is.

00:08:13.280 --> 00:08:17.390 align:middle line:84%
By the Bayes rule, this
conditional distribution is

00:08:17.390 --> 00:08:20.430 align:middle line:84%
obtained from the prior, and the
model of the measurement

00:08:20.430 --> 00:08:21.850 align:middle line:90%
process that we have.

00:08:21.850 --> 00:08:24.510 align:middle line:90%
And we get to this expression.

00:08:24.510 --> 00:08:29.520 align:middle line:84%
So in Bayesian estimation, we
want to find the most likely

00:08:29.520 --> 00:08:30.870 align:middle line:90%
value of Theta.

00:08:30.870 --> 00:08:33.070 align:middle line:84%
And we need to maximize
this quantity over

00:08:33.070 --> 00:08:34.539 align:middle line:90%
all possible Theta's.

00:08:34.539 --> 00:08:38.210 align:middle line:84%
First thing to notice is that
the denominator is a constant.

00:08:38.210 --> 00:08:40.220 align:middle line:90%
It does not involve Theta.

00:08:40.220 --> 00:08:43.250 align:middle line:84%
So when you maximize this
quantity, you don't care about

00:08:43.250 --> 00:08:44.520 align:middle line:90%
the denominator.

00:08:44.520 --> 00:08:47.800 align:middle line:84%
You just want to maximize
the numerator.

00:08:47.800 --> 00:08:52.310 align:middle line:84%
Now, here, things start to look
a little more similar.

00:08:52.310 --> 00:08:56.530 align:middle line:84%
And they would be exactly of
the same kind, if that term

00:08:56.530 --> 00:08:59.890 align:middle line:84%
here was absent, it the
prior was absent.

00:08:59.890 --> 00:09:03.860 align:middle line:84%
The two are going to become
the same if that prior was

00:09:03.860 --> 00:09:05.830 align:middle line:90%
just a constant.

00:09:05.830 --> 00:09:10.160 align:middle line:84%
So if that prior is a constant,
then maximum

00:09:10.160 --> 00:09:13.720 align:middle line:84%
likelihood estimation takes
exactly the same form as

00:09:13.720 --> 00:09:17.360 align:middle line:84%
Bayesian maximum posterior
probability estimation.

00:09:17.360 --> 00:09:21.230 align:middle line:84%
So you can give this particular
interpretation of

00:09:21.230 --> 00:09:22.680 align:middle line:90%
maximum likelihood estimation.

00:09:22.680 --> 00:09:27.400 align:middle line:84%
Maximum likelihood estimation
is essentially what you have

00:09:27.400 --> 00:09:31.380 align:middle line:84%
done, if you were in a Bayesian
world, and you had

00:09:31.380 --> 00:09:35.400 align:middle line:84%
assumed a prior on the Theta's
that's uniform, all the

00:09:35.400 --> 00:09:37.030 align:middle line:90%
Theta's being equally likely.

00:09:37.030 --> 00:09:42.620 align:middle line:90%


00:09:42.620 --> 00:09:42.725 align:middle line:90%
Okay.

00:09:42.725 --> 00:09:45.770 align:middle line:84%
So let's look at a
simple example.

00:09:45.770 --> 00:09:48.510 align:middle line:84%
Suppose that the Xi's are
independent, identically

00:09:48.510 --> 00:09:50.770 align:middle line:84%
distributed random
variables, with a

00:09:50.770 --> 00:09:52.690 align:middle line:90%
certain parameter Theta.

00:09:52.690 --> 00:09:55.910 align:middle line:84%
So the distribution of each
one of the Xi's is this

00:09:55.910 --> 00:09:57.950 align:middle line:90%
particular term.

00:09:57.950 --> 00:09:59.840 align:middle line:90%
So Theta is one-dimensional.

00:09:59.840 --> 00:10:01.280 align:middle line:84%
It's a one-dimensional
parameter.

00:10:01.280 --> 00:10:03.180 align:middle line:90%
But we have several data.

00:10:03.180 --> 00:10:07.020 align:middle line:84%
We write down the formula
for the probability of a

00:10:07.020 --> 00:10:12.360 align:middle line:84%
particular X vector, given a
particular value of Theta.

00:10:12.360 --> 00:10:14.950 align:middle line:84%
But again, when I use the word,
given, here it's not in

00:10:14.950 --> 00:10:16.080 align:middle line:90%
the conditioning sense.

00:10:16.080 --> 00:10:18.770 align:middle line:84%
It's the value of the
density for a

00:10:18.770 --> 00:10:21.710 align:middle line:90%
particular choice of Theta.

00:10:21.710 --> 00:10:24.890 align:middle line:84%
Here, I wrote down, I defined
maximum likelihood estimation

00:10:24.890 --> 00:10:26.190 align:middle line:90%
in terms of PMFs.

00:10:26.190 --> 00:10:28.050 align:middle line:84%
That's what you would
do if the X's were

00:10:28.050 --> 00:10:29.950 align:middle line:90%
discrete random variables.

00:10:29.950 --> 00:10:32.770 align:middle line:84%
Here, the X's are continuous
random variables, so instead

00:10:32.770 --> 00:10:36.220 align:middle line:84%
of I'm using the PDF
instead of the PMF.

00:10:36.220 --> 00:10:39.530 align:middle line:84%
So this a definition, here,
generalizes to the case of

00:10:39.530 --> 00:10:40.900 align:middle line:90%
continuous random variables.

00:10:40.900 --> 00:10:44.620 align:middle line:84%
And you use F's instead of
X's, our usual recipe.

00:10:44.620 --> 00:10:47.560 align:middle line:84%
So the maximum likelihood
estimate is defined.

00:10:47.560 --> 00:10:51.880 align:middle line:84%
Now, since the Xi's are
independent, the joint density

00:10:51.880 --> 00:10:54.410 align:middle line:84%
of all the X's together
is the product of

00:10:54.410 --> 00:10:57.680 align:middle line:90%
the individual densities.

00:10:57.680 --> 00:10:59.170 align:middle line:90%
So you look at this quantity.

00:10:59.170 --> 00:11:03.310 align:middle line:84%
This is the density or sort of
probability of observing a

00:11:03.310 --> 00:11:05.340 align:middle line:90%
particular sequence of X's.

00:11:05.340 --> 00:11:08.230 align:middle line:84%
And we ask the question, what's
the value of Theta that

00:11:08.230 --> 00:11:10.940 align:middle line:84%
makes the X's that we
observe most likely?

00:11:10.940 --> 00:11:13.160 align:middle line:84%
So we want to carry out
this maximization.

00:11:13.160 --> 00:11:17.430 align:middle line:84%
Now this maximization is just
a calculational problem.

00:11:17.430 --> 00:11:19.920 align:middle line:84%
We're going to do this
maximization by taking the

00:11:19.920 --> 00:11:21.910 align:middle line:90%
logarithm of this expression.

00:11:21.910 --> 00:11:23.880 align:middle line:84%
Maximizing an expression
is the same as

00:11:23.880 --> 00:11:25.790 align:middle line:90%
maximizing the logarithm.

00:11:25.790 --> 00:11:28.790 align:middle line:84%
So the logarithm of this
expression, the logarithm of a

00:11:28.790 --> 00:11:31.290 align:middle line:84%
product is the sum of
the logarithms.

00:11:31.290 --> 00:11:34.390 align:middle line:84%
You get contributions from
this Theta term.

00:11:34.390 --> 00:11:37.660 align:middle line:84%
There's n of these, so we
get an n log Theta.

00:11:37.660 --> 00:11:40.430 align:middle line:84%
And then we have the sum of the
logarithms of these terms.

00:11:40.430 --> 00:11:43.060 align:middle line:90%
It gives us minus Theta.

00:11:43.060 --> 00:11:45.020 align:middle line:90%
And then the sum of the X's.

00:11:45.020 --> 00:11:47.060 align:middle line:84%
So we need to maximize
this expression

00:11:47.060 --> 00:11:48.630 align:middle line:90%
with respect to Theta.

00:11:48.630 --> 00:11:51.130 align:middle line:84%
The way to do this maximization
is you take the

00:11:51.130 --> 00:11:53.320 align:middle line:84%
derivative, with respect
to Theta.

00:11:53.320 --> 00:11:58.520 align:middle line:84%
And you get n over Theta equals
to the sum of the X's.

00:11:58.520 --> 00:12:00.280 align:middle line:90%
And then you solve for Theta.

00:12:00.280 --> 00:12:02.040 align:middle line:84%
And you find that the
maximum likelihood

00:12:02.040 --> 00:12:04.680 align:middle line:90%
estimate is this quantity.

00:12:04.680 --> 00:12:13.160 align:middle line:84%
Which sort of makes sense,
because this is the reciprocal

00:12:13.160 --> 00:12:16.700 align:middle line:90%
of the sample mean of X's.

00:12:16.700 --> 00:12:19.520 align:middle line:84%
Theta, in an exponential
distribution, we know that

00:12:19.520 --> 00:12:23.380 align:middle line:84%
it's 1 over (the mean of the
exponential distribution).

00:12:23.380 --> 00:12:26.960 align:middle line:84%
So it looks like a reasonable
estimate.

00:12:26.960 --> 00:12:29.570 align:middle line:84%
So in any case, this is the
estimates that the maximum

00:12:29.570 --> 00:12:33.420 align:middle line:84%
likelihood estimation procedure
tells us that we

00:12:33.420 --> 00:12:35.780 align:middle line:90%
should report.

00:12:35.780 --> 00:12:39.790 align:middle line:84%
This formula here, of course,
tells you what to do if you

00:12:39.790 --> 00:12:42.640 align:middle line:84%
have already observed
specific numbers.

00:12:42.640 --> 00:12:46.020 align:middle line:84%
If you have observed specific
numbers, then you observe this

00:12:46.020 --> 00:12:49.110 align:middle line:84%
particular number as your
estimate of Theta.

00:12:49.110 --> 00:12:52.000 align:middle line:84%
If you want to describe your
estimation procedure more

00:12:52.000 --> 00:12:55.900 align:middle line:84%
abstractly, what you have
constructed is an estimator,

00:12:55.900 --> 00:12:59.690 align:middle line:84%
which is a box that's takes in
the random variables, capital

00:12:59.690 --> 00:13:05.430 align:middle line:84%
X1 up to Capital Xn, and
produces out your estimate,

00:13:05.430 --> 00:13:07.440 align:middle line:84%
which is also a random
variable.

00:13:07.440 --> 00:13:10.760 align:middle line:84%
Because it's a function of these
random variables and is

00:13:10.760 --> 00:13:14.750 align:middle line:84%
denoted by an upper case Theta,
to indicate that this

00:13:14.750 --> 00:13:17.470 align:middle line:90%
is now a random variable.

00:13:17.470 --> 00:13:21.040 align:middle line:84%
So this is an equality
about numbers.

00:13:21.040 --> 00:13:23.860 align:middle line:84%
This is a description of the
general procedure, which is an

00:13:23.860 --> 00:13:25.745 align:middle line:84%
equality between two
random variables.

00:13:25.745 --> 00:13:28.360 align:middle line:90%


00:13:28.360 --> 00:13:31.920 align:middle line:84%
And this gives you the more
abstract view of what we're

00:13:31.920 --> 00:13:35.040 align:middle line:90%
doing here.

00:13:35.040 --> 00:13:35.352 align:middle line:90%
All right.

00:13:35.352 --> 00:13:37.970 align:middle line:84%
So what can we tell about
our estimate?

00:13:37.970 --> 00:13:40.090 align:middle line:90%
Is it good or is it bad?

00:13:40.090 --> 00:13:42.960 align:middle line:84%
So we should look at this
particular random variable and

00:13:42.960 --> 00:13:46.220 align:middle line:84%
talk about the statistical
properties that it has.

00:13:46.220 --> 00:13:49.930 align:middle line:84%
What we would like is this
random variable to be close to

00:13:49.930 --> 00:13:55.810 align:middle line:84%
the true value of Theta, with
high probability, no matter

00:13:55.810 --> 00:13:59.470 align:middle line:84%
what Theta is, since we don't
know what Theta is.

00:13:59.470 --> 00:14:01.400 align:middle line:84%
Let's make a little
more specific the

00:14:01.400 --> 00:14:05.100 align:middle line:90%
properties that we want.

00:14:05.100 --> 00:14:08.470 align:middle line:84%
So we cook up the estimator
somehow.

00:14:08.470 --> 00:14:11.850 align:middle line:84%
So this estimator corresponds,
again, to a box that takes

00:14:11.850 --> 00:14:15.400 align:middle line:84%
data in, the capital X's,
and produces an

00:14:15.400 --> 00:14:17.470 align:middle line:90%
estimate Theta hat.

00:14:17.470 --> 00:14:18.710 align:middle line:90%
This estimate is random.

00:14:18.710 --> 00:14:23.070 align:middle line:84%
Sometimes it will be above
the true value of Theta.

00:14:23.070 --> 00:14:25.660 align:middle line:90%
Sometimes it will be below.

00:14:25.660 --> 00:14:30.220 align:middle line:84%
Ideally, we would like it to not
have a systematic error,

00:14:30.220 --> 00:14:32.810 align:middle line:84%
on the positive side or
the negative side.

00:14:32.810 --> 00:14:37.310 align:middle line:84%
So a reasonable wish to have,
for a good estimator, is that,

00:14:37.310 --> 00:14:41.700 align:middle line:84%
on the average, it gives
you the correct value.

00:14:41.700 --> 00:14:45.850 align:middle line:84%
Now here, let's be a little more
specific about what that

00:14:45.850 --> 00:14:47.740 align:middle line:90%
expectation is.

00:14:47.740 --> 00:14:51.270 align:middle line:84%
This is an expectation, with
respect to the probability

00:14:51.270 --> 00:14:54.240 align:middle line:90%
distribution of Theta hat.

00:14:54.240 --> 00:14:58.780 align:middle line:84%
The probability distribution
of Theta hat is affected by

00:14:58.780 --> 00:15:01.410 align:middle line:84%
the probability distribution
of the X's.

00:15:01.410 --> 00:15:03.760 align:middle line:84%
Because Theta hat is a
function of the X's.

00:15:03.760 --> 00:15:05.930 align:middle line:84%
And the probability distribution
of the X's is

00:15:05.930 --> 00:15:09.220 align:middle line:84%
affected by the true
value of Theta.

00:15:09.220 --> 00:15:13.710 align:middle line:84%
So depending on which one is the
true value of Theta, this

00:15:13.710 --> 00:15:16.650 align:middle line:84%
is going to be a different
expectation.

00:15:16.650 --> 00:15:20.830 align:middle line:84%
So if you were to write this
expectation out in more

00:15:20.830 --> 00:15:25.840 align:middle line:84%
detail, it would look
something like this.

00:15:25.840 --> 00:15:28.690 align:middle line:84%
You need to write down
the probability

00:15:28.690 --> 00:15:30.260 align:middle line:90%
distribution of Theta hat.

00:15:30.260 --> 00:15:32.890 align:middle line:90%


00:15:32.890 --> 00:15:36.470 align:middle line:84%
And this is going to
be some function.

00:15:36.470 --> 00:15:41.200 align:middle line:84%
But this function depends on the
true Theta, is affected by

00:15:41.200 --> 00:15:42.800 align:middle line:90%
the true Theta.

00:15:42.800 --> 00:15:48.300 align:middle line:84%
And then you integrate this
with respect to Theta hat.

00:15:48.300 --> 00:15:49.430 align:middle line:90%
What's the point here?

00:15:49.430 --> 00:15:53.280 align:middle line:84%
Again, Theta hat is a
function of the X's.

00:15:53.280 --> 00:15:57.000 align:middle line:84%
So the density of Theta
hat is affected by the

00:15:57.000 --> 00:15:58.400 align:middle line:90%
density of the X's.

00:15:58.400 --> 00:16:00.730 align:middle line:84%
The density of the X's
is affected by the

00:16:00.730 --> 00:16:02.380 align:middle line:90%
true value of Theta.

00:16:02.380 --> 00:16:05.420 align:middle line:84%
So the distribution of Theta
hat is affected by

00:16:05.420 --> 00:16:07.680 align:middle line:90%
the value of Theta.

00:16:07.680 --> 00:16:10.500 align:middle line:84%
Another way to put it is, as
I've mentioned a few minutes

00:16:10.500 --> 00:16:14.550 align:middle line:84%
ago, in this business, it's
as if we are considering

00:16:14.550 --> 00:16:17.880 align:middle line:84%
different possible probabilistic
models, one

00:16:17.880 --> 00:16:20.470 align:middle line:84%
probabilistic model for
each choice of Theta.

00:16:20.470 --> 00:16:22.560 align:middle line:84%
And we're trying to guess
which one of these

00:16:22.560 --> 00:16:25.200 align:middle line:84%
probabilistic models
is the true one.

00:16:25.200 --> 00:16:28.420 align:middle line:84%
One way of emphasizing the
fact that this expression

00:16:28.420 --> 00:16:31.780 align:middle line:84%
depends on the true Theta is
to put a little subscript

00:16:31.780 --> 00:16:36.790 align:middle line:84%
here, expectation, under the
particular value of the

00:16:36.790 --> 00:16:38.300 align:middle line:90%
parameter Theta.

00:16:38.300 --> 00:16:42.450 align:middle line:84%
So depending on what value the
true parameter Theta takes,

00:16:42.450 --> 00:16:45.000 align:middle line:84%
this expectation will have
a different value.

00:16:45.000 --> 00:16:49.730 align:middle line:84%
And what we would like is that
no matter what the true value

00:16:49.730 --> 00:16:55.300 align:middle line:84%
is, that our estimate will not
have a bias on the positive or

00:16:55.300 --> 00:16:57.140 align:middle line:90%
the negative sides.

00:16:57.140 --> 00:17:00.150 align:middle line:84%
So this is a property
that's desirable.

00:17:00.150 --> 00:17:02.160 align:middle line:90%
Is it always going to be true?

00:17:02.160 --> 00:17:05.218 align:middle line:84%
Not necessarily, it depends on
what estimator we construct.

00:17:05.218 --> 00:17:09.160 align:middle line:90%


00:17:09.160 --> 00:17:12.400 align:middle line:84%
Is it true for our exponential
example?

00:17:12.400 --> 00:17:14.770 align:middle line:84%
Unfortunately not, the estimate
that we have in the

00:17:14.770 --> 00:17:18.300 align:middle line:84%
exponential example turns
out to be biased.

00:17:18.300 --> 00:17:22.900 align:middle line:84%
And one extreme way of seeing
this is to consider the case

00:17:22.900 --> 00:17:25.160 align:middle line:90%
where our sample size is 1.

00:17:25.160 --> 00:17:27.020 align:middle line:84%
We're trying to estimate
Theta.

00:17:27.020 --> 00:17:30.370 align:middle line:84%
And the estimator from the
previous slide, in that case,

00:17:30.370 --> 00:17:33.410 align:middle line:90%
is just 1/X1.

00:17:33.410 --> 00:17:37.990 align:middle line:84%
Now X1 has a fair amount of
density in the vicinity of 0,

00:17:37.990 --> 00:17:41.360 align:middle line:84%
which means that 1/X1 has
significant probability of

00:17:41.360 --> 00:17:42.810 align:middle line:90%
being very large.

00:17:42.810 --> 00:17:46.140 align:middle line:84%
And if you do the calculation,
this ultimately makes the

00:17:46.140 --> 00:17:49.170 align:middle line:84%
expected value of 1/X1
to be infinite.

00:17:49.170 --> 00:17:52.870 align:middle line:84%
Now infinity is definitely
not the correct value.

00:17:52.870 --> 00:17:56.330 align:middle line:84%
So our estimate is
biased upwards.

00:17:56.330 --> 00:18:00.130 align:middle line:84%
And it's actually biased
a lot upwards.

00:18:00.130 --> 00:18:01.800 align:middle line:90%
So that's how things are.

00:18:01.800 --> 00:18:06.690 align:middle line:84%
Maximum likelihood estimates,
in general, will be biased.

00:18:06.690 --> 00:18:10.870 align:middle line:84%
But under some conditions,
they will turn out to be

00:18:10.870 --> 00:18:12.780 align:middle line:90%
asymptotically unbiased.

00:18:12.780 --> 00:18:16.810 align:middle line:84%
That is, as you get more and
more data, as your X vector is

00:18:16.810 --> 00:18:21.750 align:middle line:84%
longer and longer, with
independent data, the estimate

00:18:21.750 --> 00:18:25.010 align:middle line:84%
that you're going to have, the
expected value of your

00:18:25.010 --> 00:18:26.860 align:middle line:84%
estimator is going
to get closer and

00:18:26.860 --> 00:18:28.370 align:middle line:90%
closer to the true value.

00:18:28.370 --> 00:18:31.330 align:middle line:84%
So you do have some nice
asymptotic properties, but

00:18:31.330 --> 00:18:34.000 align:middle line:84%
we're not going to prove
anything like this.

00:18:34.000 --> 00:18:37.680 align:middle line:84%
Speaking of asymptotic
properties, in general, what

00:18:37.680 --> 00:18:40.950 align:middle line:84%
we would like to have is that,
as you collect more and more

00:18:40.950 --> 00:18:46.550 align:middle line:84%
data, you get the correct
answer, in some sense.

00:18:46.550 --> 00:18:49.360 align:middle line:84%
And the sense that we're going
to use here is the limiting

00:18:49.360 --> 00:18:52.560 align:middle line:84%
sense of convergence in
probability, since this is the

00:18:52.560 --> 00:18:55.270 align:middle line:84%
only notion of convergence of
random variables that we have

00:18:55.270 --> 00:18:56.540 align:middle line:90%
in our hands.

00:18:56.540 --> 00:18:59.600 align:middle line:84%
This is similar to what
we had in the pollster

00:18:59.600 --> 00:19:01.180 align:middle line:90%
problem, for example.

00:19:01.180 --> 00:19:04.900 align:middle line:84%
If we had a bigger and bigger
sample size, we could be more

00:19:04.900 --> 00:19:08.360 align:middle line:84%
and more confident that the
estimate that we obtained is

00:19:08.360 --> 00:19:11.970 align:middle line:84%
close to the unknown true
parameter of the distribution

00:19:11.970 --> 00:19:13.320 align:middle line:90%
that we have.

00:19:13.320 --> 00:19:16.420 align:middle line:84%
So this is a desirable
property.

00:19:16.420 --> 00:19:20.720 align:middle line:84%
If you have an infinitely large
amount of data, you

00:19:20.720 --> 00:19:25.070 align:middle line:84%
should be able to estimate
an unknown parameter

00:19:25.070 --> 00:19:26.890 align:middle line:90%
more or less exactly.

00:19:26.890 --> 00:19:32.280 align:middle line:84%
So this is it desirable property
of estimators.

00:19:32.280 --> 00:19:35.560 align:middle line:84%
It turns out that maximum
likelihood estimation, given

00:19:35.560 --> 00:19:39.330 align:middle line:84%
independent data, does have
this property, under mild

00:19:39.330 --> 00:19:40.280 align:middle line:90%
conditions.

00:19:40.280 --> 00:19:43.100 align:middle line:84%
So maximum likelihood
estimation, in this respect,

00:19:43.100 --> 00:19:45.180 align:middle line:90%
is a good approach.

00:19:45.180 --> 00:19:48.520 align:middle line:84%
So let's see, do we have this
consistency property in our

00:19:48.520 --> 00:19:50.150 align:middle line:90%
exponential example?

00:19:50.150 --> 00:19:56.560 align:middle line:84%
In our exponential example, we
used this quantity to estimate

00:19:56.560 --> 00:19:59.040 align:middle line:90%
the unknown parameter Theta.

00:19:59.040 --> 00:20:01.000 align:middle line:84%
What properties does
this quantity have

00:20:01.000 --> 00:20:03.160 align:middle line:90%
as n goes to infinity?

00:20:03.160 --> 00:20:06.580 align:middle line:84%
Well this quantity is the
reciprocal of that quantity up

00:20:06.580 --> 00:20:09.890 align:middle line:84%
here, which is the
sample mean.

00:20:09.890 --> 00:20:12.950 align:middle line:84%
We know from the weak law of
large numbers, that the sample

00:20:12.950 --> 00:20:16.350 align:middle line:84%
mean converges to
the expectation.

00:20:16.350 --> 00:20:19.250 align:middle line:84%
So this property here
comes from the weak

00:20:19.250 --> 00:20:21.660 align:middle line:90%
law of large numbers.

00:20:21.660 --> 00:20:24.670 align:middle line:84%
In probability, this quantity
converges to the expected

00:20:24.670 --> 00:20:29.830 align:middle line:84%
value, which, for exponential
distributions, is 1/Theta.

00:20:29.830 --> 00:20:33.460 align:middle line:84%
Now, if something converges to
something, then the reciprocal

00:20:33.460 --> 00:20:37.680 align:middle line:84%
of that should converge to
the reciprocal of that.

00:20:37.680 --> 00:20:41.520 align:middle line:84%
That's a property that's
certainly correct for numbers.

00:20:41.520 --> 00:20:44.000 align:middle line:84%
But you're not talking about
convergence of numbers.

00:20:44.000 --> 00:20:46.420 align:middle line:84%
We're talking about convergence
in probability,

00:20:46.420 --> 00:20:48.820 align:middle line:84%
which is a more complicated
notion.

00:20:48.820 --> 00:20:52.370 align:middle line:84%
Fortunately, it turns out that
the same thing is true, when

00:20:52.370 --> 00:20:54.660 align:middle line:84%
we deal with convergence
in probability.

00:20:54.660 --> 00:20:58.690 align:middle line:84%
One can show, although we will
not bother doing this, that

00:20:58.690 --> 00:21:01.840 align:middle line:84%
indeed, the reciprocal of this,
which is our estimate,

00:21:01.840 --> 00:21:05.880 align:middle line:84%
converges in probability to
the reciprocal of that.

00:21:05.880 --> 00:21:08.880 align:middle line:84%
And that reciprocal is the
true parameter Theta.

00:21:08.880 --> 00:21:11.570 align:middle line:84%
So for this particular
exponential example, we do

00:21:11.570 --> 00:21:15.250 align:middle line:84%
have the desirable property,
that as the number of data

00:21:15.250 --> 00:21:18.230 align:middle line:84%
becomes larger and larger,
the estimate that we have

00:21:18.230 --> 00:21:20.970 align:middle line:84%
constructed will get closer
and closer to the true

00:21:20.970 --> 00:21:22.510 align:middle line:90%
parameter value.

00:21:22.510 --> 00:21:27.050 align:middle line:84%
And this is true no matter
what Theta is.

00:21:27.050 --> 00:21:30.130 align:middle line:84%
No matter what the true
parameter Theta is, we're

00:21:30.130 --> 00:21:33.240 align:middle line:84%
going to get close to it as
we collect more data.

00:21:33.240 --> 00:21:35.780 align:middle line:90%


00:21:35.780 --> 00:21:35.950 align:middle line:90%
Okay.

00:21:35.950 --> 00:21:39.100 align:middle line:84%
So these are two rough
qualitative properties that

00:21:39.100 --> 00:21:42.350 align:middle line:90%
would be nice to have.

00:21:42.350 --> 00:21:47.340 align:middle line:84%
If you want to get a little
more quantitative, you can

00:21:47.340 --> 00:21:50.210 align:middle line:84%
start looking at the mean
squared error that your

00:21:50.210 --> 00:21:52.000 align:middle line:90%
estimator gives.

00:21:52.000 --> 00:21:56.600 align:middle line:84%
Now, once more, the comment I
was making up there applies.

00:21:56.600 --> 00:22:00.540 align:middle line:84%
Namely, that this expectation
here is an expectation with

00:22:00.540 --> 00:22:04.600 align:middle line:84%
respect to the probability
distribution of Theta hat that

00:22:04.600 --> 00:22:07.280 align:middle line:84%
corresponds to a particular
value of little theta.

00:22:07.280 --> 00:22:09.840 align:middle line:90%
So fix a little theta.

00:22:09.840 --> 00:22:11.910 align:middle line:90%
Write down this expression.

00:22:11.910 --> 00:22:14.550 align:middle line:84%
Look at the probability
distribution of Theta hat,

00:22:14.550 --> 00:22:16.380 align:middle line:90%
under that little theta.

00:22:16.380 --> 00:22:18.220 align:middle line:90%
And do this calculation.

00:22:18.220 --> 00:22:20.610 align:middle line:84%
You're going to get some
quantity that depends on the

00:22:20.610 --> 00:22:21.860 align:middle line:90%
little theta.

00:22:21.860 --> 00:22:24.200 align:middle line:90%


00:22:24.200 --> 00:22:28.450 align:middle line:84%
And so all quantities in this
equality here should be

00:22:28.450 --> 00:22:33.360 align:middle line:84%
interpreted as quantities under
that particular value of

00:22:33.360 --> 00:22:34.490 align:middle line:90%
little theta.

00:22:34.490 --> 00:22:38.640 align:middle line:84%
So if you wanted to make this
more explicit, you could start

00:22:38.640 --> 00:22:41.870 align:middle line:84%
throwing little subscripts
everywhere in those

00:22:41.870 --> 00:22:44.430 align:middle line:90%
expressions.

00:22:44.430 --> 00:22:49.190 align:middle line:84%
And let's see what those
expressions tell us.

00:22:49.190 --> 00:22:55.430 align:middle line:84%
The expected value squared of
a random variable, we know

00:22:55.430 --> 00:22:59.210 align:middle line:84%
that it's always equal to the
variance of this random

00:22:59.210 --> 00:23:03.790 align:middle line:84%
variable, plus the expectation
of the

00:23:03.790 --> 00:23:05.860 align:middle line:90%
random variable squared.

00:23:05.860 --> 00:23:08.465 align:middle line:84%
So the expectation value of that
random variable, squared.

00:23:08.465 --> 00:23:12.020 align:middle line:90%


00:23:12.020 --> 00:23:17.030 align:middle line:84%
This equality here is just our
familiar formula, that the

00:23:17.030 --> 00:23:23.250 align:middle line:84%
expected value of X squared is
the variance of X plus the

00:23:23.250 --> 00:23:26.350 align:middle line:90%
expected value of X squared.

00:23:26.350 --> 00:23:30.040 align:middle line:84%
So we apply this formula
to X equal to

00:23:30.040 --> 00:23:34.024 align:middle line:90%
Theta hat minus Theta.

00:23:34.024 --> 00:23:37.180 align:middle line:90%


00:23:37.180 --> 00:23:41.220 align:middle line:84%
Now, remember that, in this
classical setting, theta is

00:23:41.220 --> 00:23:42.140 align:middle line:90%
just a constant.

00:23:42.140 --> 00:23:43.450 align:middle line:90%
We have fixed Theta.

00:23:43.450 --> 00:23:45.850 align:middle line:84%
We want to calculate the
variance of this quantity,

00:23:45.850 --> 00:23:47.760 align:middle line:90%
under that particular Theta.

00:23:47.760 --> 00:23:51.000 align:middle line:84%
When you add or subtract a
constant to a random variable,

00:23:51.000 --> 00:23:54.070 align:middle line:90%
the variance doesn't change.

00:23:54.070 --> 00:23:56.860 align:middle line:84%
This is the same as the variance
of our estimator.

00:23:56.860 --> 00:24:00.300 align:middle line:84%
And what we've got here is
the bias of our estimate.

00:24:00.300 --> 00:24:02.580 align:middle line:84%
It tells us, on the average,
whether we

00:24:02.580 --> 00:24:04.470 align:middle line:90%
fall above or below.

00:24:04.470 --> 00:24:06.850 align:middle line:84%
And we're taking the bias
to be b squared.

00:24:06.850 --> 00:24:10.110 align:middle line:84%
If we have an unbiased
estimator, the bias

00:24:10.110 --> 00:24:13.690 align:middle line:90%
term will be 0.

00:24:13.690 --> 00:24:18.250 align:middle line:84%
So ideally we want Theta hat
to be very close to Theta.

00:24:18.250 --> 00:24:21.840 align:middle line:84%
And since Theta is a constant,
if that happens, the variance

00:24:21.840 --> 00:24:25.650 align:middle line:84%
of Theta hat would
be very small.

00:24:25.650 --> 00:24:26.870 align:middle line:90%
So Theta is a constant.

00:24:26.870 --> 00:24:30.180 align:middle line:84%
If Theta hat has a distribution
that's

00:24:30.180 --> 00:24:33.610 align:middle line:84%
concentrated just around own
little theta, then Theta hat

00:24:33.610 --> 00:24:35.250 align:middle line:90%
would have a small variance.

00:24:35.250 --> 00:24:37.690 align:middle line:84%
So this is one desire
that have.

00:24:37.690 --> 00:24:39.740 align:middle line:84%
We're going to have
a small variance.

00:24:39.740 --> 00:24:43.710 align:middle line:84%
But we also want to have a small
bias at the same time.

00:24:43.710 --> 00:24:47.370 align:middle line:84%
So the general form of the mean
squared error has two

00:24:47.370 --> 00:24:48.240 align:middle line:90%
contributions.

00:24:48.240 --> 00:24:50.530 align:middle line:84%
One is the variance
of our estimator.

00:24:50.530 --> 00:24:52.350 align:middle line:90%
The other is the bias.

00:24:52.350 --> 00:24:54.990 align:middle line:84%
And one usually wants to design
an estimator that

00:24:54.990 --> 00:24:58.900 align:middle line:84%
simultaneously keeps both
of these terms small.

00:24:58.900 --> 00:25:03.250 align:middle line:84%
So here's an estimation method
that would do very well with

00:25:03.250 --> 00:25:05.080 align:middle line:84%
respect to this term,
but badly with

00:25:05.080 --> 00:25:06.680 align:middle line:90%
respect to that term.

00:25:06.680 --> 00:25:09.410 align:middle line:84%
So suppose that my distribution
is, let's say,

00:25:09.410 --> 00:25:13.700 align:middle line:84%
normal with an unknown mean
Theta and variance 1.

00:25:13.700 --> 00:25:17.580 align:middle line:84%
And I use as my estimator
something very dumb.

00:25:17.580 --> 00:25:23.330 align:middle line:84%
I always produce an estimate
that says my estimate is 100.

00:25:23.330 --> 00:25:26.430 align:middle line:84%
So I'm just ignoring the
data and report 100.

00:25:26.430 --> 00:25:27.750 align:middle line:90%
What does this do?

00:25:27.750 --> 00:25:30.950 align:middle line:84%
The variance of my
estimator is 0.

00:25:30.950 --> 00:25:33.690 align:middle line:84%
There's no randomness in the
estimate that I report.

00:25:33.690 --> 00:25:37.020 align:middle line:84%
But the bias is going
to be pretty bad.

00:25:37.020 --> 00:25:44.180 align:middle line:84%
The bias is going to be Theta
hat, which is 100 minus the

00:25:44.180 --> 00:25:46.770 align:middle line:90%
true value of Theta.

00:25:46.770 --> 00:25:50.340 align:middle line:84%
And for some Theta's, my bias
is going to be horrible.

00:25:50.340 --> 00:25:54.600 align:middle line:84%
If my true Theta happens
to be 0, my bias

00:25:54.600 --> 00:25:56.200 align:middle line:90%
squared is a huge term.

00:25:56.200 --> 00:25:57.810 align:middle line:90%
And I get a large error.

00:25:57.810 --> 00:26:00.220 align:middle line:84%
So what's the moral
of this example?

00:26:00.220 --> 00:26:03.700 align:middle line:84%
There are ways of making that
variance very small, but, in

00:26:03.700 --> 00:26:07.360 align:middle line:84%
those cases, you pay a
price in the bias.

00:26:07.360 --> 00:26:10.340 align:middle line:84%
So you want to do something a
little more delicate, where

00:26:10.340 --> 00:26:14.640 align:middle line:84%
you try to keep both terms
small at the same time.

00:26:14.640 --> 00:26:16.720 align:middle line:84%
So these types of considerations
become

00:26:16.720 --> 00:26:20.280 align:middle line:84%
important when you start to try
to design sophisticated

00:26:20.280 --> 00:26:22.840 align:middle line:84%
estimators for more complicated
problems.

00:26:22.840 --> 00:26:24.800 align:middle line:84%
But we will not do this
in this class.

00:26:24.800 --> 00:26:26.720 align:middle line:84%
This belongs to further
classes on

00:26:26.720 --> 00:26:28.750 align:middle line:90%
statistics and inference.

00:26:28.750 --> 00:26:31.960 align:middle line:84%
For this class, for parameter
estimation, we will basically

00:26:31.960 --> 00:26:34.400 align:middle line:84%
stick to two very
simple methods.

00:26:34.400 --> 00:26:37.930 align:middle line:84%
One is the maximum likelihood
method we've just discussed.

00:26:37.930 --> 00:26:41.300 align:middle line:84%
And the other method is what you
would do if you were still

00:26:41.300 --> 00:26:44.010 align:middle line:84%
in high school and didn't
know any probability.

00:26:44.010 --> 00:26:46.610 align:middle line:90%
You get data.

00:26:46.610 --> 00:26:50.430 align:middle line:84%
And these data come from
some distribution

00:26:50.430 --> 00:26:51.850 align:middle line:90%
with an unknown mean.

00:26:51.850 --> 00:26:53.930 align:middle line:84%
And you want to estimate
that the unknown mean.

00:26:53.930 --> 00:26:54.810 align:middle line:90%
What would you do?

00:26:54.810 --> 00:26:57.990 align:middle line:84%
You would just take those data
and average them out.

00:26:57.990 --> 00:27:00.440 align:middle line:84%
So let's make this a little
more specific.

00:27:00.440 --> 00:27:04.770 align:middle line:84%
We have X's that come from
a given distribution.

00:27:04.770 --> 00:27:07.775 align:middle line:84%
We know the general form of
the distribution, perhaps.

00:27:07.775 --> 00:27:10.570 align:middle line:90%


00:27:10.570 --> 00:27:15.180 align:middle line:84%
We do know, perhaps, the
variance of that distribution,

00:27:15.180 --> 00:27:17.050 align:middle line:90%
or, perhaps, we don't know it.

00:27:17.050 --> 00:27:19.030 align:middle line:90%
But we do not know the mean.

00:27:19.030 --> 00:27:22.700 align:middle line:84%
And we want to estimate the
mean of that distribution.

00:27:22.700 --> 00:27:25.370 align:middle line:84%
Now, we can write
this situation.

00:27:25.370 --> 00:27:27.710 align:middle line:84%
We can represent it in
a different form.

00:27:27.710 --> 00:27:30.120 align:middle line:90%
The Xi's are equal to Theta.

00:27:30.120 --> 00:27:31.380 align:middle line:90%
This is the mean.

00:27:31.380 --> 00:27:34.310 align:middle line:84%
Plus a 0 mean random
variable, that you

00:27:34.310 --> 00:27:36.000 align:middle line:90%
can think of as noise.

00:27:36.000 --> 00:27:39.380 align:middle line:84%
So this corresponds to the usual
situation you would have

00:27:39.380 --> 00:27:41.950 align:middle line:84%
in a lab, where you
go and try to

00:27:41.950 --> 00:27:43.870 align:middle line:90%
measure an unknown quantity.

00:27:43.870 --> 00:27:45.260 align:middle line:90%
You get lots of measurements.

00:27:45.260 --> 00:27:49.490 align:middle line:84%
But each time that you measure
them, your measurements have

00:27:49.490 --> 00:27:51.920 align:middle line:90%
some extra noise in there.

00:27:51.920 --> 00:27:54.510 align:middle line:84%
And you want to kind of
get rid of that noise.

00:27:54.510 --> 00:27:57.860 align:middle line:84%
The way to try to get rid of
the measurement noise is to

00:27:57.860 --> 00:28:01.170 align:middle line:84%
collect lots of data and
average them out.

00:28:01.170 --> 00:28:02.930 align:middle line:90%
This is the sample mean.

00:28:02.930 --> 00:28:07.380 align:middle line:84%
And this is a very, very
reasonable way of trying to

00:28:07.380 --> 00:28:10.130 align:middle line:84%
estimate the unknown
mean of the X's.

00:28:10.130 --> 00:28:12.700 align:middle line:90%
So this is the sample mean.

00:28:12.700 --> 00:28:17.840 align:middle line:84%
It's a reasonable, plausible,
in general, pretty good

00:28:17.840 --> 00:28:22.390 align:middle line:84%
estimator of the unknown mean
of a certain distribution.

00:28:22.390 --> 00:28:26.910 align:middle line:84%
We can apply this estimator
without really knowing a lot

00:28:26.910 --> 00:28:28.810 align:middle line:84%
about the distribution
of the X's.

00:28:28.810 --> 00:28:31.010 align:middle line:84%
Actually, we don't need to
know anything about the

00:28:31.010 --> 00:28:32.320 align:middle line:90%
distribution.

00:28:32.320 --> 00:28:35.840 align:middle line:84%
We can still apply it, because
the variance, for example,

00:28:35.840 --> 00:28:37.130 align:middle line:90%
does not show up here.

00:28:37.130 --> 00:28:38.660 align:middle line:84%
We don't need to know
the variance to

00:28:38.660 --> 00:28:40.520 align:middle line:90%
calculate that quantity.

00:28:40.520 --> 00:28:43.520 align:middle line:84%
Does this estimator have
good properties?

00:28:43.520 --> 00:28:45.110 align:middle line:90%
Yes, it does.

00:28:45.110 --> 00:28:48.110 align:middle line:84%
What's the expected value
of the sample mean?

00:28:48.110 --> 00:28:51.910 align:middle line:84%
If the expectation of this, it's
the expectation of this

00:28:51.910 --> 00:28:53.600 align:middle line:90%
sum divided by n.

00:28:53.600 --> 00:28:56.410 align:middle line:84%
The expected value for each
one of the X's is Theta.

00:28:56.410 --> 00:28:58.290 align:middle line:84%
So the expected value
of the sample mean

00:28:58.290 --> 00:29:00.010 align:middle line:90%
is just Theta itself.

00:29:00.010 --> 00:29:03.310 align:middle line:90%
So our estimator is unbiased.

00:29:03.310 --> 00:29:06.410 align:middle line:84%
No matter what Theta is, our
estimator does not have a

00:29:06.410 --> 00:29:11.130 align:middle line:84%
systematic error in
either direction.

00:29:11.130 --> 00:29:13.870 align:middle line:84%
Furthermore, the weak law of
large numbers tells us that

00:29:13.870 --> 00:29:18.140 align:middle line:84%
this quantity converges to the
true parameter in probability.

00:29:18.140 --> 00:29:20.700 align:middle line:84%
So it's a consistent
estimator.

00:29:20.700 --> 00:29:21.920 align:middle line:90%
This is good.

00:29:21.920 --> 00:29:26.740 align:middle line:84%
And if you want to calculate
the mean squared error

00:29:26.740 --> 00:29:28.780 align:middle line:84%
corresponding to
this estimator.

00:29:28.780 --> 00:29:31.550 align:middle line:84%
Remember how we defined the
mean squared error?

00:29:31.550 --> 00:29:35.300 align:middle line:90%
It's this quantity.

00:29:35.300 --> 00:29:38.680 align:middle line:84%
Then it's a calculation that we
have done a fair number of

00:29:38.680 --> 00:29:40.080 align:middle line:90%
times by now.

00:29:40.080 --> 00:29:43.640 align:middle line:84%
The mean squared error is the
variance of the distribution

00:29:43.640 --> 00:29:46.000 align:middle line:90%
of the X's divided by n.

00:29:46.000 --> 00:29:49.370 align:middle line:84%
So as we get more and more data,
the mean squared error

00:29:49.370 --> 00:29:52.170 align:middle line:90%
goes down to 0.

00:29:52.170 --> 00:29:56.420 align:middle line:84%
In some examples, it turns out
that the sample mean is also

00:29:56.420 --> 00:29:58.930 align:middle line:84%
the same as the maximum
likelihood estimate.

00:29:58.930 --> 00:30:02.790 align:middle line:84%
For example, if the X's are
coming from a normal

00:30:02.790 --> 00:30:07.700 align:middle line:84%
distribution, you can write down
the likelihood, do the

00:30:07.700 --> 00:30:10.240 align:middle line:84%
maximization with respect to
Theta, you'll find that the

00:30:10.240 --> 00:30:15.190 align:middle line:84%
maximum likelihood estimate is
the same as the sample mean.

00:30:15.190 --> 00:30:18.730 align:middle line:84%
In other cases, the sample mean
will be different from

00:30:18.730 --> 00:30:20.850 align:middle line:90%
the maximum likelihood.

00:30:20.850 --> 00:30:23.990 align:middle line:84%
And then you have a choice
about which one of the

00:30:23.990 --> 00:30:24.860 align:middle line:90%
two you would use.

00:30:24.860 --> 00:30:27.890 align:middle line:84%
Probably, in most reasonable
situations, you would just use

00:30:27.890 --> 00:30:31.460 align:middle line:84%
the sample mean, because it's
simple, easy to compute, and

00:30:31.460 --> 00:30:33.830 align:middle line:90%
has nice properties.

00:30:33.830 --> 00:30:33.936 align:middle line:90%
All right.

00:30:33.936 --> 00:30:35.110 align:middle line:90%
So you go to your boss.

00:30:35.110 --> 00:30:38.120 align:middle line:84%
And you report and say,
OK, I did all my

00:30:38.120 --> 00:30:39.910 align:middle line:90%
experiments in the lab.

00:30:39.910 --> 00:30:49.820 align:middle line:84%
And the average value that I got
is a certain number, 2.37.

00:30:49.820 --> 00:30:52.490 align:middle line:84%
So is that the informative
to your boss?

00:30:52.490 --> 00:30:55.470 align:middle line:84%
Well your boss would like to
know how much they can trust

00:30:55.470 --> 00:30:58.280 align:middle line:90%
this number, 2.37.

00:30:58.280 --> 00:31:00.630 align:middle line:84%
Well, I know that the true
value is not going to be

00:31:00.630 --> 00:31:02.270 align:middle line:90%
exactly that.

00:31:02.270 --> 00:31:07.410 align:middle line:90%
But how close should it be?

00:31:07.410 --> 00:31:09.820 align:middle line:84%
So give me a range of
what you think are

00:31:09.820 --> 00:31:12.080 align:middle line:90%
possible values of Theta.

00:31:12.080 --> 00:31:16.220 align:middle line:90%
So the situation is like this.

00:31:16.220 --> 00:31:20.370 align:middle line:84%
So suppose that we observe X's
that are coming from a certain

00:31:20.370 --> 00:31:22.070 align:middle line:90%
distribution.

00:31:22.070 --> 00:31:24.230 align:middle line:84%
And we're trying to
estimate the mean.

00:31:24.230 --> 00:31:25.480 align:middle line:90%
We get our data.

00:31:25.480 --> 00:31:27.880 align:middle line:90%


00:31:27.880 --> 00:31:32.090 align:middle line:84%
Maybe our data looks something
like this.

00:31:32.090 --> 00:31:34.090 align:middle line:90%
You calculate the mean.

00:31:34.090 --> 00:31:36.140 align:middle line:90%
You find the sample mean.

00:31:36.140 --> 00:31:40.120 align:middle line:84%
So let's suppose that the sample
mean is a number, for

00:31:40.120 --> 00:31:45.570 align:middle line:90%
some reason take to be 2.37.

00:31:45.570 --> 00:31:48.300 align:middle line:84%
But you want to convey something
to your boss about

00:31:48.300 --> 00:31:51.450 align:middle line:84%
how spread out these
data were.

00:31:51.450 --> 00:31:56.690 align:middle line:84%
So the boss asks you to give
him or her some kind of

00:31:56.690 --> 00:32:05.340 align:middle line:84%
interval on which Theta, the
true parameter, might lie.

00:32:05.340 --> 00:32:07.540 align:middle line:84%
So the boss asked you
for an interval.

00:32:07.540 --> 00:32:11.740 align:middle line:84%
So what you do is you end up
reporting an interval.

00:32:11.740 --> 00:32:14.990 align:middle line:84%
And you somehow use the data
that you have seen to

00:32:14.990 --> 00:32:17.580 align:middle line:90%
construct this interval.

00:32:17.580 --> 00:32:19.900 align:middle line:84%
And you report to your
boss also the

00:32:19.900 --> 00:32:21.420 align:middle line:90%
endpoints of this interval.

00:32:21.420 --> 00:32:24.020 align:middle line:84%
Let's give names to
these endpoints,

00:32:24.020 --> 00:32:27.710 align:middle line:90%
Theta_n- and Theta_n+.

00:32:27.710 --> 00:32:31.000 align:middle line:84%
The ends here just play the role
of keeping track of how

00:32:31.000 --> 00:32:33.000 align:middle line:90%
many data we're using.

00:32:33.000 --> 00:32:39.320 align:middle line:84%
So what you report to your boss
is this interval as well.

00:32:39.320 --> 00:32:42.340 align:middle line:84%
Are these Theta's here, the
endpoints of the interval,

00:32:42.340 --> 00:32:44.220 align:middle line:90%
lowercase or uppercase?

00:32:44.220 --> 00:32:45.750 align:middle line:90%
What should they be?

00:32:45.750 --> 00:32:48.180 align:middle line:84%
Well you construct these
intervals after

00:32:48.180 --> 00:32:49.430 align:middle line:90%
you see your data.

00:32:49.430 --> 00:32:53.830 align:middle line:84%
You take the data into account
to construct your interval.

00:32:53.830 --> 00:32:57.020 align:middle line:84%
So these definitely should
depend on the data.

00:32:57.020 --> 00:32:59.460 align:middle line:84%
And therefore they are
random variables.

00:32:59.460 --> 00:33:03.240 align:middle line:84%
Same thing with your estimator,
in general, it's

00:33:03.240 --> 00:33:05.120 align:middle line:90%
going to be a random variable.

00:33:05.120 --> 00:33:07.930 align:middle line:84%
Although, when you go and report
numbers to your boss,

00:33:07.930 --> 00:33:10.580 align:middle line:84%
you give the specific
realizations of the random

00:33:10.580 --> 00:33:15.450 align:middle line:84%
variables, given the
data that you got.

00:33:15.450 --> 00:33:21.500 align:middle line:84%
So instead of having
just a single box

00:33:21.500 --> 00:33:25.050 align:middle line:90%
that produces estimates.

00:33:25.050 --> 00:33:29.540 align:middle line:84%
So our previous picture was that
you have your estimator

00:33:29.540 --> 00:33:34.130 align:middle line:84%
that takes X's and produces
Theta hats.

00:33:34.130 --> 00:33:40.960 align:middle line:84%
Now our box will also be
producing Theta hats minus and

00:33:40.960 --> 00:33:42.570 align:middle line:90%
Theta hats plus.

00:33:42.570 --> 00:33:45.180 align:middle line:84%
It's going to produce
an interval as well.

00:33:45.180 --> 00:33:48.670 align:middle line:84%
The X's are random, therefore
these quantities are random.

00:33:48.670 --> 00:33:52.340 align:middle line:84%
Once you go and do the
experiment and obtain your

00:33:52.340 --> 00:33:55.930 align:middle line:84%
data, then your data
will be some

00:33:55.930 --> 00:33:58.810 align:middle line:90%
lowercase x, specific numbers.

00:33:58.810 --> 00:34:00.950 align:middle line:84%
And then your estimates
and estimator

00:34:00.950 --> 00:34:05.110 align:middle line:90%
become also lower case.

00:34:05.110 --> 00:34:08.010 align:middle line:84%
What would we like this
interval to do?

00:34:08.010 --> 00:34:11.760 align:middle line:84%
We would like it to be highly
likely to contain the true

00:34:11.760 --> 00:34:13.810 align:middle line:90%
value of the parameter.

00:34:13.810 --> 00:34:17.800 align:middle line:84%
So we might impose some specs
of the following kind.

00:34:17.800 --> 00:34:19.170 align:middle line:90%
I pick a number, alpha.

00:34:19.170 --> 00:34:21.170 align:middle line:84%
Usually that alpha,
think of it as a

00:34:21.170 --> 00:34:23.050 align:middle line:90%
probability of a large error.

00:34:23.050 --> 00:34:27.449 align:middle line:84%
Typical value of alpha might
be 0.05, in which case this

00:34:27.449 --> 00:34:30.360 align:middle line:90%
number here is point 0.95.

00:34:30.360 --> 00:34:33.989 align:middle line:84%
And you're given specs that
say something like this.

00:34:33.989 --> 00:34:41.110 align:middle line:84%
I would like, with probability
at least 0.95, this to happen,

00:34:41.110 --> 00:34:44.739 align:middle line:84%
which says that the true
parameter lies inside the

00:34:44.739 --> 00:34:47.100 align:middle line:90%
confidence interval.

00:34:47.100 --> 00:34:50.840 align:middle line:84%
Now let's try to interpret
this statement.

00:34:50.840 --> 00:34:53.560 align:middle line:84%
Suppose that you did the
experiment, and that you ended

00:34:53.560 --> 00:34:56.230 align:middle line:84%
up reporting to your boss
a confidence interval

00:34:56.230 --> 00:35:01.520 align:middle line:90%
from 1.97 to 2.56.

00:35:01.520 --> 00:35:03.170 align:middle line:84%
That's what you report
to your boss.

00:35:03.170 --> 00:35:06.790 align:middle line:90%


00:35:06.790 --> 00:35:08.300 align:middle line:90%
And suppose that the confidence

00:35:08.300 --> 00:35:10.280 align:middle line:90%
interval has this property.

00:35:10.280 --> 00:35:16.400 align:middle line:84%
Can you go to your boss and say,
with probability 95%, the

00:35:16.400 --> 00:35:20.090 align:middle line:84%
true value of Theta is between
these two numbers?

00:35:20.090 --> 00:35:22.630 align:middle line:84%
Is that a meaningful
statement?

00:35:22.630 --> 00:35:26.100 align:middle line:84%
So the statement is, the
tentative statement is, with

00:35:26.100 --> 00:35:30.200 align:middle line:84%
probability 95%, the true
value of Theta is

00:35:30.200 --> 00:35:34.930 align:middle line:90%
between 1.97 and 2.56.

00:35:34.930 --> 00:35:38.910 align:middle line:84%
Well, what is random
in that statement?

00:35:38.910 --> 00:35:40.460 align:middle line:90%
There's nothing random.

00:35:40.460 --> 00:35:43.070 align:middle line:84%
The true value of theta
is a constant.

00:35:43.070 --> 00:35:44.720 align:middle line:90%
1.97 is a number.

00:35:44.720 --> 00:35:46.740 align:middle line:90%
2.56 is a number.

00:35:46.740 --> 00:35:52.960 align:middle line:84%
So it doesn't make any sense to
talk about the probability

00:35:52.960 --> 00:35:54.920 align:middle line:84%
that theta is in
this interval.

00:35:54.920 --> 00:35:57.540 align:middle line:84%
Either theta happens to be
in that interval, or it

00:35:57.540 --> 00:35:58.760 align:middle line:90%
happens to not be.

00:35:58.760 --> 00:36:01.560 align:middle line:84%
But there are no probabilities
associated with this.

00:36:01.560 --> 00:36:04.700 align:middle line:90%
Because theta is not random.

00:36:04.700 --> 00:36:06.690 align:middle line:84%
Syntactically, you
can see this.

00:36:06.690 --> 00:36:09.210 align:middle line:84%
Because theta here
is a lower case.

00:36:09.210 --> 00:36:11.930 align:middle line:84%
So what kind of probabilities
are we talking about here?

00:36:11.930 --> 00:36:13.460 align:middle line:90%
Where's the randomness?

00:36:13.460 --> 00:36:15.880 align:middle line:84%
Well the random thing
is the interval.

00:36:15.880 --> 00:36:17.560 align:middle line:90%
It's not theta.

00:36:17.560 --> 00:36:21.090 align:middle line:84%
So the statement that is being
made here is that the

00:36:21.090 --> 00:36:24.290 align:middle line:84%
interval, that's being
constructed by our procedure,

00:36:24.290 --> 00:36:28.410 align:middle line:84%
should have the property that,
with probability 95%, it's

00:36:28.410 --> 00:36:33.280 align:middle line:84%
going to fall on top of the
true value of theta.

00:36:33.280 --> 00:36:37.680 align:middle line:84%
So the right way of interpreting
what the 95%

00:36:37.680 --> 00:36:42.270 align:middle line:84%
confidence interval is, is
something like the following.

00:36:42.270 --> 00:36:45.390 align:middle line:84%
We have the true value of theta
that we don't know.

00:36:45.390 --> 00:36:46.750 align:middle line:90%
I get data.

00:36:46.750 --> 00:36:50.150 align:middle line:84%
Based on the data, I construct
a confidence interval.

00:36:50.150 --> 00:36:51.950 align:middle line:90%
I get my confidence interval.

00:36:51.950 --> 00:36:52.790 align:middle line:90%
I got lucky.

00:36:52.790 --> 00:36:54.850 align:middle line:84%
And the true value of
theta is in here.

00:36:54.850 --> 00:36:57.790 align:middle line:84%
Next day, I do the same
experiment, take my data,

00:36:57.790 --> 00:37:00.500 align:middle line:84%
construct a confidence
interval.

00:37:00.500 --> 00:37:04.040 align:middle line:84%
And I get this confidence
interval, lucky once more.

00:37:04.040 --> 00:37:06.320 align:middle line:90%
Next day I get data.

00:37:06.320 --> 00:37:09.620 align:middle line:84%
I use my data to come up with
an estimate of theta and the

00:37:09.620 --> 00:37:10.660 align:middle line:90%
confidence interval.

00:37:10.660 --> 00:37:12.340 align:middle line:90%
That day, I was unlucky.

00:37:12.340 --> 00:37:15.000 align:middle line:84%
And I got a confidence
interval out there.

00:37:15.000 --> 00:37:20.890 align:middle line:84%
What the requirement here is, is
that 95% of the days, where

00:37:20.890 --> 00:37:25.270 align:middle line:84%
we use this certain procedure
for constructing confidence

00:37:25.270 --> 00:37:29.180 align:middle line:84%
intervals, 95% of those days,
we will be lucky.

00:37:29.180 --> 00:37:33.750 align:middle line:84%
And we will capture the correct
value of theta by your

00:37:33.750 --> 00:37:35.160 align:middle line:90%
confidence interval.

00:37:35.160 --> 00:37:39.390 align:middle line:84%
So it's a statement about the
distribution of these random

00:37:39.390 --> 00:37:42.820 align:middle line:84%
confidence intervals, how likely
are they to fall on top

00:37:42.820 --> 00:37:45.210 align:middle line:84%
of the true theta, as opposed
to how likely

00:37:45.210 --> 00:37:47.060 align:middle line:90%
they are to fall outside.

00:37:47.060 --> 00:37:50.770 align:middle line:84%
So it's a statement about
probabilities associated with

00:37:50.770 --> 00:37:52.380 align:middle line:90%
a confidence interval.

00:37:52.380 --> 00:37:55.080 align:middle line:84%
They're not probabilities about
theta, because theta,

00:37:55.080 --> 00:37:58.370 align:middle line:90%
itself, is not random.

00:37:58.370 --> 00:38:02.080 align:middle line:84%
So this is what the confidence
interval is, in general, and

00:38:02.080 --> 00:38:03.470 align:middle line:90%
how we interpret it.

00:38:03.470 --> 00:38:07.470 align:middle line:84%
How do we construct a 95%
confidence interval?

00:38:07.470 --> 00:38:09.320 align:middle line:84%
Let's go through this
exercise, in

00:38:09.320 --> 00:38:10.980 align:middle line:90%
a particular example.

00:38:10.980 --> 00:38:13.970 align:middle line:84%
The calculations are exactly the
same as the ones that you

00:38:13.970 --> 00:38:17.770 align:middle line:84%
did when we talked about laws
of large numbers and the

00:38:17.770 --> 00:38:19.240 align:middle line:90%
central limit theorem.

00:38:19.240 --> 00:38:22.600 align:middle line:84%
So there's nothing new
calculationally but it's,

00:38:22.600 --> 00:38:25.440 align:middle line:84%
perhaps, new in terms of the
language that we use and the

00:38:25.440 --> 00:38:26.800 align:middle line:90%
interpretation.

00:38:26.800 --> 00:38:30.890 align:middle line:84%
So we got our sample mean
from some distribution.

00:38:30.890 --> 00:38:34.650 align:middle line:84%
And we would like to calculate
a 95% confidence interval.

00:38:34.650 --> 00:38:39.590 align:middle line:90%


00:38:39.590 --> 00:38:42.650 align:middle line:84%
We know from the normal tables,
that the standard

00:38:42.650 --> 00:38:54.011 align:middle line:84%
normal has 2.5% on the tail,
that's after 1.96.

00:38:54.011 --> 00:38:58.060 align:middle line:84%
Yes, by this time,
the number 1.96

00:38:58.060 --> 00:39:00.600 align:middle line:90%
should be pretty familiar.

00:39:00.600 --> 00:39:05.880 align:middle line:84%
So if this probability
here is 2.5%, this

00:39:05.880 --> 00:39:09.510 align:middle line:90%
number here is 1.96.

00:39:09.510 --> 00:39:12.310 align:middle line:84%
Now look at this random
variable here.

00:39:12.310 --> 00:39:15.000 align:middle line:90%
This is the sample mean.

00:39:15.000 --> 00:39:17.950 align:middle line:84%
Difference, from the true mean,
normalized by the usual

00:39:17.950 --> 00:39:18.940 align:middle line:90%
normalizing factor.

00:39:18.940 --> 00:39:22.090 align:middle line:84%
By the central limit theorem,
this is approximately normal.

00:39:22.090 --> 00:39:26.790 align:middle line:84%
So it has probability 0.95
of being less than 1.96.

00:39:26.790 --> 00:39:31.050 align:middle line:84%
Now take this event here
and rewrite it.

00:39:31.050 --> 00:39:36.240 align:middle line:84%
This the event, well, that
Theta hat minus theta is

00:39:36.240 --> 00:39:40.350 align:middle line:84%
bigger than this number and
smaller than that number.

00:39:40.350 --> 00:39:45.650 align:middle line:84%
This event here is equivalent
to that event here.

00:39:45.650 --> 00:39:50.670 align:middle line:84%
And so this suggests a way of
constructing our 95% percent

00:39:50.670 --> 00:39:52.130 align:middle line:90%
confidence interval.

00:39:52.130 --> 00:39:56.330 align:middle line:84%
I'm going to report the
interval, which gives this as

00:39:56.330 --> 00:40:00.350 align:middle line:84%
the lower end of the confidence
interval, and gives

00:40:00.350 --> 00:40:05.720 align:middle line:84%
this as the upper end of
the confidence interval

00:40:05.720 --> 00:40:09.180 align:middle line:84%
In other words, at the end of
the experiment, we report the

00:40:09.180 --> 00:40:12.170 align:middle line:84%
sample mean, which
is our estimate.

00:40:12.170 --> 00:40:14.230 align:middle line:90%
And we report also, an interval

00:40:14.230 --> 00:40:16.080 align:middle line:90%
around the sample mean.

00:40:16.080 --> 00:40:20.510 align:middle line:84%
And this is our 95% confidence
interval.

00:40:20.510 --> 00:40:22.800 align:middle line:90%
The confidence interval becomes

00:40:22.800 --> 00:40:26.050 align:middle line:90%
smaller, when n is larger.

00:40:26.050 --> 00:40:28.950 align:middle line:84%
In some sense, we're more
certain that we're doing a

00:40:28.950 --> 00:40:32.390 align:middle line:84%
good estimation job, so we can
have a small interval and

00:40:32.390 --> 00:40:36.000 align:middle line:84%
still be quite confident that
our interval captures the true

00:40:36.000 --> 00:40:37.520 align:middle line:90%
value of the parameter.

00:40:37.520 --> 00:40:41.890 align:middle line:84%
Also, if our data have very
little noise, when you have

00:40:41.890 --> 00:40:45.060 align:middle line:84%
more accurate measurements,
you're more confident that

00:40:45.060 --> 00:40:47.220 align:middle line:90%
your estimate is pretty good.

00:40:47.220 --> 00:40:51.120 align:middle line:84%
And that results in a smaller
confidence interval, smaller

00:40:51.120 --> 00:40:52.610 align:middle line:84%
length of the confidence
interval.

00:40:52.610 --> 00:40:56.040 align:middle line:84%
And still you have 95%
probability of capturing the

00:40:56.040 --> 00:40:57.650 align:middle line:90%
true value of theta.

00:40:57.650 --> 00:41:01.660 align:middle line:84%
So we did this exercise by
taking 95% confidence

00:41:01.660 --> 00:41:04.010 align:middle line:84%
intervals and the corresponding
value from the

00:41:04.010 --> 00:41:06.670 align:middle line:90%
normal tables, which is 1.96.

00:41:06.670 --> 00:41:11.390 align:middle line:84%
Of course, you can do it more
generally, if you set your

00:41:11.390 --> 00:41:13.730 align:middle line:90%
alpha to be some other number.

00:41:13.730 --> 00:41:16.590 align:middle line:84%
Again, you look at the
normal tables.

00:41:16.590 --> 00:41:20.460 align:middle line:84%
And you find the value here,
so that the tail has

00:41:20.460 --> 00:41:22.640 align:middle line:90%
probability alpha over 2.

00:41:22.640 --> 00:41:26.790 align:middle line:84%
And instead of using these 1.96,
you use whatever number

00:41:26.790 --> 00:41:31.380 align:middle line:84%
you get from the
normal tables.

00:41:31.380 --> 00:41:33.520 align:middle line:84%
And this tells you
how to construct

00:41:33.520 --> 00:41:36.680 align:middle line:90%
a confidence interval.

00:41:36.680 --> 00:41:42.060 align:middle line:84%
Well, to be exact, this
is not necessarily a

00:41:42.060 --> 00:41:44.640 align:middle line:90%
95% confidence interval.

00:41:44.640 --> 00:41:47.540 align:middle line:84%
It's approximately a 95%
confidence interval.

00:41:47.540 --> 00:41:48.950 align:middle line:90%
Why is this?

00:41:48.950 --> 00:41:51.060 align:middle line:84%
Because we've done
an approximation.

00:41:51.060 --> 00:41:53.890 align:middle line:84%
We have used the central
limit theorem.

00:41:53.890 --> 00:41:59.990 align:middle line:84%
So it might turn out to be a
95.5% confidence interval

00:41:59.990 --> 00:42:03.220 align:middle line:84%
instead of 95%, because
our calculations are

00:42:03.220 --> 00:42:04.740 align:middle line:90%
not entirely accurate.

00:42:04.740 --> 00:42:08.230 align:middle line:84%
But for reasonable values of
n, using the central limit

00:42:08.230 --> 00:42:10.190 align:middle line:84%
theorem is a good
approximation.

00:42:10.190 --> 00:42:13.330 align:middle line:84%
And that's what people
almost always do.

00:42:13.330 --> 00:42:17.350 align:middle line:84%
So just take the value from
the normal tables.

00:42:17.350 --> 00:42:18.600 align:middle line:90%
Okay, except for one catch.

00:42:18.600 --> 00:42:22.830 align:middle line:90%


00:42:22.830 --> 00:42:24.590 align:middle line:90%
I used the data.

00:42:24.590 --> 00:42:26.440 align:middle line:90%
I obtained my estimate.

00:42:26.440 --> 00:42:29.830 align:middle line:84%
And I want to go to my boss and
report this theta minus

00:42:29.830 --> 00:42:33.010 align:middle line:84%
and theta hat, which is the
confidence interval.

00:42:33.010 --> 00:42:35.720 align:middle line:90%
What's the difficulty?

00:42:35.720 --> 00:42:37.540 align:middle line:90%
I know what n is.

00:42:37.540 --> 00:42:40.790 align:middle line:84%
But I don't know what sigma
is, in general.

00:42:40.790 --> 00:42:44.750 align:middle line:84%
So if I don't know sigma,
what am I going to do?

00:42:44.750 --> 00:42:48.980 align:middle line:84%
Here, there's a few options
for what you can do.

00:42:48.980 --> 00:42:52.910 align:middle line:84%
And the first option is familiar
from what we did when

00:42:52.910 --> 00:42:55.020 align:middle line:84%
we talked about the
pollster problem.

00:42:55.020 --> 00:42:58.480 align:middle line:84%
We don't know what sigma is,
but maybe we have an upper

00:42:58.480 --> 00:43:00.030 align:middle line:90%
bound on sigma.

00:43:00.030 --> 00:43:03.540 align:middle line:84%
For example, if the Xi's
Bernoulli random variables, we

00:43:03.540 --> 00:43:06.910 align:middle line:84%
have seen that the standard
deviation is at most 1/2.

00:43:06.910 --> 00:43:10.220 align:middle line:84%
So use the most conservative
value for sigma.

00:43:10.220 --> 00:43:13.520 align:middle line:84%
Using the most conservative
value means that you take

00:43:13.520 --> 00:43:17.890 align:middle line:84%
bigger confidence intervals
than necessary.

00:43:17.890 --> 00:43:20.780 align:middle line:90%
So that's one option.

00:43:20.780 --> 00:43:25.480 align:middle line:84%
Another option is to try to
estimate sigma from the data.

00:43:25.480 --> 00:43:27.630 align:middle line:90%
How do you do this estimation?

00:43:27.630 --> 00:43:31.140 align:middle line:84%
In special cases, for special
types of distributions, you

00:43:31.140 --> 00:43:34.180 align:middle line:84%
can think of heuristic ways
of doing this estimation.

00:43:34.180 --> 00:43:38.390 align:middle line:84%
For example, in the case of
Bernoulli random variables, we

00:43:38.390 --> 00:43:42.420 align:middle line:84%
know that the true value of
sigma, the standard deviation

00:43:42.420 --> 00:43:45.120 align:middle line:84%
of a Bernoulli random variable,
is the square root

00:43:45.120 --> 00:43:47.670 align:middle line:84%
of theta1 minus theta,
where theta is

00:43:47.670 --> 00:43:50.290 align:middle line:90%
the mean of the Bernoulli.

00:43:50.290 --> 00:43:51.900 align:middle line:90%
Try to use this formula.

00:43:51.900 --> 00:43:54.140 align:middle line:84%
But theta is the thing we're
trying to estimate in the

00:43:54.140 --> 00:43:54.760 align:middle line:90%
first place.

00:43:54.760 --> 00:43:55.880 align:middle line:90%
We don't know it.

00:43:55.880 --> 00:43:57.150 align:middle line:90%
What do we do?

00:43:57.150 --> 00:44:00.850 align:middle line:84%
Well, we have an estimate for
theta, the estimate, produced

00:44:00.850 --> 00:44:04.195 align:middle line:84%
by our estimation procedure,
the sample mean.

00:44:04.195 --> 00:44:05.670 align:middle line:90%
So I obtain my data.

00:44:05.670 --> 00:44:06.540 align:middle line:90%
I get my data.

00:44:06.540 --> 00:44:09.030 align:middle line:84%
I produce the estimate
theta hat.

00:44:09.030 --> 00:44:10.740 align:middle line:90%
It's an estimate of the mean.

00:44:10.740 --> 00:44:14.770 align:middle line:84%
Use that estimate in this
formula to come up with an

00:44:14.770 --> 00:44:17.290 align:middle line:84%
estimate of my standard
deviation.

00:44:17.290 --> 00:44:20.210 align:middle line:84%
And then use that standard
deviation, in the construction

00:44:20.210 --> 00:44:22.510 align:middle line:84%
of the confidence interval,
pretending

00:44:22.510 --> 00:44:24.180 align:middle line:90%
that this is correct.

00:44:24.180 --> 00:44:29.050 align:middle line:84%
Well the number of your data is
large, then we know, from

00:44:29.050 --> 00:44:31.870 align:middle line:84%
the law of large numbers, that
theta hat is a pretty good

00:44:31.870 --> 00:44:33.130 align:middle line:90%
estimate of theta.

00:44:33.130 --> 00:44:36.670 align:middle line:84%
So sigma hat is going to be a
pretty good estimate of sigma.

00:44:36.670 --> 00:44:42.380 align:middle line:84%
So we're not making large errors
by using this approach.

00:44:42.380 --> 00:44:47.980 align:middle line:84%
So in this scenario here, things
were simple, because we

00:44:47.980 --> 00:44:49.890 align:middle line:90%
had an analytical formula.

00:44:49.890 --> 00:44:52.210 align:middle line:90%
Sigma was determined by theta.

00:44:52.210 --> 00:44:54.420 align:middle line:84%
So we could come up
with a quick and

00:44:54.420 --> 00:44:57.340 align:middle line:90%
dirty estimate of sigma.

00:44:57.340 --> 00:45:00.940 align:middle line:84%
In general, if you do not have
any nice formulas of this

00:45:00.940 --> 00:45:03.000 align:middle line:90%
kind, what could you do?

00:45:03.000 --> 00:45:04.920 align:middle line:84%
Well, you still need
to come up with an

00:45:04.920 --> 00:45:07.110 align:middle line:90%
estimate of sigma somehow.

00:45:07.110 --> 00:45:08.950 align:middle line:90%
What is a generic method for

00:45:08.950 --> 00:45:11.300 align:middle line:90%
estimating a standard deviation?

00:45:11.300 --> 00:45:14.440 align:middle line:84%
Equivalently, what could be a
generic method for estimating

00:45:14.440 --> 00:45:16.920 align:middle line:90%
a variance?

00:45:16.920 --> 00:45:19.360 align:middle line:84%
Well the variance is
an expected value

00:45:19.360 --> 00:45:20.940 align:middle line:90%
of some random variable.

00:45:20.940 --> 00:45:25.610 align:middle line:84%
The variance is the mean of the
random variable inside of

00:45:25.610 --> 00:45:28.200 align:middle line:90%
those brackets.

00:45:28.200 --> 00:45:33.160 align:middle line:84%
How does one estimate the mean
of some random variable?

00:45:33.160 --> 00:45:36.140 align:middle line:84%
You obtain lots of measurements
of that random

00:45:36.140 --> 00:45:40.210 align:middle line:90%
variable and average them out.

00:45:40.210 --> 00:45:45.170 align:middle line:84%
So this would be a reasonable
way of estimating the variance

00:45:45.170 --> 00:45:47.310 align:middle line:90%
of a distribution.

00:45:47.310 --> 00:45:50.590 align:middle line:84%
And again, the weak law of large
numbers tells us that

00:45:50.590 --> 00:45:55.370 align:middle line:84%
this average converges to the
expected value of this, which

00:45:55.370 --> 00:45:58.590 align:middle line:84%
is just the variance of
the distribution.

00:45:58.590 --> 00:46:01.700 align:middle line:84%
So we got a nice and
consistent way

00:46:01.700 --> 00:46:03.940 align:middle line:90%
of estimating variances.

00:46:03.940 --> 00:46:08.100 align:middle line:84%
But now, we seem to be getting
in a vicious circle here,

00:46:08.100 --> 00:46:10.580 align:middle line:84%
because to estimate
the variance, we

00:46:10.580 --> 00:46:12.910 align:middle line:90%
need to know the mean.

00:46:12.910 --> 00:46:16.075 align:middle line:84%
And the mean is something we're
trying to estimate in

00:46:16.075 --> 00:46:18.250 align:middle line:90%
the first place.

00:46:18.250 --> 00:46:18.400 align:middle line:90%
Okay.

00:46:18.400 --> 00:46:20.880 align:middle line:84%
But we do have an estimate
from the mean.

00:46:20.880 --> 00:46:24.640 align:middle line:84%
So a reasonable approximation,
once more, is to plug-in,

00:46:24.640 --> 00:46:27.620 align:middle line:84%
here, since we don't
know the mean, the

00:46:27.620 --> 00:46:29.270 align:middle line:90%
estimate of the mean.

00:46:29.270 --> 00:46:32.370 align:middle line:84%
And so you get that expression,
but with a theta

00:46:32.370 --> 00:46:35.130 align:middle line:90%
hat instead of theta itself.

00:46:35.130 --> 00:46:37.980 align:middle line:84%
And this is another
reasonable way of

00:46:37.980 --> 00:46:40.180 align:middle line:90%
estimating the variance.

00:46:40.180 --> 00:46:42.940 align:middle line:84%
It does have the same
consistency properties.

00:46:42.940 --> 00:46:44.050 align:middle line:90%
Why?

00:46:44.050 --> 00:46:51.100 align:middle line:84%
When n is large, this is going
to behave the same as that,

00:46:51.100 --> 00:46:53.640 align:middle line:84%
because theta hat converges
to theta.

00:46:53.640 --> 00:46:57.890 align:middle line:84%
And when n is large, this is
approximately the same as

00:46:57.890 --> 00:46:58.820 align:middle line:90%
sigma squared.

00:46:58.820 --> 00:47:02.220 align:middle line:84%
So for a large n, this quantity
also converges to

00:47:02.220 --> 00:47:03.350 align:middle line:90%
sigma squared.

00:47:03.350 --> 00:47:05.500 align:middle line:84%
And we have a consistent
estimate of

00:47:05.500 --> 00:47:07.000 align:middle line:90%
the variance as well.

00:47:07.000 --> 00:47:09.490 align:middle line:84%
And we can take that consistent
estimate and use it

00:47:09.490 --> 00:47:12.360 align:middle line:84%
back in the construction
of confidence interval.

00:47:12.360 --> 00:47:16.310 align:middle line:84%
One little detail, here,
we're dividing by n.

00:47:16.310 --> 00:47:19.590 align:middle line:90%
Here, we're dividing by n-1.

00:47:19.590 --> 00:47:21.050 align:middle line:90%
Why do we do this?

00:47:21.050 --> 00:47:24.630 align:middle line:84%
Well, it turns out that's what
you need to do for these

00:47:24.630 --> 00:47:28.590 align:middle line:84%
estimates to be an unbiased
estimate of the variance.

00:47:28.590 --> 00:47:32.080 align:middle line:84%
One has to do a little bit of
a calculation, and one finds

00:47:32.080 --> 00:47:36.650 align:middle line:84%
that that's the factor that you
need to have here in order

00:47:36.650 --> 00:47:37.770 align:middle line:90%
to be unbiased.

00:47:37.770 --> 00:47:42.280 align:middle line:84%
Of course, if you get 100 data
points, whether you divide by

00:47:42.280 --> 00:47:46.070 align:middle line:84%
100 or divided by 99, it's
going to make only a tiny

00:47:46.070 --> 00:47:48.620 align:middle line:84%
difference in your estimate
of your variance.

00:47:48.620 --> 00:47:50.740 align:middle line:84%
So it's going to make only
a tiny difference in your

00:47:50.740 --> 00:47:52.670 align:middle line:84%
estimate of the standard
deviation.

00:47:52.670 --> 00:47:54.180 align:middle line:90%
It's not a big deal.

00:47:54.180 --> 00:47:56.550 align:middle line:90%
And it doesn't really matter.

00:47:56.550 --> 00:48:00.720 align:middle line:84%
But if you want to show off
about your deeper knowledge of

00:48:00.720 --> 00:48:06.810 align:middle line:84%
statistics, you throw in the
1 over n-1 factor in there.

00:48:06.810 --> 00:48:11.350 align:middle line:84%
So now one basically needs to
put together this story here,

00:48:11.350 --> 00:48:15.260 align:middle line:90%
how you estimate the variance.

00:48:15.260 --> 00:48:18.370 align:middle line:84%
You first estimate
the sample mean.

00:48:18.370 --> 00:48:21.010 align:middle line:84%
And then you do some extra
work to come up with a

00:48:21.010 --> 00:48:23.020 align:middle line:84%
reasonable estimate of
the variance and

00:48:23.020 --> 00:48:24.640 align:middle line:90%
the standard deviation.

00:48:24.640 --> 00:48:27.510 align:middle line:84%
And then you use your estimate,
of the standard

00:48:27.510 --> 00:48:32.960 align:middle line:84%
deviation, to come up with a
confidence interval, which has

00:48:32.960 --> 00:48:35.150 align:middle line:90%
these two endpoints.

00:48:35.150 --> 00:48:39.130 align:middle line:84%
In doing this procedure, there's
basically a number of

00:48:39.130 --> 00:48:41.810 align:middle line:84%
approximations that
are involved.

00:48:41.810 --> 00:48:43.570 align:middle line:84%
There are two types
of approximations.

00:48:43.570 --> 00:48:46.170 align:middle line:84%
One approximation is that we're
pretending that the

00:48:46.170 --> 00:48:48.720 align:middle line:84%
sample mean has a normal
distribution.

00:48:48.720 --> 00:48:51.080 align:middle line:84%
That's something we're justified
to do, by the

00:48:51.080 --> 00:48:52.470 align:middle line:90%
central limit theorem.

00:48:52.470 --> 00:48:53.550 align:middle line:90%
But it's not exact.

00:48:53.550 --> 00:48:54.910 align:middle line:90%
It's an approximation.

00:48:54.910 --> 00:48:58.080 align:middle line:84%
And the second approximation
that comes in is that, instead

00:48:58.080 --> 00:49:01.260 align:middle line:84%
of using the correct standard
deviation, in general, you

00:49:01.260 --> 00:49:04.850 align:middle line:84%
will have to use some
approximation of

00:49:04.850 --> 00:49:06.100 align:middle line:90%
the standard deviation.

00:49:06.100 --> 00:49:08.390 align:middle line:90%


00:49:08.390 --> 00:49:11.200 align:middle line:84%
Okay so you will be getting a
little bit of practice with

00:49:11.200 --> 00:49:14.550 align:middle line:84%
these concepts in recitation
and tutorial.

00:49:14.550 --> 00:49:18.070 align:middle line:84%
And we will move on to
new topics next week.

00:49:18.070 --> 00:49:20.930 align:middle line:84%
But the material that's going
to be covered in the final

00:49:20.930 --> 00:49:23.570 align:middle line:90%
exam is only up to this point.

00:49:23.570 --> 00:49:28.220 align:middle line:84%
So next week is just
general education.

00:49:28.220 --> 00:49:30.550 align:middle line:84%
Hopefully useful, but it's
not in the exam.

00:49:30.550 --> 00:49:31.800 align:middle line:90%