WEBVTT

00:00:00.000 --> 00:00:00.040 align:middle line:90%


00:00:00.040 --> 00:00:02.460 align:middle line:84%
The following content is
provided under a Creative

00:00:02.460 --> 00:00:03.870 align:middle line:90%
Commons license.

00:00:03.870 --> 00:00:06.910 align:middle line:84%
Your support will help MIT
OpenCourseWare continue to

00:00:06.910 --> 00:00:10.560 align:middle line:84%
offer high quality educational
resources for free.

00:00:10.560 --> 00:00:13.460 align:middle line:84%
To make a donation or view
additional materials from

00:00:13.460 --> 00:00:19.290 align:middle line:84%
hundreds of MIT courses, visit
MIT OpenCourseWare at

00:00:19.290 --> 00:00:22.004 align:middle line:90%
ocw.mit.edu

00:00:22.004 --> 00:00:24.966 align:middle line:84%
JOHN TSISIKLIS: So here's
the agenda for today.

00:00:24.966 --> 00:00:26.848 align:middle line:84%
We're going to do a
very quick review.

00:00:26.848 --> 00:00:28.936 align:middle line:84%
And then we're going
to introduce some

00:00:28.936 --> 00:00:30.560 align:middle line:90%
very important concepts.

00:00:30.560 --> 00:00:34.060 align:middle line:84%
The idea is that all
information is--

00:00:34.060 --> 00:00:36.450 align:middle line:90%
Information is always partial.

00:00:36.450 --> 00:00:40.260 align:middle line:84%
And the question is what do we
do to probabilities if we have

00:00:40.260 --> 00:00:43.340 align:middle line:84%
some partial information about
the random experiments.

00:00:43.340 --> 00:00:45.770 align:middle line:84%
We're going to introduce the
important concept of

00:00:45.770 --> 00:00:47.530 align:middle line:90%
conditional probability.

00:00:47.530 --> 00:00:50.860 align:middle line:84%
And then we will see three
very useful ways

00:00:50.860 --> 00:00:52.670 align:middle line:90%
in which it is used.

00:00:52.670 --> 00:00:55.410 align:middle line:84%
And these ways basically
correspond to divide and

00:00:55.410 --> 00:00:58.070 align:middle line:84%
conquer methods for breaking
up problems

00:00:58.070 --> 00:01:00.120 align:middle line:90%
into simpler pieces.

00:01:00.120 --> 00:01:04.010 align:middle line:84%
And also one more fundamental
tool which allows us to use

00:01:04.010 --> 00:01:07.420 align:middle line:84%
conditional probabilities to do
inference, that is, if we

00:01:07.420 --> 00:01:09.440 align:middle line:84%
get a little bit of information
about some

00:01:09.440 --> 00:01:12.620 align:middle line:84%
phenomenon, what can we
infer about the things

00:01:12.620 --> 00:01:14.640 align:middle line:90%
that we have not seen?

00:01:14.640 --> 00:01:17.050 align:middle line:90%
So our quick review.

00:01:17.050 --> 00:01:22.100 align:middle line:84%
In setting up a model of a
random experiment, the first

00:01:22.100 --> 00:01:25.930 align:middle line:84%
thing to do is to come up with
a list of all the possible

00:01:25.930 --> 00:01:27.870 align:middle line:90%
outcomes of the experiment.

00:01:27.870 --> 00:01:31.120 align:middle line:84%
So that list is what we
call the sample space.

00:01:31.120 --> 00:01:32.480 align:middle line:90%
It's a set.

00:01:32.480 --> 00:01:34.580 align:middle line:84%
And the elements of the
sample space are all

00:01:34.580 --> 00:01:35.720 align:middle line:90%
the possible outcomes.

00:01:35.720 --> 00:01:37.560 align:middle line:90%
Those possible outcomes must be

00:01:37.560 --> 00:01:39.690 align:middle line:90%
distinguishable from each other.

00:01:39.690 --> 00:01:41.020 align:middle line:90%
They're mutually exclusive.

00:01:41.020 --> 00:01:44.900 align:middle line:84%
Either one happens or the other
happens, but not both.

00:01:44.900 --> 00:01:47.440 align:middle line:84%
And they are collectively
exhaustive, that is no matter

00:01:47.440 --> 00:01:50.480 align:middle line:84%
what the outcome of the
experiment is going to be an

00:01:50.480 --> 00:01:52.130 align:middle line:90%
element of the sample space.

00:01:52.130 --> 00:01:54.200 align:middle line:84%
And then we discussed last
time that there's also an

00:01:54.200 --> 00:01:57.510 align:middle line:84%
element of art in how to choose
your sample space,

00:01:57.510 --> 00:02:01.440 align:middle line:84%
depending on how much detail
you want to capture.

00:02:01.440 --> 00:02:03.130 align:middle line:90%
This is usually the easy part.

00:02:03.130 --> 00:02:06.980 align:middle line:84%
Then the more interesting part
is to assign probabilities to

00:02:06.980 --> 00:02:10.660 align:middle line:84%
our model, that is to make some
statements about what we

00:02:10.660 --> 00:02:14.610 align:middle line:84%
believe to be likely and what
we believe to be unlikely.

00:02:14.610 --> 00:02:17.720 align:middle line:84%
The way we do that is by
assigning probabilities to

00:02:17.720 --> 00:02:20.510 align:middle line:90%
subsets of the sample space.

00:02:20.510 --> 00:02:26.120 align:middle line:84%
So as we have our sample space
here, we may have a subset A.

00:02:26.120 --> 00:02:31.090 align:middle line:84%
And we assign a number to that
subset P(A), which is the

00:02:31.090 --> 00:02:33.910 align:middle line:84%
probability that this
event happens.

00:02:33.910 --> 00:02:37.080 align:middle line:84%
Or this is the probability that
when we do the experiment

00:02:37.080 --> 00:02:39.860 align:middle line:84%
and we get an outcome it's the
probability that the outcome

00:02:39.860 --> 00:02:41.850 align:middle line:84%
happens to fall inside
that event.

00:02:41.850 --> 00:02:44.500 align:middle line:84%
We have certain rules that
probabilities should satisfy.

00:02:44.500 --> 00:02:46.210 align:middle line:90%
They're non-negative.

00:02:46.210 --> 00:02:49.780 align:middle line:84%
The probability of the overall
sample space is equal to one,

00:02:49.780 --> 00:02:52.900 align:middle line:84%
which expresses the fact that
we're are certain, no matter

00:02:52.900 --> 00:02:55.480 align:middle line:84%
what, the outcome is going
to be an element

00:02:55.480 --> 00:02:56.830 align:middle line:90%
of the sample space.

00:02:56.830 --> 00:02:59.760 align:middle line:84%
Well, if we set the top right
so that it exhausts all

00:02:59.760 --> 00:03:03.190 align:middle line:84%
possibilities, this should
be the case.

00:03:03.190 --> 00:03:05.480 align:middle line:84%
And then there's another
interesting property of

00:03:05.480 --> 00:03:09.240 align:middle line:84%
probabilities that says that,
if we have two events or two

00:03:09.240 --> 00:03:11.910 align:middle line:84%
subsets that are disjoint, and
we're interested in the

00:03:11.910 --> 00:03:17.670 align:middle line:84%
probability, that one or the
other happens, that is the

00:03:17.670 --> 00:03:21.870 align:middle line:84%
outcome belongs to A or belongs
to B. For disjoint

00:03:21.870 --> 00:03:25.320 align:middle line:84%
events the total probability of
these two, taken together,

00:03:25.320 --> 00:03:28.030 align:middle line:84%
is just the sum of their
individual probabilities.

00:03:28.030 --> 00:03:30.270 align:middle line:84%
So probabilities behave
like masses.

00:03:30.270 --> 00:03:34.760 align:middle line:84%
The mass of the object
consisting of A and B is the

00:03:34.760 --> 00:03:37.230 align:middle line:84%
sum of the masses of
these two objects.

00:03:37.230 --> 00:03:39.720 align:middle line:84%
Or you can think of
probabilities as areas.

00:03:39.720 --> 00:03:41.240 align:middle line:84%
They have, again, the
same property.

00:03:41.240 --> 00:03:45.490 align:middle line:84%
The area of A together with B is
the area of A plus the area

00:03:45.490 --> 00:03:46.410 align:middle line:90%
B.

00:03:46.410 --> 00:03:50.290 align:middle line:84%
But as we discussed at the end
of last lecture, it's useful

00:03:50.290 --> 00:03:53.970 align:middle line:84%
to have in our hands a more
general version of this

00:03:53.970 --> 00:03:58.990 align:middle line:84%
additivity property, which says
the following, if we take

00:03:58.990 --> 00:04:00.982 align:middle line:90%
a sequence of sets--

00:04:00.982 --> 00:04:07.480 align:middle line:90%
A1, A2, A3, A4, and so on.

00:04:07.480 --> 00:04:09.630 align:middle line:84%
And we put all of those
sets together.

00:04:09.630 --> 00:04:11.410 align:middle line:90%
It's an infinite sequence.

00:04:11.410 --> 00:04:14.950 align:middle line:84%
And we ask for the probability
that the outcome falls

00:04:14.950 --> 00:04:19.170 align:middle line:84%
somewhere in this infinite
union, that is we are asking

00:04:19.170 --> 00:04:22.640 align:middle line:84%
for the probability that the
outcome belongs to one of

00:04:22.640 --> 00:04:27.950 align:middle line:84%
these sets, and assuming that
the sets are disjoint, we can

00:04:27.950 --> 00:04:32.820 align:middle line:84%
again find the probability for
the overall set by adding up

00:04:32.820 --> 00:04:36.000 align:middle line:84%
the probabilities of the
individual sets.

00:04:36.000 --> 00:04:38.910 align:middle line:84%
So this is a nice and
simple property.

00:04:38.910 --> 00:04:43.130 align:middle line:84%
But it's a little more subtle
than you might think.

00:04:43.130 --> 00:04:45.820 align:middle line:84%
And let's see what's going
on by considering

00:04:45.820 --> 00:04:47.770 align:middle line:90%
the following example.

00:04:47.770 --> 00:04:51.850 align:middle line:84%
We had an example last time
where we take our sample space

00:04:51.850 --> 00:04:53.800 align:middle line:90%
to be the unit square.

00:04:53.800 --> 00:04:58.110 align:middle line:84%
And we said let's consider a
probability law that says that

00:04:58.110 --> 00:05:04.190 align:middle line:84%
the probability of a subset is
just the area of that subset.

00:05:04.190 --> 00:05:07.630 align:middle line:84%
So let's consider this
probability law.

00:05:07.630 --> 00:05:08.530 align:middle line:90%
OK.

00:05:08.530 --> 00:05:13.990 align:middle line:84%
Now the unit square is
the set --let me just

00:05:13.990 --> 00:05:15.210 align:middle line:90%
draw it this way--

00:05:15.210 --> 00:05:20.520 align:middle line:84%
the unit square is the union of
one element set consisting

00:05:20.520 --> 00:05:21.680 align:middle line:90%
all of the points.

00:05:21.680 --> 00:05:28.280 align:middle line:84%
So the unit square is made up
by the union of the various

00:05:28.280 --> 00:05:30.740 align:middle line:90%
points inside the square.

00:05:30.740 --> 00:05:33.830 align:middle line:90%
So union over all x's and y's.

00:05:33.830 --> 00:05:34.770 align:middle line:90%
OK?

00:05:34.770 --> 00:05:36.690 align:middle line:84%
So the square is made
up out of all the

00:05:36.690 --> 00:05:38.400 align:middle line:90%
points that this contains.

00:05:38.400 --> 00:05:41.140 align:middle line:84%
And now let's do
a calculation.

00:05:41.140 --> 00:05:45.060 align:middle line:84%
One is the probability of our
overall sample space, which is

00:05:45.060 --> 00:05:47.260 align:middle line:90%
the unit square.

00:05:47.260 --> 00:06:02.000 align:middle line:84%
Now the unit square is the union
of these things, which,

00:06:02.000 --> 00:06:06.810 align:middle line:84%
according to our additivity
axiom, is the sum of the

00:06:06.810 --> 00:06:10.595 align:middle line:84%
probabilities of all of these
one element sets.

00:06:10.595 --> 00:06:16.830 align:middle line:90%


00:06:16.830 --> 00:06:20.580 align:middle line:84%
Now what is the probability
of a one element set?

00:06:20.580 --> 00:06:23.520 align:middle line:84%
What is the probability of
this one element set?

00:06:23.520 --> 00:06:26.100 align:middle line:84%
What's the probability that our
outcome is exactly that

00:06:26.100 --> 00:06:27.490 align:middle line:90%
particular point?

00:06:27.490 --> 00:06:31.460 align:middle line:84%
Well, it's the area of that
set, which is zero.

00:06:31.460 --> 00:06:33.990 align:middle line:90%
So it's just the sum of zeros.

00:06:33.990 --> 00:06:35.950 align:middle line:84%
And by any reasonable
definition the

00:06:35.950 --> 00:06:38.370 align:middle line:90%
sum of zeros is zero.

00:06:38.370 --> 00:06:42.220 align:middle line:84%
So we just proved that
one is equal to zero.

00:06:42.220 --> 00:06:42.680 align:middle line:90%
OK.

00:06:42.680 --> 00:06:48.340 align:middle line:84%
Either probability theory is
dead or there is some mistake

00:06:48.340 --> 00:06:51.030 align:middle line:90%
in the derivation that I did.

00:06:51.030 --> 00:06:54.580 align:middle line:84%
OK, the mistake is quite
subtle and it

00:06:54.580 --> 00:06:57.300 align:middle line:90%
comes at this step.

00:06:57.300 --> 00:07:00.640 align:middle line:84%
We're sort of applied the
additivity axiom by saying

00:07:00.640 --> 00:07:04.040 align:middle line:84%
that the unit square is the
union of all those sets.

00:07:04.040 --> 00:07:06.500 align:middle line:84%
Can we really apply our
additivity axiom.

00:07:06.500 --> 00:07:07.260 align:middle line:90%
Here's the catch.

00:07:07.260 --> 00:07:11.470 align:middle line:84%
The additivity axiom applies
to the case where we have a

00:07:11.470 --> 00:07:17.180 align:middle line:84%
sequence of disjoint events
and we take their union.

00:07:17.180 --> 00:07:21.740 align:middle line:90%
Is this a sequence of sets?

00:07:21.740 --> 00:07:27.780 align:middle line:84%
Can you make up the whole unit
square by taking a sequence of

00:07:27.780 --> 00:07:31.310 align:middle line:84%
elements inside it and cover
the whole unit square?

00:07:31.310 --> 00:07:34.900 align:middle line:84%
Well if you try, if you start
looking at the sequence of one

00:07:34.900 --> 00:07:40.910 align:middle line:84%
element points, that sequence
will never be able to exhaust

00:07:40.910 --> 00:07:43.100 align:middle line:90%
the whole unit square.

00:07:43.100 --> 00:07:45.680 align:middle line:84%
So there's a deeper reason
behind that.

00:07:45.680 --> 00:07:48.790 align:middle line:84%
And the reason is that infinite
sets are not all of

00:07:48.790 --> 00:07:50.130 align:middle line:90%
the same size.

00:07:50.130 --> 00:07:52.620 align:middle line:84%
The integers are an
infinite set.

00:07:52.620 --> 00:07:55.510 align:middle line:84%
And you can arrange the integers
in a sequence.

00:07:55.510 --> 00:07:57.630 align:middle line:84%
But the continuous set
like the units

00:07:57.630 --> 00:08:00.205 align:middle line:90%
square is a bigger set.

00:08:00.205 --> 00:08:02.050 align:middle line:90%
It's so-called uncountable.

00:08:02.050 --> 00:08:06.160 align:middle line:84%
It has more elements than
any sequence could have.

00:08:06.160 --> 00:08:13.610 align:middle line:84%
So this union here is not of
this kind, where we would have

00:08:13.610 --> 00:08:16.930 align:middle line:90%
a sequence of events.

00:08:16.930 --> 00:08:18.370 align:middle line:84%
It's a different
kind of union.

00:08:18.370 --> 00:08:23.070 align:middle line:84%
It's a Union that involves a
union of many, many more sets.

00:08:23.070 --> 00:08:25.420 align:middle line:84%
So the countable additivity
axiom does not

00:08:25.420 --> 00:08:27.360 align:middle line:90%
apply in this case.

00:08:27.360 --> 00:08:30.230 align:middle line:84%
Because, we're not dealing
with a sequence of sets.

00:08:30.230 --> 00:08:33.780 align:middle line:84%
And so this is the
incorrect step.

00:08:33.780 --> 00:08:37.240 align:middle line:84%
So at some level you might think
that this is puzzling

00:08:37.240 --> 00:08:38.580 align:middle line:90%
and awfully confusing.

00:08:38.580 --> 00:08:41.070 align:middle line:84%
On the other hand, if you think
about areas of the way

00:08:41.070 --> 00:08:43.520 align:middle line:84%
you're used to them from
calculus, there's nothing

00:08:43.520 --> 00:08:44.940 align:middle line:90%
mysterious about it.

00:08:44.940 --> 00:08:47.460 align:middle line:84%
Every point on the unit
square has zero area.

00:08:47.460 --> 00:08:50.140 align:middle line:84%
When you put all the points
together, they make up

00:08:50.140 --> 00:08:52.330 align:middle line:84%
something that has
finite area.

00:08:52.330 --> 00:08:55.470 align:middle line:84%
So there shouldn't be any
mystery behind it.

00:08:55.470 --> 00:09:00.230 align:middle line:84%
Now, one interesting thing that
this discussion tells us,

00:09:00.230 --> 00:09:03.670 align:middle line:84%
especially the fact that the
single elements set has zero

00:09:03.670 --> 00:09:05.790 align:middle line:90%
area, is the following--

00:09:05.790 --> 00:09:08.960 align:middle line:84%
Individual points have
zero probability.

00:09:08.960 --> 00:09:12.390 align:middle line:84%
After you do the experiment and
you observe the outcome,

00:09:12.390 --> 00:09:14.660 align:middle line:84%
it's going to be an
individual point.

00:09:14.660 --> 00:09:18.160 align:middle line:84%
So what happened in that
experiment is something that

00:09:18.160 --> 00:09:21.820 align:middle line:84%
initially you thought had zero
probability of occurring.

00:09:21.820 --> 00:09:25.420 align:middle line:84%
So if you happen to get some
particular numbers and you

00:09:25.420 --> 00:09:28.290 align:middle line:84%
say, "Well, in the beginning,
what did I think about those

00:09:28.290 --> 00:09:29.280 align:middle line:90%
specific numbers?

00:09:29.280 --> 00:09:31.290 align:middle line:84%
I thought they had
zero probability.

00:09:31.290 --> 00:09:36.250 align:middle line:84%
But yet those particular
numbers did occur."

00:09:36.250 --> 00:09:41.640 align:middle line:84%
So one moral from this is that
zero probability does not mean

00:09:41.640 --> 00:09:42.890 align:middle line:90%
impossible.

00:09:42.890 --> 00:09:46.920 align:middle line:84%
It just means extremely,
extremely unlikely by itself.

00:09:46.920 --> 00:09:49.420 align:middle line:84%
So zero probability
things do happen.

00:09:49.420 --> 00:09:53.340 align:middle line:84%
In such continuous models,
actually zero probability

00:09:53.340 --> 00:09:56.930 align:middle line:84%
outcomes are everything
that happens.

00:09:56.930 --> 00:10:00.790 align:middle line:84%
And the bumper sticker version
of this is to always expect

00:10:00.790 --> 00:10:02.220 align:middle line:90%
the unexpected.

00:10:02.220 --> 00:10:05.095 align:middle line:90%
Yes?

00:10:05.095 --> 00:10:06.345 align:middle line:90%
AUDIENCE: [INAUDIBLE].

00:10:06.345 --> 00:10:08.532 align:middle line:90%


00:10:08.532 --> 00:10:11.800 align:middle line:84%
JOHN TSISIKLIS: Well,
probability is supposed to be

00:10:11.800 --> 00:10:12.530 align:middle line:90%
a real number.

00:10:12.530 --> 00:10:16.220 align:middle line:84%
So it's either zero or it's
a positive number.

00:10:16.220 --> 00:10:21.350 align:middle line:84%
So you can think of the
probability of things just

00:10:21.350 --> 00:10:25.040 align:middle line:84%
close to that point and those
probabilities are tiny and

00:10:25.040 --> 00:10:26.390 align:middle line:90%
close to zero.

00:10:26.390 --> 00:10:28.780 align:middle line:84%
So that's how we're going to
interpret probabilities in

00:10:28.780 --> 00:10:29.810 align:middle line:90%
continuous models.

00:10:29.810 --> 00:10:31.340 align:middle line:84%
But this is two chapters
ahead.

00:10:31.340 --> 00:10:33.950 align:middle line:90%


00:10:33.950 --> 00:10:34.230 align:middle line:90%
Yeah?

00:10:34.230 --> 00:10:36.198 align:middle line:84%
AUDIENCE: How do we interpret
probability of zero?

00:10:36.198 --> 00:10:37.674 align:middle line:84%
If we can use models that
way, then how about

00:10:37.674 --> 00:10:38.658 align:middle line:90%
probability of one?

00:10:38.658 --> 00:10:40.462 align:middle line:84%
That it it's extremely
likely but not

00:10:40.462 --> 00:10:42.110 align:middle line:90%
necessarily for certain?

00:10:42.110 --> 00:10:43.320 align:middle line:84%
JOHN TSISIKLIS: That's
also the case.

00:10:43.320 --> 00:10:47.450 align:middle line:84%
For example, if you ask in this
continuous model, if you

00:10:47.450 --> 00:10:52.190 align:middle line:84%
ask me for the probability that
x, y, is different than

00:10:52.190 --> 00:10:55.840 align:middle line:84%
the zero, zero this is
the whole square,

00:10:55.840 --> 00:10:57.220 align:middle line:90%
except for one point.

00:10:57.220 --> 00:11:01.150 align:middle line:84%
So the area of this is
going to be one.

00:11:01.150 --> 00:11:06.330 align:middle line:84%
But this event is not entirely
certain because the zero, zero

00:11:06.330 --> 00:11:08.210 align:middle line:90%
outcome is also possible.

00:11:08.210 --> 00:11:12.330 align:middle line:84%
So again, probability of one
means essential certainty.

00:11:12.330 --> 00:11:16.450 align:middle line:84%
But it still allows the
possibility that the outcome

00:11:16.450 --> 00:11:18.320 align:middle line:90%
might be outside that set.

00:11:18.320 --> 00:11:20.910 align:middle line:84%
So these are some of the weird
things that are happening when

00:11:20.910 --> 00:11:22.680 align:middle line:90%
you have continuous models.

00:11:22.680 --> 00:11:25.240 align:middle line:84%
And that's why we start to
this class with discrete

00:11:25.240 --> 00:11:27.050 align:middle line:84%
models, on which would
be spending the

00:11:27.050 --> 00:11:30.400 align:middle line:90%
next couple of weeks.

00:11:30.400 --> 00:11:30.820 align:middle line:90%
OK.

00:11:30.820 --> 00:11:35.650 align:middle line:84%
So now once we have set up our
probability model and we have

00:11:35.650 --> 00:11:39.160 align:middle line:84%
a legitimate probability law
that has these properties,

00:11:39.160 --> 00:11:43.070 align:middle line:84%
then the rest is
usually simple.

00:11:43.070 --> 00:11:45.950 align:middle line:84%
Somebody asks you a question of
calculating the probability

00:11:45.950 --> 00:11:47.520 align:middle line:90%
of some event.

00:11:47.520 --> 00:11:50.270 align:middle line:84%
While you were told something
about the probability law,

00:11:50.270 --> 00:11:52.520 align:middle line:84%
such as for example the
probabilities are equal to

00:11:52.520 --> 00:11:55.460 align:middle line:84%
areas, and then you just
need to calculate.

00:11:55.460 --> 00:11:58.730 align:middle line:84%
In these type of examples
somebody would give you a set

00:11:58.730 --> 00:12:00.230 align:middle line:84%
and you would have
to calculate the

00:12:00.230 --> 00:12:01.500 align:middle line:90%
area of that set.

00:12:01.500 --> 00:12:06.060 align:middle line:84%
So the rest is just calculation
and simple.

00:12:06.060 --> 00:12:09.390 align:middle line:84%
Alright, so now it's time
to start with our main

00:12:09.390 --> 00:12:12.600 align:middle line:90%
business for today.

00:12:12.600 --> 00:12:16.880 align:middle line:84%
And the starting point
is the following--

00:12:16.880 --> 00:12:18.920 align:middle line:84%
You know something
about the world.

00:12:18.920 --> 00:12:21.690 align:middle line:84%
And based on what you know when
you set up a probability

00:12:21.690 --> 00:12:23.820 align:middle line:84%
model and you write down
probabilities for the

00:12:23.820 --> 00:12:26.000 align:middle line:90%
different outcomes.

00:12:26.000 --> 00:12:28.950 align:middle line:84%
Then something happens, and
somebody tells you a little

00:12:28.950 --> 00:12:33.620 align:middle line:84%
more about the world, gives
you some new information.

00:12:33.620 --> 00:12:37.430 align:middle line:84%
This new information, in
general, should change your

00:12:37.430 --> 00:12:41.240 align:middle line:84%
beliefs about what happened
or what may happen.

00:12:41.240 --> 00:12:44.550 align:middle line:84%
So whenever we're given new
information, some partial

00:12:44.550 --> 00:12:47.400 align:middle line:84%
information about the outcome
of the experiment, we should

00:12:47.400 --> 00:12:49.750 align:middle line:90%
revise our beliefs.

00:12:49.750 --> 00:12:54.470 align:middle line:84%
And conditional probabilities
are just the probabilities

00:12:54.470 --> 00:12:58.820 align:middle line:84%
that apply after the revision
of our beliefs, when we're

00:12:58.820 --> 00:13:00.580 align:middle line:90%
given some information.

00:13:00.580 --> 00:13:04.510 align:middle line:84%
So lets make this into
a numerical example.

00:13:04.510 --> 00:13:07.870 align:middle line:84%
So inside the sample space, this
part of the sample space,

00:13:07.870 --> 00:13:12.580 align:middle line:84%
let's say has probability 3/6,
this part has 2/6, and that

00:13:12.580 --> 00:13:14.550 align:middle line:90%
part has 1/6.

00:13:14.550 --> 00:13:17.940 align:middle line:84%
I guess that means that out here
we have zero probability.

00:13:17.940 --> 00:13:21.900 align:middle line:84%
So these were our initial
beliefs about the outcome of

00:13:21.900 --> 00:13:23.270 align:middle line:90%
the experiment.

00:13:23.270 --> 00:13:27.160 align:middle line:84%
Suppose now that someone
comes and tells you

00:13:27.160 --> 00:13:30.960 align:middle line:90%
that event B occurred.

00:13:30.960 --> 00:13:33.560 align:middle line:84%
So they don't tell you the
full outcome with the

00:13:33.560 --> 00:13:34.440 align:middle line:90%
experiment.

00:13:34.440 --> 00:13:38.960 align:middle line:84%
But they just tell you that the
outcome is known to lie

00:13:38.960 --> 00:13:41.060 align:middle line:90%
inside this set B.

00:13:41.060 --> 00:13:44.320 align:middle line:84%
Well then, you should certainly
change your beliefs

00:13:44.320 --> 00:13:45.560 align:middle line:90%
in some way.

00:13:45.560 --> 00:13:48.420 align:middle line:84%
And your new beliefs about what
is likely to occur and

00:13:48.420 --> 00:13:51.770 align:middle line:84%
what is not is going to be
denoted by this notation.

00:13:51.770 --> 00:13:55.330 align:middle line:84%
This is the conditional
probability that the event A

00:13:55.330 --> 00:13:57.970 align:middle line:84%
is going to occur, the
probability that the outcome

00:13:57.970 --> 00:14:01.580 align:middle line:84%
is going to fall inside the set
A given that we are told

00:14:01.580 --> 00:14:05.890 align:middle line:84%
and we're sure that the event
lies inside the event B Now

00:14:05.890 --> 00:14:09.000 align:middle line:84%
once you're told that the
outcome lies inside the event

00:14:09.000 --> 00:14:13.740 align:middle line:84%
B, then our old sample space
in some ways is irrelevant.

00:14:13.740 --> 00:14:16.975 align:middle line:84%
We have then you sample space,
which is just the set B. We

00:14:16.975 --> 00:14:21.020 align:middle line:84%
are certain that the outcome
is going to be inside B.

00:14:21.020 --> 00:14:25.465 align:middle line:84%
For example, what is this
conditional probability?

00:14:25.465 --> 00:14:29.120 align:middle line:90%


00:14:29.120 --> 00:14:30.160 align:middle line:90%
It should be one.

00:14:30.160 --> 00:14:33.250 align:middle line:84%
Given that I told you that B
occurred, you're certain that

00:14:33.250 --> 00:14:36.380 align:middle line:84%
B occurred, so this has
unit probability.

00:14:36.380 --> 00:14:40.340 align:middle line:84%
So here we see an instance of
revision of our beliefs.

00:14:40.340 --> 00:14:44.880 align:middle line:84%
Initially, event B had the
probability of (2+1)/6 --

00:14:44.880 --> 00:14:46.300 align:middle line:90%
that's 1/2.

00:14:46.300 --> 00:14:49.500 align:middle line:84%
Initially, we thought B
had probability 1/2.

00:14:49.500 --> 00:14:52.370 align:middle line:84%
Once we're told that B occurred,
the new probability

00:14:52.370 --> 00:14:54.250 align:middle line:90%
of B is equal to one.

00:14:54.250 --> 00:14:55.160 align:middle line:90%
OK.

00:14:55.160 --> 00:15:00.860 align:middle line:84%
How do we revise the probability
that A occurs?

00:15:00.860 --> 00:15:03.950 align:middle line:84%
So we are going to have the
outcome of the experiment.

00:15:03.950 --> 00:15:07.330 align:middle line:84%
We know that it's inside B. So
we will either get something

00:15:07.330 --> 00:15:09.200 align:middle line:90%
here, and A does not occur.

00:15:09.200 --> 00:15:12.570 align:middle line:84%
Or something inside here,
and A does occur.

00:15:12.570 --> 00:15:16.280 align:middle line:84%
What's the likelihood that,
given that we're inside B, the

00:15:16.280 --> 00:15:18.160 align:middle line:90%
outcome is inside here?

00:15:18.160 --> 00:15:21.380 align:middle line:84%
Here's how we're going
to think about.

00:15:21.380 --> 00:15:26.110 align:middle line:84%
This part of this set B, in
which A also occurs, in our

00:15:26.110 --> 00:15:31.280 align:middle line:84%
initial model was twice as
likely as that part of B. So

00:15:31.280 --> 00:15:36.220 align:middle line:84%
outcomes inside here
collectively were twice as

00:15:36.220 --> 00:15:38.950 align:middle line:90%
likely as outcomes out there.

00:15:38.950 --> 00:15:43.240 align:middle line:84%
So we're going to keep the same
proportions and say, that

00:15:43.240 --> 00:15:47.280 align:middle line:84%
given that we are inside the set
B, we still want outcomes

00:15:47.280 --> 00:15:51.120 align:middle line:84%
inside here to be twice as
likely outcomes there.

00:15:51.120 --> 00:15:55.800 align:middle line:84%
So the proportion of the
probabilities should be two

00:15:55.800 --> 00:15:57.570 align:middle line:90%
versus one.

00:15:57.570 --> 00:16:01.210 align:middle line:84%
And these probabilities should
add up to one because together

00:16:01.210 --> 00:16:04.340 align:middle line:84%
they make the conditional
probability of B. So the

00:16:04.340 --> 00:16:09.260 align:middle line:84%
conditional probabilities should
be 2/3 probability of

00:16:09.260 --> 00:16:13.080 align:middle line:84%
being here and 1/3 probability
of being there.

00:16:13.080 --> 00:16:16.860 align:middle line:84%
That's how we revise
our probabilities.

00:16:16.860 --> 00:16:20.740 align:middle line:84%
That's a reasonable, intuitively
reasonable, way of

00:16:20.740 --> 00:16:22.230 align:middle line:90%
doing this revision.

00:16:22.230 --> 00:16:26.650 align:middle line:84%
Let's translate what we
did into a definition.

00:16:26.650 --> 00:16:29.490 align:middle line:84%
The definition says the
following, that the

00:16:29.490 --> 00:16:33.410 align:middle line:84%
conditional probability of A
given that B occurred is

00:16:33.410 --> 00:16:35.270 align:middle line:90%
calculated as follows.

00:16:35.270 --> 00:16:39.430 align:middle line:84%
We look at the total probability
of B. And out of

00:16:39.430 --> 00:16:43.190 align:middle line:84%
that probability that was inside
here, what fraction of

00:16:43.190 --> 00:16:48.310 align:middle line:84%
that probability is assigned to
points for which the event

00:16:48.310 --> 00:16:49.780 align:middle line:90%
A also occurs?

00:16:49.780 --> 00:16:54.480 align:middle line:90%


00:16:54.480 --> 00:16:56.860 align:middle line:84%
Does it give us the same numbers
as we got with this

00:16:56.860 --> 00:16:58.420 align:middle line:90%
heuristic argument?

00:16:58.420 --> 00:17:01.530 align:middle line:84%
Well in this example,
probability of A intersection

00:17:01.530 --> 00:17:06.359 align:middle line:84%
B is 2/6, divided by total
probability of B, which is

00:17:06.359 --> 00:17:12.369 align:middle line:84%
3/6, and so it's 2/3, which
agrees with this answer that's

00:17:12.369 --> 00:17:13.589 align:middle line:90%
we got before.

00:17:13.589 --> 00:17:18.280 align:middle line:84%
So the former indeed matches
what we were trying to do.

00:17:18.280 --> 00:17:21.040 align:middle line:90%
One little technical detail.

00:17:21.040 --> 00:17:24.970 align:middle line:84%
If the event B has zero
probability, and then here we

00:17:24.970 --> 00:17:27.770 align:middle line:84%
have a ratio that doesn't
make sense.

00:17:27.770 --> 00:17:30.470 align:middle line:84%
So in this case, we say that
conditional probabilities are

00:17:30.470 --> 00:17:31.720 align:middle line:90%
not defined.

00:17:31.720 --> 00:17:34.780 align:middle line:90%


00:17:34.780 --> 00:17:38.980 align:middle line:84%
Now you can take this definition
and unravel it and

00:17:38.980 --> 00:17:40.260 align:middle line:90%
write it in this form.

00:17:40.260 --> 00:17:43.510 align:middle line:84%
The probability of A
intersection B is the

00:17:43.510 --> 00:17:46.780 align:middle line:84%
probability of B times the
conditional probability.

00:17:46.780 --> 00:17:50.350 align:middle line:90%


00:17:50.350 --> 00:17:53.820 align:middle line:84%
So this is just consequence of
the definition but it has a

00:17:53.820 --> 00:17:55.370 align:middle line:90%
nice interpretation.

00:17:55.370 --> 00:17:57.930 align:middle line:84%
Think of probabilities
as frequencies.

00:17:57.930 --> 00:18:01.480 align:middle line:84%
If I do the experiment over and
over, what fraction of the

00:18:01.480 --> 00:18:05.300 align:middle line:84%
time is it going to be the case
that both A and B occur?

00:18:05.300 --> 00:18:08.490 align:middle line:84%
Well, there's going to be a
certain fraction of the time

00:18:08.490 --> 00:18:10.820 align:middle line:90%
at which B occurs.

00:18:10.820 --> 00:18:14.760 align:middle line:84%
And out of those times when B
occurs, there's going to be a

00:18:14.760 --> 00:18:17.270 align:middle line:84%
further fraction of
the experiments in

00:18:17.270 --> 00:18:19.410 align:middle line:90%
which A also occurs.

00:18:19.410 --> 00:18:21.930 align:middle line:84%
So interpret the conditional
probability as follows.

00:18:21.930 --> 00:18:24.320 align:middle line:84%
You only look at those
experiments at which

00:18:24.320 --> 00:18:26.050 align:middle line:90%
B happens to occur.

00:18:26.050 --> 00:18:29.820 align:middle line:84%
And look at what fraction of
those experiments where B

00:18:29.820 --> 00:18:33.670 align:middle line:84%
already occurred, event
A also occurs.

00:18:33.670 --> 00:18:39.610 align:middle line:84%
And there's a symmetrical
version of this equality.

00:18:39.610 --> 00:18:44.660 align:middle line:84%
There's symmetry between the
events B and A. So you also

00:18:44.660 --> 00:18:48.890 align:middle line:84%
have this relation that
goes the other way.

00:18:48.890 --> 00:18:53.950 align:middle line:84%
OK, so what do we use these
conditional probabilities for?

00:18:53.950 --> 00:18:55.120 align:middle line:90%
First, one comment.

00:18:55.120 --> 00:18:58.100 align:middle line:84%
Conditional probabilities
are just like ordinary

00:18:58.100 --> 00:18:59.170 align:middle line:90%
probabilities.

00:18:59.170 --> 00:19:02.820 align:middle line:84%
They're the new probabilities
that apply in a new universe

00:19:02.820 --> 00:19:07.300 align:middle line:84%
where event B is known
to have occurred.

00:19:07.300 --> 00:19:10.620 align:middle line:84%
So we had an original
probability model.

00:19:10.620 --> 00:19:12.210 align:middle line:90%
We are told that B occurs.

00:19:12.210 --> 00:19:13.840 align:middle line:90%
We revise our model.

00:19:13.840 --> 00:19:16.690 align:middle line:84%
Our new model should still be
legitimate probability model.

00:19:16.690 --> 00:19:20.770 align:middle line:84%
So it should satisfy all sorts
of properties that ordinary

00:19:20.770 --> 00:19:23.210 align:middle line:90%
probabilities do satisfy.

00:19:23.210 --> 00:19:29.230 align:middle line:84%
So for example, if A and B are
disjoint events, then we know

00:19:29.230 --> 00:19:33.830 align:middle line:84%
that the probability of A
union B is equal to the

00:19:33.830 --> 00:19:39.230 align:middle line:84%
probability of A plus
probability of B. And now if I

00:19:39.230 --> 00:19:42.770 align:middle line:84%
tell you that a certain event C
occurred, we're placed in a

00:19:42.770 --> 00:19:45.220 align:middle line:84%
new universe where
event C occurred.

00:19:45.220 --> 00:19:47.515 align:middle line:84%
We have new probabilities
for that universe.

00:19:47.515 --> 00:19:49.880 align:middle line:84%
These are the conditional
probabilities.

00:19:49.880 --> 00:19:52.960 align:middle line:84%
And conditional probabilities
also satisfy

00:19:52.960 --> 00:19:54.820 align:middle line:90%
this kind of property.

00:19:54.820 --> 00:19:58.380 align:middle line:84%
So this is just our usual
additivity axiom but the

00:19:58.380 --> 00:20:02.290 align:middle line:84%
applied in a new model, in which
we were told that event

00:20:02.290 --> 00:20:03.250 align:middle line:90%
C occurred.

00:20:03.250 --> 00:20:06.580 align:middle line:84%
So conditional probabilities
do not taste or smell any

00:20:06.580 --> 00:20:09.970 align:middle line:84%
different than ordinary
probabilities do.

00:20:09.970 --> 00:20:14.350 align:middle line:84%
Conditional probabilities, given
a specific event B, just

00:20:14.350 --> 00:20:19.480 align:middle line:84%
form a probability law
on our sample space.

00:20:19.480 --> 00:20:22.460 align:middle line:84%
It's a different probability
law but it's still a

00:20:22.460 --> 00:20:26.430 align:middle line:84%
probability law that has all
of the desired properties.

00:20:26.430 --> 00:20:30.360 align:middle line:84%
OK, so where do conditional
probabilities come up?

00:20:30.360 --> 00:20:32.450 align:middle line:84%
They do come up in quizzes
and they do

00:20:32.450 --> 00:20:34.070 align:middle line:90%
come up in silly problems.

00:20:34.070 --> 00:20:35.680 align:middle line:90%
So let's start with this.

00:20:35.680 --> 00:20:37.790 align:middle line:84%
We have this example
from last time.

00:20:37.790 --> 00:20:42.220 align:middle line:84%
Two rolls of a die, all possible
pairs of roles are

00:20:42.220 --> 00:20:46.410 align:middle line:84%
equally likely, so every element
in this square has

00:20:46.410 --> 00:20:47.660 align:middle line:90%
probability of 1/16.

00:20:47.660 --> 00:20:50.300 align:middle line:90%


00:20:50.300 --> 00:20:52.330 align:middle line:84%
So all elements are
equally likely.

00:20:52.330 --> 00:20:54.280 align:middle line:90%
That's our original model.

00:20:54.280 --> 00:20:57.210 align:middle line:84%
Then somebody comes and tells us
that the minimum of the two

00:20:57.210 --> 00:20:59.530 align:middle line:90%
rolls is equal to zero.

00:20:59.530 --> 00:21:02.060 align:middle line:90%
What's that event?

00:21:02.060 --> 00:21:05.990 align:middle line:84%
The minimum equal to zero can
happen in many ways, if we get

00:21:05.990 --> 00:21:08.990 align:middle line:84%
two zeros or if we
get a zero and--

00:21:08.990 --> 00:21:13.140 align:middle line:84%
sorry, if we get two
two's, or get a two

00:21:13.140 --> 00:21:14.830 align:middle line:90%
and something larger.

00:21:14.830 --> 00:21:21.400 align:middle line:84%
And so the is our new event B.
The red event is the event B.

00:21:21.400 --> 00:21:23.500 align:middle line:84%
And now we want to calculate
probabilities

00:21:23.500 --> 00:21:25.310 align:middle line:90%
inside this new universe.

00:21:25.310 --> 00:21:28.770 align:middle line:84%
For example, you may be
interested in the question,

00:21:28.770 --> 00:21:31.960 align:middle line:84%
questions about the maximum
of the two rolls.

00:21:31.960 --> 00:21:34.310 align:middle line:84%
In the new universe, what's
the probability that the

00:21:34.310 --> 00:21:37.550 align:middle line:90%
maximum is equal to one?

00:21:37.550 --> 00:21:44.320 align:middle line:84%
The maximum being equal to
one is this black event.

00:21:44.320 --> 00:21:49.240 align:middle line:84%
And given that we're told that
B occurred, this black events

00:21:49.240 --> 00:21:50.300 align:middle line:90%
cannot happen.

00:21:50.300 --> 00:21:53.240 align:middle line:84%
So this probability
is equal to zero.

00:21:53.240 --> 00:21:56.500 align:middle line:84%
How about the maximum
being equal to two,

00:21:56.500 --> 00:21:59.110 align:middle line:90%
given that event B?

00:21:59.110 --> 00:22:01.760 align:middle line:84%
OK, we can use the
definition here.

00:22:01.760 --> 00:22:05.730 align:middle line:84%
It's going to be the probability
that the maximum

00:22:05.730 --> 00:22:10.590 align:middle line:84%
is equal to two and B occurs
divided by the probability of

00:22:10.590 --> 00:22:16.020 align:middle line:84%
B. The probability that the
maximum is equal to two.

00:22:16.020 --> 00:22:19.470 align:middle line:84%
OK, what's the event that the
maximum is equal to two?

00:22:19.470 --> 00:22:20.340 align:middle line:90%
Let's draw it.

00:22:20.340 --> 00:22:22.300 align:middle line:84%
This is going to be
the blue event.

00:22:22.300 --> 00:22:25.950 align:middle line:84%
The maximum is equal to
two if we get any

00:22:25.950 --> 00:22:28.520 align:middle line:90%
of those blue points.

00:22:28.520 --> 00:22:32.310 align:middle line:84%
So the intersection of the two
events is the intersection of

00:22:32.310 --> 00:22:35.170 align:middle line:84%
the red event and
the blue event.

00:22:35.170 --> 00:22:37.770 align:middle line:84%
There's only one point in
their intersection.

00:22:37.770 --> 00:22:39.640 align:middle line:84%
So the probability of
that intersection

00:22:39.640 --> 00:22:41.080 align:middle line:90%
happening is 1/16.

00:22:41.080 --> 00:22:43.740 align:middle line:90%


00:22:43.740 --> 00:22:45.160 align:middle line:90%
That's the numerator.

00:22:45.160 --> 00:22:47.110 align:middle line:90%
How about the denominator?

00:22:47.110 --> 00:22:50.610 align:middle line:84%
The event B consists of five
elements, each one of which

00:22:50.610 --> 00:22:52.270 align:middle line:90%
had probability of 1/16.

00:22:52.270 --> 00:22:54.570 align:middle line:90%
So that's 5/16.

00:22:54.570 --> 00:22:58.340 align:middle line:90%
And so the answer is 1/5.

00:22:58.340 --> 00:23:02.830 align:middle line:84%
Could we have gotten this
answer in a faster way?

00:23:02.830 --> 00:23:04.190 align:middle line:90%
Yes.

00:23:04.190 --> 00:23:05.560 align:middle line:90%
Here's how it goes.

00:23:05.560 --> 00:23:09.060 align:middle line:84%
We're trying to find the
conditional probability that

00:23:09.060 --> 00:23:13.210 align:middle line:84%
we get this point, given
that B occurred.

00:23:13.210 --> 00:23:15.570 align:middle line:90%
B consist of five elements.

00:23:15.570 --> 00:23:18.250 align:middle line:84%
All of those five elements were
equally likely when we

00:23:18.250 --> 00:23:22.720 align:middle line:84%
started, so they remain equally
likely afterwards.

00:23:22.720 --> 00:23:25.180 align:middle line:84%
Because when we define
conditional probabilities, we

00:23:25.180 --> 00:23:28.110 align:middle line:84%
keep the same proportions
inside the set.

00:23:28.110 --> 00:23:31.940 align:middle line:84%
So the five red elements
were equally likely.

00:23:31.940 --> 00:23:35.050 align:middle line:84%
They remain equally likely
in the conditional world.

00:23:35.050 --> 00:23:39.080 align:middle line:84%
So conditional event B having
happened, each one of these

00:23:39.080 --> 00:23:41.580 align:middle line:84%
five elements has the
same probability.

00:23:41.580 --> 00:23:44.300 align:middle line:84%
So the probability that we
actually get this point is

00:23:44.300 --> 00:23:46.210 align:middle line:90%
going to be 1/5.

00:23:46.210 --> 00:23:48.280 align:middle line:90%
And so that's the shortcut.

00:23:48.280 --> 00:23:53.070 align:middle line:84%
More generally, whenever you
have a uniform distribution on

00:23:53.070 --> 00:23:56.470 align:middle line:84%
your initial sample space,
when you condition on an

00:23:56.470 --> 00:24:01.000 align:middle line:84%
event, your new distribution is
still going to be uniform,

00:24:01.000 --> 00:24:05.010 align:middle line:84%
but on the smaller events
of that we considered.

00:24:05.010 --> 00:24:09.780 align:middle line:84%
So we started with a uniform
distribution on the big square

00:24:09.780 --> 00:24:13.730 align:middle line:84%
and we ended up with a
uniform distribution

00:24:13.730 --> 00:24:17.230 align:middle line:90%
just on the red point.

00:24:17.230 --> 00:24:19.850 align:middle line:84%
Now besides silly problems,
however, conditional

00:24:19.850 --> 00:24:25.070 align:middle line:84%
probabilities show up in real
and interesting situations.

00:24:25.070 --> 00:24:27.390 align:middle line:84%
And this example is going
to give you some

00:24:27.390 --> 00:24:30.430 align:middle line:90%
idea of how that happens.

00:24:30.430 --> 00:24:32.250 align:middle line:90%
OK.

00:24:32.250 --> 00:24:35.450 align:middle line:84%
Actually, in this example,
instead of starting with a

00:24:35.450 --> 00:24:39.480 align:middle line:84%
probability model in terms of
regular probabilities, I'm

00:24:39.480 --> 00:24:43.070 align:middle line:84%
actually going to define the
model in terms of conditional

00:24:43.070 --> 00:24:43.890 align:middle line:90%
probabilities.

00:24:43.890 --> 00:24:45.880 align:middle line:84%
And we'll see how
this is done.

00:24:45.880 --> 00:24:48.330 align:middle line:90%
So here's the story.

00:24:48.330 --> 00:24:52.210 align:middle line:84%
There may be an airplane flying
up in the sky, in a

00:24:52.210 --> 00:24:55.400 align:middle line:84%
particular sector of the sky
that you're watching.

00:24:55.400 --> 00:24:57.950 align:middle line:84%
Sometimes there is one sometimes
there isn't.

00:24:57.950 --> 00:25:01.760 align:middle line:84%
And from experience you know
that when you look up, there's

00:25:01.760 --> 00:25:04.400 align:middle line:84%
five percent probability that
the plane is flying above

00:25:04.400 --> 00:25:09.670 align:middle line:84%
there and 95% probability that
there's no plane up there.

00:25:09.670 --> 00:25:14.930 align:middle line:84%
So event A is the event that the
plane is flying out there.

00:25:14.930 --> 00:25:19.140 align:middle line:84%
Now you bought this wonderful
radar that's looks up.

00:25:19.140 --> 00:25:23.300 align:middle line:84%
And you're told in the
manufacturer's specs that, if

00:25:23.300 --> 00:25:27.310 align:middle line:84%
there is a plane out there,
your radar is going to

00:25:27.310 --> 00:25:30.090 align:middle line:84%
register something, a
blip on the screen

00:25:30.090 --> 00:25:32.940 align:middle line:90%
with probability 99%.

00:25:32.940 --> 00:25:35.540 align:middle line:84%
And it will not register
anything with

00:25:35.540 --> 00:25:37.500 align:middle line:90%
probability one percent.

00:25:37.500 --> 00:25:43.890 align:middle line:84%
So this particular part of the
picture is a self-contained

00:25:43.890 --> 00:25:50.280 align:middle line:84%
probability model of what your
radar does in a world where a

00:25:50.280 --> 00:25:52.530 align:middle line:90%
plane is out there.

00:25:52.530 --> 00:25:55.380 align:middle line:84%
So I'm telling you that the
plane is out there.

00:25:55.380 --> 00:25:58.240 align:middle line:84%
So we're now dealing with
conditional probabilities

00:25:58.240 --> 00:26:00.920 align:middle line:84%
because I gave you some
particular information.

00:26:00.920 --> 00:26:04.120 align:middle line:84%
Given this information that the
plane is out there, that's

00:26:04.120 --> 00:26:07.770 align:middle line:84%
how your radar is going to
behave with probability 99% is

00:26:07.770 --> 00:26:10.320 align:middle line:84%
going to detect it, with
probability one percent is

00:26:10.320 --> 00:26:11.620 align:middle line:90%
going to miss it.

00:26:11.620 --> 00:26:14.100 align:middle line:84%
So this piece of the picture
is a self-contained

00:26:14.100 --> 00:26:15.060 align:middle line:90%
probability model.

00:26:15.060 --> 00:26:17.130 align:middle line:84%
The probabilities
add up to one.

00:26:17.130 --> 00:26:20.300 align:middle line:84%
But it's a piece of
a larger model.

00:26:20.300 --> 00:26:22.820 align:middle line:84%
Similarly, there's the
other possibility.

00:26:22.820 --> 00:26:27.980 align:middle line:84%
Maybe a plane is not up there
and the manufacturer specs

00:26:27.980 --> 00:26:32.630 align:middle line:84%
tell you something about
false alarms.

00:26:32.630 --> 00:26:37.490 align:middle line:84%
A false alarm is the situation
where the plane is not there,

00:26:37.490 --> 00:26:41.190 align:middle line:84%
but for some reason your radar
picked up some noise or

00:26:41.190 --> 00:26:43.700 align:middle line:84%
whatever and shows a
blip on the screen.

00:26:43.700 --> 00:26:46.790 align:middle line:84%
And suppose that this happens
with probability ten percent.

00:26:46.790 --> 00:26:49.170 align:middle line:84%
Whereas with probability
90% your radar

00:26:49.170 --> 00:26:51.220 align:middle line:90%
gives the correct answer.

00:26:51.220 --> 00:26:55.430 align:middle line:84%
So this is sort of a model of
what's going to happen with

00:26:55.430 --> 00:26:59.430 align:middle line:84%
respect to both the plane --
we're given probabilities

00:26:59.430 --> 00:27:02.000 align:middle line:84%
about this -- and we're given
probabilities about how the

00:27:02.000 --> 00:27:04.120 align:middle line:90%
radar behaves.

00:27:04.120 --> 00:27:07.740 align:middle line:84%
So here I have indirectly
specified the probability law

00:27:07.740 --> 00:27:10.810 align:middle line:84%
in our model by starting with
conditional probabilities as

00:27:10.810 --> 00:27:13.670 align:middle line:84%
opposed to starting with
ordinary probabilities.

00:27:13.670 --> 00:27:17.160 align:middle line:84%
Can we derive ordinary
probabilities starting from

00:27:17.160 --> 00:27:18.740 align:middle line:90%
the conditional number ones?

00:27:18.740 --> 00:27:20.340 align:middle line:90%
Yeah, we certainly can.

00:27:20.340 --> 00:27:25.810 align:middle line:84%
Let's look at this event, A
intersection B, which is the

00:27:25.810 --> 00:27:31.160 align:middle line:84%
event up here, that there
is a plane and our

00:27:31.160 --> 00:27:33.750 align:middle line:90%
radar picks it up.

00:27:33.750 --> 00:27:35.760 align:middle line:84%
How can we calculate
this probability?

00:27:35.760 --> 00:27:38.600 align:middle line:84%
Well we use the definition of
conditional probabilities and

00:27:38.600 --> 00:27:41.430 align:middle line:84%
this is the probability of
A times the conditional

00:27:41.430 --> 00:27:50.260 align:middle line:84%
probability of B given A.
So it's 0.05 times 0.99.

00:27:50.260 --> 00:27:53.290 align:middle line:84%
And the answer, in
case you care--

00:27:53.290 --> 00:27:56.730 align:middle line:90%
It's 0.0495.

00:27:56.730 --> 00:27:57.650 align:middle line:90%
OK.

00:27:57.650 --> 00:28:01.370 align:middle line:84%
So we can calculate the
probabilities of final

00:28:01.370 --> 00:28:05.120 align:middle line:84%
outcomes, which are the leaves
of the tree, by using the

00:28:05.120 --> 00:28:07.250 align:middle line:84%
probabilities that
we have along the

00:28:07.250 --> 00:28:09.000 align:middle line:90%
branches of the tree.

00:28:09.000 --> 00:28:11.950 align:middle line:84%
So essentially, what we ended
up doing was to multiply the

00:28:11.950 --> 00:28:13.700 align:middle line:84%
probability of this
branch times the

00:28:13.700 --> 00:28:17.220 align:middle line:90%
probability of that branch.

00:28:17.220 --> 00:28:20.690 align:middle line:84%
Now, how about the answer
to this question.

00:28:20.690 --> 00:28:25.350 align:middle line:84%
What is the probability
that our radar is

00:28:25.350 --> 00:28:28.660 align:middle line:90%
going to register something?

00:28:28.660 --> 00:28:32.800 align:middle line:84%
OK, this is an event that can
happen in multiple ways.

00:28:32.800 --> 00:28:38.020 align:middle line:84%
It's the event that consists
of this outcome.

00:28:38.020 --> 00:28:41.640 align:middle line:84%
There is a plane and the radar
registers something together

00:28:41.640 --> 00:28:46.440 align:middle line:84%
with this outcome, there is no
plane but the radar still

00:28:46.440 --> 00:28:48.470 align:middle line:90%
registers something.

00:28:48.470 --> 00:28:52.650 align:middle line:84%
So to find the probability of
this event, we need the

00:28:52.650 --> 00:28:56.940 align:middle line:84%
individual probabilities
of the two outcomes.

00:28:56.940 --> 00:29:00.780 align:middle line:84%
For the first outcome, we
already calculated it.

00:29:00.780 --> 00:29:03.870 align:middle line:84%
For the second outcome, the
probability that this happens

00:29:03.870 --> 00:29:08.480 align:middle line:84%
is going to be this probability
95% times 0.10,

00:29:08.480 --> 00:29:11.280 align:middle line:84%
which is the conditional
probability for taking this

00:29:11.280 --> 00:29:15.070 align:middle line:84%
branch, given that there
was no plane out there.

00:29:15.070 --> 00:29:18.080 align:middle line:90%
So we just add the numbers.

00:29:18.080 --> 00:29:26.950 align:middle line:84%
0.05 times 0.99 plus 0.95
times 0.1 and the

00:29:26.950 --> 00:29:31.720 align:middle line:90%
final answer is 0.1445.

00:29:31.720 --> 00:29:32.410 align:middle line:90%
OK.

00:29:32.410 --> 00:29:35.730 align:middle line:84%
And now here's the interesting
question.

00:29:35.730 --> 00:29:41.480 align:middle line:84%
Given that your radar recorded
something, how likely is it

00:29:41.480 --> 00:29:45.070 align:middle line:84%
that there is an airplane
up there?

00:29:45.070 --> 00:29:46.810 align:middle line:84%
Your radar registering
something --

00:29:46.810 --> 00:29:48.730 align:middle line:84%
that can be caused
by two things.

00:29:48.730 --> 00:29:52.390 align:middle line:84%
Either there's a plane there,
and your radar did its job.

00:29:52.390 --> 00:29:57.400 align:middle line:84%
Or there was nothing, but your
radar fired a false alarm.

00:29:57.400 --> 00:30:01.690 align:middle line:84%
What's the probability that this
is the case as opposed to

00:30:01.690 --> 00:30:05.370 align:middle line:90%
that being the case?

00:30:05.370 --> 00:30:06.460 align:middle line:90%
OK.

00:30:06.460 --> 00:30:10.510 align:middle line:84%
The intuitive shortcut would
be that it should be the

00:30:10.510 --> 00:30:12.930 align:middle line:90%
probability--

00:30:12.930 --> 00:30:15.820 align:middle line:84%
you look at their relative odds
of these two elements and

00:30:15.820 --> 00:30:19.570 align:middle line:84%
you use them to find out how
much more likely it is to be

00:30:19.570 --> 00:30:21.730 align:middle line:84%
there as opposed
to being there.

00:30:21.730 --> 00:30:24.240 align:middle line:84%
But instead of doing this,
let's just write down the

00:30:24.240 --> 00:30:26.570 align:middle line:90%
definition and just use it.

00:30:26.570 --> 00:30:30.480 align:middle line:84%
It's the probability of A and
B happening, divided by the

00:30:30.480 --> 00:30:34.250 align:middle line:84%
probability of B. This is just
our definition of conditional

00:30:34.250 --> 00:30:35.540 align:middle line:90%
probabilities.

00:30:35.540 --> 00:30:39.300 align:middle line:84%
Now we have already found
the numerator.

00:30:39.300 --> 00:30:42.450 align:middle line:84%
We have already calculated
the denominator.

00:30:42.450 --> 00:30:46.440 align:middle line:84%
So we take the ratio of these
two numbers and we find the

00:30:46.440 --> 00:30:47.650 align:middle line:90%
final answer --

00:30:47.650 --> 00:30:54.490 align:middle line:90%
which is 0.34.

00:30:54.490 --> 00:30:55.980 align:middle line:90%
OK.

00:30:55.980 --> 00:30:59.040 align:middle line:84%
There's this slightly
curious thing that's

00:30:59.040 --> 00:31:02.270 align:middle line:90%
happened in this example.

00:31:02.270 --> 00:31:08.380 align:middle line:84%
Doesn't this number feel
a little too low?

00:31:08.380 --> 00:31:10.700 align:middle line:90%
My radar --

00:31:10.700 --> 00:31:13.820 align:middle line:84%
So this is a conditional
probability, given that my

00:31:13.820 --> 00:31:17.110 align:middle line:84%
radar said there is something
out there, that there is

00:31:17.110 --> 00:31:19.200 align:middle line:90%
indeed something there.

00:31:19.200 --> 00:31:21.960 align:middle line:84%
So it's sort of the probability
that our radar

00:31:21.960 --> 00:31:24.560 align:middle line:90%
gave the correct answer.

00:31:24.560 --> 00:31:28.580 align:middle line:84%
Now, the specs of our radar
we're pretty good.

00:31:28.580 --> 00:31:31.460 align:middle line:84%
In this situation, it gives
you the correct

00:31:31.460 --> 00:31:34.160 align:middle line:90%
answer 99% of the time.

00:31:34.160 --> 00:31:36.020 align:middle line:84%
In this situation, it gives
you the correct

00:31:36.020 --> 00:31:38.400 align:middle line:90%
answer 90% of the time.

00:31:38.400 --> 00:31:39.730 align:middle line:84%
So you would think
that your radar

00:31:39.730 --> 00:31:41.870 align:middle line:90%
there is really reliable.

00:31:41.870 --> 00:31:47.730 align:middle line:84%
But yet here the radar recorded
something, but the

00:31:47.730 --> 00:31:51.900 align:middle line:84%
chance that the answer that
you get out of this is the

00:31:51.900 --> 00:31:55.180 align:middle line:84%
right one, given that it
recorded something, the chance

00:31:55.180 --> 00:31:58.970 align:middle line:84%
that there is an airplane
out there is only 30%.

00:31:58.970 --> 00:32:01.980 align:middle line:84%
So you cannot really rely on
the measurements from your

00:32:01.980 --> 00:32:06.650 align:middle line:84%
radar, even though the specs of
the radar were really good.

00:32:06.650 --> 00:32:08.620 align:middle line:90%
What's the reason for this?

00:32:08.620 --> 00:32:17.730 align:middle line:84%
Well, the reason is that false
alarms are pretty common.

00:32:17.730 --> 00:32:20.110 align:middle line:84%
Most of the time there's
nothing.

00:32:20.110 --> 00:32:23.750 align:middle line:84%
And there's a ten percent
probability of false alarms.

00:32:23.750 --> 00:32:26.640 align:middle line:84%
So there's roughly a ten percent
probability that in

00:32:26.640 --> 00:32:29.730 align:middle line:84%
any given experiment, you
have a false alarm.

00:32:29.730 --> 00:32:33.450 align:middle line:84%
And there is about the five
percent probability that

00:32:33.450 --> 00:32:37.090 align:middle line:84%
something out there and
your radar gets it.

00:32:37.090 --> 00:32:41.350 align:middle line:84%
So when your radar records
something, it's actually more

00:32:41.350 --> 00:32:44.980 align:middle line:84%
likely to be a false
alarm rather than

00:32:44.980 --> 00:32:46.860 align:middle line:90%
being an actual airplane.

00:32:46.860 --> 00:32:49.100 align:middle line:84%
This has probability ten
percent roughly.

00:32:49.100 --> 00:32:52.000 align:middle line:84%
This has probability roughly
five percent

00:32:52.000 --> 00:32:55.130 align:middle line:84%
So conditional probabilities
are sometimes

00:32:55.130 --> 00:32:58.250 align:middle line:84%
counter-intuitive in terms of
the answers that they get.

00:32:58.250 --> 00:33:01.210 align:middle line:84%
And you can make similar
stories about doctors

00:33:01.210 --> 00:33:04.370 align:middle line:84%
interpreting the results
of tests.

00:33:04.370 --> 00:33:07.560 align:middle line:84%
So you tested positive for
a certain disease.

00:33:07.560 --> 00:33:11.260 align:middle line:84%
Does it mean that you have
the disease necessarily?

00:33:11.260 --> 00:33:14.590 align:middle line:84%
Well if that disease has been
eradicated from the face of

00:33:14.590 --> 00:33:17.900 align:middle line:84%
the earth, testing positive
doesn't mean that you have the

00:33:17.900 --> 00:33:21.740 align:middle line:84%
disease, even if the test
was designed to be

00:33:21.740 --> 00:33:23.320 align:middle line:90%
a pretty good one.

00:33:23.320 --> 00:33:28.190 align:middle line:84%
So unfortunately, doctors do get
it wrong also sometimes.

00:33:28.190 --> 00:33:29.990 align:middle line:84%
And the reasoning that
comes in such

00:33:29.990 --> 00:33:32.290 align:middle line:90%
situations is pretty subtle.

00:33:32.290 --> 00:33:34.890 align:middle line:84%
Now for the rest of the lecture,
what we're going to

00:33:34.890 --> 00:33:40.710 align:middle line:84%
do is to take this example where
we did three things and

00:33:40.710 --> 00:33:41.880 align:middle line:90%
abstract them.

00:33:41.880 --> 00:33:44.540 align:middle line:84%
These three trivial calculations
that's we just

00:33:44.540 --> 00:33:50.190 align:middle line:84%
did are three very important,
very basic tools that you use

00:33:50.190 --> 00:33:53.350 align:middle line:84%
to solve more general
probability problems.

00:33:53.350 --> 00:33:55.040 align:middle line:90%
So what's the first one?

00:33:55.040 --> 00:33:58.040 align:middle line:84%
We find the probability of a
composite event, two things

00:33:58.040 --> 00:34:01.300 align:middle line:84%
happening, by multiplying
probabilities and conditional

00:34:01.300 --> 00:34:03.130 align:middle line:90%
probabilities.

00:34:03.130 --> 00:34:08.639 align:middle line:84%
More general version of this,
look at any situation, maybe

00:34:08.639 --> 00:34:10.860 align:middle line:84%
involving lots and
lots of events.

00:34:10.860 --> 00:34:15.510 align:middle line:84%
So here's a story that event A
may happen or may not happen.

00:34:15.510 --> 00:34:19.440 align:middle line:84%
Given that A occurred, it's
possible that B happens or

00:34:19.440 --> 00:34:21.360 align:middle line:90%
that B does not happen.

00:34:21.360 --> 00:34:25.280 align:middle line:84%
Given that B also happens, it's
possible that the event C

00:34:25.280 --> 00:34:29.770 align:middle line:84%
also happens or that event
C does not happen.

00:34:29.770 --> 00:34:33.400 align:middle line:84%
And somebody specifies for you
a model by giving you all

00:34:33.400 --> 00:34:36.230 align:middle line:84%
these conditional probabilities
along the way.

00:34:36.230 --> 00:34:39.570 align:middle line:84%
Notice what we move along
the branches as the tree

00:34:39.570 --> 00:34:40.690 align:middle line:90%
progresses.

00:34:40.690 --> 00:34:45.110 align:middle line:84%
Any point in the tree
corresponds to certain events

00:34:45.110 --> 00:34:47.050 align:middle line:90%
having happened.

00:34:47.050 --> 00:34:50.980 align:middle line:84%
And then, given that this
has happened, we specify

00:34:50.980 --> 00:34:52.360 align:middle line:90%
conditional probabilities.

00:34:52.360 --> 00:34:55.989 align:middle line:84%
Given that this has happened,
how likely is it for that C

00:34:55.989 --> 00:34:57.900 align:middle line:90%
also occurs?

00:34:57.900 --> 00:35:00.890 align:middle line:84%
Given a model of this kind, how
do we find the probability

00:35:00.890 --> 00:35:02.660 align:middle line:90%
or for this event?

00:35:02.660 --> 00:35:05.310 align:middle line:84%
The answer is extremely
simple.

00:35:05.310 --> 00:35:09.930 align:middle line:84%
All that you do is move along
with the tree and multiply

00:35:09.930 --> 00:35:12.950 align:middle line:84%
conditional probabilities
along the way.

00:35:12.950 --> 00:35:16.900 align:middle line:84%
So in terms of frequencies, how
often do all three things

00:35:16.900 --> 00:35:19.310 align:middle line:90%
happen, A, B, and C?

00:35:19.310 --> 00:35:22.450 align:middle line:84%
You first see how often
does A occur.

00:35:22.450 --> 00:35:24.860 align:middle line:84%
Out of the times that
A occurs, how

00:35:24.860 --> 00:35:26.710 align:middle line:90%
often does B occur?

00:35:26.710 --> 00:35:29.630 align:middle line:84%
And out of the times where both
A and B have occurred,

00:35:29.630 --> 00:35:31.660 align:middle line:90%
how often does C occur?

00:35:31.660 --> 00:35:34.390 align:middle line:84%
And you can just multiply those
three frequencies with

00:35:34.390 --> 00:35:36.440 align:middle line:90%
each other.

00:35:36.440 --> 00:35:39.740 align:middle line:84%
What is the formal
proof of this?

00:35:39.740 --> 00:35:43.000 align:middle line:84%
Well, the only thing we have in
our hands is the definition

00:35:43.000 --> 00:35:44.890 align:middle line:90%
of conditional probabilities.

00:35:44.890 --> 00:35:49.660 align:middle line:90%
So let's just use this.

00:35:49.660 --> 00:35:50.910 align:middle line:90%
And--

00:35:50.910 --> 00:35:54.370 align:middle line:90%


00:35:54.370 --> 00:35:55.000 align:middle line:90%
OK.

00:35:55.000 --> 00:35:58.210 align:middle line:84%
Now, the definition of
conditional probabilities

00:35:58.210 --> 00:36:00.770 align:middle line:84%
tells us that the probability
of two things is the

00:36:00.770 --> 00:36:03.660 align:middle line:84%
probability of one of them
times a conditional

00:36:03.660 --> 00:36:04.620 align:middle line:90%
probability.

00:36:04.620 --> 00:36:05.850 align:middle line:90%
Unfortunately, here we have the

00:36:05.850 --> 00:36:07.310 align:middle line:90%
probability of three things.

00:36:07.310 --> 00:36:09.000 align:middle line:90%
What can I do?

00:36:09.000 --> 00:36:13.570 align:middle line:84%
I can put a parenthesis in here
and think of this as the

00:36:13.570 --> 00:36:18.640 align:middle line:84%
probability of this and that
and apply our definition of

00:36:18.640 --> 00:36:20.300 align:middle line:84%
conditional probabilities
here.

00:36:20.300 --> 00:36:23.920 align:middle line:84%
The probability of two things
happening is the probability

00:36:23.920 --> 00:36:28.430 align:middle line:84%
that the first happens times
the conditional probability

00:36:28.430 --> 00:36:34.070 align:middle line:84%
that the second happens, given
A and B, given that the first

00:36:34.070 --> 00:36:35.330 align:middle line:90%
one happened.

00:36:35.330 --> 00:36:38.850 align:middle line:84%
So this is just the definition
of the conditional probability

00:36:38.850 --> 00:36:41.980 align:middle line:84%
of an event, given
another event.

00:36:41.980 --> 00:36:44.270 align:middle line:84%
That other event is a
composite one, but

00:36:44.270 --> 00:36:45.330 align:middle line:90%
that's not an issue.

00:36:45.330 --> 00:36:47.300 align:middle line:90%
It's just an event.

00:36:47.300 --> 00:36:50.040 align:middle line:84%
And then we use the definition
of conditional probabilities

00:36:50.040 --> 00:36:56.290 align:middle line:84%
once more to break this apart
and make it P(A), P(B given A)

00:36:56.290 --> 00:36:58.260 align:middle line:84%
and then finally,
the last term.

00:36:58.260 --> 00:37:00.930 align:middle line:90%


00:37:00.930 --> 00:37:01.270 align:middle line:90%
OK.

00:37:01.270 --> 00:37:03.680 align:middle line:84%
So this proves the formula
that I have up

00:37:03.680 --> 00:37:05.290 align:middle line:90%
there on the slides.

00:37:05.290 --> 00:37:07.470 align:middle line:84%
And if you wish to calculate
any other

00:37:07.470 --> 00:37:09.330 align:middle line:90%
probability in this diagram.

00:37:09.330 --> 00:37:12.590 align:middle line:84%
For example, if you want to
calculate this probability,

00:37:12.590 --> 00:37:15.580 align:middle line:84%
you would still multiply the
conditional probabilities

00:37:15.580 --> 00:37:18.560 align:middle line:84%
along the different branches
of the tree.

00:37:18.560 --> 00:37:22.360 align:middle line:84%
In particular, here in this
branch, you would have the

00:37:22.360 --> 00:37:26.670 align:middle line:84%
conditional probability of
C complement, given A

00:37:26.670 --> 00:37:29.790 align:middle line:84%
intersection B complement,
and so on.

00:37:29.790 --> 00:37:32.070 align:middle line:84%
So you write down probabilities
along all those

00:37:32.070 --> 00:37:35.940 align:middle line:84%
tree branches and just multiply
them as you go.

00:37:35.940 --> 00:37:38.510 align:middle line:90%


00:37:38.510 --> 00:37:44.450 align:middle line:84%
So this was the first skill
that we are covering.

00:37:44.450 --> 00:37:46.690 align:middle line:90%
What was the second one?

00:37:46.690 --> 00:37:53.240 align:middle line:84%
What we did was to calculate
the total probability of a

00:37:53.240 --> 00:37:58.520 align:middle line:84%
certain event B that
consisted of--

00:37:58.520 --> 00:38:02.820 align:middle line:84%
was made up from different
possibilities, which

00:38:02.820 --> 00:38:05.580 align:middle line:84%
corresponded to different
scenarios.

00:38:05.580 --> 00:38:08.870 align:middle line:84%
So we wanted to calculate the
probability of this event B

00:38:08.870 --> 00:38:12.030 align:middle line:84%
that consisted of those
two elements.

00:38:12.030 --> 00:38:13.280 align:middle line:90%
Let's generalize.

00:38:13.280 --> 00:38:18.600 align:middle line:90%


00:38:18.600 --> 00:38:23.080 align:middle line:90%
So we have our big model.

00:38:23.080 --> 00:38:26.110 align:middle line:84%
And this sample space
is partitioned

00:38:26.110 --> 00:38:27.410 align:middle line:90%
in a number of sets.

00:38:27.410 --> 00:38:30.620 align:middle line:84%
In our radar example, we had
a partition in two sets.

00:38:30.620 --> 00:38:33.600 align:middle line:84%
Either a plane is there, or
a plane is not there.

00:38:33.600 --> 00:38:35.850 align:middle line:84%
Since we're trying to
generalize, now I'm going to

00:38:35.850 --> 00:38:39.410 align:middle line:84%
give you a picture for the case
of three possibilities or

00:38:39.410 --> 00:38:41.360 align:middle line:90%
three possible scenarios.

00:38:41.360 --> 00:38:45.160 align:middle line:84%
So whatever happens in the
world, there are three

00:38:45.160 --> 00:38:49.660 align:middle line:84%
possible scenarios,
A1, A2, A3.

00:38:49.660 --> 00:38:54.695 align:middle line:84%
So think of these as there's
nothing in the air, there's an

00:38:54.695 --> 00:38:58.190 align:middle line:84%
airplane in the air, or there's
a flock of geese

00:38:58.190 --> 00:38:59.490 align:middle line:90%
flying in the air.

00:38:59.490 --> 00:39:03.050 align:middle line:84%
So there's three possible
scenarios.

00:39:03.050 --> 00:39:08.972 align:middle line:84%
And then there's a certain event
B of interest, such as a

00:39:08.972 --> 00:39:12.800 align:middle line:84%
radar records something or
doesn't record something.

00:39:12.800 --> 00:39:15.870 align:middle line:90%
We specify this model by giving

00:39:15.870 --> 00:39:18.040 align:middle line:90%
probabilities for the Ai's--

00:39:18.040 --> 00:39:20.690 align:middle line:90%


00:39:20.690 --> 00:39:23.420 align:middle line:84%
That's the probability of
the different scenarios.

00:39:23.420 --> 00:39:27.180 align:middle line:84%
And somebody also gives us the
probabilities that this event

00:39:27.180 --> 00:39:31.010 align:middle line:84%
B is going to occur, given
that the Ai-th

00:39:31.010 --> 00:39:33.480 align:middle line:90%
scenario has occurred.

00:39:33.480 --> 00:39:36.230 align:middle line:84%
Think of the Ai's
as scenarios.

00:39:36.230 --> 00:39:39.130 align:middle line:90%


00:39:39.130 --> 00:39:43.110 align:middle line:84%
And we want to calculate the
overall probability of the

00:39:43.110 --> 00:39:47.210 align:middle line:84%
event B. What's happening
in this example?

00:39:47.210 --> 00:39:49.640 align:middle line:84%
Perhaps, instead of this
picture, it's easier to

00:39:49.640 --> 00:39:54.970 align:middle line:84%
visualize if I go back to the
picture I was using before.

00:39:54.970 --> 00:39:59.990 align:middle line:84%
We have three possible
scenarios, A1, A2, A3.

00:39:59.990 --> 00:40:05.150 align:middle line:84%
And under each scenario, B may
happen or B may not happen.

00:40:05.150 --> 00:40:11.360 align:middle line:90%


00:40:11.360 --> 00:40:12.250 align:middle line:90%
And so on.

00:40:12.250 --> 00:40:16.060 align:middle line:84%
So here we have A2 intersection
B. And here we

00:40:16.060 --> 00:40:22.110 align:middle line:84%
have A3 intersection B. In the
previous slide, we found how

00:40:22.110 --> 00:40:25.350 align:middle line:84%
to calculate the probability
of any event of this kind,

00:40:25.350 --> 00:40:28.870 align:middle line:84%
which is done by multiplying
probabilities here and

00:40:28.870 --> 00:40:31.100 align:middle line:84%
conditional probabilities
there.

00:40:31.100 --> 00:40:34.320 align:middle line:84%
Now we are asked to calculate
the total probability of the

00:40:34.320 --> 00:40:38.410 align:middle line:84%
event B. The event B can happen
in three possible ways.

00:40:38.410 --> 00:40:39.900 align:middle line:90%
It can happen here.

00:40:39.900 --> 00:40:41.700 align:middle line:90%
It can happen there.

00:40:41.700 --> 00:40:43.780 align:middle line:90%
And it can happen here.

00:40:43.780 --> 00:40:50.020 align:middle line:84%
So this is our event B. It
consists of three elements.

00:40:50.020 --> 00:40:53.370 align:middle line:84%
To calculate the total
probability of our event B,

00:40:53.370 --> 00:40:56.730 align:middle line:84%
all we need to do is to add
these three probabilities.

00:40:56.730 --> 00:40:59.440 align:middle line:90%


00:40:59.440 --> 00:41:03.510 align:middle line:84%
So B is an event that consists
of these three elements.

00:41:03.510 --> 00:41:06.450 align:middle line:84%
There are three ways
that B can happen.

00:41:06.450 --> 00:41:10.390 align:middle line:84%
Either B happens together with
A1, or B happens together with

00:41:10.390 --> 00:41:13.030 align:middle line:84%
A2, or B happens together
with A3.

00:41:13.030 --> 00:41:15.340 align:middle line:84%
So we need to add the
probabilities of these three

00:41:15.340 --> 00:41:16.630 align:middle line:90%
contingencies.

00:41:16.630 --> 00:41:18.980 align:middle line:84%
For each one of those
contingencies, we can

00:41:18.980 --> 00:41:23.020 align:middle line:84%
calculate its probability by
using the multiplication rule.

00:41:23.020 --> 00:41:27.580 align:middle line:84%
So the probability of A1 and
B happening is this--

00:41:27.580 --> 00:41:30.030 align:middle line:84%
It's the probability of A1
and then B happening

00:41:30.030 --> 00:41:32.020 align:middle line:90%
given that A1 happens.

00:41:32.020 --> 00:41:36.140 align:middle line:84%
The probability of this
contingency is found by taking

00:41:36.140 --> 00:41:39.470 align:middle line:84%
the probability that A2 happens
times the conditional

00:41:39.470 --> 00:41:42.350 align:middle line:84%
probability of A2, given
that B happened.

00:41:42.350 --> 00:41:44.640 align:middle line:84%
And similarly for
the third one.

00:41:44.640 --> 00:41:48.030 align:middle line:84%
So this is the general rule
that we have here.

00:41:48.030 --> 00:41:50.830 align:middle line:84%
The rule is written for the
case of three scenarios.

00:41:50.830 --> 00:41:54.020 align:middle line:84%
But obviously, it has a
generalization for the case of

00:41:54.020 --> 00:41:57.440 align:middle line:84%
four or five or more
scenarios.

00:41:57.440 --> 00:42:02.050 align:middle line:84%
It gives you a way of breaking
up the calculation of an event

00:42:02.050 --> 00:42:06.740 align:middle line:84%
that can happen in multiple ways
by considering individual

00:42:06.740 --> 00:42:09.720 align:middle line:84%
probabilities for the different
ways that the event

00:42:09.720 --> 00:42:10.970 align:middle line:90%
can happen.

00:42:10.970 --> 00:42:12.950 align:middle line:90%


00:42:12.950 --> 00:42:14.640 align:middle line:90%
OK.

00:42:14.640 --> 00:42:16.300 align:middle line:90%
So--

00:42:16.300 --> 00:42:16.656 align:middle line:90%
Yes?

00:42:16.656 --> 00:42:18.180 align:middle line:84%
AUDIENCE: Does this
have to change for

00:42:18.180 --> 00:42:19.800 align:middle line:90%
infinite sample space?

00:42:19.800 --> 00:42:20.760 align:middle line:90%
JOHN TSISIKLIS: No.

00:42:20.760 --> 00:42:23.050 align:middle line:84%
This is true whether
your sample space

00:42:23.050 --> 00:42:25.450 align:middle line:90%
is infinite or finite.

00:42:25.450 --> 00:42:28.410 align:middle line:84%
What I'm using in this argument
that we have a

00:42:28.410 --> 00:42:33.670 align:middle line:84%
partition into just three
scenarios, three events.

00:42:33.670 --> 00:42:36.720 align:middle line:84%
So it's a partition to a finite
number of events.

00:42:36.720 --> 00:42:41.100 align:middle line:84%
It's also true if it's a
partition into an infinite

00:42:41.100 --> 00:42:43.670 align:middle line:90%
sequence of events.

00:42:43.670 --> 00:42:47.550 align:middle line:84%
But that's, I think, one of the
theoretical problems at

00:42:47.550 --> 00:42:49.430 align:middle line:90%
the end of the chapter.

00:42:49.430 --> 00:42:54.350 align:middle line:84%
You probably may not
need it for now.

00:42:54.350 --> 00:42:57.550 align:middle line:84%
OK, going back to
the story here.

00:42:57.550 --> 00:43:00.410 align:middle line:84%
There are three possible
scenarios about what could

00:43:00.410 --> 00:43:03.390 align:middle line:84%
happen in the world that
are captured here.

00:43:03.390 --> 00:43:08.660 align:middle line:84%
Event, under each scenario,
event B may or may not happen.

00:43:08.660 --> 00:43:11.850 align:middle line:84%
And so these probabilities tell
us the likelihoods of the

00:43:11.850 --> 00:43:13.270 align:middle line:90%
different scenarios.

00:43:13.270 --> 00:43:17.640 align:middle line:84%
These conditional probabilities
tell us how

00:43:17.640 --> 00:43:21.030 align:middle line:84%
likely is it for B to happen
under one scenario, or the

00:43:21.030 --> 00:43:23.760 align:middle line:84%
other scenario, or the
other scenario.

00:43:23.760 --> 00:43:28.510 align:middle line:84%
The overall probability of
B is found by taking some

00:43:28.510 --> 00:43:32.380 align:middle line:84%
combination of the probabilities
of B in the

00:43:32.380 --> 00:43:34.250 align:middle line:84%
different possible
worlds, in the

00:43:34.250 --> 00:43:36.230 align:middle line:90%
different possible scenarios.

00:43:36.230 --> 00:43:38.690 align:middle line:84%
Under some scenario, B
may be very likely.

00:43:38.690 --> 00:43:42.280 align:middle line:84%
Under another scenario, it
may be very unlikely.

00:43:42.280 --> 00:43:45.740 align:middle line:84%
We take all of these into
account and weigh them

00:43:45.740 --> 00:43:48.590 align:middle line:84%
according to the likelihood
of the scenarios.

00:43:48.590 --> 00:43:53.040 align:middle line:84%
Now notice that since A1, A2,
and three form a partition,

00:43:53.040 --> 00:43:58.530 align:middle line:84%
these three probabilities
have what property?

00:43:58.530 --> 00:44:00.810 align:middle line:90%
Add to what?

00:44:00.810 --> 00:44:03.640 align:middle line:90%
They add to one.

00:44:03.640 --> 00:44:06.020 align:middle line:84%
So it's the probability of this
branch, plus this branch,

00:44:06.020 --> 00:44:07.240 align:middle line:90%
plus this branch.

00:44:07.240 --> 00:44:11.660 align:middle line:84%
So what we have here is a
weighted average of the

00:44:11.660 --> 00:44:15.120 align:middle line:84%
probabilities of the B's into
the different worlds, or in

00:44:15.120 --> 00:44:16.690 align:middle line:90%
the different scenarios.

00:44:16.690 --> 00:44:17.860 align:middle line:90%
Special case.

00:44:17.860 --> 00:44:20.370 align:middle line:84%
Suppose the three scenarios
are equally likely.

00:44:20.370 --> 00:44:25.300 align:middle line:84%
So P of A1 equals 1/3, equals
to P of A2, P of A3.

00:44:25.300 --> 00:44:27.320 align:middle line:90%
what are we saying here?

00:44:27.320 --> 00:44:31.750 align:middle line:84%
In that case of equally likely
scenarios, the probability of

00:44:31.750 --> 00:44:35.920 align:middle line:84%
B is the average of the
probabilities of B in the

00:44:35.920 --> 00:44:38.835 align:middle line:84%
three different words, or in the
three different scenarios.

00:44:38.835 --> 00:44:42.950 align:middle line:90%


00:44:42.950 --> 00:44:43.450 align:middle line:90%
OK.

00:44:43.450 --> 00:44:46.630 align:middle line:90%
So to finally, the last step.

00:44:46.630 --> 00:44:53.800 align:middle line:84%
If we go back again two slides,
the last thing that we

00:44:53.800 --> 00:44:57.510 align:middle line:84%
did was to calculate a
conditional probability of

00:44:57.510 --> 00:45:01.760 align:middle line:84%
this kind, probability of
A given B, which is a

00:45:01.760 --> 00:45:04.080 align:middle line:84%
probability associated
essentially with

00:45:04.080 --> 00:45:05.630 align:middle line:90%
an inference problem.

00:45:05.630 --> 00:45:09.840 align:middle line:84%
Given that our radar recorded
something, how likely is it

00:45:09.840 --> 00:45:12.060 align:middle line:90%
that the plane was up there?

00:45:12.060 --> 00:45:15.240 align:middle line:84%
So we're trying to infer whether
a plane was up there

00:45:15.240 --> 00:45:18.610 align:middle line:84%
or not, based on the information
that we've got.

00:45:18.610 --> 00:45:20.770 align:middle line:90%
So let's generalize once more.

00:45:20.770 --> 00:45:24.560 align:middle line:90%


00:45:24.560 --> 00:45:28.250 align:middle line:84%
And we're just going to rewrite
what we did in that

00:45:28.250 --> 00:45:32.190 align:middle line:84%
example, but in terms of general
symbols instead of the

00:45:32.190 --> 00:45:33.650 align:middle line:90%
specific numbers.

00:45:33.650 --> 00:45:38.180 align:middle line:84%
So once more, the model that we
have involves probabilities

00:45:38.180 --> 00:45:40.480 align:middle line:90%
of the different scenarios.

00:45:40.480 --> 00:45:42.830 align:middle line:84%
These we call them prior
probabilities.

00:45:42.830 --> 00:45:46.690 align:middle line:84%
They're are our initial beliefs
about how likely each

00:45:46.690 --> 00:45:49.360 align:middle line:90%
scenario is to occur.

00:45:49.360 --> 00:45:54.500 align:middle line:84%
We also have a model of our
measuring device that tells us

00:45:54.500 --> 00:45:58.110 align:middle line:84%
under that scenario how likely
is it that our radar will

00:45:58.110 --> 00:46:00.140 align:middle line:90%
register something or not.

00:46:00.140 --> 00:46:03.220 align:middle line:84%
So we're given again these
conditional probabilities.

00:46:03.220 --> 00:46:04.330 align:middle line:90%
We're given the conditional

00:46:04.330 --> 00:46:06.950 align:middle line:90%
probabilities for these branches.

00:46:06.950 --> 00:46:11.050 align:middle line:84%
Then we are told that
event B occurred.

00:46:11.050 --> 00:46:15.330 align:middle line:84%
And on the basis of this new
information, we want to form

00:46:15.330 --> 00:46:18.510 align:middle line:84%
some new beliefs about the
relative likelihood of the

00:46:18.510 --> 00:46:20.110 align:middle line:90%
different scenarios.

00:46:20.110 --> 00:46:23.790 align:middle line:84%
Going back again to our radar
example, an airplane was

00:46:23.790 --> 00:46:26.340 align:middle line:90%
present with probability 5%.

00:46:26.340 --> 00:46:29.180 align:middle line:84%
Given that the radar recorded
something, we're going to

00:46:29.180 --> 00:46:30.540 align:middle line:90%
change our beliefs.

00:46:30.540 --> 00:46:34.870 align:middle line:84%
Now, a plane is present
with probability 34%.

00:46:34.870 --> 00:46:38.270 align:middle line:84%
The radar, since we saw
something, we are going to

00:46:38.270 --> 00:46:41.880 align:middle line:84%
revise our beliefs as to whether
the plane is out there

00:46:41.880 --> 00:46:43.130 align:middle line:90%
or is not there.

00:46:43.130 --> 00:46:46.040 align:middle line:90%


00:46:46.040 --> 00:46:52.660 align:middle line:84%
And so what we need to do is to
calculate the conditional

00:46:52.660 --> 00:46:57.290 align:middle line:84%
probabilities of the different
scenarios, given the

00:46:57.290 --> 00:46:59.340 align:middle line:90%
information that we got.

00:46:59.340 --> 00:47:02.330 align:middle line:84%
So initially, we have these
probabilities for the

00:47:02.330 --> 00:47:04.000 align:middle line:90%
different scenarios.

00:47:04.000 --> 00:47:06.870 align:middle line:84%
Once we get the information,
we update them and we

00:47:06.870 --> 00:47:09.760 align:middle line:84%
calculate our revised
probabilities or conditional

00:47:09.760 --> 00:47:14.130 align:middle line:84%
probabilities given the
observation that we made.

00:47:14.130 --> 00:47:14.730 align:middle line:90%
OK.

00:47:14.730 --> 00:47:15.760 align:middle line:90%
So what do we do?

00:47:15.760 --> 00:47:17.620 align:middle line:84%
We just use the definition
of conditional

00:47:17.620 --> 00:47:19.360 align:middle line:90%
probabilities twice.

00:47:19.360 --> 00:47:22.490 align:middle line:84%
By definition the conditional
probability is the probability

00:47:22.490 --> 00:47:25.740 align:middle line:84%
of two things happening divided
by the probability of

00:47:25.740 --> 00:47:27.960 align:middle line:90%
the conditioning event.

00:47:27.960 --> 00:47:30.480 align:middle line:84%
Now, I'm using the definition
of conditional probabilities

00:47:30.480 --> 00:47:33.550 align:middle line:84%
once more, or rather I use
the multiplication rule.

00:47:33.550 --> 00:47:35.970 align:middle line:84%
The probability of two things
happening is the probability

00:47:35.970 --> 00:47:38.740 align:middle line:90%
of the first and the second.

00:47:38.740 --> 00:47:41.190 align:middle line:84%
So these are things that
are given to us.

00:47:41.190 --> 00:47:43.430 align:middle line:84%
They're the probabilities of
the different scenarios.

00:47:43.430 --> 00:47:47.750 align:middle line:84%
And it's the model of our
measuring device, which we

00:47:47.750 --> 00:47:51.810 align:middle line:90%
assume to be available.

00:47:51.810 --> 00:47:53.450 align:middle line:90%
And how about the denominator?

00:47:53.450 --> 00:47:57.780 align:middle line:84%
This is total probability of the
event B. But we just found

00:47:57.780 --> 00:48:01.140 align:middle line:84%
that's it's easy to calculate
using the formula in the

00:48:01.140 --> 00:48:02.400 align:middle line:90%
previous slide.

00:48:02.400 --> 00:48:04.750 align:middle line:84%
To find the overall probability
of event B

00:48:04.750 --> 00:48:08.260 align:middle line:84%
occurring, we look at the
probabilities of B occurring

00:48:08.260 --> 00:48:11.560 align:middle line:84%
under the different scenario
and weigh them according to

00:48:11.560 --> 00:48:13.710 align:middle line:84%
the probabilities of
all the scenarios.

00:48:13.710 --> 00:48:17.370 align:middle line:84%
So in the end, we have a formula
for the conditional

00:48:17.370 --> 00:48:22.730 align:middle line:84%
probability, A's given B,
based on the data of the

00:48:22.730 --> 00:48:25.090 align:middle line:84%
problem, which were
probabilities of the different

00:48:25.090 --> 00:48:27.360 align:middle line:84%
scenarios and conditional
probabilities of

00:48:27.360 --> 00:48:29.490 align:middle line:90%
B, given the A's.

00:48:29.490 --> 00:48:33.320 align:middle line:84%
So what this calculation does
is, basically, it reverses the

00:48:33.320 --> 00:48:35.310 align:middle line:90%
order of conditioning.

00:48:35.310 --> 00:48:39.000 align:middle line:84%
We are given conditional
probabilities of these kind,

00:48:39.000 --> 00:48:42.950 align:middle line:84%
where it's B given A and we
produce new conditional

00:48:42.950 --> 00:48:46.630 align:middle line:84%
probabilities, where things
go the other way.

00:48:46.630 --> 00:48:53.530 align:middle line:84%
So schematically, what's
happening here is that we have

00:48:53.530 --> 00:48:59.995 align:middle line:84%
model of cause and
effect and--

00:48:59.995 --> 00:49:02.550 align:middle line:90%


00:49:02.550 --> 00:49:09.840 align:middle line:84%
So a scenario occurs and that
may cause B to happen or may

00:49:09.840 --> 00:49:11.880 align:middle line:90%
not cause it to happen.

00:49:11.880 --> 00:49:14.495 align:middle line:84%
So this is a cause/effect
model.

00:49:14.495 --> 00:49:17.300 align:middle line:90%


00:49:17.300 --> 00:49:20.090 align:middle line:84%
And it's modeled using
probabilities, such as

00:49:20.090 --> 00:49:23.350 align:middle line:90%
probability of B given Ai.

00:49:23.350 --> 00:49:28.710 align:middle line:84%
And what we want to do is
inference where we are told

00:49:28.710 --> 00:49:35.910 align:middle line:84%
that B occurs, and we want
to infer whether Ai

00:49:35.910 --> 00:49:38.580 align:middle line:90%
also occurred or not.

00:49:38.580 --> 00:49:42.050 align:middle line:84%
And the appropriate
probabilities for that are the

00:49:42.050 --> 00:49:45.010 align:middle line:84%
conditional probabilities
that A occurred,

00:49:45.010 --> 00:49:48.110 align:middle line:90%
given that B occurred.

00:49:48.110 --> 00:49:52.250 align:middle line:84%
So we're starting with a causal
model of our situation.

00:49:52.250 --> 00:49:57.220 align:middle line:84%
It models from a given cause how
likely is a certain effect

00:49:57.220 --> 00:49:58.830 align:middle line:90%
to be observed.

00:49:58.830 --> 00:50:02.920 align:middle line:84%
And then we do inference, which
answers the question,

00:50:02.920 --> 00:50:06.730 align:middle line:84%
given that the effect was
observed, how likely is it

00:50:06.730 --> 00:50:10.870 align:middle line:84%
that the world was in this
particular situation or state

00:50:10.870 --> 00:50:12.940 align:middle line:90%
or scenario.

00:50:12.940 --> 00:50:17.260 align:middle line:84%
So the name of the Bayes rule
comes from Thomas Bayes, a

00:50:17.260 --> 00:50:20.750 align:middle line:84%
British theologian back
in the 1700s.

00:50:20.750 --> 00:50:21.530 align:middle line:90%
It actually--

00:50:21.530 --> 00:50:25.000 align:middle line:84%
This calculation addresses
a basic problem, a basic

00:50:25.000 --> 00:50:30.230 align:middle line:84%
philosophical problem, how one
can learn from experience or

00:50:30.230 --> 00:50:33.300 align:middle line:84%
from experimental data and
some systematic way.

00:50:33.300 --> 00:50:35.840 align:middle line:84%
So the British at that time
were preoccupied with this

00:50:35.840 --> 00:50:36.710 align:middle line:90%
type of question.

00:50:36.710 --> 00:50:41.200 align:middle line:84%
Is there a basic theory that
about how we can incorporate

00:50:41.200 --> 00:50:44.280 align:middle line:84%
new knowledge to previous
knowledge.

00:50:44.280 --> 00:50:47.600 align:middle line:84%
And this calculation made an
argument that, yes, it is

00:50:47.600 --> 00:50:50.100 align:middle line:84%
possible to do that in
a systematic way.

00:50:50.100 --> 00:50:53.040 align:middle line:84%
So the philosophical
underpinnings of this have a

00:50:53.040 --> 00:50:57.050 align:middle line:84%
very long history and a lot
of discussion around them.

00:50:57.050 --> 00:51:00.560 align:middle line:84%
But for our purposes, it's just
an extremely useful tool.

00:51:00.560 --> 00:51:03.550 align:middle line:84%
And it's the foundation of
almost everything that gets

00:51:03.550 --> 00:51:07.190 align:middle line:84%
done when you try to do
inference based on partial

00:51:07.190 --> 00:51:08.860 align:middle line:90%
observations.

00:51:08.860 --> 00:51:09.690 align:middle line:90%
Very well.

00:51:09.690 --> 00:51:10.940 align:middle line:90%
Till next time.

00:51:10.940 --> 00:51:11.760 align:middle line:90%