WEBVTT

00:00:00.000 --> 00:00:00.040 align:middle line:90%


00:00:00.040 --> 00:00:02.460 align:middle line:84%
The following content is
provided under a Creative

00:00:02.460 --> 00:00:03.870 align:middle line:90%
Commons license.

00:00:03.870 --> 00:00:06.910 align:middle line:84%
Your support will help MIT
OpenCourseWare continue to

00:00:06.910 --> 00:00:10.560 align:middle line:84%
offer high quality educational
resources for free.

00:00:10.560 --> 00:00:13.460 align:middle line:84%
To make a donation or view
additional materials from

00:00:13.460 --> 00:00:19.290 align:middle line:84%
hundreds of MIT courses, visit
MIT OpenCourseWare at

00:00:19.290 --> 00:00:20.540 align:middle line:90%
ocw.mit.edu.

00:00:20.540 --> 00:00:23.750 align:middle line:90%


00:00:23.750 --> 00:00:26.970 align:middle line:90%
PROFESSOR: Let us start.

00:00:26.970 --> 00:00:30.080 align:middle line:84%
So as always, we're to have
a quick review of what we

00:00:30.080 --> 00:00:31.240 align:middle line:90%
discussed last time.

00:00:31.240 --> 00:00:34.240 align:middle line:84%
And then today we're going
to introduce just one new

00:00:34.240 --> 00:00:38.120 align:middle line:84%
concept, the notion of
independence of two events.

00:00:38.120 --> 00:00:41.030 align:middle line:84%
And we will play with
that concept.

00:00:41.030 --> 00:00:43.110 align:middle line:84%
So what did we talk
about last time?

00:00:43.110 --> 00:00:46.410 align:middle line:84%
The idea is that we have an
experiment, and the experiment

00:00:46.410 --> 00:00:48.800 align:middle line:90%
has a sample space omega.

00:00:48.800 --> 00:00:52.300 align:middle line:84%
And then somebody comes and
tells us you know the outcome

00:00:52.300 --> 00:00:56.840 align:middle line:84%
of the experiments happens to
lie inside this particular

00:00:56.840 --> 00:01:00.470 align:middle line:84%
event B. Given this information,
it kind of

00:01:00.470 --> 00:01:03.070 align:middle line:84%
changes what we know about
the situation.

00:01:03.070 --> 00:01:05.510 align:middle line:84%
It tells us that the outcome
is going to be somewhere

00:01:05.510 --> 00:01:06.630 align:middle line:90%
inside here.

00:01:06.630 --> 00:01:09.800 align:middle line:84%
So this is essentially
our new sample space.

00:01:09.800 --> 00:01:13.130 align:middle line:84%
And now we need to we reassign
probabilities to the various

00:01:13.130 --> 00:01:16.550 align:middle line:84%
possible outcomes, because, for
example, these outcomes,

00:01:16.550 --> 00:01:20.340 align:middle line:84%
even if they had positive
probability beforehand, now

00:01:20.340 --> 00:01:22.890 align:middle line:84%
that we're told that B occurred,
those outcomes out

00:01:22.890 --> 00:01:25.220 align:middle line:84%
there are going to have
zero probability.

00:01:25.220 --> 00:01:27.670 align:middle line:84%
So we need to revise
our probabilities.

00:01:27.670 --> 00:01:29.740 align:middle line:84%
The new probabilities are
called conditional

00:01:29.740 --> 00:01:33.390 align:middle line:84%
probabilities, and they're
defined this way.

00:01:33.390 --> 00:01:37.000 align:middle line:84%
The conditional probability that
A occurs given that we're

00:01:37.000 --> 00:01:40.670 align:middle line:84%
told that B occurred is
calculated by this formula,

00:01:40.670 --> 00:01:42.880 align:middle line:90%
which tells us the following--

00:01:42.880 --> 00:01:45.750 align:middle line:84%
out of the total probability
that was initially assigned to

00:01:45.750 --> 00:01:49.760 align:middle line:84%
the event B, what fraction of
that probability is assigned

00:01:49.760 --> 00:01:54.310 align:middle line:84%
to outcomes that also
make A to happen?

00:01:54.310 --> 00:01:58.650 align:middle line:84%
So out of the total probability
assigned to B, we

00:01:58.650 --> 00:02:03.080 align:middle line:84%
see what fraction of that total
probability is assigned

00:02:03.080 --> 00:02:06.650 align:middle line:84%
to those elements here that
will also make A happen.

00:02:06.650 --> 00:02:09.360 align:middle line:84%
Conditional probabilities are
left undefined if the

00:02:09.360 --> 00:02:12.860 align:middle line:90%
denominator here is zero.

00:02:12.860 --> 00:02:16.030 align:middle line:84%
An easy consequence of the
definition is if we bring that

00:02:16.030 --> 00:02:18.670 align:middle line:84%
term to the other side, then we
can find the probability of

00:02:18.670 --> 00:02:22.040 align:middle line:84%
two things happening by taking
the probability that the first

00:02:22.040 --> 00:02:25.170 align:middle line:84%
thing happens, and then, given
that the first thing happened,

00:02:25.170 --> 00:02:28.230 align:middle line:84%
the conditional probability that
the second one happens.

00:02:28.230 --> 00:02:31.680 align:middle line:84%
Then we saw last time that we
can divide and conquer in

00:02:31.680 --> 00:02:36.010 align:middle line:84%
calculating probabilities of
mildly complicated events by

00:02:36.010 --> 00:02:39.070 align:middle line:84%
breaking it down into
different scenarios.

00:02:39.070 --> 00:02:41.350 align:middle line:84%
So event B can happen
in two ways.

00:02:41.350 --> 00:02:44.810 align:middle line:84%
It can happen either together
with A, which is this

00:02:44.810 --> 00:02:48.700 align:middle line:84%
probability, or it can happen
together with A complement,

00:02:48.700 --> 00:02:49.880 align:middle line:90%
which is this probability.

00:02:49.880 --> 00:02:53.100 align:middle line:84%
So basically what we're saying
that the total probability of

00:02:53.100 --> 00:02:57.960 align:middle line:84%
B is the probability of this,
which is A intersection B,

00:02:57.960 --> 00:03:02.620 align:middle line:84%
plus the probability of that,
which is A complement

00:03:02.620 --> 00:03:07.530 align:middle line:90%
intersection B.

00:03:07.530 --> 00:03:11.355 align:middle line:84%
So these two facts here,
multiplication rule and the

00:03:11.355 --> 00:03:14.760 align:middle line:84%
total probability theorem, are
basic tools that one uses to

00:03:14.760 --> 00:03:16.990 align:middle line:84%
break down probability
calculations

00:03:16.990 --> 00:03:18.930 align:middle line:90%
into a simpler parts.

00:03:18.930 --> 00:03:21.600 align:middle line:84%
So we find probabilities of
two things happening by

00:03:21.600 --> 00:03:24.300 align:middle line:90%
looking at each one at a time.

00:03:24.300 --> 00:03:27.830 align:middle line:84%
And this is what we do to break
up a situation with two

00:03:27.830 --> 00:03:29.760 align:middle line:90%
different possible scenarios.

00:03:29.760 --> 00:03:32.110 align:middle line:84%
Then we also have
the Bayes rule,

00:03:32.110 --> 00:03:33.600 align:middle line:90%
which does the following.

00:03:33.600 --> 00:03:36.640 align:middle line:84%
Given a model that has
conditional probabilities of

00:03:36.640 --> 00:03:38.970 align:middle line:84%
this kind, the Bayes rule
allows us to calculate

00:03:38.970 --> 00:03:41.570 align:middle line:84%
conditional probabilities in
which the events appear in

00:03:41.570 --> 00:03:43.020 align:middle line:90%
different order.

00:03:43.020 --> 00:03:45.740 align:middle line:84%
You can think of these
probabilities as describing a

00:03:45.740 --> 00:03:49.270 align:middle line:84%
causal model of a certain
situation, whereas these are

00:03:49.270 --> 00:03:52.670 align:middle line:84%
the probabilities that you get
after you do some inference

00:03:52.670 --> 00:03:55.480 align:middle line:84%
based on the information that
you have available.

00:03:55.480 --> 00:03:59.200 align:middle line:84%
Now the Bayes rule, we derived
it, and it's a trivial

00:03:59.200 --> 00:04:01.040 align:middle line:90%
half-line calculation.

00:04:01.040 --> 00:04:03.670 align:middle line:84%
But it underlies lots
and lots of useful

00:04:03.670 --> 00:04:05.290 align:middle line:90%
things in the real world.

00:04:05.290 --> 00:04:07.650 align:middle line:84%
We had the radar example
last time.

00:04:07.650 --> 00:04:10.410 align:middle line:84%
You can think of more
complicated situations in

00:04:10.410 --> 00:04:14.570 align:middle line:84%
which there's a bunch or lots of
different hypotheses about

00:04:14.570 --> 00:04:15.920 align:middle line:90%
the environment.

00:04:15.920 --> 00:04:18.899 align:middle line:84%
Given any particular setting in
the environment, you have a

00:04:18.899 --> 00:04:21.140 align:middle line:84%
measuring device that
can produce

00:04:21.140 --> 00:04:23.670 align:middle line:90%
many different outcomes.

00:04:23.670 --> 00:04:29.210 align:middle line:84%
And you observe the final
outcome out of your measuring

00:04:29.210 --> 00:04:31.820 align:middle line:84%
device, and you're trying
to guess which

00:04:31.820 --> 00:04:34.210 align:middle line:90%
particular branch occurred.

00:04:34.210 --> 00:04:36.640 align:middle line:84%
That is, you're trying to guess
the state of the world

00:04:36.640 --> 00:04:38.500 align:middle line:84%
based on a particular
measurement.

00:04:38.500 --> 00:04:40.770 align:middle line:84%
That's what inference
is all about.

00:04:40.770 --> 00:04:44.450 align:middle line:84%
So real world problems only
differ from the simple example

00:04:44.450 --> 00:04:48.610 align:middle line:84%
that we saw last time in that
this kind of tree is a little

00:04:48.610 --> 00:04:50.040 align:middle line:90%
more complicated.

00:04:50.040 --> 00:04:52.150 align:middle line:84%
You might have infinitely
many possible

00:04:52.150 --> 00:04:54.450 align:middle line:90%
outcomes here and so on.

00:04:54.450 --> 00:04:57.960 align:middle line:84%
So setting up the model may be
more elaborate, but the basic

00:04:57.960 --> 00:05:01.170 align:middle line:84%
calculation that's done based
on the Bayes rule is

00:05:01.170 --> 00:05:04.430 align:middle line:84%
essentially the same as
the one that we saw.

00:05:04.430 --> 00:05:07.190 align:middle line:84%
Now something that we discuss
is that sometimes we use

00:05:07.190 --> 00:05:11.050 align:middle line:84%
conditional probabilities to
describe models, and let's do

00:05:11.050 --> 00:05:14.030 align:middle line:84%
this by looking at a
model where we toss

00:05:14.030 --> 00:05:16.150 align:middle line:90%
a coin three times.

00:05:16.150 --> 00:05:19.090 align:middle line:84%
And how do we use conditional
probabilities to

00:05:19.090 --> 00:05:20.630 align:middle line:90%
describe the situation?

00:05:20.630 --> 00:05:22.950 align:middle line:90%
So we have one experiment.

00:05:22.950 --> 00:05:26.590 align:middle line:84%
But that one experiment consists
of three consecutive

00:05:26.590 --> 00:05:27.880 align:middle line:90%
coin tosses.

00:05:27.880 --> 00:05:32.380 align:middle line:84%
So the possible outcomes, our
sample space, consists of

00:05:32.380 --> 00:05:37.070 align:middle line:84%
strings of length 3 that tell
us whether we had heads,

00:05:37.070 --> 00:05:39.200 align:middle line:90%
tails, and in what sequence.

00:05:39.200 --> 00:05:43.110 align:middle line:84%
So three heads in a row is
one particular outcome.

00:05:43.110 --> 00:05:46.460 align:middle line:84%
So what is the meaning
of those labels in

00:05:46.460 --> 00:05:48.030 align:middle line:90%
front of the branches?

00:05:48.030 --> 00:05:51.850 align:middle line:84%
So this P here, of course,
stands for the probability

00:05:51.850 --> 00:05:55.640 align:middle line:84%
that the first toss
resulted in heads.

00:05:55.640 --> 00:05:59.270 align:middle line:84%
And let me use this notation
to denote that

00:05:59.270 --> 00:06:01.170 align:middle line:90%
the first was heads.

00:06:01.170 --> 00:06:04.570 align:middle line:90%
I put an H in toss one.

00:06:04.570 --> 00:06:08.350 align:middle line:84%
How about the meaning of
this probability here?

00:06:08.350 --> 00:06:10.570 align:middle line:84%
Well the meaning of this
probability is

00:06:10.570 --> 00:06:11.670 align:middle line:90%
a conditional one.

00:06:11.670 --> 00:06:14.170 align:middle line:84%
It's the conditional probability
that the second

00:06:14.170 --> 00:06:18.340 align:middle line:84%
toss resulted in heads,
given that the first

00:06:18.340 --> 00:06:21.440 align:middle line:90%
one resulted in heads.

00:06:21.440 --> 00:06:26.830 align:middle line:84%
And similarly this label here
corresponds to the probability

00:06:26.830 --> 00:06:31.550 align:middle line:84%
that the third toss resulted in
heads, given that the first

00:06:31.550 --> 00:06:35.010 align:middle line:84%
one and the second one
resulted in heads.

00:06:35.010 --> 00:06:39.610 align:middle line:84%
So in this particular model that
I wrote down here, those

00:06:39.610 --> 00:06:44.740 align:middle line:84%
probabilities, P, of obtaining
heads remain the same no

00:06:44.740 --> 00:06:47.570 align:middle line:84%
matter what happened in
the previous toss.

00:06:47.570 --> 00:06:52.020 align:middle line:84%
For example, even if the first
toss was tails, we still have

00:06:52.020 --> 00:06:56.920 align:middle line:84%
the same probability, P, that
the second one is heads, given

00:06:56.920 --> 00:06:59.100 align:middle line:90%
that the first one was tails.

00:06:59.100 --> 00:07:01.820 align:middle line:84%
So we're assuming that no matter
what happened in the

00:07:01.820 --> 00:07:05.550 align:middle line:84%
first toss, the second toss will
still have a conditional

00:07:05.550 --> 00:07:08.960 align:middle line:84%
probability equal to P. So that
conditional probability

00:07:08.960 --> 00:07:12.800 align:middle line:84%
does not depend on what happened
in the first toss.

00:07:12.800 --> 00:07:16.040 align:middle line:84%
And we will see that this is a
very special situation, and

00:07:16.040 --> 00:07:19.240 align:middle line:84%
that's really the concept of
independence that we are going

00:07:19.240 --> 00:07:20.850 align:middle line:90%
to introduce shortly.

00:07:20.850 --> 00:07:25.540 align:middle line:84%
But before we get to
independence, let's practice

00:07:25.540 --> 00:07:29.060 align:middle line:84%
once more the three skills that
we covered last time in

00:07:29.060 --> 00:07:30.490 align:middle line:90%
this example.

00:07:30.490 --> 00:07:33.470 align:middle line:84%
So first skill was
multiplication rule.

00:07:33.470 --> 00:07:35.660 align:middle line:84%
How do you find the
probability of

00:07:35.660 --> 00:07:38.000 align:middle line:90%
several things happening?

00:07:38.000 --> 00:07:41.390 align:middle line:84%
That is the probability that
we have tails followed by

00:07:41.390 --> 00:07:44.140 align:middle line:90%
heads followed by tails.

00:07:44.140 --> 00:07:50.350 align:middle line:84%
So here we're talking about this
particular outcome here,

00:07:50.350 --> 00:07:53.130 align:middle line:84%
tails followed by heads
followed by tails.

00:07:53.130 --> 00:07:57.070 align:middle line:84%
And the way we calculate such
a probability is by

00:07:57.070 --> 00:08:01.480 align:middle line:84%
multiplying conditional
probabilities along the path

00:08:01.480 --> 00:08:03.560 align:middle line:90%
that takes us to this outcome.

00:08:03.560 --> 00:08:05.160 align:middle line:84%
And so these conditional
probabilities

00:08:05.160 --> 00:08:06.220 align:middle line:90%
are recorded here.

00:08:06.220 --> 00:08:11.840 align:middle line:84%
So it's going to be (1 minus P)
times P times (1 minus P).

00:08:11.840 --> 00:08:14.480 align:middle line:84%
So this is the multiplication
rule.

00:08:14.480 --> 00:08:17.990 align:middle line:84%
Second question is how do we
find the probability of a

00:08:17.990 --> 00:08:20.510 align:middle line:90%
mildly complicated event?

00:08:20.510 --> 00:08:23.520 align:middle line:84%
So the event of interest here
that I wrote down is the

00:08:23.520 --> 00:08:25.850 align:middle line:84%
probability that in the
three tosses, we had a

00:08:25.850 --> 00:08:28.650 align:middle line:90%
total of one head.

00:08:28.650 --> 00:08:30.470 align:middle line:90%
Exactly one head.

00:08:30.470 --> 00:08:33.450 align:middle line:84%
This is an event that can
happen in multiple ways.

00:08:33.450 --> 00:08:35.940 align:middle line:90%
It happens here.

00:08:35.940 --> 00:08:38.200 align:middle line:90%
It happens here.

00:08:38.200 --> 00:08:41.380 align:middle line:90%
And it also happens here.

00:08:41.380 --> 00:08:44.480 align:middle line:84%
So we want to find the total
probability of the event

00:08:44.480 --> 00:08:46.290 align:middle line:84%
consisting of these
three outcomes.

00:08:46.290 --> 00:08:47.370 align:middle line:90%
What do we do?

00:08:47.370 --> 00:08:51.100 align:middle line:84%
We just add the probabilities
of each individual outcome.

00:08:51.100 --> 00:08:53.850 align:middle line:84%
How do we find the probability
of an individual outcome?

00:08:53.850 --> 00:08:56.250 align:middle line:90%
Well, that's what we just did.

00:08:56.250 --> 00:09:00.300 align:middle line:84%
Now notice that this outcome
has probability P times (1

00:09:00.300 --> 00:09:01.550 align:middle line:90%
minus P) squared.

00:09:01.550 --> 00:09:04.260 align:middle line:90%


00:09:04.260 --> 00:09:07.000 align:middle line:90%
That one should not be there.

00:09:07.000 --> 00:09:08.984 align:middle line:90%
So where is it?

00:09:08.984 --> 00:09:09.750 align:middle line:90%
Ah.

00:09:09.750 --> 00:09:11.000 align:middle line:90%
It's this one.

00:09:11.000 --> 00:09:13.830 align:middle line:90%


00:09:13.830 --> 00:09:18.610 align:middle line:84%
OK, so the probability of this
outcome is (1 minus P times P)

00:09:18.610 --> 00:09:20.970 align:middle line:84%
times (1 minus P), the
same probability.

00:09:20.970 --> 00:09:25.470 align:middle line:84%
And finally, this one is again
(1 minus P) squared times P.

00:09:25.470 --> 00:09:29.240 align:middle line:84%
So this event of one head can
happen in three ways.

00:09:29.240 --> 00:09:32.270 align:middle line:84%
And each one of those three ways
has the same probability

00:09:32.270 --> 00:09:33.380 align:middle line:90%
of occurring.

00:09:33.380 --> 00:09:36.440 align:middle line:90%
And this is the answer.

00:09:36.440 --> 00:09:40.110 align:middle line:84%
And finally, the last thing that
we learned how to do is

00:09:40.110 --> 00:09:41.980 align:middle line:90%
to use the Bayes rule to

00:09:41.980 --> 00:09:44.230 align:middle line:90%
calculate and make an inference.

00:09:44.230 --> 00:09:47.045 align:middle line:84%
So somebody tells you that there
was exactly one head in

00:09:47.045 --> 00:09:49.350 align:middle line:90%
your three tosses.

00:09:49.350 --> 00:09:52.610 align:middle line:84%
What is the probability
that the first

00:09:52.610 --> 00:09:55.110 align:middle line:90%
toss resulted in heads?

00:09:55.110 --> 00:09:59.980 align:middle line:84%
OK, I guess you can guess the
answer here if I tell you that

00:09:59.980 --> 00:10:01.710 align:middle line:90%
there were three tosses.

00:10:01.710 --> 00:10:03.590 align:middle line:90%
One of them was heads.

00:10:03.590 --> 00:10:05.670 align:middle line:84%
Where was that head
in the first, the

00:10:05.670 --> 00:10:07.300 align:middle line:90%
second, or the third?

00:10:07.300 --> 00:10:10.410 align:middle line:84%
Well, by symmetry, they should
all be equally likely.

00:10:10.410 --> 00:10:13.770 align:middle line:84%
So there should be probably
just 1/3 that that head

00:10:13.770 --> 00:10:16.070 align:middle line:90%
occurred in the first toss.

00:10:16.070 --> 00:10:19.230 align:middle line:84%
Let's check our intuition
using the definitions.

00:10:19.230 --> 00:10:21.280 align:middle line:84%
So the definition of conditional
probability tells

00:10:21.280 --> 00:10:26.030 align:middle line:84%
us the conditional probability
is the probability of both

00:10:26.030 --> 00:10:27.310 align:middle line:90%
things happening.

00:10:27.310 --> 00:10:33.890 align:middle line:84%
First toss is heads, and we have
exactly one head divided

00:10:33.890 --> 00:10:36.430 align:middle line:84%
by the probability
of one head.

00:10:36.430 --> 00:10:40.720 align:middle line:90%


00:10:40.720 --> 00:10:44.860 align:middle line:84%
What is the probability that the
first toss is heads, and

00:10:44.860 --> 00:10:47.340 align:middle line:90%
we have exactly one head?

00:10:47.340 --> 00:10:51.810 align:middle line:84%
This is the same as the event
heads, tails, tails.

00:10:51.810 --> 00:10:54.280 align:middle line:84%
If I tell you that the first is
heads, and there's only one

00:10:54.280 --> 00:10:57.030 align:middle line:84%
head, it means that the
others are tails.

00:10:57.030 --> 00:11:03.080 align:middle line:84%
So this is the probability of
heads, tails, tails divided by

00:11:03.080 --> 00:11:06.080 align:middle line:90%
the probability of one head.

00:11:06.080 --> 00:11:08.660 align:middle line:84%
And we know all of these
quantities probability of

00:11:08.660 --> 00:11:12.080 align:middle line:84%
heads, tails, tails is P times
(1 minus P) squared.

00:11:12.080 --> 00:11:14.806 align:middle line:84%
Probability of one
head is 3 times P

00:11:14.806 --> 00:11:17.680 align:middle line:90%
times (1 minus P) squared.

00:11:17.680 --> 00:11:22.820 align:middle line:84%
So the final answer is 1/3,
which is what you should have

00:11:22.820 --> 00:11:27.280 align:middle line:84%
a guessed on intuitive
grounds.

00:11:27.280 --> 00:11:27.740 align:middle line:90%
Very good.

00:11:27.740 --> 00:11:31.110 align:middle line:84%
So we got our practice on
the material that we

00:11:31.110 --> 00:11:33.040 align:middle line:90%
did cover last time.

00:11:33.040 --> 00:11:33.700 align:middle line:90%
Again, think.

00:11:33.700 --> 00:11:38.050 align:middle line:84%
There's basically three basic
skills that we are practicing

00:11:38.050 --> 00:11:40.210 align:middle line:90%
and exercising here.

00:11:40.210 --> 00:11:43.870 align:middle line:84%
In the problems, quizzes, and in
the real life, you may have

00:11:43.870 --> 00:11:47.560 align:middle line:84%
to apply those three skills in
somewhat more complicated

00:11:47.560 --> 00:11:49.590 align:middle line:84%
settings, but in the
end that's what it

00:11:49.590 --> 00:11:51.860 align:middle line:90%
boils down to usually.

00:11:51.860 --> 00:11:55.240 align:middle line:84%
Now let's focus on this special
feature of this

00:11:55.240 --> 00:11:59.610 align:middle line:84%
particular model that I
discussed a little earlier.

00:11:59.610 --> 00:12:03.010 align:middle line:84%
Think of the event heads
in the second toss.

00:12:03.010 --> 00:12:05.690 align:middle line:90%


00:12:05.690 --> 00:12:09.750 align:middle line:84%
Initially, the probability of
heads in the second toss, you

00:12:09.750 --> 00:12:12.460 align:middle line:84%
know, that it's P, the
probability of

00:12:12.460 --> 00:12:14.290 align:middle line:90%
success of your coin.

00:12:14.290 --> 00:12:19.100 align:middle line:84%
If I tell you that the first
toss resulted in heads, what's

00:12:19.100 --> 00:12:21.240 align:middle line:84%
the probability that the
second toss is heads?

00:12:21.240 --> 00:12:24.870 align:middle line:84%
It's again P. If I tell you that
the first toss was tails,

00:12:24.870 --> 00:12:27.510 align:middle line:84%
what's the probability that
the second toss is heads?

00:12:27.510 --> 00:12:33.290 align:middle line:84%
It's again P. So whether I tell
you the result of the

00:12:33.290 --> 00:12:37.280 align:middle line:84%
first toss, or I don't tell
you, it doesn't make any

00:12:37.280 --> 00:12:38.490 align:middle line:90%
difference to you.

00:12:38.490 --> 00:12:40.690 align:middle line:84%
You would always say the
probability of heads in the

00:12:40.690 --> 00:12:44.970 align:middle line:84%
second toss is going to P, no
matter what happened in the

00:12:44.970 --> 00:12:46.400 align:middle line:90%
first toss.

00:12:46.400 --> 00:12:49.550 align:middle line:84%
This is a special situation to
which we're going to give a

00:12:49.550 --> 00:12:53.540 align:middle line:84%
name, and we're going to call
that property independence.

00:12:53.540 --> 00:12:58.520 align:middle line:84%
Basically independence between
two things stands for the fact

00:12:58.520 --> 00:13:02.690 align:middle line:84%
that the first thing, whether
it occurred or not, doesn't

00:13:02.690 --> 00:13:05.630 align:middle line:84%
give you any information, does
not cause you to change your

00:13:05.630 --> 00:13:08.980 align:middle line:84%
beliefs about the
second event.

00:13:08.980 --> 00:13:11.600 align:middle line:90%
This is the intuition.

00:13:11.600 --> 00:13:16.130 align:middle line:84%
Let's try to translate this
into mathematics.

00:13:16.130 --> 00:13:19.510 align:middle line:84%
We have two events, and we're
going to say that they're

00:13:19.510 --> 00:13:26.010 align:middle line:84%
independent if your initial
beliefs about B are not going

00:13:26.010 --> 00:13:30.140 align:middle line:84%
to change if I tell you
that A occurred.

00:13:30.140 --> 00:13:34.700 align:middle line:84%
So you believe something
how likely B is.

00:13:34.700 --> 00:13:37.640 align:middle line:84%
Then somebody comes and tells
you, you know, A has happened.

00:13:37.640 --> 00:13:39.790 align:middle line:84%
Are you going to change
your beliefs?

00:13:39.790 --> 00:13:42.200 align:middle line:84%
No, I'm not going
to change them.

00:13:42.200 --> 00:13:45.020 align:middle line:84%
Whenever you are in such a
situation, then you say that

00:13:45.020 --> 00:13:47.040 align:middle line:84%
the two events are
independent.

00:13:47.040 --> 00:13:51.470 align:middle line:84%
Intuitively, the fact that A
occurred does not convey any

00:13:51.470 --> 00:13:55.720 align:middle line:84%
information to you about the
likelihood of event B. The

00:13:55.720 --> 00:13:58.480 align:middle line:84%
information that A provides
is not so

00:13:58.480 --> 00:14:00.780 align:middle line:90%
useful, is not relevant.

00:14:00.780 --> 00:14:03.010 align:middle line:84%
A has to do with
something else.

00:14:03.010 --> 00:14:06.250 align:middle line:84%
It's not useful for your
guessing whether B is going to

00:14:06.250 --> 00:14:07.780 align:middle line:90%
occur or not.

00:14:07.780 --> 00:14:13.650 align:middle line:84%
So we can take this as a first
attempt into a definition of

00:14:13.650 --> 00:14:15.870 align:middle line:90%
independence.

00:14:15.870 --> 00:14:23.130 align:middle line:84%
Now remember that we have this
property, the probability of

00:14:23.130 --> 00:14:25.690 align:middle line:84%
two things happening is the
probability of the first times

00:14:25.690 --> 00:14:27.920 align:middle line:84%
the conditional probability
of the second.

00:14:27.920 --> 00:14:31.390 align:middle line:84%
If we have independence, this
conditional probability is the

00:14:31.390 --> 00:14:33.840 align:middle line:84%
same as the unconditional
probability.

00:14:33.840 --> 00:14:38.040 align:middle line:84%
So if we have independence
according to that definition,

00:14:38.040 --> 00:14:41.190 align:middle line:84%
we get this property that you
can find the probability of

00:14:41.190 --> 00:14:44.440 align:middle line:84%
two things happening by just
multiplying their individual

00:14:44.440 --> 00:14:45.640 align:middle line:90%
probabilities.

00:14:45.640 --> 00:14:48.070 align:middle line:84%
Probability of heads in
the first toss is 1/2.

00:14:48.070 --> 00:14:50.900 align:middle line:84%
Probability of heads in the
second toss is 1/2.

00:14:50.900 --> 00:14:54.200 align:middle line:84%
Probability of heads
heads is 1/4.

00:14:54.200 --> 00:14:57.590 align:middle line:84%
That's what happens if your two
tosses are independent of

00:14:57.590 --> 00:14:58.730 align:middle line:90%
each other.

00:14:58.730 --> 00:15:03.110 align:middle line:84%
So this property here is
a consequence of this

00:15:03.110 --> 00:15:08.470 align:middle line:84%
definition, but it's actually
nicer, better, simpler,

00:15:08.470 --> 00:15:12.880 align:middle line:84%
cleaner, more beautiful to take
this as our definition

00:15:12.880 --> 00:15:14.380 align:middle line:90%
instead of that one.

00:15:14.380 --> 00:15:17.180 align:middle line:84%
Are the two definitions
equivalent?

00:15:17.180 --> 00:15:21.040 align:middle line:84%
Well, they're are almost the
same, except for one thing.

00:15:21.040 --> 00:15:24.250 align:middle line:84%
Conditional probabilities are
only defined if you condition

00:15:24.250 --> 00:15:26.900 align:middle line:84%
on an event that has positive
probability.

00:15:26.900 --> 00:15:31.090 align:middle line:84%
So this definition would be
limited to cases where event A

00:15:31.090 --> 00:15:34.080 align:middle line:84%
has positive probability,
whereas this definition is

00:15:34.080 --> 00:15:38.140 align:middle line:84%
something that you can
write down always.

00:15:38.140 --> 00:15:43.280 align:middle line:84%
We will say that two events are
independent if and only if

00:15:43.280 --> 00:15:46.940 align:middle line:84%
their probability of happening
simultaneously is equal to the

00:15:46.940 --> 00:15:50.470 align:middle line:84%
product of their two individual
probabilities.

00:15:50.470 --> 00:15:54.690 align:middle line:84%
And in particular, we can have
events of zero probability.

00:15:54.690 --> 00:15:56.220 align:middle line:84%
There's nothing wrong
with that.

00:15:56.220 --> 00:16:01.450 align:middle line:84%
If A has 0 probability, then A
intersection B will also have

00:16:01.450 --> 00:16:04.990 align:middle line:84%
zero probability, because it's
an even smaller event.

00:16:04.990 --> 00:16:09.200 align:middle line:84%
And so we're going to get
zero is equal to zero.

00:16:09.200 --> 00:16:13.920 align:middle line:84%
A corollary of what I just said,
if an event A has zero

00:16:13.920 --> 00:16:17.700 align:middle line:84%
probability, it's actually
independent of any other event

00:16:17.700 --> 00:16:20.220 align:middle line:84%
in our model, because
we're going to get

00:16:20.220 --> 00:16:21.810 align:middle line:90%
zero is equal to zero.

00:16:21.810 --> 00:16:24.140 align:middle line:84%
And the definition is going
to be satisfied.

00:16:24.140 --> 00:16:27.560 align:middle line:84%
This is a little bit harder to
reconcile with the intuition

00:16:27.560 --> 00:16:32.800 align:middle line:84%
we have about independence, but
then again, it's part of

00:16:32.800 --> 00:16:35.610 align:middle line:90%
the mathematical definition.

00:16:35.610 --> 00:16:40.450 align:middle line:84%
So what I want you to retain
is this notion that the

00:16:40.450 --> 00:16:46.300 align:middle line:84%
independence is something that
you can check formally using

00:16:46.300 --> 00:16:50.420 align:middle line:84%
this definition, but also you
can check intuitively by if,

00:16:50.420 --> 00:16:54.280 align:middle line:84%
in some cases, you can reason
that whatever happens and

00:16:54.280 --> 00:16:58.310 align:middle line:84%
determines whether A is going
to occur or not, has nothing

00:16:58.310 --> 00:17:01.850 align:middle line:84%
absolutely to do with whatever
happens and determines whether

00:17:01.850 --> 00:17:04.369 align:middle line:90%
B is going to occur or not.

00:17:04.369 --> 00:17:08.440 align:middle line:84%
So if I'm doing a science
experiment in this room, and

00:17:08.440 --> 00:17:12.569 align:middle line:84%
it gets hit by some noise that's
causes randomness.

00:17:12.569 --> 00:17:16.040 align:middle line:84%
And then five years later,
somebody somewhere else does

00:17:16.040 --> 00:17:19.069 align:middle line:84%
the same science experiment
somewhere else, it gets hit by

00:17:19.069 --> 00:17:23.230 align:middle line:84%
other noise, you would usually
say that these experiments are

00:17:23.230 --> 00:17:23.940 align:middle line:90%
independent.

00:17:23.940 --> 00:17:30.230 align:middle line:84%
So what events happen in one
experiment are not going to

00:17:30.230 --> 00:17:33.290 align:middle line:84%
change your beliefs about what
might be happening in the

00:17:33.290 --> 00:17:36.610 align:middle line:84%
other, because the sources of
noise in these two experiments

00:17:36.610 --> 00:17:38.350 align:middle line:90%
are completely unrelated.

00:17:38.350 --> 00:17:40.110 align:middle line:84%
They have nothing to
do with each other.

00:17:40.110 --> 00:17:43.470 align:middle line:84%
So if I flip a coin here today,
and I flip a coin in my

00:17:43.470 --> 00:17:47.890 align:middle line:84%
office tomorrow, one shouldn't
affect the other.

00:17:47.890 --> 00:17:52.690 align:middle line:84%
So the events that I get from
these should be independent.

00:17:52.690 --> 00:17:55.700 align:middle line:84%
So that's usually how
independence arises.

00:17:55.700 --> 00:17:57.580 align:middle line:90%
By having distinct physical

00:17:57.580 --> 00:17:59.940 align:middle line:90%
phenomena that do not interact.

00:17:59.940 --> 00:18:03.690 align:middle line:84%
Sometimes you also get
independence even though there

00:18:03.690 --> 00:18:06.590 align:middle line:84%
is a physical interaction, but
you just happen to have a

00:18:06.590 --> 00:18:08.930 align:middle line:90%
numerical accident.

00:18:08.930 --> 00:18:13.340 align:middle line:84%
A and B might be physically
related very tightly, but a

00:18:13.340 --> 00:18:16.820 align:middle line:84%
numerical accident happens and
you get equality here, that's

00:18:16.820 --> 00:18:20.070 align:middle line:84%
another case where we
do get independence.

00:18:20.070 --> 00:18:24.350 align:middle line:84%
Now suppose that we have
two events that are

00:18:24.350 --> 00:18:27.380 align:middle line:90%
laid out like this.

00:18:27.380 --> 00:18:30.240 align:middle line:84%
Are these two events
independent or not?

00:18:30.240 --> 00:18:34.570 align:middle line:90%


00:18:34.570 --> 00:18:36.620 align:middle line:84%
The picture kind of tells
you that one is

00:18:36.620 --> 00:18:38.140 align:middle line:90%
separate from the other.

00:18:38.140 --> 00:18:41.170 align:middle line:84%
But separate has nothing
to do with independent.

00:18:41.170 --> 00:18:45.340 align:middle line:84%
In fact, these two events are as
dependent as Siamese twins.

00:18:45.340 --> 00:18:46.480 align:middle line:90%
Why is that?

00:18:46.480 --> 00:18:51.560 align:middle line:84%
If I tell you that A occurred,
then you are certain that B

00:18:51.560 --> 00:18:53.060 align:middle line:90%
did not occur.

00:18:53.060 --> 00:18:57.780 align:middle line:84%
So information about the
occurrence of A definitely

00:18:57.780 --> 00:19:01.090 align:middle line:84%
affects your beliefs about the
possible occurrence or

00:19:01.090 --> 00:19:05.490 align:middle line:84%
non-occurrence of B. When the
picture is like that, knowing

00:19:05.490 --> 00:19:09.480 align:middle line:84%
that A occurred will change
drastically my beliefs about

00:19:09.480 --> 00:19:13.030 align:middle line:84%
B, because now I suddenly
become certain

00:19:13.030 --> 00:19:14.870 align:middle line:90%
that B did not occur.

00:19:14.870 --> 00:19:18.260 align:middle line:84%
So a picture like this is a
case actually of extreme

00:19:18.260 --> 00:19:19.360 align:middle line:90%
dependence.

00:19:19.360 --> 00:19:23.440 align:middle line:84%
So don't confuse independence
with disjointness.

00:19:23.440 --> 00:19:26.406 align:middle line:84%
They're very different
types of properties.

00:19:26.406 --> 00:19:27.080 align:middle line:90%
AUDIENCE: Question.

00:19:27.080 --> 00:19:27.520 align:middle line:90%
PROFESSOR: Yes?

00:19:27.520 --> 00:19:29.400 align:middle line:84%
AUDIENCE: So I understand
the explanation, but the

00:19:29.400 --> 00:19:31.954 align:middle line:84%
probability of A intersect B
[INAUDIBLE] to zero, because

00:19:31.954 --> 00:19:32.910 align:middle line:90%
they're disjoint.

00:19:32.910 --> 00:19:33.388 align:middle line:90%
PROFESSOR: Yes.

00:19:33.388 --> 00:19:35.539 align:middle line:84%
AUDIENCE: But then the product
of probability A and

00:19:35.539 --> 00:19:37.690 align:middle line:84%
probability B, one of them
is going to be 1.

00:19:37.690 --> 00:19:39.602 align:middle line:90%
[INAUDIBLE]

00:19:39.602 --> 00:19:42.690 align:middle line:84%
PROFESSOR: No, suppose that
the probabilities are 1/3,

00:19:42.690 --> 00:19:46.610 align:middle line:84%
1/4, and the rest
is out there.

00:19:46.610 --> 00:19:48.560 align:middle line:84%
You check the definition
of independence.

00:19:48.560 --> 00:19:52.440 align:middle line:84%
Probability of A intersection
B is zero.

00:19:52.440 --> 00:19:58.520 align:middle line:84%
Probability of A times the
probability of B is 1/12.

00:19:58.520 --> 00:20:00.630 align:middle line:90%
The two are not equal.

00:20:00.630 --> 00:20:02.710 align:middle line:84%
Therefore we do not
have independence.

00:20:02.710 --> 00:20:03.199 align:middle line:90%
AUDIENCE: Right.

00:20:03.199 --> 00:20:05.644 align:middle line:84%
So what's wrong with the
intuition of the probability

00:20:05.644 --> 00:20:09.556 align:middle line:84%
of A being 1, and the
other one being 0?

00:20:09.556 --> 00:20:12.490 align:middle line:90%
[INAUDIBLE].

00:20:12.490 --> 00:20:12.610 align:middle line:90%
PROFESSOR: No.

00:20:12.610 --> 00:20:19.340 align:middle line:84%
The probability of A given
B is equal to 0.

00:20:19.340 --> 00:20:23.870 align:middle line:84%
Probability of A is
equal to 1/3.

00:20:23.870 --> 00:20:26.650 align:middle line:84%
So again, these two
are different.

00:20:26.650 --> 00:20:30.210 align:middle line:84%
So we had some initial beliefs
about A, but as soon as we are

00:20:30.210 --> 00:20:34.440 align:middle line:84%
told that B occurred, our
beliefs about A changed.

00:20:34.440 --> 00:20:37.770 align:middle line:84%
And so since our beliefs
changed, that means that B

00:20:37.770 --> 00:20:40.666 align:middle line:90%
conveys information about A.

00:20:40.666 --> 00:20:42.931 align:middle line:84%
AUDIENCE: So can you not draw
independent [INAUDIBLE] on a

00:20:42.931 --> 00:20:43.390 align:middle line:90%
Venn diagram?

00:20:43.390 --> 00:20:44.430 align:middle line:90%
PROFESSOR: I can't hear you.

00:20:44.430 --> 00:20:45.352 align:middle line:90%
AUDIENCE: Can you draw

00:20:45.352 --> 00:20:46.735 align:middle line:90%
independence on a Venn diagram?

00:20:46.735 --> 00:20:51.320 align:middle line:84%
PROFESSOR: No, the Venn diagram
is never enough to

00:20:51.320 --> 00:20:53.400 align:middle line:90%
decide independence.

00:20:53.400 --> 00:20:56.350 align:middle line:84%
So the typical picture in which
you're going to have

00:20:56.350 --> 00:21:00.120 align:middle line:84%
independence would be one event
this way, and another

00:21:00.120 --> 00:21:01.760 align:middle line:90%
event this way.

00:21:01.760 --> 00:21:03.800 align:middle line:84%
You need to take the probability
of this times the

00:21:03.800 --> 00:21:07.795 align:middle line:84%
probability of that, and check
that, numerically, it's equal

00:21:07.795 --> 00:21:11.350 align:middle line:84%
to the probability of
this intersection.

00:21:11.350 --> 00:21:14.330 align:middle line:84%
So it's more than
a Venn diagram.

00:21:14.330 --> 00:21:16.138 align:middle line:84%
Numbers need to come
out right.

00:21:16.138 --> 00:21:19.730 align:middle line:90%


00:21:19.730 --> 00:21:23.570 align:middle line:84%
Now we did say some time ago
that conditional probabilities

00:21:23.570 --> 00:21:27.680 align:middle line:84%
are just like ordinary
probabilities, and whatever we

00:21:27.680 --> 00:21:31.870 align:middle line:84%
do in probability theory
can also be done

00:21:31.870 --> 00:21:34.350 align:middle line:90%
in conditional universes.

00:21:34.350 --> 00:21:37.680 align:middle line:84%
Talking about conditional
probabilities.

00:21:37.680 --> 00:21:42.870 align:middle line:84%
So since we have a notion of
independence, then there

00:21:42.870 --> 00:21:47.470 align:middle line:84%
should be also a notion of
conditional independence.

00:21:47.470 --> 00:21:55.070 align:middle line:84%
So independence was defined
by the probability that A

00:21:55.070 --> 00:21:59.070 align:middle line:84%
intersection B is equal to the
probability of A times the

00:21:59.070 --> 00:22:01.920 align:middle line:90%
probability of B.

00:22:01.920 --> 00:22:05.670 align:middle line:84%
What would be a reasonable
definition of conditional

00:22:05.670 --> 00:22:06.840 align:middle line:90%
independence?

00:22:06.840 --> 00:22:09.355 align:middle line:84%
Conditional independence would
mean that this same property

00:22:09.355 --> 00:22:13.210 align:middle line:84%
could be true, but in a
conditional universe where we

00:22:13.210 --> 00:22:15.660 align:middle line:84%
are told that the certain
event happens.

00:22:15.660 --> 00:22:19.060 align:middle line:84%
So if we're told that the event
C has happened, then

00:22:19.060 --> 00:22:22.240 align:middle line:84%
were transported in a
conditional universe where the

00:22:22.240 --> 00:22:26.460 align:middle line:84%
only thing that matters are
conditional probabilities.

00:22:26.460 --> 00:22:31.320 align:middle line:84%
And this is just the same plain,
previous definition of

00:22:31.320 --> 00:22:35.190 align:middle line:84%
independence, but applied in
a conditional universe.

00:22:35.190 --> 00:22:40.020 align:middle line:84%
So this is the definition of
conditional independence.

00:22:40.020 --> 00:22:43.390 align:middle line:90%


00:22:43.390 --> 00:22:46.940 align:middle line:84%
So it's independence, but with
reference to the conditional

00:22:46.940 --> 00:22:48.830 align:middle line:90%
probabilities.

00:22:48.830 --> 00:22:51.830 align:middle line:84%
And intuitively it has, again,
the same meaning, that in the

00:22:51.830 --> 00:22:56.410 align:middle line:84%
conditional world, if I tell you
that A occurred, then that

00:22:56.410 --> 00:22:58.940 align:middle line:84%
doesn't change your
beliefs about B.

00:22:58.940 --> 00:23:01.100 align:middle line:84%
So suppose you had a
picture like this.

00:23:01.100 --> 00:23:06.880 align:middle line:84%
And somebody told you that
events A and B are independent

00:23:06.880 --> 00:23:09.630 align:middle line:90%
unconditionally.

00:23:09.630 --> 00:23:14.320 align:middle line:84%
Then somebody comes and tells
you that event C actually has

00:23:14.320 --> 00:23:18.150 align:middle line:84%
occurred, so we now live
in this new universe.

00:23:18.150 --> 00:23:22.450 align:middle line:84%
In this new universe, is the
independence of A and B going

00:23:22.450 --> 00:23:25.180 align:middle line:90%
to be preserved or not?

00:23:25.180 --> 00:23:29.300 align:middle line:84%
Are A and B independent
in this new universe?

00:23:29.300 --> 00:23:34.780 align:middle line:84%
The answer is no, because in the
new universe, whatever is

00:23:34.780 --> 00:23:36.790 align:middle line:90%
left of event A is this piece.

00:23:36.790 --> 00:23:39.630 align:middle line:84%
Whatever is left of event
B is this piece.

00:23:39.630 --> 00:23:42.310 align:middle line:84%
And these two pieces
are disjoint.

00:23:42.310 --> 00:23:45.490 align:middle line:84%
So we are back in a situation
of this kind.

00:23:45.490 --> 00:23:46.450 align:middle line:90%
So in the conditional

00:23:46.450 --> 00:23:49.620 align:middle line:90%
universe, A and B are disjoint.

00:23:49.620 --> 00:23:53.380 align:middle line:84%
And therefore, generically,
they're not going to be

00:23:53.380 --> 00:23:54.730 align:middle line:90%
independent.

00:23:54.730 --> 00:23:58.030 align:middle line:84%
What's the moral of
this example?

00:23:58.030 --> 00:24:01.870 align:middle line:84%
Having independence in the
original model does not imply

00:24:01.870 --> 00:24:05.930 align:middle line:84%
independence in a conditional
model.

00:24:05.930 --> 00:24:08.490 align:middle line:90%
The opposite is also possible.

00:24:08.490 --> 00:24:12.160 align:middle line:84%
And let's illustrate
by another example.

00:24:12.160 --> 00:24:17.960 align:middle line:84%
So I have two coins, and both
of them are badly biased.

00:24:17.960 --> 00:24:21.680 align:middle line:84%
One coin is much biased
in favor of heads.

00:24:21.680 --> 00:24:25.320 align:middle line:84%
The other coin is much biased
in favor of tails.

00:24:25.320 --> 00:24:28.050 align:middle line:84%
So the probabilities
being 90%.

00:24:28.050 --> 00:24:33.050 align:middle line:84%
Let's consider independent flips
of coin A. This is the

00:24:33.050 --> 00:24:34.980 align:middle line:90%
relevant model.

00:24:34.980 --> 00:24:39.600 align:middle line:84%
This is a model of two
independent flips

00:24:39.600 --> 00:24:41.240 align:middle line:90%
of the first coin.

00:24:41.240 --> 00:24:43.850 align:middle line:84%
There's going to be two flips,
and each one has probability

00:24:43.850 --> 00:24:46.080 align:middle line:90%
0.9 of being heads.

00:24:46.080 --> 00:24:49.330 align:middle line:84%
So that's a model that describes
coin A. You can

00:24:49.330 --> 00:24:52.540 align:middle line:84%
think of this as a conditional
model which is a model of the

00:24:52.540 --> 00:24:55.940 align:middle line:84%
coin flips conditioned on the
fact that they have chosen

00:24:55.940 --> 00:24:57.460 align:middle line:90%
coin A.

00:24:57.460 --> 00:25:01.460 align:middle line:84%
Alternatively we could be
dealing with coin B In a

00:25:01.460 --> 00:25:05.260 align:middle line:84%
conditional world where we
chose coin B and flip it

00:25:05.260 --> 00:25:08.130 align:middle line:84%
twice, this is the
relevant model.

00:25:08.130 --> 00:25:10.660 align:middle line:84%
The probability of two heads,
for example, is the

00:25:10.660 --> 00:25:13.280 align:middle line:84%
probability of heads the first
time, heads the second time,

00:25:13.280 --> 00:25:16.070 align:middle line:90%
and each one is 0.1.

00:25:16.070 --> 00:25:19.960 align:middle line:84%
Now I'm building this into a
bigger experiment in which I

00:25:19.960 --> 00:25:25.160 align:middle line:84%
first start by choosing one of
the two coins at random.

00:25:25.160 --> 00:25:26.620 align:middle line:90%
So I have these two coins.

00:25:26.620 --> 00:25:28.610 align:middle line:90%
I blindly pick one of them.

00:25:28.610 --> 00:25:32.610 align:middle line:84%
And then I start
flipping them.

00:25:32.610 --> 00:25:36.620 align:middle line:84%
So the question now is, are the
coin flips, or the coin

00:25:36.620 --> 00:25:39.730 align:middle line:84%
tosses, are they independent
of each other?

00:25:39.730 --> 00:25:46.370 align:middle line:84%
If we just stay inside this
sub-model here, are the coin

00:25:46.370 --> 00:25:47.620 align:middle line:90%
flips independent?

00:25:47.620 --> 00:25:52.240 align:middle line:90%


00:25:52.240 --> 00:25:56.540 align:middle line:84%
They are independent, because
the probability of heads in

00:25:56.540 --> 00:26:01.780 align:middle line:84%
the second toss is the same,
0.9, no matter what happened

00:26:01.780 --> 00:26:03.550 align:middle line:90%
in the first toss.

00:26:03.550 --> 00:26:06.550 align:middle line:84%
So the conditional probabilities
of what happens

00:26:06.550 --> 00:26:10.050 align:middle line:84%
in the second toss are not
affected by the outcome of the

00:26:10.050 --> 00:26:11.180 align:middle line:90%
first toss.

00:26:11.180 --> 00:26:14.620 align:middle line:84%
So the second toss and the first
toss are independent.

00:26:14.620 --> 00:26:17.800 align:middle line:84%
So here we're just dealing
with plain,

00:26:17.800 --> 00:26:19.990 align:middle line:90%
independent coin flips.

00:26:19.990 --> 00:26:24.940 align:middle line:84%
Similarity the coin flips within
this sub-model are also

00:26:24.940 --> 00:26:26.190 align:middle line:90%
independent.

00:26:26.190 --> 00:26:28.840 align:middle line:90%


00:26:28.840 --> 00:26:33.410 align:middle line:84%
Now the question is, if we look
at the big model as just

00:26:33.410 --> 00:26:38.955 align:middle line:84%
one probability model, instead
of looking at the conditional

00:26:38.955 --> 00:26:44.530 align:middle line:84%
sub-models, are the coin flips
independent of each other?

00:26:44.530 --> 00:26:49.590 align:middle line:84%
Does the outcome of a few coin
flips give you information

00:26:49.590 --> 00:26:53.610 align:middle line:90%
about subsequent coin flips?

00:26:53.610 --> 00:27:02.570 align:middle line:84%
Well if I observe ten
heads in a row--

00:27:02.570 --> 00:27:05.960 align:middle line:84%
So instead of two coin flips,
now let's think of doing more

00:27:05.960 --> 00:27:10.070 align:middle line:84%
of them so that the tree
gets expanded.

00:27:10.070 --> 00:27:13.800 align:middle line:90%
So let's start with this.

00:27:13.800 --> 00:27:16.020 align:middle line:90%
I don't know which coin it is.

00:27:16.020 --> 00:27:18.970 align:middle line:84%
What's the probability that
the 11th coin toss

00:27:18.970 --> 00:27:20.220 align:middle line:90%
is going to be heads?

00:27:20.220 --> 00:27:25.570 align:middle line:90%


00:27:25.570 --> 00:27:29.370 align:middle line:84%
There's complete symmetry here,
so the answer could not

00:27:29.370 --> 00:27:32.330 align:middle line:90%
be anything other than 1/2.

00:27:32.330 --> 00:27:36.950 align:middle line:84%
So let's justify it,
why is it 1/2?

00:27:36.950 --> 00:27:40.560 align:middle line:84%
Well, the probability that the
11th toss is heads, how can

00:27:40.560 --> 00:27:42.380 align:middle line:90%
that outcome happen?

00:27:42.380 --> 00:27:43.840 align:middle line:90%
It can happen in two ways.

00:27:43.840 --> 00:27:50.480 align:middle line:84%
You can choose coin A, which
happens with probability 1/2.

00:27:50.480 --> 00:27:54.370 align:middle line:84%
And having chosen coin A,
there's probability 0.9 that

00:27:54.370 --> 00:27:58.500 align:middle line:84%
it results in that you get
heads in the 11th toss.

00:27:58.500 --> 00:28:03.540 align:middle line:84%
Or you can choose coin B. And
if it's coin B when you flip

00:28:03.540 --> 00:28:06.710 align:middle line:84%
it, there's probably 0.1
that you have heads.

00:28:06.710 --> 00:28:08.860 align:middle line:90%
So the final answer is 1/2.

00:28:08.860 --> 00:28:11.370 align:middle line:90%


00:28:11.370 --> 00:28:14.820 align:middle line:84%
So each one of the coins is
biased, but they're biased in

00:28:14.820 --> 00:28:16.190 align:middle line:90%
different ways.

00:28:16.190 --> 00:28:20.340 align:middle line:84%
If I don't know which coin it
is, their two biases kind of

00:28:20.340 --> 00:28:23.740 align:middle line:84%
cancel out, and the probability
of obtaining heads

00:28:23.740 --> 00:28:27.880 align:middle line:84%
is just in the middle,
then it's 1/2.

00:28:27.880 --> 00:28:31.720 align:middle line:84%
Now if someone tells you that
the first ten tosses were

00:28:31.720 --> 00:28:34.940 align:middle line:84%
heads, is that going to
change your beliefs

00:28:34.940 --> 00:28:37.300 align:middle line:90%
about the 11th toss?

00:28:37.300 --> 00:28:41.820 align:middle line:84%
Here's how a reasonable person
would think about it.

00:28:41.820 --> 00:28:49.480 align:middle line:84%
If it's coin B the probability
of obtaining 10 heads in a row

00:28:49.480 --> 00:28:51.510 align:middle line:90%
is negligible.

00:28:51.510 --> 00:28:55.270 align:middle line:84%
It's going to be 0.1
to the 10th.

00:28:55.270 --> 00:28:59.110 align:middle line:84%
If it's coin A. The probability
of 10 heads in a

00:28:59.110 --> 00:29:01.380 align:middle line:84%
row is a more reasonable
number.

00:29:01.380 --> 00:29:03.850 align:middle line:90%
It's 0.9 to the 10th.

00:29:03.850 --> 00:29:10.320 align:middle line:84%
So this event is a lot more
likely to occur with coin A,

00:29:10.320 --> 00:29:13.910 align:middle line:90%
rather than coin B.

00:29:13.910 --> 00:29:18.820 align:middle line:84%
The plausible explanation of
having seen ten heads in a row

00:29:18.820 --> 00:29:25.730 align:middle line:84%
is that I actually chose coin A.
When you see ten heads in a

00:29:25.730 --> 00:29:29.690 align:middle line:84%
row, you are pretty certain that
it's coin A that we're

00:29:29.690 --> 00:29:30.940 align:middle line:90%
dealing with.

00:29:30.940 --> 00:29:33.800 align:middle line:84%
And once you're pretty certain
that it's coin A that we're

00:29:33.800 --> 00:29:36.350 align:middle line:84%
dealing with, what's the
probability that the

00:29:36.350 --> 00:29:38.246 align:middle line:90%
next toss is heads?

00:29:38.246 --> 00:29:40.960 align:middle line:90%
It's going to be 0.9.

00:29:40.960 --> 00:29:45.270 align:middle line:84%
So essentially here I'm doing
an inference calculation.

00:29:45.270 --> 00:29:48.990 align:middle line:84%
Given this information, I'm
making an inference about

00:29:48.990 --> 00:29:50.700 align:middle line:90%
which coin I'm dealing with.

00:29:50.700 --> 00:29:53.540 align:middle line:90%


00:29:53.540 --> 00:29:57.240 align:middle line:84%
I become pretty certain that
it's coin A, and given that

00:29:57.240 --> 00:30:00.640 align:middle line:84%
it's coin A, this probability
is going to be 0.9.

00:30:00.640 --> 00:30:04.070 align:middle line:84%
And I'm putting an approximate
sign here, because the

00:30:04.070 --> 00:30:06.220 align:middle line:84%
inference that I did
is approximate.

00:30:06.220 --> 00:30:09.850 align:middle line:84%
I'm pretty certain it's coin A.
I'm not 100% certain that

00:30:09.850 --> 00:30:11.200 align:middle line:90%
it's coin A.

00:30:11.200 --> 00:30:15.430 align:middle line:84%
But in any case what happens
here is that the unconditional

00:30:15.430 --> 00:30:19.440 align:middle line:84%
probability is different from
the conditional probability.

00:30:19.440 --> 00:30:23.710 align:middle line:84%
This information here makes
me change my beliefs

00:30:23.710 --> 00:30:25.590 align:middle line:90%
about the 11th toss.

00:30:25.590 --> 00:30:30.560 align:middle line:84%
And this means that the 11th
toss is dependent on the

00:30:30.560 --> 00:30:31.530 align:middle line:90%
previous tosses.

00:30:31.530 --> 00:30:35.590 align:middle line:84%
So the coin tosses have
now become dependent.

00:30:35.590 --> 00:30:38.790 align:middle line:84%
What is the physical link that
causes this dependence?

00:30:38.790 --> 00:30:42.710 align:middle line:84%
Well, the physical link is
the choice of the coin.

00:30:42.710 --> 00:30:46.580 align:middle line:84%
By choosing a particular coin,
I'm introducing a pattern in

00:30:46.580 --> 00:30:48.200 align:middle line:90%
the future coin tosses.

00:30:48.200 --> 00:30:52.740 align:middle line:84%
And that pattern is what
causes dependence.

00:30:52.740 --> 00:30:55.670 align:middle line:84%
OK, so I've been playing a
little bit too loose with the

00:30:55.670 --> 00:30:59.810 align:middle line:84%
language here, because we
defined the concept of

00:30:59.810 --> 00:31:01.810 align:middle line:90%
independence of two events.

00:31:01.810 --> 00:31:06.180 align:middle line:84%
But here I have been referring
to independent coin tosses,

00:31:06.180 --> 00:31:08.380 align:middle line:84%
where I'm thinking about
many coin tosses,

00:31:08.380 --> 00:31:11.200 align:middle line:90%
like 10 or 11 of them.

00:31:11.200 --> 00:31:15.160 align:middle line:84%
So to be proper, I should have
defined for you also the

00:31:15.160 --> 00:31:18.710 align:middle line:84%
notion of independence of
multiple events, not just two.

00:31:18.710 --> 00:31:21.970 align:middle line:84%
We don't want to just say coin
toss one is independent from

00:31:21.970 --> 00:31:23.170 align:middle line:90%
coin toss two.

00:31:23.170 --> 00:31:26.250 align:middle line:84%
We want to be able to say
something like, these 10 then

00:31:26.250 --> 00:31:29.690 align:middle line:84%
coin tosses are all independent
of each other.

00:31:29.690 --> 00:31:33.800 align:middle line:84%
Intuitively what that means
should be the same thing--

00:31:33.800 --> 00:31:37.450 align:middle line:84%
that information about some of
the coin tosses doesn't change

00:31:37.450 --> 00:31:40.220 align:middle line:84%
your beliefs about the remaining
coin tosses.

00:31:40.220 --> 00:31:43.580 align:middle line:84%
How do we translate that into
a mathematical definition?

00:31:43.580 --> 00:31:48.600 align:middle line:84%
Well, an ugly attempt
would be to impose

00:31:48.600 --> 00:31:51.800 align:middle line:90%
requirements such as this.

00:31:51.800 --> 00:31:56.780 align:middle line:84%
Think of A1 being the event that
the first flip was heads.

00:31:56.780 --> 00:32:00.980 align:middle line:84%
A2 is the event of that the
second flip was heads.

00:32:00.980 --> 00:32:04.320 align:middle line:84%
A3, the third flip, was
heads, and so on.

00:32:04.320 --> 00:32:08.310 align:middle line:84%
Here is an event whose
occurrence is not determined

00:32:08.310 --> 00:32:10.860 align:middle line:90%
by the first three coin flips.

00:32:10.860 --> 00:32:13.400 align:middle line:84%
And here's an event whose
occurrence or not is

00:32:13.400 --> 00:32:16.680 align:middle line:84%
determined by the fifth
and sixth coin flip.

00:32:16.680 --> 00:32:19.080 align:middle line:84%
If we think physically that
all those coin flips have

00:32:19.080 --> 00:32:22.220 align:middle line:84%
nothing to do with each other,
information about the fifth

00:32:22.220 --> 00:32:26.420 align:middle line:84%
and sixth coin flip are not
going to change what we expect

00:32:26.420 --> 00:32:27.960 align:middle line:90%
from the first three.

00:32:27.960 --> 00:32:30.780 align:middle line:84%
So the probability of this
event, the conditional

00:32:30.780 --> 00:32:33.050 align:middle line:84%
probability, should be the
same as the unconditional

00:32:33.050 --> 00:32:34.430 align:middle line:90%
probability.

00:32:34.430 --> 00:32:38.850 align:middle line:84%
And we would like a relation
of this kind to be true, no

00:32:38.850 --> 00:32:43.480 align:middle line:84%
matter what kind of formula you
write down, as long as the

00:32:43.480 --> 00:32:47.230 align:middle line:84%
events that show up here are
different from the events that

00:32:47.230 --> 00:32:49.250 align:middle line:90%
show up there.

00:32:49.250 --> 00:32:49.770 align:middle line:90%
OK.

00:32:49.770 --> 00:32:52.150 align:middle line:84%
That's sort of an
ugly definition.

00:32:52.150 --> 00:32:55.350 align:middle line:84%
The mathematical definition that
actually does the job,

00:32:55.350 --> 00:32:59.530 align:middle line:84%
and leads to all the
formulas of this

00:32:59.530 --> 00:33:01.130 align:middle line:90%
kind, is the following.

00:33:01.130 --> 00:33:03.610 align:middle line:84%
We're going to say that the
collection of events are

00:33:03.610 --> 00:33:07.090 align:middle line:84%
independent if we can find the
probability of their joint

00:33:07.090 --> 00:33:11.780 align:middle line:84%
occurrence by just multiplying
probabilities.

00:33:11.780 --> 00:33:17.380 align:middle line:84%
And that will be true even if
you look at sub-collections of

00:33:17.380 --> 00:33:18.640 align:middle line:90%
these events.

00:33:18.640 --> 00:33:20.670 align:middle line:90%
Let's make that more precise.

00:33:20.670 --> 00:33:24.310 align:middle line:84%
If we have three events, the
definition tells us that the

00:33:24.310 --> 00:33:27.560 align:middle line:84%
three events are independent
if the following are true.

00:33:27.560 --> 00:33:31.830 align:middle line:84%
Probability A1 and A2 and A3,
you can calculate this

00:33:31.830 --> 00:33:34.840 align:middle line:84%
probability by multiplying
individual probabilities.

00:33:34.840 --> 00:33:38.370 align:middle line:90%


00:33:38.370 --> 00:33:44.320 align:middle line:84%
But the same is true even if
you take fewer events.

00:33:44.320 --> 00:33:46.740 align:middle line:84%
Just a few indices out
of the indices

00:33:46.740 --> 00:33:48.340 align:middle line:90%
that we have available.

00:33:48.340 --> 00:33:54.970 align:middle line:84%
So we also require P(A1
intersection A2) is P(A1)

00:33:54.970 --> 00:33:57.600 align:middle line:90%
times P(A2).

00:33:57.600 --> 00:34:01.250 align:middle line:84%
And similarly for the other
possibilities of

00:34:01.250 --> 00:34:02.500 align:middle line:90%
choosing the indices.

00:34:02.500 --> 00:34:10.900 align:middle line:90%


00:34:10.900 --> 00:34:14.659 align:middle line:84%
OK, so independence,
mathematical definition,

00:34:14.659 --> 00:34:18.860 align:middle line:84%
requires that calculating
probabilities of any

00:34:18.860 --> 00:34:22.370 align:middle line:84%
intersection of the events we
have in our hands, that

00:34:22.370 --> 00:34:25.590 align:middle line:84%
calculation can be done by just
multiplying individual

00:34:25.590 --> 00:34:27.000 align:middle line:90%
probabilities.

00:34:27.000 --> 00:34:30.230 align:middle line:84%
And this has to apply to the
case where we consider all of

00:34:30.230 --> 00:34:33.300 align:middle line:90%
the events in our hands or just

00:34:33.300 --> 00:34:36.900 align:middle line:90%
sub-collections of those events.

00:34:36.900 --> 00:34:42.130 align:middle line:84%
Now these relations just by
themselves are called pairwise

00:34:42.130 --> 00:34:44.389 align:middle line:90%
independence.

00:34:44.389 --> 00:34:47.179 align:middle line:84%
So this relation, for example,
tells us that A1 is

00:34:47.179 --> 00:34:48.710 align:middle line:90%
independent from A2.

00:34:48.710 --> 00:34:51.130 align:middle line:84%
This tells us that A2 is
independent from A3.

00:34:51.130 --> 00:34:54.670 align:middle line:84%
This will tell us that A1
is independent from A3.

00:34:54.670 --> 00:34:58.990 align:middle line:84%
But independence of all the
events together actually

00:34:58.990 --> 00:35:01.020 align:middle line:90%
requires a little more.

00:35:01.020 --> 00:35:05.080 align:middle line:84%
One more equality that has to do
with all three events being

00:35:05.080 --> 00:35:07.000 align:middle line:90%
considered at the same time.

00:35:07.000 --> 00:35:10.562 align:middle line:84%
And this extra equality
is not redundant.

00:35:10.562 --> 00:35:13.020 align:middle line:84%
It actually does make
a difference.

00:35:13.020 --> 00:35:15.390 align:middle line:84%
Independence and pairwise
independence

00:35:15.390 --> 00:35:17.310 align:middle line:90%
are different things.

00:35:17.310 --> 00:35:20.320 align:middle line:84%
So let's illustrate the
situation with an example.

00:35:20.320 --> 00:35:22.790 align:middle line:84%
Suppose we have two
coin flips.

00:35:22.790 --> 00:35:28.390 align:middle line:84%
The coin tosses are independent,
so the bias is

00:35:28.390 --> 00:35:32.910 align:middle line:84%
1/2, so all possible outcomes
have a probability of 1/2

00:35:32.910 --> 00:35:36.100 align:middle line:90%
times 1/2, which is 1/4.

00:35:36.100 --> 00:35:40.520 align:middle line:84%
And let's consider now a bunch
of different events.

00:35:40.520 --> 00:35:46.290 align:middle line:84%
One event is that the
first toss is heads.

00:35:46.290 --> 00:35:48.950 align:middle line:90%
This is this blue set here.

00:35:48.950 --> 00:35:54.990 align:middle line:84%
Another event is the second
toss is heads.

00:35:54.990 --> 00:35:57.970 align:middle line:84%
And this is this black
event here.

00:35:57.970 --> 00:36:00.770 align:middle line:90%


00:36:00.770 --> 00:36:01.850 align:middle line:90%
OK.

00:36:01.850 --> 00:36:04.500 align:middle line:84%
Are these two events
independent?

00:36:04.500 --> 00:36:06.660 align:middle line:84%
If you check it mathematically,
yes.

00:36:06.660 --> 00:36:09.270 align:middle line:84%
Probability of A is probability
of B is 1/2.

00:36:09.270 --> 00:36:13.170 align:middle line:84%
Probability of A times
probability of B is 1/4, which

00:36:13.170 --> 00:36:16.700 align:middle line:84%
is the same as the probability
of A intersection B,

00:36:16.700 --> 00:36:18.070 align:middle line:90%
which is this set.

00:36:18.070 --> 00:36:20.680 align:middle line:84%
So we have just checked
mathematically that A and B

00:36:20.680 --> 00:36:22.180 align:middle line:90%
are independent.

00:36:22.180 --> 00:36:26.210 align:middle line:84%
Now lets consider a third event
which is that the first

00:36:26.210 --> 00:36:30.080 align:middle line:84%
and second toss give
the same result.

00:36:30.080 --> 00:36:32.270 align:middle line:90%
I'll use a different color.

00:36:32.270 --> 00:36:35.400 align:middle line:84%
First and second toss to
give the same result.

00:36:35.400 --> 00:36:38.350 align:middle line:84%
This is the event that
we obtain heads,

00:36:38.350 --> 00:36:40.700 align:middle line:90%
heads or tails, tails.

00:36:40.700 --> 00:36:43.030 align:middle line:84%
So this is the probability
of C. What's the

00:36:43.030 --> 00:36:44.280 align:middle line:90%
probability of C?

00:36:44.280 --> 00:36:47.790 align:middle line:90%


00:36:47.790 --> 00:36:51.520 align:middle line:84%
Well, C is made up of two
outcomes, each one of which

00:36:51.520 --> 00:36:55.500 align:middle line:84%
has probability 1/4, so the
probability of C is 1/2.

00:36:55.500 --> 00:36:58.600 align:middle line:84%
What is the probability
of C intersection A?

00:36:58.600 --> 00:37:02.760 align:middle line:84%
C intersection A is just this
one outcome, and has

00:37:02.760 --> 00:37:06.030 align:middle line:90%
probability 1/4.

00:37:06.030 --> 00:37:10.040 align:middle line:84%
What's the probability of A
intersection B intersection C?

00:37:10.040 --> 00:37:13.650 align:middle line:84%
The three events intersect just
this outcome, so this

00:37:13.650 --> 00:37:15.620 align:middle line:90%
probability is also 1/4.

00:37:15.620 --> 00:37:18.610 align:middle line:90%


00:37:18.610 --> 00:37:19.860 align:middle line:90%
OK.

00:37:19.860 --> 00:37:24.130 align:middle line:90%


00:37:24.130 --> 00:37:27.060 align:middle line:84%
What's the probability
of C given A and B?

00:37:27.060 --> 00:37:29.800 align:middle line:90%


00:37:29.800 --> 00:37:34.840 align:middle line:84%
If A has occurred, and B has
occurred, you are certain that

00:37:34.840 --> 00:37:36.980 align:middle line:90%
this outcome here happened.

00:37:36.980 --> 00:37:40.160 align:middle line:84%
If the first toss is H and the
second toss is H, then you're

00:37:40.160 --> 00:37:41.970 align:middle line:84%
certain of the first
and second toss

00:37:41.970 --> 00:37:43.760 align:middle line:90%
gave the same result.

00:37:43.760 --> 00:37:46.500 align:middle line:84%
So the conditional probability
of C given A and

00:37:46.500 --> 00:37:49.050 align:middle line:90%
B is equal to 1.

00:37:49.050 --> 00:37:51.640 align:middle line:84%
So do we have independence
in this example?

00:37:51.640 --> 00:37:54.310 align:middle line:90%


00:37:54.310 --> 00:37:55.970 align:middle line:90%
We don't.

00:37:55.970 --> 00:38:00.210 align:middle line:84%
C, that we obtain the same
result in the first and the

00:38:00.210 --> 00:38:04.020 align:middle line:84%
second toss, has probability
1/2.

00:38:04.020 --> 00:38:08.480 align:middle line:84%
Half of the possible outcomes
give us two coin flips with

00:38:08.480 --> 00:38:10.700 align:middle line:84%
the same result-- heads,
heads or tails, tails.

00:38:10.700 --> 00:38:12.970 align:middle line:84%
So the probability
of C is 1/2.

00:38:12.970 --> 00:38:17.590 align:middle line:84%
But if I tell you that the
events A and B both occurred,

00:38:17.590 --> 00:38:20.900 align:middle line:84%
then you're certain
that C occurred.

00:38:20.900 --> 00:38:23.190 align:middle line:84%
If I tell you that we had heads
and heads, then you're

00:38:23.190 --> 00:38:25.460 align:middle line:84%
certain the outcomes
were the same.

00:38:25.460 --> 00:38:28.830 align:middle line:84%
So the conditional probability
is different from the

00:38:28.830 --> 00:38:31.400 align:middle line:90%
unconditional probability.

00:38:31.400 --> 00:38:37.050 align:middle line:84%
So by combining these two
relations together, we get

00:38:37.050 --> 00:38:39.235 align:middle line:84%
that the three events
are not independent.

00:38:39.235 --> 00:38:42.260 align:middle line:90%


00:38:42.260 --> 00:38:45.520 align:middle line:84%
But are they pairwise
independent?

00:38:45.520 --> 00:38:49.020 align:middle line:90%
Is A independent from B?

00:38:49.020 --> 00:38:53.400 align:middle line:84%
Yes, because probability of A
times probability of B is 1/4,

00:38:53.400 --> 00:38:58.670 align:middle line:84%
which is probability of
A intersection B. Is C

00:38:58.670 --> 00:39:02.350 align:middle line:90%
independent from A?

00:39:02.350 --> 00:39:05.780 align:middle line:84%
Well, the probability
of C and A is 1/4.

00:39:05.780 --> 00:39:07.620 align:middle line:90%
The probability of C is 1/2.

00:39:07.620 --> 00:39:09.830 align:middle line:90%
The probability of A is 1/2.

00:39:09.830 --> 00:39:11.150 align:middle line:90%
So it checks.

00:39:11.150 --> 00:39:17.960 align:middle line:84%
1/4 is equal to 1/2 and 1/2,
so event C and event A are

00:39:17.960 --> 00:39:19.410 align:middle line:90%
independent.

00:39:19.410 --> 00:39:24.490 align:middle line:84%
Knowing that the first toss was
heads does not change your

00:39:24.490 --> 00:39:28.600 align:middle line:84%
beliefs about whether the two
tosses are going to have the

00:39:28.600 --> 00:39:31.380 align:middle line:90%
same outcome or not.

00:39:31.380 --> 00:39:34.200 align:middle line:84%
Knowing that the first was
heads, well, the second is

00:39:34.200 --> 00:39:36.520 align:middle line:84%
equally likely to be
heads or tails.

00:39:36.520 --> 00:39:39.710 align:middle line:84%
So event C has just the
same probability,

00:39:39.710 --> 00:39:42.140 align:middle line:90%
again, 1/2, to occur.

00:39:42.140 --> 00:39:46.110 align:middle line:84%
To put it the opposite way,
if I tell you that the two

00:39:46.110 --> 00:39:47.860 align:middle line:90%
results were the same--

00:39:47.860 --> 00:39:51.130 align:middle line:84%
so it's either heads, heads
or tails, tails--

00:39:51.130 --> 00:39:53.070 align:middle line:84%
what does that tell you
about the first toss?

00:39:53.070 --> 00:39:54.800 align:middle line:90%
Is it heads, or is it tails?

00:39:54.800 --> 00:39:56.570 align:middle line:84%
Well, it doesn't tell
you anything.

00:39:56.570 --> 00:39:59.700 align:middle line:84%
It could be either over the
two, so the probability of

00:39:59.700 --> 00:40:04.490 align:middle line:84%
heads in the first toss is equal
to 1/2, and telling you

00:40:04.490 --> 00:40:07.460 align:middle line:84%
C occurred does not
change anything.

00:40:07.460 --> 00:40:10.830 align:middle line:84%
So this is an example that
illustrates the case where we

00:40:10.830 --> 00:40:14.650 align:middle line:84%
have three events in which
we check that pairwise

00:40:14.650 --> 00:40:18.140 align:middle line:84%
independence holds for
any combination of

00:40:18.140 --> 00:40:19.250 align:middle line:90%
two of these events.

00:40:19.250 --> 00:40:21.900 align:middle line:84%
We have the probability of their
intersection is equal to

00:40:21.900 --> 00:40:23.760 align:middle line:84%
the product of their
probabilities.

00:40:23.760 --> 00:40:27.930 align:middle line:84%
On the other hand, the three
events taken all together are

00:40:27.930 --> 00:40:29.500 align:middle line:90%
not independent.

00:40:29.500 --> 00:40:32.780 align:middle line:84%
A doesn't tell me anything
useful, whether C is going to

00:40:32.780 --> 00:40:34.710 align:middle line:90%
occur or not.

00:40:34.710 --> 00:40:36.730 align:middle line:84%
B doesn't tell me
anything useful.

00:40:36.730 --> 00:40:40.840 align:middle line:84%
But if I tell you that both A
and B occurred, the two of

00:40:40.840 --> 00:40:44.150 align:middle line:84%
them together tell me something
useful about C.

00:40:44.150 --> 00:40:47.165 align:middle line:84%
Namely, they tell me that C
certainly has occurred.

00:40:47.165 --> 00:40:49.750 align:middle line:90%


00:40:49.750 --> 00:40:51.000 align:middle line:90%
Very good.

00:40:51.000 --> 00:40:53.900 align:middle line:90%


00:40:53.900 --> 00:40:56.890 align:middle line:84%
So independence is this somewhat
subtle concept.

00:40:56.890 --> 00:40:59.710 align:middle line:84%
Once you grasp the intuition of
what it really means, then

00:40:59.710 --> 00:41:02.910 align:middle line:90%
things perhaps fall in place.

00:41:02.910 --> 00:41:06.630 align:middle line:84%
But it's a concept where
it's easy to get some

00:41:06.630 --> 00:41:07.430 align:middle line:90%
misunderstanding.

00:41:07.430 --> 00:41:11.370 align:middle line:84%
So just take some
time to digest.

00:41:11.370 --> 00:41:14.810 align:middle line:84%
So to lighten things up, I'm
going to spend the remaining

00:41:14.810 --> 00:41:18.810 align:middle line:84%
four minutes talking about the
very nice, simple problem that

00:41:18.810 --> 00:41:23.240 align:middle line:84%
involves conditional
probabilities and the like.

00:41:23.240 --> 00:41:28.140 align:middle line:84%
So here's the problem,
formulated exactly as it shows

00:41:28.140 --> 00:41:30.250 align:middle line:90%
up in various textbooks.

00:41:30.250 --> 00:41:31.780 align:middle line:84%
And the formulation says
the following.

00:41:31.780 --> 00:41:35.310 align:middle line:84%
Well, consider one of those
anachronistic places where

00:41:35.310 --> 00:41:40.050 align:middle line:84%
they still have kings or queens,
and where actually

00:41:40.050 --> 00:41:43.090 align:middle line:84%
boys take precedence
over girls.

00:41:43.090 --> 00:41:44.600 align:middle line:90%
So if there is a boy--

00:41:44.600 --> 00:41:47.280 align:middle line:90%


00:41:47.280 --> 00:41:52.400 align:middle line:84%
if the royal family has a boy,
then he will become the king

00:41:52.400 --> 00:41:58.080 align:middle line:84%
even if he has an older sister
who might be the queen.

00:41:58.080 --> 00:42:02.930 align:middle line:84%
So we have one of those
royal families.

00:42:02.930 --> 00:42:06.810 align:middle line:84%
That royal family had two
children, and we know that

00:42:06.810 --> 00:42:08.060 align:middle line:90%
there is a king.

00:42:08.060 --> 00:42:11.370 align:middle line:90%


00:42:11.370 --> 00:42:14.250 align:middle line:84%
There is a king, which means
that at least one of the two

00:42:14.250 --> 00:42:16.030 align:middle line:90%
children was a boy.

00:42:16.030 --> 00:42:18.970 align:middle line:84%
Otherwise we wouldn't
have a king.

00:42:18.970 --> 00:42:21.885 align:middle line:84%
What is the probability that the
king's sibling is female?

00:42:21.885 --> 00:42:24.920 align:middle line:90%


00:42:24.920 --> 00:42:26.170 align:middle line:90%
OK.

00:42:26.170 --> 00:42:28.260 align:middle line:90%


00:42:28.260 --> 00:42:30.830 align:middle line:84%
I guess we need to make some
assumptions about genetics.

00:42:30.830 --> 00:42:33.910 align:middle line:84%
Let's assume that every child
is a boy or a girl with

00:42:33.910 --> 00:42:39.440 align:middle line:84%
probability 1/2, and that
different children, what they

00:42:39.440 --> 00:42:42.900 align:middle line:84%
are is independent from what
the other children were.

00:42:42.900 --> 00:42:47.660 align:middle line:84%
So every childbirth is basically
a coin flip.

00:42:47.660 --> 00:42:50.740 align:middle line:84%
OK, so if you take that,
you say, well,

00:42:50.740 --> 00:42:52.980 align:middle line:90%
the king is a child.

00:42:52.980 --> 00:42:55.890 align:middle line:90%
His sibling is another child.

00:42:55.890 --> 00:42:58.450 align:middle line:84%
Children are independent
of each other.

00:42:58.450 --> 00:43:05.860 align:middle line:84%
So the probability that the
sibling is a girl is 1/2.

00:43:05.860 --> 00:43:07.620 align:middle line:90%
That's the naive answer.

00:43:07.620 --> 00:43:09.270 align:middle line:84%
Now let's try to
do it formally.

00:43:09.270 --> 00:43:12.410 align:middle line:84%
Let's set up a model
of the experiment.

00:43:12.410 --> 00:43:15.650 align:middle line:84%
The royal family had two
children, as we we're told, so

00:43:15.650 --> 00:43:17.020 align:middle line:90%
there's four outcomes--

00:43:17.020 --> 00:43:22.040 align:middle line:84%
boy boy, boy girl, girl
boy, and girl girl.

00:43:22.040 --> 00:43:26.520 align:middle line:84%
Now, we are told that there is
a king, which means what?

00:43:26.520 --> 00:43:29.530 align:middle line:84%
This outcome here
did not happen.

00:43:29.530 --> 00:43:30.760 align:middle line:90%
It is not possible.

00:43:30.760 --> 00:43:33.810 align:middle line:84%
There are three outcomes
that remain possible.

00:43:33.810 --> 00:43:37.940 align:middle line:84%
So this is our conditional
sample space given

00:43:37.940 --> 00:43:40.500 align:middle line:90%
that there is king.

00:43:40.500 --> 00:43:43.170 align:middle line:84%
What are the probabilities
for the original model?

00:43:43.170 --> 00:43:46.420 align:middle line:84%
Well with the model that we
assume that every child is a

00:43:46.420 --> 00:43:50.885 align:middle line:84%
boy or a girl independently with
probability 1/2, then the

00:43:50.885 --> 00:43:54.950 align:middle line:84%
four outcomes would be equally
likely, and they're like this.

00:43:54.950 --> 00:43:57.110 align:middle line:84%
These are the original
probabilities.

00:43:57.110 --> 00:44:00.810 align:middle line:84%
But once we are told that this
outcome did not happen,

00:44:00.810 --> 00:44:03.900 align:middle line:84%
because we have a king, then
we are transported to the

00:44:03.900 --> 00:44:05.830 align:middle line:90%
smaller sample space.

00:44:05.830 --> 00:44:08.380 align:middle line:84%
In this sample space, what's
the probability that the

00:44:08.380 --> 00:44:10.360 align:middle line:90%
sibling is a girl?

00:44:10.360 --> 00:44:15.160 align:middle line:84%
Well the sibling is a girl in
two out of the three outcomes.

00:44:15.160 --> 00:44:17.290 align:middle line:84%
So the probability that
the sibling is a

00:44:17.290 --> 00:44:21.880 align:middle line:90%
girl is actually 2/3.

00:44:21.880 --> 00:44:25.780 align:middle line:84%
So that's supposed to
be the right answer.

00:44:25.780 --> 00:44:29.620 align:middle line:84%
Maybe a little
counter-intuitive.

00:44:29.620 --> 00:44:32.960 align:middle line:84%
So you can play smart and say,
oh I understand such problems

00:44:32.960 --> 00:44:35.990 align:middle line:84%
better than you, here is a trick
problem and here's why

00:44:35.990 --> 00:44:37.800 align:middle line:90%
the answer is 2/3.

00:44:37.800 --> 00:44:41.300 align:middle line:84%
But actually I'm not fully
justified in saying that the

00:44:41.300 --> 00:44:42.930 align:middle line:90%
answer is 2/3.

00:44:42.930 --> 00:44:46.520 align:middle line:84%
I made lots of hidden
assumptions when I put this

00:44:46.520 --> 00:44:50.040 align:middle line:84%
model down, which I
didn't yet state.

00:44:50.040 --> 00:44:54.960 align:middle line:84%
So to reverse engineer this
answer, let's actually think

00:44:54.960 --> 00:44:57.960 align:middle line:84%
what's the probability model for
which this would have been

00:44:57.960 --> 00:44:59.320 align:middle line:90%
the right answer.

00:44:59.320 --> 00:45:01.300 align:middle line:84%
And here's the probability
model.

00:45:01.300 --> 00:45:02.800 align:middle line:90%
The royal family--

00:45:02.800 --> 00:45:07.050 align:middle line:84%
the royal parents decided to
have exactly two children.

00:45:07.050 --> 00:45:08.960 align:middle line:90%
They went and had them.

00:45:08.960 --> 00:45:11.670 align:middle line:84%
It turned out that at
least one was a boy

00:45:11.670 --> 00:45:13.390 align:middle line:90%
and became a king.

00:45:13.390 --> 00:45:15.580 align:middle line:90%
Under this scenario--

00:45:15.580 --> 00:45:18.070 align:middle line:84%
that they decide to have
exactly two children--

00:45:18.070 --> 00:45:20.840 align:middle line:84%
then this is the big
sample space.

00:45:20.840 --> 00:45:23.350 align:middle line:84%
It turned out that
one was a boy.

00:45:23.350 --> 00:45:25.560 align:middle line:90%
That eliminates this outcome.

00:45:25.560 --> 00:45:27.410 align:middle line:84%
And then this picture
is correct and this

00:45:27.410 --> 00:45:28.750 align:middle line:90%
is the right answer.

00:45:28.750 --> 00:45:31.680 align:middle line:84%
But there's hidden assumptions
being there.

00:45:31.680 --> 00:45:35.170 align:middle line:84%
How about if the royal
family had followed

00:45:35.170 --> 00:45:37.230 align:middle line:90%
the following strategy?

00:45:37.230 --> 00:45:41.760 align:middle line:84%
We're going to have children
until we get a boy, so that we

00:45:41.760 --> 00:45:45.700 align:middle line:84%
get a king, and then
we'll stop.

00:45:45.700 --> 00:45:47.660 align:middle line:84%
OK, given they have two
children, what's the

00:45:47.660 --> 00:45:50.660 align:middle line:84%
probability that the
sibling is a girl?

00:45:50.660 --> 00:45:51.880 align:middle line:90%
It's 1.

00:45:51.880 --> 00:45:55.260 align:middle line:84%
The reason that they had two
children was because the first

00:45:55.260 --> 00:45:57.800 align:middle line:84%
was a girl, so they had
to have a second.

00:45:57.800 --> 00:46:00.820 align:middle line:84%
So assumptions about
reproductive practices

00:46:00.820 --> 00:46:03.130 align:middle line:84%
actually need to come in,
and they're going

00:46:03.130 --> 00:46:04.630 align:middle line:90%
to affect the decisions.

00:46:04.630 --> 00:46:08.010 align:middle line:84%
Or, if it's one of those ancient
kingdoms where a king

00:46:08.010 --> 00:46:11.790 align:middle line:84%
would always make sure too
strangle any of his brothers,

00:46:11.790 --> 00:46:15.560 align:middle line:84%
then the probability that the
sibling is a girl is actually

00:46:15.560 --> 00:46:17.570 align:middle line:90%
1 again, and so on.

00:46:17.570 --> 00:46:20.590 align:middle line:84%
So it means that one needs to be
careful when you start with

00:46:20.590 --> 00:46:24.330 align:middle line:84%
loosely worded problems to
make sure exactly what it

00:46:24.330 --> 00:46:26.950 align:middle line:84%
means and what assumptions
you're making.

00:46:26.950 --> 00:46:28.880 align:middle line:90%
All right, see you next week.

00:46:28.880 --> 00:46:30.130 align:middle line:90%