WEBVTT

00:00:00.000 --> 00:00:00.040 align:middle line:90%


00:00:00.040 --> 00:00:02.460 align:middle line:84%
The following content is
provided under a Creative

00:00:02.460 --> 00:00:03.870 align:middle line:90%
Commons license.

00:00:03.870 --> 00:00:06.320 align:middle line:84%
Your support will help
MIT OpenCourseWare

00:00:06.320 --> 00:00:10.560 align:middle line:84%
continue to offer high quality
educational resources for free.

00:00:10.560 --> 00:00:13.300 align:middle line:84%
To make a donation or
view additional materials

00:00:13.300 --> 00:00:17.210 align:middle line:84%
from hundreds of MIT courses,
visit MIT OpenCourseWare

00:00:17.210 --> 00:00:17.865 align:middle line:90%
at ocw.mit.edu.

00:00:17.865 --> 00:00:22.440 align:middle line:90%


00:00:22.440 --> 00:00:28.500 align:middle line:84%
PROFESSOR: OK, so
welcome to 6.041/6.431,

00:00:28.500 --> 00:00:31.750 align:middle line:84%
the class on probability
models and the like.

00:00:31.750 --> 00:00:32.740 align:middle line:90%
I'm John Tsitsiklis.

00:00:32.740 --> 00:00:35.480 align:middle line:84%
I will be teaching
this class, and I'm

00:00:35.480 --> 00:00:39.760 align:middle line:84%
looking forward to this being
an enjoyable and also useful

00:00:39.760 --> 00:00:41.060 align:middle line:90%
experience.

00:00:41.060 --> 00:00:43.370 align:middle line:84%
We have a fair amount
of staff involved

00:00:43.370 --> 00:00:46.280 align:middle line:84%
in this course, your
recitation instructors

00:00:46.280 --> 00:00:50.140 align:middle line:84%
and also a bunch of TAs, but
I want to single out our head

00:00:50.140 --> 00:00:54.450 align:middle line:84%
TA, Uzoma, who is the
key person in this class.

00:00:54.450 --> 00:00:56.550 align:middle line:84%
Everything has to
go through him.

00:00:56.550 --> 00:00:59.080 align:middle line:84%
If he doesn't know in
which recitation section

00:00:59.080 --> 00:01:03.600 align:middle line:84%
you are, then simply you do not
exist, so keep that in mind.

00:01:03.600 --> 00:01:04.099 align:middle line:90%
All right.

00:01:04.099 --> 00:01:07.770 align:middle line:84%
So we want to jump
right into the subject,

00:01:07.770 --> 00:01:09.630 align:middle line:84%
but I'm going to take
just a few minutes

00:01:09.630 --> 00:01:12.600 align:middle line:84%
to talk about a few
administrative details

00:01:12.600 --> 00:01:14.580 align:middle line:90%
and how the course is run.

00:01:14.580 --> 00:01:17.650 align:middle line:84%
So we're going to have
lectures twice a week

00:01:17.650 --> 00:01:20.300 align:middle line:84%
and I'm going to use old
fashioned transparencies.

00:01:20.300 --> 00:01:23.270 align:middle line:84%
Now, you get copies of these
slides with plenty of space

00:01:23.270 --> 00:01:25.760 align:middle line:90%
for you to keep notes on them.

00:01:25.760 --> 00:01:29.670 align:middle line:84%
A useful way of making
good use of the slides

00:01:29.670 --> 00:01:33.920 align:middle line:84%
is to use them as a sort
of mnemonic summary of what

00:01:33.920 --> 00:01:35.720 align:middle line:90%
happens in lecture.

00:01:35.720 --> 00:01:37.510 align:middle line:84%
Not everything that
I'm going to say

00:01:37.510 --> 00:01:40.490 align:middle line:84%
is, of course, on the
slides, but by looking them

00:01:40.490 --> 00:01:42.760 align:middle line:84%
you get the sense of
what's happening right now.

00:01:42.760 --> 00:01:45.070 align:middle line:84%
And it may be a good
idea to review them

00:01:45.070 --> 00:01:47.061 align:middle line:90%
before you go to recitation.

00:01:47.061 --> 00:01:48.310 align:middle line:90%
So what happens in recitation?

00:01:48.310 --> 00:01:51.090 align:middle line:84%
In recitation, your
recitation instructor

00:01:51.090 --> 00:01:53.990 align:middle line:84%
is going to maybe review
some of the theory

00:01:53.990 --> 00:01:57.150 align:middle line:84%
and then solve some
problems for you.

00:01:57.150 --> 00:01:59.260 align:middle line:84%
And then you have
tutorials where

00:01:59.260 --> 00:02:02.750 align:middle line:84%
you meet in very small
groups together with your TA.

00:02:02.750 --> 00:02:05.330 align:middle line:84%
And what happens in tutorials
is that you actually

00:02:05.330 --> 00:02:07.960 align:middle line:84%
do the problem solving
with the help of your TA

00:02:07.960 --> 00:02:12.290 align:middle line:84%
and the help of your classmates
in your tutorial section.

00:02:12.290 --> 00:02:14.340 align:middle line:84%
Now probability is
a tricky subject.

00:02:14.340 --> 00:02:16.750 align:middle line:84%
You may be reading the
text, listening to lectures,

00:02:16.750 --> 00:02:19.630 align:middle line:84%
everything makes perfect
sense, and so on,

00:02:19.630 --> 00:02:23.100 align:middle line:84%
but until you actually sit
down and try to solve problems,

00:02:23.100 --> 00:02:26.160 align:middle line:84%
you don't quite appreciate the
subtleties and the difficulties

00:02:26.160 --> 00:02:27.310 align:middle line:90%
that are involved.

00:02:27.310 --> 00:02:30.550 align:middle line:84%
So problem solving is a
key part of this class.

00:02:30.550 --> 00:02:33.916 align:middle line:84%
And tutorials are extremely
useful just for this reason

00:02:33.916 --> 00:02:35.290 align:middle line:84%
because that's
where you actually

00:02:35.290 --> 00:02:38.120 align:middle line:84%
get the practice of solving
problems on your own,

00:02:38.120 --> 00:02:42.090 align:middle line:84%
as opposed to seeing someone
else who's solving them

00:02:42.090 --> 00:02:43.510 align:middle line:90%
for you.

00:02:43.510 --> 00:02:47.270 align:middle line:84%
OK but, mechanics, a key part
of what's going to happen today

00:02:47.270 --> 00:02:51.500 align:middle line:84%
is that you will turn in
your schedule forms that

00:02:51.500 --> 00:02:55.350 align:middle line:84%
are at the end of the handout
that you have in your hands.

00:02:55.350 --> 00:02:59.820 align:middle line:84%
Then, the TAs will be working
frantically through the night,

00:02:59.820 --> 00:03:03.660 align:middle line:84%
and they're going to be
producing a list of who

00:03:03.660 --> 00:03:05.700 align:middle line:90%
goes into what section.

00:03:05.700 --> 00:03:09.330 align:middle line:84%
And when that happens,
any person in this class,

00:03:09.330 --> 00:03:11.410 align:middle line:84%
with probability
90%, is going to be

00:03:11.410 --> 00:03:15.520 align:middle line:84%
happy with their assignment
and, with probability 10%,

00:03:15.520 --> 00:03:17.670 align:middle line:90%
they're going to be unhappy.

00:03:17.670 --> 00:03:20.860 align:middle line:84%
Now, unhappy people
have an option, though.

00:03:20.860 --> 00:03:23.230 align:middle line:84%
You can resubmit
your form together

00:03:23.230 --> 00:03:25.280 align:middle line:84%
with your full schedule
and constraints,

00:03:25.280 --> 00:03:28.150 align:middle line:84%
give it back to the
head TA, who will then

00:03:28.150 --> 00:03:32.160 align:middle line:84%
do some further juggling
and reassign people,

00:03:32.160 --> 00:03:36.040 align:middle line:84%
and after that happens,
90% of those unhappy people

00:03:36.040 --> 00:03:37.570 align:middle line:90%
will become happy.

00:03:37.570 --> 00:03:42.270 align:middle line:84%
And 10% of them will
be less unhappy.

00:03:42.270 --> 00:03:42.840 align:middle line:90%
OK.

00:03:42.840 --> 00:03:46.350 align:middle line:84%
So what's the probability
that a random person

00:03:46.350 --> 00:03:49.800 align:middle line:84%
is going to be unhappy at
the end of this process?

00:03:49.800 --> 00:03:50.780 align:middle line:90%
It's 1%.

00:03:50.780 --> 00:03:51.330 align:middle line:90%
Excellent.

00:03:51.330 --> 00:03:51.829 align:middle line:90%
Good.

00:03:51.829 --> 00:03:53.200 align:middle line:90%
Maybe you don't need this class.

00:03:53.200 --> 00:03:54.340 align:middle line:90%
OK, so 1%.

00:03:54.340 --> 00:03:56.810 align:middle line:84%
We have about 100
people in this class,

00:03:56.810 --> 00:03:59.590 align:middle line:84%
so there's going to be
about one unhappy person.

00:03:59.590 --> 00:04:02.630 align:middle line:84%
I mean, anywhere you look
in life, in any group

00:04:02.630 --> 00:04:05.370 align:middle line:84%
you look at, there's always
one unhappy person, right?

00:04:05.370 --> 00:04:09.060 align:middle line:90%
So, what can we do about it?

00:04:09.060 --> 00:04:09.660 align:middle line:90%
All right.

00:04:09.660 --> 00:04:11.690 align:middle line:84%
Another important
part about mechanics

00:04:11.690 --> 00:04:13.290 align:middle line:84%
is to read carefully
the statement

00:04:13.290 --> 00:04:15.050 align:middle line:84%
that we have about
collaboration,

00:04:15.050 --> 00:04:17.019 align:middle line:90%
academic honesty, and all that.

00:04:17.019 --> 00:04:19.070 align:middle line:84%
You're encouraged,
it's a very good idea

00:04:19.070 --> 00:04:21.140 align:middle line:90%
to work with other students.

00:04:21.140 --> 00:04:24.550 align:middle line:84%
You can consult sources
that are out there, but when

00:04:24.550 --> 00:04:26.870 align:middle line:84%
you sit down and
write your solutions

00:04:26.870 --> 00:04:29.510 align:middle line:84%
you have to do that by
setting things aside

00:04:29.510 --> 00:04:32.050 align:middle line:90%
and just write them on your own.

00:04:32.050 --> 00:04:37.040 align:middle line:84%
You cannot copy something that
somebody else has given to you.

00:04:37.040 --> 00:04:39.790 align:middle line:84%
One reason is that
we're not going

00:04:39.790 --> 00:04:43.750 align:middle line:84%
to like it when it happens,
and then another reason

00:04:43.750 --> 00:04:46.212 align:middle line:84%
is that you're not going
to do yourself any favor.

00:04:46.212 --> 00:04:48.045 align:middle line:84%
Really the only way to
do well in this class

00:04:48.045 --> 00:04:51.620 align:middle line:84%
is to get a lot of practice by
solving problems yourselves.

00:04:51.620 --> 00:04:53.470 align:middle line:84%
So if you don't do
that on your own,

00:04:53.470 --> 00:04:56.880 align:middle line:84%
then when quiz and
exam time comes,

00:04:56.880 --> 00:04:59.070 align:middle line:84%
things are going
to be difficult.

00:04:59.070 --> 00:05:01.100 align:middle line:84%
So, as I mentioned
here, we're going

00:05:01.100 --> 00:05:04.520 align:middle line:84%
to have recitation
sections, that some of them

00:05:04.520 --> 00:05:08.220 align:middle line:84%
are for 6.041 students,
some are for 6.431 students,

00:05:08.220 --> 00:05:10.270 align:middle line:84%
the graduate section
of the class.

00:05:10.270 --> 00:05:14.170 align:middle line:84%
Now undergraduates can sit
in the graduate recitation

00:05:14.170 --> 00:05:14.690 align:middle line:90%
sections.

00:05:14.690 --> 00:05:16.850 align:middle line:84%
What's going to happen
there is that things

00:05:16.850 --> 00:05:18.780 align:middle line:84%
may be just a little
faster and you

00:05:18.780 --> 00:05:21.930 align:middle line:84%
may be covering a problem
that's a little more advanced

00:05:21.930 --> 00:05:24.670 align:middle line:84%
and is not covered in
the undergrad sections.

00:05:24.670 --> 00:05:27.470 align:middle line:84%
But if you sit in
the graduate section,

00:05:27.470 --> 00:05:29.040 align:middle line:84%
and you're an
undergraduate, you're

00:05:29.040 --> 00:05:33.130 align:middle line:84%
still just responsible for
the undergraduate material.

00:05:33.130 --> 00:05:35.500 align:middle line:84%
That is, you can just do
the undergraduate work

00:05:35.500 --> 00:05:37.429 align:middle line:84%
in the class, but
maybe be exposed

00:05:37.429 --> 00:05:38.470 align:middle line:90%
at the different section.

00:05:38.470 --> 00:05:41.070 align:middle line:90%


00:05:41.070 --> 00:05:43.036 align:middle line:90%
OK.

00:05:43.036 --> 00:05:46.220 align:middle line:84%
A few words about the
style of this class.

00:05:46.220 --> 00:05:50.760 align:middle line:84%
We want to focus on
basic ideas and concepts.

00:05:50.760 --> 00:05:52.620 align:middle line:84%
There's going to be
lots of formulas,

00:05:52.620 --> 00:05:54.710 align:middle line:84%
but what we try to
do in this class

00:05:54.710 --> 00:05:56.530 align:middle line:84%
is to actually
have you understand

00:05:56.530 --> 00:05:58.190 align:middle line:90%
what those formulas mean.

00:05:58.190 --> 00:06:01.260 align:middle line:84%
And, in a year from now when
almost all of the formulas

00:06:01.260 --> 00:06:03.510 align:middle line:84%
have been wiped out
from your memory,

00:06:03.510 --> 00:06:05.610 align:middle line:84%
you still have the
basic concepts.

00:06:05.610 --> 00:06:08.690 align:middle line:84%
You can understand them, so
when you look things up again,

00:06:08.690 --> 00:06:12.820 align:middle line:90%
they will still make sense.

00:06:12.820 --> 00:06:15.650 align:middle line:84%
It's not the plug and
chug kind of class

00:06:15.650 --> 00:06:19.210 align:middle line:84%
where you're given a list of
formulas, you're given numbers,

00:06:19.210 --> 00:06:21.470 align:middle line:84%
and you plug in and
you get answers.

00:06:21.470 --> 00:06:24.950 align:middle line:84%
The really hard part is
usually to choose which

00:06:24.950 --> 00:06:26.280 align:middle line:90%
formulas you're going to use.

00:06:26.280 --> 00:06:28.900 align:middle line:84%
You need judgment,
you need intuition.

00:06:28.900 --> 00:06:32.400 align:middle line:84%
Lots of probability problems,
at least the interesting ones,

00:06:32.400 --> 00:06:34.450 align:middle line:84%
often have lots of
different solutions.

00:06:34.450 --> 00:06:37.440 align:middle line:84%
Some are extremely long,
some are extremely short.

00:06:37.440 --> 00:06:40.430 align:middle line:84%
The extremely short ones
usually involve some kind

00:06:40.430 --> 00:06:43.860 align:middle line:84%
of deeper understanding of
what's going on so that you

00:06:43.860 --> 00:06:46.350 align:middle line:90%
can pick a shortcut and use it.

00:06:46.350 --> 00:06:49.110 align:middle line:84%
And hopefully you are
going to develop this skill

00:06:49.110 --> 00:06:51.630 align:middle line:90%
during this class.

00:06:51.630 --> 00:06:56.360 align:middle line:84%
Now, I could spend a lot of
time in this lecture talking

00:06:56.360 --> 00:06:58.570 align:middle line:84%
about why the
subject is important.

00:06:58.570 --> 00:07:02.270 align:middle line:84%
I'll keep it short because
I think it's almost obvious.

00:07:02.270 --> 00:07:05.650 align:middle line:84%
Anything that happens
in life is uncertain.

00:07:05.650 --> 00:07:08.910 align:middle line:84%
There's uncertainty anywhere,
so whatever you try to do,

00:07:08.910 --> 00:07:11.830 align:middle line:84%
you need to have some way
of dealing or thinking

00:07:11.830 --> 00:07:13.930 align:middle line:90%
about this uncertainty.

00:07:13.930 --> 00:07:16.280 align:middle line:84%
And the way to do that
in a systematic way

00:07:16.280 --> 00:07:17.955 align:middle line:84%
is by using the
models that are given

00:07:17.955 --> 00:07:20.110 align:middle line:90%
to us by probability theory.

00:07:20.110 --> 00:07:21.800 align:middle line:84%
So if you're an
engineer and you're

00:07:21.800 --> 00:07:24.630 align:middle line:84%
dealing with a communication
system or signal processing,

00:07:24.630 --> 00:07:28.440 align:middle line:84%
basically you're facing
a fight against noise.

00:07:28.440 --> 00:07:30.380 align:middle line:90%
Noise is random, is uncertain.

00:07:30.380 --> 00:07:31.450 align:middle line:90%
How do you model it?

00:07:31.450 --> 00:07:33.120 align:middle line:90%
How do you deal with it?

00:07:33.120 --> 00:07:35.210 align:middle line:84%
If you're a manager,
I guess you're

00:07:35.210 --> 00:07:38.410 align:middle line:84%
dealing with customer demand,
which is, of course, random.

00:07:38.410 --> 00:07:40.950 align:middle line:84%
Or you're dealing
with the stock market,

00:07:40.950 --> 00:07:42.820 align:middle line:90%
which is definitely random.

00:07:42.820 --> 00:07:48.190 align:middle line:84%
Or you play the casino, which
is, again, random, and so on.

00:07:48.190 --> 00:07:50.800 align:middle line:84%
And the same goes for
pretty much any other field

00:07:50.800 --> 00:07:52.880 align:middle line:90%
that you can think of.

00:07:52.880 --> 00:07:57.190 align:middle line:84%
But, independent of which
field you're coming from,

00:07:57.190 --> 00:08:00.630 align:middle line:84%
the basic concepts and tools
are really all the same.

00:08:00.630 --> 00:08:03.100 align:middle line:84%
So you may see in
bookstores that there

00:08:03.100 --> 00:08:05.630 align:middle line:84%
are books, probability
for scientists,

00:08:05.630 --> 00:08:07.670 align:middle line:84%
probability for
engineers, probability

00:08:07.670 --> 00:08:11.440 align:middle line:84%
for social scientists,
probability for astrologists.

00:08:11.440 --> 00:08:13.960 align:middle line:84%
Well, what all those
books have inside

00:08:13.960 --> 00:08:16.790 align:middle line:84%
them is exactly the same
models, the same equations,

00:08:16.790 --> 00:08:18.040 align:middle line:90%
the same problems.

00:08:18.040 --> 00:08:21.510 align:middle line:84%
They just make them somewhat
different word problems.

00:08:21.510 --> 00:08:24.780 align:middle line:84%
The basic concepts are
just one and the same,

00:08:24.780 --> 00:08:28.420 align:middle line:84%
and we'll take this
as an excuse for not

00:08:28.420 --> 00:08:31.960 align:middle line:84%
going too much into specific
domain applications.

00:08:31.960 --> 00:08:33.830 align:middle line:84%
We will have
problems and examples

00:08:33.830 --> 00:08:36.150 align:middle line:84%
that are motivated,
in some loose sense,

00:08:36.150 --> 00:08:38.140 align:middle line:90%
from real world situations.

00:08:38.140 --> 00:08:40.510 align:middle line:84%
But we're not really
trying in this class

00:08:40.510 --> 00:08:46.220 align:middle line:84%
to develop the skills for
domain-specific problems.

00:08:46.220 --> 00:08:49.660 align:middle line:84%
Rather, we're going to try to
stick to general understanding

00:08:49.660 --> 00:08:52.260 align:middle line:90%
of the subject.

00:08:52.260 --> 00:08:52.760 align:middle line:90%
OK.

00:08:52.760 --> 00:08:57.280 align:middle line:84%
So the next slide, of which
you do have in your handout,

00:08:57.280 --> 00:09:01.080 align:middle line:84%
gives you a few more
details about the class.

00:09:01.080 --> 00:09:03.590 align:middle line:84%
Maybe one thing to
comment here is that you

00:09:03.590 --> 00:09:06.370 align:middle line:90%
do need to read the text.

00:09:06.370 --> 00:09:08.560 align:middle line:84%
And with calculus
books, perhaps you

00:09:08.560 --> 00:09:11.100 align:middle line:84%
can live with a just a
two page summary of all

00:09:11.100 --> 00:09:13.190 align:middle line:84%
of the interesting
formulas in calculus,

00:09:13.190 --> 00:09:18.050 align:middle line:84%
and you can get by just
with those formulas.

00:09:18.050 --> 00:09:21.080 align:middle line:84%
But here, because we want to
develop concepts and intuition,

00:09:21.080 --> 00:09:24.700 align:middle line:84%
actually reading words, as
opposed to just browsing

00:09:24.700 --> 00:09:27.430 align:middle line:84%
through equations,
does make a difference.

00:09:27.430 --> 00:09:30.250 align:middle line:84%
In the beginning, the
class is kind of easy.

00:09:30.250 --> 00:09:32.420 align:middle line:84%
When we deal with
discrete probability,

00:09:32.420 --> 00:09:36.420 align:middle line:84%
that's the material until our
first quiz, and some of you

00:09:36.420 --> 00:09:39.790 align:middle line:84%
may get by without being too
systematic about following

00:09:39.790 --> 00:09:40.710 align:middle line:90%
the material.

00:09:40.710 --> 00:09:43.970 align:middle line:84%
But it does get substantially
harder afterwards.

00:09:43.970 --> 00:09:47.020 align:middle line:84%
And I would keep
restating that you

00:09:47.020 --> 00:09:52.460 align:middle line:84%
do have to read the text to
really understand the material.

00:09:52.460 --> 00:09:52.980 align:middle line:90%
OK.

00:09:52.980 --> 00:09:57.850 align:middle line:84%
So now we can start with the
real part of the lecture.

00:09:57.850 --> 00:10:01.670 align:middle line:90%
Let us set the goals for today.

00:10:01.670 --> 00:10:04.320 align:middle line:84%
So probability, or
probability theory,

00:10:04.320 --> 00:10:08.130 align:middle line:84%
is a framework for
dealing with uncertainty,

00:10:08.130 --> 00:10:10.250 align:middle line:84%
for dealing with
situations in which we

00:10:10.250 --> 00:10:12.200 align:middle line:90%
have some kind of randomness.

00:10:12.200 --> 00:10:15.860 align:middle line:84%
So what we want to do is, by
the end of today's lecture,

00:10:15.860 --> 00:10:18.780 align:middle line:84%
to give you anything
that you need

00:10:18.780 --> 00:10:21.980 align:middle line:84%
to know how to set up what
does it take to set up

00:10:21.980 --> 00:10:23.970 align:middle line:90%
a probabilistic model.

00:10:23.970 --> 00:10:28.150 align:middle line:84%
And what are the basic rules
of the game for dealing

00:10:28.150 --> 00:10:30.520 align:middle line:90%
with probabilistic models?

00:10:30.520 --> 00:10:32.230 align:middle line:84%
So, by the end of
this lecture, you

00:10:32.230 --> 00:10:35.290 align:middle line:84%
will have essentially recovered
half of this semester's

00:10:35.290 --> 00:10:36.860 align:middle line:90%
tuition, right?

00:10:36.860 --> 00:10:40.100 align:middle line:84%
So we're going to talk about
probabilistic models in more

00:10:40.100 --> 00:10:40.820 align:middle line:90%
detail--

00:10:40.820 --> 00:10:42.740 align:middle line:84%
the sample space,
which is basically

00:10:42.740 --> 00:10:44.500 align:middle line:84%
a description of
all the things that

00:10:44.500 --> 00:10:47.410 align:middle line:84%
may happen during a
random experiment,

00:10:47.410 --> 00:10:50.590 align:middle line:84%
and the probability law,
which describes our beliefs

00:10:50.590 --> 00:10:54.130 align:middle line:84%
about which outcomes are
more likely to occur compared

00:10:54.130 --> 00:10:56.080 align:middle line:90%
to other outcomes.

00:10:56.080 --> 00:10:59.130 align:middle line:84%
Probability laws have to obey
certain properties that we

00:10:59.130 --> 00:11:00.640 align:middle line:90%
call the axioms of probability.

00:11:00.640 --> 00:11:03.570 align:middle line:84%
So the main part
of today's lecture

00:11:03.570 --> 00:11:05.380 align:middle line:84%
is to describe
those axioms, which

00:11:05.380 --> 00:11:08.540 align:middle line:84%
are the rules of the
game, and consider

00:11:08.540 --> 00:11:12.770 align:middle line:90%
a few really trivial examples.

00:11:12.770 --> 00:11:15.370 align:middle line:84%
OK, so let's start
with our agenda.

00:11:15.370 --> 00:11:17.600 align:middle line:84%
The first piece in a
probabilistic model

00:11:17.600 --> 00:11:21.850 align:middle line:84%
is a description of the
sample space of an experiment.

00:11:21.850 --> 00:11:25.960 align:middle line:84%
So we do an experiment,
and by experiment we

00:11:25.960 --> 00:11:30.270 align:middle line:84%
just mean that just
something happens out there.

00:11:30.270 --> 00:11:34.200 align:middle line:84%
And that something that happens,
it could be flipping a coin,

00:11:34.200 --> 00:11:38.910 align:middle line:84%
or it could be rolling
a dice, or it could be

00:11:38.910 --> 00:11:41.550 align:middle line:90%
doing something in a card game.

00:11:41.550 --> 00:11:44.190 align:middle line:84%
So we fix a
particular experiment.

00:11:44.190 --> 00:11:48.780 align:middle line:84%
And we come up with a list of
all the possible things that

00:11:48.780 --> 00:11:51.090 align:middle line:84%
may happen during
this experiment.

00:11:51.090 --> 00:11:54.880 align:middle line:84%
So we write down a list of
all the possible outcomes.

00:11:54.880 --> 00:11:57.610 align:middle line:84%
So here's a list of all
the possible outcomes

00:11:57.610 --> 00:11:59.050 align:middle line:90%
of the experiment.

00:11:59.050 --> 00:12:01.720 align:middle line:84%
I use the word
"list," but, if you

00:12:01.720 --> 00:12:04.160 align:middle line:84%
want to be a little
more formal, it's better

00:12:04.160 --> 00:12:06.730 align:middle line:90%
to think of that list as a set.

00:12:06.730 --> 00:12:08.630 align:middle line:90%
So we have a set.

00:12:08.630 --> 00:12:11.000 align:middle line:90%
That set is our sample space.

00:12:11.000 --> 00:12:14.530 align:middle line:84%
And it's a set whose elements
are the possible outcomes

00:12:14.530 --> 00:12:15.920 align:middle line:90%
of the experiment.

00:12:15.920 --> 00:12:18.530 align:middle line:84%
So, for example, if you're
dealing with flipping a coin,

00:12:18.530 --> 00:12:22.380 align:middle line:84%
your sample space would be
heads, this is one outcome,

00:12:22.380 --> 00:12:24.450 align:middle line:90%
tails is one outcome.

00:12:24.450 --> 00:12:26.460 align:middle line:84%
And this set, which
has two elements,

00:12:26.460 --> 00:12:29.170 align:middle line:84%
is the sample space
of the experiment.

00:12:29.170 --> 00:12:29.670 align:middle line:90%
OK.

00:12:29.670 --> 00:12:32.570 align:middle line:84%
What do we need to
think about when we're

00:12:32.570 --> 00:12:34.430 align:middle line:90%
setting up the sample space?

00:12:34.430 --> 00:12:36.690 align:middle line:84%
First, the list should
be mutually exclusive,

00:12:36.690 --> 00:12:37.830 align:middle line:90%
collectively exhaustive.

00:12:37.830 --> 00:12:39.150 align:middle line:90%
What does that mean?

00:12:39.150 --> 00:12:41.070 align:middle line:84%
Collectively
exhaustive means that,

00:12:41.070 --> 00:12:43.740 align:middle line:84%
no matter what happens
in the experiment,

00:12:43.740 --> 00:12:47.700 align:middle line:84%
you're going to get one of
the outcomes inside here.

00:12:47.700 --> 00:12:51.010 align:middle line:84%
So you have not forgotten any
of the possibilities of what

00:12:51.010 --> 00:12:53.020 align:middle line:90%
may happen in the experiment.

00:12:53.020 --> 00:12:56.730 align:middle line:84%
Mutually exclusive means
that if this happens,

00:12:56.730 --> 00:12:58.870 align:middle line:90%
then that cannot happen.

00:12:58.870 --> 00:13:00.660 align:middle line:84%
So at the end of
the experiment, you

00:13:00.660 --> 00:13:05.610 align:middle line:84%
should be able to point out
to me just one, exactly one,

00:13:05.610 --> 00:13:10.540 align:middle line:84%
of these outcomes and say, this
is the outcome that happened.

00:13:10.540 --> 00:13:11.040 align:middle line:90%
OK.

00:13:11.040 --> 00:13:13.690 align:middle line:84%
So these are sort of
basic requirements.

00:13:13.690 --> 00:13:16.540 align:middle line:84%
There's another requirement
which is a little more loose.

00:13:16.540 --> 00:13:18.130 align:middle line:84%
When you set up
your sample space,

00:13:18.130 --> 00:13:19.930 align:middle line:84%
sometimes you do
have some freedom

00:13:19.930 --> 00:13:24.900 align:middle line:84%
about the details of how
you're going to describe it.

00:13:24.900 --> 00:13:27.020 align:middle line:84%
And the question
is, how much detail

00:13:27.020 --> 00:13:28.730 align:middle line:90%
are you going to include?

00:13:28.730 --> 00:13:31.410 align:middle line:84%
So let's take this coin
flipping experiment

00:13:31.410 --> 00:13:34.070 align:middle line:84%
and think of the
following sample space.

00:13:34.070 --> 00:13:37.550 align:middle line:84%
One possible outcome is heads,
a second possible outcome

00:13:37.550 --> 00:13:43.750 align:middle line:84%
is tails and it's raining,
and the third possible outcome

00:13:43.750 --> 00:13:45.500 align:middle line:90%
is tails and it's not raining.

00:13:45.500 --> 00:13:49.180 align:middle line:90%


00:13:49.180 --> 00:13:53.530 align:middle line:84%
So this is another possible
sample space for the experiment

00:13:53.530 --> 00:13:56.910 align:middle line:90%
where I flip a coin just once.

00:13:56.910 --> 00:13:58.330 align:middle line:90%
It's a legitimate one.

00:13:58.330 --> 00:14:01.430 align:middle line:84%
These three possibilities
are mutually exclusive

00:14:01.430 --> 00:14:03.470 align:middle line:90%
and collectively exhaustive.

00:14:03.470 --> 00:14:05.410 align:middle line:84%
Which one is the
right sample space?

00:14:05.410 --> 00:14:08.440 align:middle line:90%
Is it this one or that one?

00:14:08.440 --> 00:14:12.020 align:middle line:84%
Well, if you think that my
coin flipping inside this room

00:14:12.020 --> 00:14:15.230 align:middle line:84%
is completely unrelated
to the weather outside,

00:14:15.230 --> 00:14:18.470 align:middle line:84%
then you're going to stick
with this sample space.

00:14:18.470 --> 00:14:22.080 align:middle line:84%
If, on the other hand, you
have some superstitious belief

00:14:22.080 --> 00:14:25.510 align:middle line:84%
that maybe rain has
an effect on my coins,

00:14:25.510 --> 00:14:29.520 align:middle line:84%
you might work with the
sample space of this kind.

00:14:29.520 --> 00:14:32.040 align:middle line:84%
So you probably
wouldn't do that,

00:14:32.040 --> 00:14:35.370 align:middle line:84%
but it's a legitimate
option, strictly speaking.

00:14:35.370 --> 00:14:38.650 align:middle line:84%
Now this example is a little
bit on the frivolous side,

00:14:38.650 --> 00:14:40.600 align:middle line:84%
but the issue that
comes up here is

00:14:40.600 --> 00:14:43.270 align:middle line:84%
a basic one that
shows up anywhere

00:14:43.270 --> 00:14:44.700 align:middle line:90%
in science and engineering.

00:14:44.700 --> 00:14:48.150 align:middle line:84%
Whenever you're dealing with
a model or with a situation,

00:14:48.150 --> 00:14:50.645 align:middle line:84%
there are zillions of
details in that situation.

00:14:50.645 --> 00:14:52.550 align:middle line:84%
And when you come
up with a model,

00:14:52.550 --> 00:14:57.170 align:middle line:84%
you choose some of those details
that you keep in your model,

00:14:57.170 --> 00:15:00.060 align:middle line:84%
and some that you say,
well, these are irrelevant.

00:15:00.060 --> 00:15:03.780 align:middle line:84%
Or maybe there are small
effects, I can neglect them,

00:15:03.780 --> 00:15:05.970 align:middle line:84%
and you keep them
outside your model.

00:15:05.970 --> 00:15:08.660 align:middle line:84%
So when you go to
the real world,

00:15:08.660 --> 00:15:11.860 align:middle line:84%
there's definitely an element
of art and some judgment

00:15:11.860 --> 00:15:15.430 align:middle line:84%
that you need to do in order
to set up an appropriate sample

00:15:15.430 --> 00:15:15.930 align:middle line:90%
space.

00:15:15.930 --> 00:15:20.270 align:middle line:90%


00:15:20.270 --> 00:15:23.310 align:middle line:90%
So, an easy example now.

00:15:23.310 --> 00:15:25.770 align:middle line:84%
So of course, the
elementary examples

00:15:25.770 --> 00:15:29.420 align:middle line:90%
are coins, cards, and dice.

00:15:29.420 --> 00:15:30.840 align:middle line:90%
So let's deal with dice.

00:15:30.840 --> 00:15:34.550 align:middle line:84%
But to keep the diagram small,
instead of a six-sided die,

00:15:34.550 --> 00:15:38.270 align:middle line:84%
we're going to think about the
die that only has four faces.

00:15:38.270 --> 00:15:40.220 align:middle line:84%
So you can do that
with a tetrahedron,

00:15:40.220 --> 00:15:41.150 align:middle line:90%
doesn't really matter.

00:15:41.150 --> 00:15:43.520 align:middle line:84%
Basically, it's a die
that when you roll it,

00:15:43.520 --> 00:15:47.360 align:middle line:84%
you get a result which is
one, two, three or four.

00:15:47.360 --> 00:15:50.590 align:middle line:84%
However, the experiment that
I'm going to think about

00:15:50.590 --> 00:15:55.770 align:middle line:84%
will consist of two
rolls of a dice.

00:15:55.770 --> 00:15:57.600 align:middle line:90%
A crucial point here--

00:15:57.600 --> 00:15:59.890 align:middle line:84%
I'm rolling the
die twice, but I'm

00:15:59.890 --> 00:16:03.680 align:middle line:84%
thinking of this as
just one experiment, not

00:16:03.680 --> 00:16:08.540 align:middle line:84%
two different experiments,
not a repetition twice

00:16:08.540 --> 00:16:10.110 align:middle line:90%
of the same experiment.

00:16:10.110 --> 00:16:12.040 align:middle line:90%
So it's one big experiment.

00:16:12.040 --> 00:16:14.210 align:middle line:84%
During that big
experiment various things

00:16:14.210 --> 00:16:17.150 align:middle line:84%
could happen, such as
I'm rolling the die once,

00:16:17.150 --> 00:16:20.384 align:middle line:84%
and then I'm rolling
the die twice.

00:16:20.384 --> 00:16:22.450 align:middle line:90%
OK.

00:16:22.450 --> 00:16:25.280 align:middle line:84%
So what's the sample
space for that experiment?

00:16:25.280 --> 00:16:28.700 align:middle line:84%
Well, the sample space consists
of the possible outcomes.

00:16:28.700 --> 00:16:32.170 align:middle line:84%
One possible outcome
is that your first roll

00:16:32.170 --> 00:16:36.670 align:middle line:84%
resulted in two and the
second roll resulted in three.

00:16:36.670 --> 00:16:40.690 align:middle line:84%
In which case, the outcome
that you get is this one,

00:16:40.690 --> 00:16:42.840 align:middle line:90%
a two followed by three.

00:16:42.840 --> 00:16:45.840 align:middle line:90%
This is one possible outcome.

00:16:45.840 --> 00:16:49.280 align:middle line:84%
The way I'm describing
things, this outcome

00:16:49.280 --> 00:16:51.750 align:middle line:84%
is to be distinguished
from this outcome

00:16:51.750 --> 00:16:56.656 align:middle line:84%
here, where a three
is followed by two.

00:16:56.656 --> 00:16:59.680 align:middle line:84%
If you're playing
backgammon, it doesn't matter

00:16:59.680 --> 00:17:02.250 align:middle line:90%
which one of the two happened.

00:17:02.250 --> 00:17:05.540 align:middle line:84%
But if you're dealing
with a probabilistic model

00:17:05.540 --> 00:17:08.310 align:middle line:84%
that you want to keep track
of everything that happens

00:17:08.310 --> 00:17:11.359 align:middle line:84%
in this composite
experiment, there

00:17:11.359 --> 00:17:14.000 align:middle line:84%
are good reasons
for distinguishing

00:17:14.000 --> 00:17:15.859 align:middle line:90%
between these two outcomes.

00:17:15.859 --> 00:17:18.609 align:middle line:84%
I mean, when this happens,
it's definitely something

00:17:18.609 --> 00:17:20.220 align:middle line:90%
different from that happening.

00:17:20.220 --> 00:17:23.319 align:middle line:84%
A two followed by a three is
different from a three followed

00:17:23.319 --> 00:17:24.349 align:middle line:90%
by a two.

00:17:24.349 --> 00:17:27.700 align:middle line:84%
So this is the correct sample
space for this experiment

00:17:27.700 --> 00:17:29.890 align:middle line:90%
where we roll the die twice.

00:17:29.890 --> 00:17:33.240 align:middle line:84%
It has a total of 16
elements and it's, of course,

00:17:33.240 --> 00:17:35.840 align:middle line:90%
a finite set.

00:17:35.840 --> 00:17:39.850 align:middle line:84%
Sometimes, instead of
describing sample spaces

00:17:39.850 --> 00:17:43.950 align:middle line:84%
in terms of lists, or sets,
or diagrams of this kind,

00:17:43.950 --> 00:17:46.550 align:middle line:84%
it's useful to
describe the experiment

00:17:46.550 --> 00:17:48.660 align:middle line:90%
in some sequential way.

00:17:48.660 --> 00:17:50.370 align:middle line:84%
Whenever you have
an experiment that

00:17:50.370 --> 00:17:53.170 align:middle line:84%
consists of multiple
stages, it might

00:17:53.170 --> 00:17:57.860 align:middle line:84%
be useful, at least visually,
to give a diagram that shows you

00:17:57.860 --> 00:17:59.940 align:middle line:90%
how those stages evolve.

00:17:59.940 --> 00:18:03.720 align:middle line:84%
And that's what we do by
using a sequential description

00:18:03.720 --> 00:18:06.960 align:middle line:84%
or a tree-based
description by drawing

00:18:06.960 --> 00:18:09.690 align:middle line:84%
a tree of the
possible evolutions

00:18:09.690 --> 00:18:11.250 align:middle line:90%
during our experiment.

00:18:11.250 --> 00:18:14.890 align:middle line:84%
So in this tree, I'm thinking
of a first stage in which I

00:18:14.890 --> 00:18:18.655 align:middle line:84%
roll the first die, and there
are four possible results, one,

00:18:18.655 --> 00:18:20.520 align:middle line:90%
two, three and four.and 4.

00:18:20.520 --> 00:18:24.310 align:middle line:84%
And, given what happened,
let's say in the first roll,

00:18:24.310 --> 00:18:26.050 align:middle line:90%
suppose I got a one.

00:18:26.050 --> 00:18:28.310 align:middle line:84%
Then I'm rolling
the second dice,

00:18:28.310 --> 00:18:30.320 align:middle line:84%
and there are four
possibilities for what

00:18:30.320 --> 00:18:32.060 align:middle line:90%
may happen to the second die.

00:18:32.060 --> 00:18:36.010 align:middle line:84%
And the possible results are
one, tow, three and four again.

00:18:36.010 --> 00:18:38.860 align:middle line:84%
So what's the relation
between the two diagrams?

00:18:38.860 --> 00:18:41.150 align:middle line:84%
Well, for example,
the outcome two

00:18:41.150 --> 00:18:46.940 align:middle line:84%
followed by three corresponds
to this path on the tree.

00:18:46.940 --> 00:18:50.550 align:middle line:84%
So this path corresponds
to two followed by a three.

00:18:50.550 --> 00:18:53.910 align:middle line:84%
Any path is associated
to a particular outcome,

00:18:53.910 --> 00:18:57.360 align:middle line:84%
any outcome is associated
to a particular path.

00:18:57.360 --> 00:18:59.320 align:middle line:84%
And, instead of
paths, you may want

00:18:59.320 --> 00:19:01.990 align:middle line:84%
to think in terms of the
leaves of this diagram.

00:19:01.990 --> 00:19:04.870 align:middle line:84%
Same thing, think of
each one of the leaves

00:19:04.870 --> 00:19:07.980 align:middle line:90%
as being one possible outcome.

00:19:07.980 --> 00:19:10.340 align:middle line:84%
And of course we have
16 outcomes here,

00:19:10.340 --> 00:19:12.790 align:middle line:90%
we have 16 outcomes here.

00:19:12.790 --> 00:19:15.920 align:middle line:84%
Maybe you noticed the subtlety
that I used in my language.

00:19:15.920 --> 00:19:18.810 align:middle line:84%
I said I rolled the
first dice and the result

00:19:18.810 --> 00:19:20.580 align:middle line:90%
that I get is a two.

00:19:20.580 --> 00:19:22.660 align:middle line:90%
I didn't use the word "outcome."

00:19:22.660 --> 00:19:24.940 align:middle line:84%
I want to reserve
the word "outcome"

00:19:24.940 --> 00:19:28.810 align:middle line:84%
to mean the overall
outcome at the end

00:19:28.810 --> 00:19:30.570 align:middle line:90%
of the overall experiment.

00:19:30.570 --> 00:19:36.300 align:middle line:84%
So "2, 3" is the outcome
of the experiment.

00:19:36.300 --> 00:19:38.910 align:middle line:84%
The experiment
consisted of stages.

00:19:38.910 --> 00:19:40.970 align:middle line:84%
Two was the result
in the first stage,

00:19:40.970 --> 00:19:43.370 align:middle line:84%
three was the result
in the second stage.

00:19:43.370 --> 00:19:45.390 align:middle line:84%
You put all those
results together,

00:19:45.390 --> 00:19:47.520 align:middle line:90%
and you get your outcome.

00:19:47.520 --> 00:19:50.220 align:middle line:84%
OK, perhaps we are
splitting hairs here,

00:19:50.220 --> 00:19:56.470 align:middle line:84%
but it's useful to keep
the concepts right.

00:19:56.470 --> 00:19:58.540 align:middle line:84%
What's special
about this example

00:19:58.540 --> 00:20:00.840 align:middle line:84%
is that, besides
being trivial, it has

00:20:00.840 --> 00:20:03.230 align:middle line:90%
a sample space which is finite.

00:20:03.230 --> 00:20:06.000 align:middle line:84%
There's 16 possible
total outcomes.

00:20:06.000 --> 00:20:09.210 align:middle line:84%
Not every experiment has
a finite sample space.

00:20:09.210 --> 00:20:12.840 align:middle line:84%
Here's an experiment in which
the sample space is infinite.

00:20:12.840 --> 00:20:17.690 align:middle line:84%
So you are playing darts and
the target is this square.

00:20:17.690 --> 00:20:20.910 align:middle line:84%
And you're perfect at
that game, so you're

00:20:20.910 --> 00:20:26.010 align:middle line:84%
sure that your darts will
always fall inside the square.

00:20:26.010 --> 00:20:29.550 align:middle line:84%
So, but where exactly your dart
would fall inside that square,

00:20:29.550 --> 00:20:31.180 align:middle line:90%
that itself is random.

00:20:31.180 --> 00:20:32.880 align:middle line:84%
We don't know what
it's going to be.

00:20:32.880 --> 00:20:34.300 align:middle line:90%
It's uncertain.

00:20:34.300 --> 00:20:37.220 align:middle line:84%
So all the possible
points inside the square

00:20:37.220 --> 00:20:39.710 align:middle line:84%
are possible outcomes
of the experiment.

00:20:39.710 --> 00:20:43.060 align:middle line:84%
So a typical outcome of the
experiment is going to a pair

00:20:43.060 --> 00:20:47.180 align:middle line:84%
of numbers, x,y, where x and y
are real numbers between zero

00:20:47.180 --> 00:20:48.280 align:middle line:90%
and one.

00:20:48.280 --> 00:20:51.170 align:middle line:84%
Now there's infinitely
many real numbers,

00:20:51.170 --> 00:20:54.100 align:middle line:84%
there's infinitely many
points in the square,

00:20:54.100 --> 00:20:56.700 align:middle line:84%
so this is an example
in which our sample

00:20:56.700 --> 00:20:58.740 align:middle line:90%
space is an infinite set.

00:20:58.740 --> 00:21:01.670 align:middle line:90%


00:21:01.670 --> 00:21:06.910 align:middle line:84%
OK, so we're going to revisit
this example a little later.

00:21:06.910 --> 00:21:11.120 align:middle line:84%
So these are two examples of
what the sample space might

00:21:11.120 --> 00:21:13.730 align:middle line:90%
be in simple experiments.

00:21:13.730 --> 00:21:17.350 align:middle line:84%
Now, the more important
order of business

00:21:17.350 --> 00:21:20.030 align:middle line:84%
is now to look at
those possible outcomes

00:21:20.030 --> 00:21:21.800 align:middle line:84%
and to make some
statements about

00:21:21.800 --> 00:21:23.910 align:middle line:90%
their relative likelihoods.

00:21:23.910 --> 00:21:29.060 align:middle line:84%
Which outcome is more likely to
occur compared to the others?

00:21:29.060 --> 00:21:33.580 align:middle line:84%
And the way we do this is
by assigning probabilities

00:21:33.580 --> 00:21:36.210 align:middle line:90%
to the outcomes.

00:21:36.210 --> 00:21:38.590 align:middle line:90%
Well, not exactly.

00:21:38.590 --> 00:21:42.440 align:middle line:84%
Suppose that all you were to
do was to assign probabilities

00:21:42.440 --> 00:21:44.320 align:middle line:90%
to individual outcomes.

00:21:44.320 --> 00:21:49.760 align:middle line:84%
If you go back to this example,
and you consider one particular

00:21:49.760 --> 00:21:52.250 align:middle line:90%
outcome-- let's say this point--

00:21:52.250 --> 00:21:55.900 align:middle line:84%
what would be the probability
that you hit exactly this point

00:21:55.900 --> 00:21:58.640 align:middle line:90%
to infinite precision?

00:21:58.640 --> 00:22:01.070 align:middle line:84%
Intuitively, that
probability would be zero.

00:22:01.070 --> 00:22:06.090 align:middle line:84%
So any individual point in this
diagram in any reasonable model

00:22:06.090 --> 00:22:08.520 align:middle line:90%
should have zero probability.

00:22:08.520 --> 00:22:11.870 align:middle line:84%
So if you just tell me that
any individual outcome has

00:22:11.870 --> 00:22:13.920 align:middle line:84%
zero probability,
you're not really

00:22:13.920 --> 00:22:17.030 align:middle line:90%
telling me much to work with.

00:22:17.030 --> 00:22:19.890 align:middle line:84%
For that reason, what
instead we're going to do

00:22:19.890 --> 00:22:23.760 align:middle line:84%
is to assign probabilities
to subsets of the sample

00:22:23.760 --> 00:22:26.065 align:middle line:84%
space, as opposed to
assigning probabilities

00:22:26.065 --> 00:22:29.170 align:middle line:90%
to individual outcomes.

00:22:29.170 --> 00:22:32.410 align:middle line:90%
So here's the picture.

00:22:32.410 --> 00:22:36.230 align:middle line:84%
We have our sample
space, which is omega,

00:22:36.230 --> 00:22:39.690 align:middle line:84%
and we consider some
subset of the sample space.

00:22:39.690 --> 00:22:45.140 align:middle line:84%
Call it A. And I want
to assign a number,

00:22:45.140 --> 00:22:50.720 align:middle line:84%
a numerical probability, to
this particular subset which

00:22:50.720 --> 00:22:56.841 align:middle line:84%
represents my belief about how
likely this set is to occur.

00:22:56.841 --> 00:22:57.340 align:middle line:90%
OK.

00:22:57.340 --> 00:22:59.710 align:middle line:90%
What do we mean "to occur?"

00:22:59.710 --> 00:23:01.720 align:middle line:84%
And I'm introducing
here a language

00:23:01.720 --> 00:23:03.770 align:middle line:84%
that's being used in
probability theory.

00:23:03.770 --> 00:23:06.580 align:middle line:84%
When we talk about subsets
of the sample space,

00:23:06.580 --> 00:23:10.470 align:middle line:84%
we usually call them events,
as opposed to subsets.

00:23:10.470 --> 00:23:13.250 align:middle line:84%
And the reason is
because it works nicely

00:23:13.250 --> 00:23:16.710 align:middle line:84%
with the language that
describes what's going on.

00:23:16.710 --> 00:23:19.010 align:middle line:90%
So the outcome is a point.

00:23:19.010 --> 00:23:20.540 align:middle line:90%
The outcome is random.

00:23:20.540 --> 00:23:25.720 align:middle line:84%
The outcome may be inside
this set, in which case

00:23:25.720 --> 00:23:31.270 align:middle line:84%
we say that event A occurred, if
we get an outcome inside here.

00:23:31.270 --> 00:23:35.120 align:middle line:84%
Or the outcome may fall
outside the set, in which case

00:23:35.120 --> 00:23:38.530 align:middle line:84%
we say that event
A did not occur.

00:23:38.530 --> 00:23:42.310 align:middle line:84%
So we're going to assign
probabilities to events.

00:23:42.310 --> 00:23:45.630 align:middle line:84%
And now, how should
we do this assignment?

00:23:45.630 --> 00:23:48.840 align:middle line:84%
Well, probabilities are meant
to describe your beliefs

00:23:48.840 --> 00:23:52.880 align:middle line:84%
about which sets are more likely
to occur versus other sets.

00:23:52.880 --> 00:23:56.080 align:middle line:84%
So there's many ways that you
can assign those probabilities.

00:23:56.080 --> 00:23:59.290 align:middle line:84%
But there are some ground
rules for this game.

00:23:59.290 --> 00:24:03.310 align:middle line:84%
First, we want probabilities to
be numbers between zero and one

00:24:03.310 --> 00:24:06.740 align:middle line:84%
because that's the
usual convention.

00:24:06.740 --> 00:24:08.680 align:middle line:84%
So a probability
of zero means we're

00:24:08.680 --> 00:24:10.820 align:middle line:84%
certain that something
is not going to happen.

00:24:10.820 --> 00:24:12.460 align:middle line:84%
Probability of one
means that we're

00:24:12.460 --> 00:24:14.870 align:middle line:84%
essentially certain that
something's going to happen.

00:24:14.870 --> 00:24:17.450 align:middle line:84%
So we want numbers
between zero and one.

00:24:17.450 --> 00:24:19.740 align:middle line:90%
We also want a few other things.

00:24:19.740 --> 00:24:21.740 align:middle line:84%
And those few other
things are going

00:24:21.740 --> 00:24:25.060 align:middle line:84%
to be encapsulated
in a set of axioms.

00:24:25.060 --> 00:24:27.610 align:middle line:84%
What "axioms" means
in this context,

00:24:27.610 --> 00:24:31.740 align:middle line:84%
it's the ground rules that any
legitimate probabilistic model

00:24:31.740 --> 00:24:33.410 align:middle line:90%
should obey.

00:24:33.410 --> 00:24:37.080 align:middle line:84%
You have a choice of what
kind of probabilities you use.

00:24:37.080 --> 00:24:39.420 align:middle line:84%
But, no matter
what you use, they

00:24:39.420 --> 00:24:43.660 align:middle line:84%
should obey certain consistency
properties because if they

00:24:43.660 --> 00:24:45.220 align:middle line:84%
obey those properties,
then you can

00:24:45.220 --> 00:24:47.370 align:middle line:84%
go ahead and do
useful calculations

00:24:47.370 --> 00:24:49.360 align:middle line:90%
and do some useful reasoning.

00:24:49.360 --> 00:24:51.010 align:middle line:90%
So what are these properties?

00:24:51.010 --> 00:24:55.060 align:middle line:84%
First, probabilities
should be non-negative.

00:24:55.060 --> 00:24:56.590 align:middle line:90%
OK?

00:24:56.590 --> 00:24:57.530 align:middle line:90%
That's our convention.

00:24:57.530 --> 00:25:00.350 align:middle line:84%
We want probabilities to be
numbers between zero and one.

00:25:00.350 --> 00:25:02.130 align:middle line:84%
So they should certainly
be non-negative.

00:25:02.130 --> 00:25:04.170 align:middle line:84%
The probability
that event A occurs

00:25:04.170 --> 00:25:06.135 align:middle line:90%
should be a non-negative number.

00:25:06.135 --> 00:25:08.110 align:middle line:90%
What's the second axiom?

00:25:08.110 --> 00:25:13.760 align:middle line:84%
The probability of the entire
sample space is equal to one.

00:25:13.760 --> 00:25:15.590 align:middle line:90%
Why does this make sense?

00:25:15.590 --> 00:25:20.120 align:middle line:84%
Well, the outcome is certain
to be an element of the sample

00:25:20.120 --> 00:25:23.720 align:middle line:84%
space because we set up a sample
space, which is collectively

00:25:23.720 --> 00:25:24.660 align:middle line:90%
exhaustive.

00:25:24.660 --> 00:25:27.559 align:middle line:84%
No matter what the
outcome is, it's

00:25:27.559 --> 00:25:29.350 align:middle line:84%
going to be an element
of the sample space.

00:25:29.350 --> 00:25:33.710 align:middle line:84%
We're certain that event
omega is going to occur.

00:25:33.710 --> 00:25:36.700 align:middle line:84%
Therefore, we represent
this certainty

00:25:36.700 --> 00:25:41.520 align:middle line:84%
by saying that the probability
of omega is equal to one.

00:25:41.520 --> 00:25:47.180 align:middle line:90%
Pretty straightforward so far.

00:25:47.180 --> 00:25:52.240 align:middle line:84%
The more interesting
axiom is the third rule.

00:25:52.240 --> 00:25:55.580 align:middle line:84%
Before getting into it,
just a quick reminder.

00:25:55.580 --> 00:26:01.950 align:middle line:84%
If you have two sets, A and
B, the intersection of A and B

00:26:01.950 --> 00:26:07.220 align:middle line:84%
consists of those elements
that belong both to A and B.

00:26:07.220 --> 00:26:09.580 align:middle line:90%
And we denote it this way.

00:26:09.580 --> 00:26:11.060 align:middle line:84%
When you think
probabilistically,

00:26:11.060 --> 00:26:15.090 align:middle line:84%
the way to think of intersection
is by using the word "and."

00:26:15.090 --> 00:26:18.890 align:middle line:84%
This event, this
intersection, is the event

00:26:18.890 --> 00:26:22.450 align:middle line:90%
that A occurred and B occurred.

00:26:22.450 --> 00:26:25.200 align:middle line:84%
If I get an outcome inside
here, A has occurred

00:26:25.200 --> 00:26:27.950 align:middle line:84%
and B has occurred
at the same time.

00:26:27.950 --> 00:26:31.150 align:middle line:84%
So you may find the word "and"
to be a little more convenient

00:26:31.150 --> 00:26:33.680 align:middle line:90%
than the word "intersection."

00:26:33.680 --> 00:26:35.950 align:middle line:84%
And similarly, we
have some notation

00:26:35.950 --> 00:26:42.280 align:middle line:84%
for the union of two events,
which we write this way.

00:26:42.280 --> 00:26:44.930 align:middle line:84%
The union of two
sets, or two events,

00:26:44.930 --> 00:26:47.080 align:middle line:84%
is the collection
of all the elements

00:26:47.080 --> 00:26:49.990 align:middle line:84%
that belong either to the
first set, or to the second,

00:26:49.990 --> 00:26:51.400 align:middle line:90%
or to both.

00:26:51.400 --> 00:26:54.970 align:middle line:84%
When you talk about events,
you can use the word "or."

00:26:54.970 --> 00:26:59.990 align:middle line:84%
So this is the event that
A occurred or B occurred.

00:26:59.990 --> 00:27:03.700 align:middle line:84%
And this "or" means that it
could also be that both of them

00:27:03.700 --> 00:27:04.200 align:middle line:90%
occurred.

00:27:04.200 --> 00:27:08.650 align:middle line:90%


00:27:08.650 --> 00:27:09.150 align:middle line:90%
OK.

00:27:09.150 --> 00:27:10.870 align:middle line:84%
So now that we
have this notation,

00:27:10.870 --> 00:27:13.835 align:middle line:90%
what does the third axiom say?

00:27:13.835 --> 00:27:19.830 align:middle line:84%
The third axiom says that if
we have two events, A and B,

00:27:19.830 --> 00:27:23.140 align:middle line:90%
that have no common elements--

00:27:23.140 --> 00:27:29.180 align:middle line:84%
so here's A, here's
B, and perhaps this

00:27:29.180 --> 00:27:31.140 align:middle line:90%
is our big sample space.

00:27:31.140 --> 00:27:33.470 align:middle line:84%
The two events have
no common elements.

00:27:33.470 --> 00:27:36.510 align:middle line:84%
So the intersection of the
two events is the empty set.

00:27:36.510 --> 00:27:38.930 align:middle line:84%
There's nothing in
their intersection.

00:27:38.930 --> 00:27:42.780 align:middle line:84%
Then, the total probability
of A together with B

00:27:42.780 --> 00:27:46.600 align:middle line:84%
has to be equal to the sum of
the individual probabilities.

00:27:46.600 --> 00:27:50.020 align:middle line:84%
So the probability that
A occurs or B occurs

00:27:50.020 --> 00:27:52.930 align:middle line:84%
is equal to the probability that
A occurs plus the probability

00:27:52.930 --> 00:27:55.040 align:middle line:90%
that B occurs.

00:27:55.040 --> 00:27:58.860 align:middle line:84%
So think of probability
as being cream cheese.

00:27:58.860 --> 00:28:03.020 align:middle line:84%
You have one pound of cream
cheese, the total probability

00:28:03.020 --> 00:28:05.340 align:middle line:84%
assigned to the
entire sample space.

00:28:05.340 --> 00:28:12.780 align:middle line:84%
And that cream cheese is
spread out over this set.

00:28:12.780 --> 00:28:15.760 align:middle line:84%
The probability of A is
how much cream cheese

00:28:15.760 --> 00:28:20.120 align:middle line:84%
sits on top of A. Probability of
B is how much sits on top of B.

00:28:20.120 --> 00:28:25.370 align:middle line:84%
The probability of A union B
is the total amount of cream

00:28:25.370 --> 00:28:28.660 align:middle line:84%
cheese sitting on
top of this and that,

00:28:28.660 --> 00:28:31.510 align:middle line:84%
which is obviously the sum
of how much is sitting here

00:28:31.510 --> 00:28:33.220 align:middle line:90%
and how much is sitting there.

00:28:33.220 --> 00:28:35.670 align:middle line:84%
So probabilities behave
like cream cheese,

00:28:35.670 --> 00:28:38.450 align:middle line:90%
or they behave like mass.

00:28:38.450 --> 00:28:47.690 align:middle line:84%
For example, if you think
of some material object,

00:28:47.690 --> 00:28:50.510 align:middle line:84%
the mass of this set
consisting of two pieces

00:28:50.510 --> 00:28:53.120 align:middle line:84%
is obviously the sum
of the two masses.

00:28:53.120 --> 00:28:55.680 align:middle line:84%
So this property is
a very intuitive one.

00:28:55.680 --> 00:28:58.282 align:middle line:84%
It's a pretty
natural one to have.

00:28:58.282 --> 00:29:00.640 align:middle line:90%
OK.

00:29:00.640 --> 00:29:03.880 align:middle line:84%
Are these axioms enough
for what we want to do?

00:29:03.880 --> 00:29:07.670 align:middle line:84%
I mentioned a while ago that
we want probabilities to be

00:29:07.670 --> 00:29:10.069 align:middle line:90%
numbers between zero and one.

00:29:10.069 --> 00:29:12.110 align:middle line:84%
Here's an axiom that tells
you that probabilities

00:29:12.110 --> 00:29:13.710 align:middle line:90%
are non-negative.

00:29:13.710 --> 00:29:15.590 align:middle line:84%
Should we have
another axiom that

00:29:15.590 --> 00:29:21.670 align:middle line:84%
tells us that probabilities
are less than or equal to one?

00:29:21.670 --> 00:29:23.150 align:middle line:90%
It's a desirable property.

00:29:23.150 --> 00:29:26.090 align:middle line:84%
We would like to
have it in our hands.

00:29:26.090 --> 00:29:29.030 align:middle line:90%
OK, why is it not in that list?

00:29:29.030 --> 00:29:32.680 align:middle line:84%
Well, the people who are in
the axiom making business

00:29:32.680 --> 00:29:34.610 align:middle line:84%
are mathematicians
and mathematicians

00:29:34.610 --> 00:29:36.390 align:middle line:90%
tend to be pretty laconic.

00:29:36.390 --> 00:29:40.020 align:middle line:84%
You don't say something if
you don't have to say it.

00:29:40.020 --> 00:29:42.580 align:middle line:90%
And this is the case here.

00:29:42.580 --> 00:29:46.410 align:middle line:84%
We don't need that extra
axiom because we can derive it

00:29:46.410 --> 00:29:48.440 align:middle line:90%
from the existing axioms.

00:29:48.440 --> 00:29:50.590 align:middle line:90%
Here's how it goes.

00:29:50.590 --> 00:29:55.180 align:middle line:84%
One is the probability over
the entire sample space.

00:29:55.180 --> 00:29:57.450 align:middle line:84%
Here we're using
the second axiom.

00:29:57.450 --> 00:30:00.310 align:middle line:90%


00:30:00.310 --> 00:30:05.690 align:middle line:84%
Now the sample space
consists of A together

00:30:05.690 --> 00:30:07.680 align:middle line:90%
with the complement of A. OK?

00:30:07.680 --> 00:30:11.200 align:middle line:90%


00:30:11.200 --> 00:30:13.000 align:middle line:84%
When I write the
complement of A,

00:30:13.000 --> 00:30:16.800 align:middle line:84%
I mean the complement of
A inside of the set omega.

00:30:16.800 --> 00:30:21.700 align:middle line:84%
So we have omega, here's A,
here's the complement of A,

00:30:21.700 --> 00:30:24.660 align:middle line:90%
and the overall set is omega.

00:30:24.660 --> 00:30:25.350 align:middle line:90%
OK.

00:30:25.350 --> 00:30:27.520 align:middle line:90%
Now, what's the next step?

00:30:27.520 --> 00:30:28.650 align:middle line:90%
What should I do next?

00:30:28.650 --> 00:30:31.320 align:middle line:90%
Which axiom should I use?

00:30:31.320 --> 00:30:35.950 align:middle line:84%
We use axiom three because a set
and the complement of that set

00:30:35.950 --> 00:30:36.730 align:middle line:90%
are disjoint.

00:30:36.730 --> 00:30:38.770 align:middle line:84%
They don't have any
common elements.

00:30:38.770 --> 00:30:42.810 align:middle line:84%
So axiom three
applies and tells me

00:30:42.810 --> 00:30:45.530 align:middle line:84%
that this is the
probability of A

00:30:45.530 --> 00:30:48.150 align:middle line:84%
plus the probability
of A complement.

00:30:48.150 --> 00:30:52.090 align:middle line:84%
In particular, the
probability of A

00:30:52.090 --> 00:30:56.850 align:middle line:84%
is equal to one minus the
probability of A complement,

00:30:56.850 --> 00:31:00.321 align:middle line:84%
and this is less
than or equal to one.

00:31:00.321 --> 00:31:00.820 align:middle line:90%
Why?

00:31:00.820 --> 00:31:03.430 align:middle line:90%


00:31:03.430 --> 00:31:06.670 align:middle line:84%
Because probabilities
are non-negative,

00:31:06.670 --> 00:31:09.810 align:middle line:90%
by the first axiom.

00:31:09.810 --> 00:31:10.310 align:middle line:90%
OK.

00:31:10.310 --> 00:31:12.440 align:middle line:84%
So we got the conclusion
that we wanted.

00:31:12.440 --> 00:31:15.520 align:middle line:84%
Probabilities are always
less than or equal to one,

00:31:15.520 --> 00:31:18.250 align:middle line:84%
and this is a simple
consequence of the three axioms

00:31:18.250 --> 00:31:20.230 align:middle line:90%
that we have.

00:31:20.230 --> 00:31:23.130 align:middle line:84%
This is a really nice
argument because it actually

00:31:23.130 --> 00:31:26.560 align:middle line:90%
uses each one of those axioms.

00:31:26.560 --> 00:31:28.160 align:middle line:84%
The argument is
simple, but you have

00:31:28.160 --> 00:31:30.490 align:middle line:84%
to use all of these
three properties

00:31:30.490 --> 00:31:33.050 align:middle line:84%
to get the conclusion
that you want.

00:31:33.050 --> 00:31:33.720 align:middle line:90%
OK.

00:31:33.720 --> 00:31:37.140 align:middle line:84%
So we can get interesting
things out of our axioms.

00:31:37.140 --> 00:31:40.050 align:middle line:84%
Can we get some more
interesting ones?

00:31:40.050 --> 00:31:44.540 align:middle line:84%
How about the union
of three sets?

00:31:44.540 --> 00:31:47.000 align:middle line:84%
What kind of probability
should it have?

00:31:47.000 --> 00:31:52.870 align:middle line:84%
So here's an event
consisting of three pieces.

00:31:52.870 --> 00:31:55.500 align:middle line:84%
And I want to say something
about the probability

00:31:55.500 --> 00:32:01.080 align:middle line:84%
of A union B union C.
What I would like to say

00:32:01.080 --> 00:32:04.940 align:middle line:84%
is that this probability is
equal to the sum of the three

00:32:04.940 --> 00:32:07.140 align:middle line:90%
individual probabilities.

00:32:07.140 --> 00:32:08.860 align:middle line:90%
How can I do it?

00:32:08.860 --> 00:32:10.870 align:middle line:84%
I have an axiom
that tells me that I

00:32:10.870 --> 00:32:12.760 align:middle line:90%
can do it for two events.

00:32:12.760 --> 00:32:15.370 align:middle line:84%
I don't have an axiom
for three events.

00:32:15.370 --> 00:32:18.730 align:middle line:84%
Well, maybe I can
massage things and still

00:32:18.730 --> 00:32:20.620 align:middle line:90%
be able to use that axiom.

00:32:20.620 --> 00:32:22.700 align:middle line:90%
And here's the trick.

00:32:22.700 --> 00:32:26.150 align:middle line:84%
The union of three sets,
you can think of it

00:32:26.150 --> 00:32:30.500 align:middle line:84%
as forming the union
of the first two sets

00:32:30.500 --> 00:32:35.670 align:middle line:84%
and then taking the
union with the third set.

00:32:35.670 --> 00:32:36.530 align:middle line:90%
OK?

00:32:36.530 --> 00:32:39.450 align:middle line:84%
So taking unions, you can
take the unions in any order

00:32:39.450 --> 00:32:40.440 align:middle line:90%
that you want.

00:32:40.440 --> 00:32:44.580 align:middle line:84%
So here we have the
union of two sets.

00:32:44.580 --> 00:32:49.150 align:middle line:84%
Now, ABC are disjoint,
by assumption

00:32:49.150 --> 00:32:51.780 align:middle line:90%
or that's how I drew it.

00:32:51.780 --> 00:32:55.760 align:middle line:84%
So if A, B, and C are
disjoint, then A union B

00:32:55.760 --> 00:32:59.030 align:middle line:84%
is disjoint from
C. So here we have

00:32:59.030 --> 00:33:01.400 align:middle line:90%
the union of two disjoint sets.

00:33:01.400 --> 00:33:05.810 align:middle line:84%
So by the additivity axiom, the
probability of that the union

00:33:05.810 --> 00:33:08.960 align:middle line:84%
is going to be the
probability of the first set

00:33:08.960 --> 00:33:12.000 align:middle line:84%
plus the probability
of the second set.

00:33:12.000 --> 00:33:15.580 align:middle line:84%
And now I can use the
additivity axiom once more

00:33:15.580 --> 00:33:17.520 align:middle line:84%
to write that this
is probability

00:33:17.520 --> 00:33:21.240 align:middle line:84%
of A plus probability
of B plus probability

00:33:21.240 --> 00:33:26.290 align:middle line:84%
of C. So by using this axiom
which was stated for two sets,

00:33:26.290 --> 00:33:29.580 align:middle line:84%
we can actually derive
a similar property

00:33:29.580 --> 00:33:32.450 align:middle line:84%
for the union of
three disjoint sets.

00:33:32.450 --> 00:33:34.960 align:middle line:84%
And then you can repeat
this argument as many times

00:33:34.960 --> 00:33:35.940 align:middle line:90%
as you want.

00:33:35.940 --> 00:33:38.750 align:middle line:84%
It's valid for the union
of ten disjoint sets,

00:33:38.750 --> 00:33:41.240 align:middle line:84%
for the union of a
hundred disjoint sets,

00:33:41.240 --> 00:33:44.910 align:middle line:84%
for the union of any
finite number of sets.

00:33:44.910 --> 00:33:53.210 align:middle line:84%
So if A1 up to An are
disjoint, then the probability

00:33:53.210 --> 00:33:59.185 align:middle line:84%
of A1 union An is equal to
the sum of the probabilities

00:33:59.185 --> 00:34:01.500 align:middle line:90%
of the individual sets.

00:34:01.500 --> 00:34:04.180 align:middle line:90%


00:34:04.180 --> 00:34:05.740 align:middle line:90%
OK.

00:34:05.740 --> 00:34:10.790 align:middle line:84%
Special case of this is when
we're dealing with finite sets.

00:34:10.790 --> 00:34:14.300 align:middle line:84%
Suppose I have just a
finite set of outcomes.

00:34:14.300 --> 00:34:17.150 align:middle line:84%
I put them together
in a set and I'm

00:34:17.150 --> 00:34:19.630 align:middle line:84%
interested in the
probability of that set.

00:34:19.630 --> 00:34:22.050 align:middle line:90%
So here's our sample space.

00:34:22.050 --> 00:34:26.620 align:middle line:84%
There's lots of outcomes,
but I'm taking a few of these

00:34:26.620 --> 00:34:30.120 align:middle line:90%
and I form a set out of them.

00:34:30.120 --> 00:34:34.760 align:middle line:84%
This is a set consisting of, in
this picture, three elements.

00:34:34.760 --> 00:34:38.260 align:middle line:84%
In general, it
consists of k elements.

00:34:38.260 --> 00:34:41.340 align:middle line:84%
Now, a finite set,
I can write it

00:34:41.340 --> 00:34:44.889 align:middle line:84%
as a union of
single element sets.

00:34:44.889 --> 00:34:49.080 align:middle line:84%
So this set here is the union
of this one element set,

00:34:49.080 --> 00:34:51.469 align:middle line:84%
together with this one
element set together

00:34:51.469 --> 00:34:53.980 align:middle line:90%
with that one element set.

00:34:53.980 --> 00:34:55.909 align:middle line:84%
So the total
probability of this set

00:34:55.909 --> 00:35:00.180 align:middle line:84%
is going to be the sum of
the probabilities of the one

00:35:00.180 --> 00:35:02.510 align:middle line:90%
element sets.

00:35:02.510 --> 00:35:05.840 align:middle line:84%
Now, probability
of a one element

00:35:05.840 --> 00:35:09.320 align:middle line:84%
set, you need to use
the brackets here

00:35:09.320 --> 00:35:12.260 align:middle line:84%
because probabilities
are assigned to sets.

00:35:12.260 --> 00:35:15.560 align:middle line:84%
But this gets kind of
tedious, so here one abuses

00:35:15.560 --> 00:35:18.510 align:middle line:84%
notation a little bit
and we get rid of those

00:35:18.510 --> 00:35:20.740 align:middle line:84%
brackets and just
write probability

00:35:20.740 --> 00:35:24.030 align:middle line:84%
of this single,
individual outcome.

00:35:24.030 --> 00:35:26.770 align:middle line:84%
In any case, conclusion
from this exercise

00:35:26.770 --> 00:35:32.010 align:middle line:84%
is that the total probability
of a finite collection

00:35:32.010 --> 00:35:35.790 align:middle line:84%
of possible outcomes,
the total probability

00:35:35.790 --> 00:35:38.360 align:middle line:84%
is equal to the sum
of the probabilities

00:35:38.360 --> 00:35:42.190 align:middle line:90%
of individual elements.

00:35:42.190 --> 00:35:46.460 align:middle line:84%
So these are basically the
axioms of probability theory.

00:35:46.460 --> 00:35:49.970 align:middle line:84%
Or, well, they're
almost the axioms.

00:35:49.970 --> 00:35:53.060 align:middle line:84%
There are some subtleties
that are involved here.

00:35:53.060 --> 00:35:58.480 align:middle line:84%
One subtlety is that this
axiom here doesn't quite

00:35:58.480 --> 00:36:01.340 align:middle line:84%
do the job for everything
we would like to do.

00:36:01.340 --> 00:36:05.080 align:middle line:84%
And we're going to come back to
this at the end of the lecture.

00:36:05.080 --> 00:36:10.380 align:middle line:84%
A second subtlety has
to do with weird sets.

00:36:10.380 --> 00:36:13.270 align:middle line:84%
We said that an event is a
subset of the sample space

00:36:13.270 --> 00:36:16.712 align:middle line:84%
and we assign
probabilities to events.

00:36:16.712 --> 00:36:19.760 align:middle line:84%
Does this mean that we are
going to assign probability

00:36:19.760 --> 00:36:23.500 align:middle line:84%
to every possible subset
of the sample space?

00:36:23.500 --> 00:36:26.660 align:middle line:84%
Ideally, we would
wish to do that.

00:36:26.660 --> 00:36:29.580 align:middle line:84%
Unfortunately, this is
not always possible.

00:36:29.580 --> 00:36:34.510 align:middle line:84%
If you take a sample
space, such as the square,

00:36:34.510 --> 00:36:36.950 align:middle line:84%
the square has
nice subsets, those

00:36:36.950 --> 00:36:40.220 align:middle line:84%
that you can describe by
cutting it with lines and so on.

00:36:40.220 --> 00:36:45.270 align:middle line:84%
But it does have some very
ugly subsets, as well,

00:36:45.270 --> 00:36:47.210 align:middle line:84%
that are impossible
to visualize,

00:36:47.210 --> 00:36:50.030 align:middle line:84%
impossible to imagine,
but they do exist.

00:36:50.030 --> 00:36:53.130 align:middle line:84%
And those very weird sets
are such that there's

00:36:53.130 --> 00:36:56.220 align:middle line:84%
no way to assign probabilities
to them in a way that's

00:36:56.220 --> 00:36:58.500 align:middle line:84%
consistent with the
axioms of probability.

00:36:58.500 --> 00:36:59.000 align:middle line:90%
OK.

00:36:59.000 --> 00:37:01.460 align:middle line:84%
So this is a very,
very fine point

00:37:01.460 --> 00:37:05.940 align:middle line:84%
that you can immediately forget
for the rest of this class.

00:37:05.940 --> 00:37:08.040 align:middle line:84%
You will only
encounter these sets

00:37:08.040 --> 00:37:12.340 align:middle line:84%
if you end up doing doctoral
work on the theoretical aspects

00:37:12.340 --> 00:37:15.910 align:middle line:90%
of probability theory.

00:37:15.910 --> 00:37:18.350 align:middle line:84%
So it's just a
mathematical subtlety

00:37:18.350 --> 00:37:20.740 align:middle line:84%
that some very weird
sets do not have

00:37:20.740 --> 00:37:22.560 align:middle line:90%
probabilities assigned to them.

00:37:22.560 --> 00:37:24.780 align:middle line:84%
But we're not going to
encounter these sets

00:37:24.780 --> 00:37:26.885 align:middle line:84%
and they do not show
up in any applications.

00:37:26.885 --> 00:37:29.340 align:middle line:90%


00:37:29.340 --> 00:37:29.840 align:middle line:90%
OK.

00:37:29.840 --> 00:37:32.410 align:middle line:84%
So now let's revisit
our examples.

00:37:32.410 --> 00:37:34.800 align:middle line:84%
Let's go back to
the die example.

00:37:34.800 --> 00:37:36.950 align:middle line:90%
We have our sample space.

00:37:36.950 --> 00:37:40.830 align:middle line:84%
Now we need to assign
a probability law.

00:37:40.830 --> 00:37:43.260 align:middle line:84%
There's lots of possible
probability laws

00:37:43.260 --> 00:37:44.690 align:middle line:90%
that you can assign.

00:37:44.690 --> 00:37:48.200 align:middle line:84%
I'm picking one
here, arbitrarily,

00:37:48.200 --> 00:37:50.960 align:middle line:84%
in which I say that every
possible outcome has

00:37:50.960 --> 00:37:55.440 align:middle line:90%
the same probability of 1/16.

00:37:55.440 --> 00:37:56.040 align:middle line:90%
OK.

00:37:56.040 --> 00:37:58.010 align:middle line:90%
Why do I make this model?

00:37:58.010 --> 00:38:01.990 align:middle line:84%
Well, empirically, if you
have well-manufactured dice,

00:38:01.990 --> 00:38:04.540 align:middle line:90%
they tend to behave that way.

00:38:04.540 --> 00:38:06.870 align:middle line:84%
We will be coming back
to this kind of story

00:38:06.870 --> 00:38:08.500 align:middle line:90%
later in this class.

00:38:08.500 --> 00:38:12.790 align:middle line:84%
But I'm not saying that
this is the only probability

00:38:12.790 --> 00:38:13.720 align:middle line:90%
law that there can be.

00:38:13.720 --> 00:38:17.460 align:middle line:84%
You might have weird dice in
which certain outcomes are

00:38:17.460 --> 00:38:19.280 align:middle line:90%
more likely than others.

00:38:19.280 --> 00:38:21.730 align:middle line:84%
But to keep things simple,
let's take every outcome

00:38:21.730 --> 00:38:24.870 align:middle line:84%
to have the same
probability of 1/16.

00:38:24.870 --> 00:38:26.790 align:middle line:90%
OK.

00:38:26.790 --> 00:38:29.110 align:middle line:84%
Now that we have in our
hands a sample space

00:38:29.110 --> 00:38:31.350 align:middle line:84%
and the probability
law, we can actually

00:38:31.350 --> 00:38:33.250 align:middle line:90%
solve any problem there is.

00:38:33.250 --> 00:38:36.070 align:middle line:84%
We can answer any question
that could be posed to us.

00:38:36.070 --> 00:38:38.230 align:middle line:84%
For example, what's
the probability

00:38:38.230 --> 00:38:43.590 align:middle line:84%
that the outcome, which is this
pair, is either 1,1 or 1,2.

00:38:43.590 --> 00:38:50.160 align:middle line:84%
We're talking here about this
particular event, 1,1 or 1,2.

00:38:50.160 --> 00:38:53.300 align:middle line:84%
So it's an event consisting
of these two items.

00:38:53.300 --> 00:38:55.880 align:middle line:84%
According to what we
were just discussing,

00:38:55.880 --> 00:38:58.720 align:middle line:84%
the probability of a finite
collection of outcomes

00:38:58.720 --> 00:39:01.170 align:middle line:84%
is the sum of their
individual probabilities.

00:39:01.170 --> 00:39:03.710 align:middle line:84%
Each one of them has
probability of 1/16,

00:39:03.710 --> 00:39:07.720 align:middle line:84%
so the probability
of this is 2/16.

00:39:07.720 --> 00:39:12.240 align:middle line:84%
How about the probability of the
event that x is equal to one.

00:39:12.240 --> 00:39:14.500 align:middle line:84%
x is the first roll, so
that's the probability

00:39:14.500 --> 00:39:18.120 align:middle line:84%
that the first roll
is equal to one.

00:39:18.120 --> 00:39:22.340 align:middle line:84%
Notice the syntax
that's being used here.

00:39:22.340 --> 00:39:26.200 align:middle line:84%
Probabilities are assigned
to subsets, to sets,

00:39:26.200 --> 00:39:30.800 align:middle line:84%
so we think of this as meaning
the set of all outcomes

00:39:30.800 --> 00:39:33.660 align:middle line:90%
such that x is equal to one.

00:39:33.660 --> 00:39:35.210 align:middle line:90%
How do you answer this question?

00:39:35.210 --> 00:39:36.680 align:middle line:84%
You go back to the
picture and you

00:39:36.680 --> 00:39:40.810 align:middle line:84%
try to visualize or identify
this event of interest.

00:39:40.810 --> 00:39:45.570 align:middle line:84%
x is equal to one corresponds
to this event here.

00:39:45.570 --> 00:39:48.950 align:middle line:84%
These are all the outcomes
at which x is equal to one.

00:39:48.950 --> 00:39:50.100 align:middle line:90%
There's four outcomes.

00:39:50.100 --> 00:39:54.180 align:middle line:84%
Each one has probability
1/16, so the answer is 4/16.

00:39:54.180 --> 00:39:56.760 align:middle line:90%


00:39:56.760 --> 00:39:57.820 align:middle line:90%
OK.

00:39:57.820 --> 00:40:06.482 align:middle line:84%
How about the probability
that x plus y is odd?

00:40:06.482 --> 00:40:07.100 align:middle line:90%
OK.

00:40:07.100 --> 00:40:09.840 align:middle line:84%
That will take a
little bit more work.

00:40:09.840 --> 00:40:11.590 align:middle line:84%
But you go to the
sample space and you

00:40:11.590 --> 00:40:16.010 align:middle line:84%
identify all the outcomes at
which the sum is an odd number.

00:40:16.010 --> 00:40:24.480 align:middle line:84%
So that's a place where the sum
is odd, these are other places,

00:40:24.480 --> 00:40:28.610 align:middle line:84%
and I guess that exhausts
all the possible outcomes

00:40:28.610 --> 00:40:31.780 align:middle line:90%
at which we have an odd sum.

00:40:31.780 --> 00:40:32.890 align:middle line:90%
We count them.

00:40:32.890 --> 00:40:34.030 align:middle line:90%
How many are there?

00:40:34.030 --> 00:40:35.540 align:middle line:84%
There's a total
of eight of them.

00:40:35.540 --> 00:40:40.490 align:middle line:84%
Each one has probability 1/16,
total probability is 8/16.

00:40:40.490 --> 00:40:41.620 align:middle line:90%
And harder question.

00:40:41.620 --> 00:40:44.310 align:middle line:84%
What is the probability that
the minimum of the two rolls

00:40:44.310 --> 00:40:45.782 align:middle line:90%
is equal to 2?

00:40:45.782 --> 00:40:47.240 align:middle line:84%
This is something
that you probably

00:40:47.240 --> 00:40:51.640 align:middle line:84%
couldn't do in your head
without the help of a diagram.

00:40:51.640 --> 00:40:54.780 align:middle line:84%
But once you have a
diagram, things are simple.

00:40:54.780 --> 00:40:55.760 align:middle line:90%
You ask the question.

00:40:55.760 --> 00:40:59.600 align:middle line:84%
OK, this is an event, that
the minimum of the two rolls

00:40:59.600 --> 00:41:01.140 align:middle line:90%
is equal to two.

00:41:01.140 --> 00:41:03.150 align:middle line:90%
This can happen in several ways.

00:41:03.150 --> 00:41:05.250 align:middle line:84%
What are the several
ways that it can happen?

00:41:05.250 --> 00:41:07.980 align:middle line:84%
Go to the diagram and
try to identify them.

00:41:07.980 --> 00:41:11.620 align:middle line:84%
So the minimum is equal to
two if both of them are two's.

00:41:11.620 --> 00:41:14.230 align:middle line:90%


00:41:14.230 --> 00:41:18.780 align:middle line:84%
Or it could be that x is two
and y is bigger, or y is two

00:41:18.780 --> 00:41:21.900 align:middle line:90%
and x is bigger.

00:41:21.900 --> 00:41:23.150 align:middle line:90%
OK.

00:41:23.150 --> 00:41:28.670 align:middle line:84%
I guess we rediscover that
yellow and blue make green,

00:41:28.670 --> 00:41:31.670 align:middle line:84%
so we see here that
there's a total

00:41:31.670 --> 00:41:34.630 align:middle line:90%
of five possible outcomes.

00:41:34.630 --> 00:41:37.645 align:middle line:84%
The probability of
this event is 5/16.

00:41:37.645 --> 00:41:41.250 align:middle line:90%


00:41:41.250 --> 00:41:46.590 align:middle line:84%
Simple example,
but the procedure

00:41:46.590 --> 00:41:48.680 align:middle line:84%
that we followed in
this example actually

00:41:48.680 --> 00:41:54.240 align:middle line:84%
applies to any probability
model you might ever encounter.

00:41:54.240 --> 00:41:56.630 align:middle line:84%
You set up your
sample space, you

00:41:56.630 --> 00:41:58.970 align:middle line:84%
make a statement that
describes the probability

00:41:58.970 --> 00:42:01.510 align:middle line:84%
law over that sample space,
then somebody asks you

00:42:01.510 --> 00:42:03.640 align:middle line:90%
questions about various events.

00:42:03.640 --> 00:42:07.010 align:middle line:84%
You go to your pictures,
identify those events,

00:42:07.010 --> 00:42:10.075 align:middle line:84%
pin them down, and
then start kind

00:42:10.075 --> 00:42:12.990 align:middle line:84%
of counting and calculating
the total probability

00:42:12.990 --> 00:42:16.560 align:middle line:84%
for those outcomes that
you're considering.

00:42:16.560 --> 00:42:19.590 align:middle line:84%
This example is a special
case of what is called

00:42:19.590 --> 00:42:22.780 align:middle line:90%
the discrete uniform law.

00:42:22.780 --> 00:42:24.970 align:middle line:84%
The model obeys the
discrete uniform law

00:42:24.970 --> 00:42:28.340 align:middle line:84%
if all outcomes
are equally likely.

00:42:28.340 --> 00:42:30.040 align:middle line:90%
It doesn't have to be that way.

00:42:30.040 --> 00:42:33.290 align:middle line:84%
That's just one example
of a probability law.

00:42:33.290 --> 00:42:37.830 align:middle line:84%
But when things are that way, if
all outcomes are equally likely

00:42:37.830 --> 00:42:46.770 align:middle line:84%
and we have N of them, and you
have a set A that has little n

00:42:46.770 --> 00:42:50.870 align:middle line:84%
elements, then each
one of those elements

00:42:50.870 --> 00:42:53.840 align:middle line:84%
has probability
one over capital N

00:42:53.840 --> 00:42:56.450 align:middle line:84%
since all outcomes
are equally likely.

00:42:56.450 --> 00:42:58.360 align:middle line:84%
And for our probabilities
to add up to one,

00:42:58.360 --> 00:43:01.000 align:middle line:84%
each one must have
this much probability,

00:43:01.000 --> 00:43:02.620 align:middle line:90%
and there's little n elements.

00:43:02.620 --> 00:43:06.120 align:middle line:84%
That gives you the probability
of the event of interest.

00:43:06.120 --> 00:43:08.710 align:middle line:84%
So problems like the one
in the previous slide

00:43:08.710 --> 00:43:10.840 align:middle line:84%
and more generally of
the type described here

00:43:10.840 --> 00:43:13.460 align:middle line:84%
under discrete uniform
law, these problems

00:43:13.460 --> 00:43:15.270 align:middle line:90%
reduce to just counting.

00:43:15.270 --> 00:43:17.500 align:middle line:84%
How many elements are
there in my sample space?

00:43:17.500 --> 00:43:21.160 align:middle line:84%
How many elements are there
inside the event of interest?

00:43:21.160 --> 00:43:24.110 align:middle line:84%
Counting is generally
simple, but for some problems

00:43:24.110 --> 00:43:25.950 align:middle line:90%
it gets pretty complicated.

00:43:25.950 --> 00:43:28.170 align:middle line:84%
And in a couple of
weeks, we're going

00:43:28.170 --> 00:43:30.390 align:middle line:84%
to have to spend the
whole lecture just

00:43:30.390 --> 00:43:33.280 align:middle line:84%
on the subject of how
to count systematically.

00:43:33.280 --> 00:43:36.730 align:middle line:84%
Now the procedure we followed
in the previous example

00:43:36.730 --> 00:43:38.440 align:middle line:84%
is the same as the
procedure you would

00:43:38.440 --> 00:43:41.330 align:middle line:84%
follow in continuous
probability problems.

00:43:41.330 --> 00:43:43.230 align:middle line:84%
So, going back to
our dart problem,

00:43:43.230 --> 00:43:46.550 align:middle line:84%
we get the random point
inside the square.

00:43:46.550 --> 00:43:48.030 align:middle line:90%
That's our sample space.

00:43:48.030 --> 00:43:50.360 align:middle line:84%
We need to assign
a probability law.

00:43:50.360 --> 00:43:53.480 align:middle line:84%
For lack of imagination, I'm
taking the probability law

00:43:53.480 --> 00:43:56.280 align:middle line:90%
to be the area of a subset.

00:43:56.280 --> 00:44:00.260 align:middle line:84%
So if we have two subsets
of the sample space

00:44:00.260 --> 00:44:04.890 align:middle line:84%
that have equal areas, then
I'm postulating that they

00:44:04.890 --> 00:44:06.386 align:middle line:90%
are equally likely to occur.

00:44:06.386 --> 00:44:09.010 align:middle line:84%
The probably that they fall here
is the same as the probability

00:44:09.010 --> 00:44:11.430 align:middle line:90%
that they fall there.

00:44:11.430 --> 00:44:13.670 align:middle line:84%
The model doesn't
have to be that way.

00:44:13.670 --> 00:44:15.980 align:middle line:84%
But if I have sort
of complete ignorance

00:44:15.980 --> 00:44:18.230 align:middle line:84%
of which points are
more likely than others,

00:44:18.230 --> 00:44:21.430 align:middle line:84%
that might be the
reasonable model to use.

00:44:21.430 --> 00:44:24.680 align:middle line:84%
So equal areas mean
equal probabilities.

00:44:24.680 --> 00:44:26.910 align:middle line:84%
If the area is twice as
large, the probability

00:44:26.910 --> 00:44:28.830 align:middle line:90%
is going to be twice as big.

00:44:28.830 --> 00:44:32.130 align:middle line:90%
So this is our model.

00:44:32.130 --> 00:44:34.580 align:middle line:90%
We can now answer questions.

00:44:34.580 --> 00:44:35.730 align:middle line:90%
Let's answer the easy one.

00:44:35.730 --> 00:44:40.660 align:middle line:84%
What's the probability that the
outcome is exactly this point?

00:44:40.660 --> 00:44:47.500 align:middle line:84%
That of course is zero because
a single point has zero area.

00:44:47.500 --> 00:44:49.580 align:middle line:84%
And since this probability
is equal to area,

00:44:49.580 --> 00:44:51.510 align:middle line:90%
that's zero probability.

00:44:51.510 --> 00:44:54.180 align:middle line:84%
How about the
probability that the sum

00:44:54.180 --> 00:44:58.150 align:middle line:84%
of the coordinates of the
point that we got is less than

00:44:58.150 --> 00:45:00.090 align:middle line:90%
or equal to 1/2?

00:45:00.090 --> 00:45:01.570 align:middle line:90%
How do you deal with it?

00:45:01.570 --> 00:45:04.770 align:middle line:84%
Well, you look at the picture
again, at your sample space,

00:45:04.770 --> 00:45:08.130 align:middle line:84%
and try to describe the event
that you're talking about.

00:45:08.130 --> 00:45:11.480 align:middle line:84%
The sum being less
than 1/2 corresponds

00:45:11.480 --> 00:45:14.960 align:middle line:84%
to getting an outcome that's
below this line, where

00:45:14.960 --> 00:45:19.600 align:middle line:84%
this line is the line where
x plus y equals to 1/2.

00:45:19.600 --> 00:45:25.860 align:middle line:84%
So the intercepts of that line
with the axis are 1/2 and 1/2.

00:45:25.860 --> 00:45:28.630 align:middle line:84%
So you describe
the event visually

00:45:28.630 --> 00:45:30.780 align:middle line:84%
and then you use
your probability law.

00:45:30.780 --> 00:45:32.250 align:middle line:84%
The probability
law that we have is

00:45:32.250 --> 00:45:36.620 align:middle line:84%
that the probability of a set is
equal to the area of that set.

00:45:36.620 --> 00:45:39.900 align:middle line:84%
So all we need to find is the
area of this triangle, which

00:45:39.900 --> 00:45:48.881 align:middle line:84%
is 1/2 times 1/2 times
1/2, half, equals to 1/8.

00:45:48.881 --> 00:45:49.380 align:middle line:90%
OK.

00:45:49.380 --> 00:45:52.020 align:middle line:84%
Moral from these two
examples is that it's always

00:45:52.020 --> 00:45:54.960 align:middle line:84%
useful to have a picture
and work with a picture

00:45:54.960 --> 00:45:58.750 align:middle line:84%
to visualize the events
that you're talking about.

00:45:58.750 --> 00:46:01.340 align:middle line:84%
And once you have a
probability law in your hands,

00:46:01.340 --> 00:46:03.700 align:middle line:84%
then it's a matter
of calculation

00:46:03.700 --> 00:46:06.540 align:middle line:84%
to find the probabilities
of an event of interest.

00:46:06.540 --> 00:46:09.080 align:middle line:84%
The calculations we did in
these two examples, of course,

00:46:09.080 --> 00:46:10.130 align:middle line:90%
were very simple.

00:46:10.130 --> 00:46:13.320 align:middle line:84%
Sometimes calculations
may be a lot harder,

00:46:13.320 --> 00:46:15.480 align:middle line:90%
but it's a different business.

00:46:15.480 --> 00:46:18.820 align:middle line:84%
It's a business of calculus,
for example, or being

00:46:18.820 --> 00:46:20.250 align:middle line:90%
good in algebra and so on.

00:46:20.250 --> 00:46:22.950 align:middle line:84%
As far as probability
is concerned,

00:46:22.950 --> 00:46:24.950 align:middle line:84%
it's clear what
you will be doing,

00:46:24.950 --> 00:46:27.920 align:middle line:84%
and then maybe you're faced
with a harder algebraic part

00:46:27.920 --> 00:46:30.540 align:middle line:84%
to actually carry
out the calculations.

00:46:30.540 --> 00:46:32.870 align:middle line:84%
The area of a triangle
is easy to compute.

00:46:32.870 --> 00:46:35.740 align:middle line:84%
If I had put down a
very complicated shape,

00:46:35.740 --> 00:46:38.390 align:middle line:84%
then you might need to
solve a hard integration

00:46:38.390 --> 00:46:40.650 align:middle line:84%
problem to find the
area of that shape,

00:46:40.650 --> 00:46:42.960 align:middle line:84%
but that's stuff that
belongs to another class

00:46:42.960 --> 00:46:46.306 align:middle line:84%
that you have presumably
mastered by now.

00:46:46.306 --> 00:46:47.000 align:middle line:90%
Good, OK.

00:46:47.000 --> 00:46:48.580 align:middle line:84%
So now let me
spend just a couple

00:46:48.580 --> 00:46:52.170 align:middle line:84%
of minutes to return to a
point that I raised before.

00:46:52.170 --> 00:46:56.270 align:middle line:84%
I was saying that the axiom
that we had about additivity

00:46:56.270 --> 00:46:58.730 align:middle line:90%
might not quite be enough.

00:46:58.730 --> 00:47:01.730 align:middle line:84%
Let's illustrate what I mean
by the following example.

00:47:01.730 --> 00:47:04.790 align:middle line:84%
Think of the experiment where
you keep flipping a coin

00:47:04.790 --> 00:47:08.120 align:middle line:84%
and you wait until you obtain
heads for the first time.

00:47:08.120 --> 00:47:11.390 align:middle line:84%
What's the sample space
of this experiment?

00:47:11.390 --> 00:47:12.970 align:middle line:84%
It might happen
the first flip, it

00:47:12.970 --> 00:47:14.700 align:middle line:90%
might happen in the tenth flip.

00:47:14.700 --> 00:47:18.490 align:middle line:84%
Heads for the first time might
occur in the millionth flip.

00:47:18.490 --> 00:47:20.370 align:middle line:84%
So the outcome of
this experiment

00:47:20.370 --> 00:47:22.020 align:middle line:84%
is going to be an
integer and there's

00:47:22.020 --> 00:47:23.820 align:middle line:90%
no bound to that integer.

00:47:23.820 --> 00:47:26.780 align:middle line:84%
You might have to wait very
much until that happens.

00:47:26.780 --> 00:47:28.860 align:middle line:84%
So the natural sample
space is the set

00:47:28.860 --> 00:47:30.950 align:middle line:90%
of all possible integers.

00:47:30.950 --> 00:47:34.690 align:middle line:84%
Somebody tells you
some information

00:47:34.690 --> 00:47:36.250 align:middle line:90%
about the probability law.

00:47:36.250 --> 00:47:39.250 align:middle line:84%
The probability that you
have to wait for n flips

00:47:39.250 --> 00:47:41.130 align:middle line:90%
is equal to two to the minus n.

00:47:41.130 --> 00:47:42.850 align:middle line:90%
Where did this come from?

00:47:42.850 --> 00:47:44.220 align:middle line:90%
That's a separate story.

00:47:44.220 --> 00:47:45.730 align:middle line:90%
Where did it come from?

00:47:45.730 --> 00:47:49.680 align:middle line:84%
Somebody tells this to us,
and those probabilities

00:47:49.680 --> 00:47:52.150 align:middle line:84%
are plotted here
as a function of n.

00:47:52.150 --> 00:47:54.770 align:middle line:84%
And you're asked to find the
probability that the outcome is

00:47:54.770 --> 00:47:56.660 align:middle line:90%
an even number.

00:47:56.660 --> 00:47:59.920 align:middle line:84%
How do you go about
calculating that probability?

00:47:59.920 --> 00:48:01.830 align:middle line:84%
So the probability of
being an even number

00:48:01.830 --> 00:48:05.300 align:middle line:84%
is the probability of
the subset that consists

00:48:05.300 --> 00:48:08.380 align:middle line:90%
of just the even numbers.

00:48:08.380 --> 00:48:11.040 align:middle line:84%
So it would be a subset
of this kind, that

00:48:11.040 --> 00:48:13.760 align:middle line:90%
includes two, four, and so on.

00:48:13.760 --> 00:48:17.080 align:middle line:84%
So any reasonable
person would say,

00:48:17.080 --> 00:48:19.960 align:middle line:84%
well the probability of
obtaining an outcome that's

00:48:19.960 --> 00:48:23.160 align:middle line:84%
either two or four
or six and so on

00:48:23.160 --> 00:48:25.770 align:middle line:84%
is equal to the probability
of obtaining a two,

00:48:25.770 --> 00:48:28.180 align:middle line:84%
plus the probability
of obtaining a four,

00:48:28.180 --> 00:48:31.130 align:middle line:84%
plus the probability of
obtaining a six, and so on.

00:48:31.130 --> 00:48:33.640 align:middle line:84%
These probabilities
are given to us.

00:48:33.640 --> 00:48:35.990 align:middle line:90%
So here I have to do my algebra.

00:48:35.990 --> 00:48:40.840 align:middle line:84%
I add this geometric series
and I get an answer of 1/3.

00:48:40.840 --> 00:48:43.430 align:middle line:84%
That's what any reasonable
person would do.

00:48:43.430 --> 00:48:46.460 align:middle line:84%
But the person who
only knows the axioms

00:48:46.460 --> 00:48:51.880 align:middle line:84%
that they posted just a
little earlier may get stuck.

00:48:51.880 --> 00:48:53.610 align:middle line:84%
They would get
stuck at this point.

00:48:53.610 --> 00:48:55.700 align:middle line:90%
How do we justify this?

00:48:55.700 --> 00:48:59.000 align:middle line:90%


00:48:59.000 --> 00:49:03.740 align:middle line:84%
We had this property for
the union of disjoint sets

00:49:03.740 --> 00:49:05.780 align:middle line:84%
and the corresponding
property that tells us

00:49:05.780 --> 00:49:10.950 align:middle line:84%
that the total probability of
finitely many things, outcomes,

00:49:10.950 --> 00:49:13.740 align:middle line:84%
is the sum of their
individual probabilities.

00:49:13.740 --> 00:49:17.940 align:middle line:84%
But here we're using it
on an infinite collection.

00:49:17.940 --> 00:49:21.400 align:middle line:84%
The probability of
infinitely many points

00:49:21.400 --> 00:49:24.350 align:middle line:84%
is equal to the sum
of the probabilities

00:49:24.350 --> 00:49:26.070 align:middle line:90%
of each one of these.

00:49:26.070 --> 00:49:30.840 align:middle line:84%
To justify this step we need to
introduce one additional rule,

00:49:30.840 --> 00:49:35.060 align:middle line:84%
an additional axiom, that tells
us that this step is actually

00:49:35.060 --> 00:49:36.160 align:middle line:90%
legitimate.

00:49:36.160 --> 00:49:38.850 align:middle line:84%
And this is the countable
additivity axiom,

00:49:38.850 --> 00:49:42.560 align:middle line:84%
which is a little stronger,
or quite a bit stronger,

00:49:42.560 --> 00:49:45.140 align:middle line:84%
than the additivity
axiom we had before.

00:49:45.140 --> 00:49:48.830 align:middle line:84%
It tells us that if we
have a sequence of sets

00:49:48.830 --> 00:49:53.760 align:middle line:84%
that are disjoint and we want
to find their total probability,

00:49:53.760 --> 00:49:58.230 align:middle line:84%
then we are allowed to add
their individual probabilities.

00:49:58.230 --> 00:50:01.000 align:middle line:84%
So the picture might
be such as follows.

00:50:01.000 --> 00:50:07.420 align:middle line:84%
We have a sequence of sets,
A1, A2, A3, and so on.

00:50:07.420 --> 00:50:09.950 align:middle line:84%
I guess in order to fit them
inside the sample space,

00:50:09.950 --> 00:50:13.920 align:middle line:84%
the sets need to get
smaller and smaller perhaps.

00:50:13.920 --> 00:50:15.340 align:middle line:90%
They are disjoint.

00:50:15.340 --> 00:50:17.330 align:middle line:90%
We have a sequence of such sets.

00:50:17.330 --> 00:50:20.180 align:middle line:84%
The total probability
of falling anywhere

00:50:20.180 --> 00:50:23.140 align:middle line:84%
inside one of those
sets is the sum

00:50:23.140 --> 00:50:25.740 align:middle line:84%
of their individual
probabilities.

00:50:25.740 --> 00:50:28.510 align:middle line:84%
A key subtlety
that's involved here

00:50:28.510 --> 00:50:33.710 align:middle line:84%
is that we're talking
about a sequence of events.

00:50:33.710 --> 00:50:36.330 align:middle line:84%
By "sequence" we mean
that these events

00:50:36.330 --> 00:50:38.450 align:middle line:90%
can be arranged in order.

00:50:38.450 --> 00:50:41.660 align:middle line:84%
I can tell you the first
event, the second event,

00:50:41.660 --> 00:50:43.530 align:middle line:90%
the third event, and so on.

00:50:43.530 --> 00:50:45.860 align:middle line:84%
So if you have such a
collection of events

00:50:45.860 --> 00:50:50.060 align:middle line:84%
that can be ordered as first,
second, third, and so on,

00:50:50.060 --> 00:50:53.200 align:middle line:84%
then you can add
their probabilities

00:50:53.200 --> 00:50:55.790 align:middle line:84%
to find the probability
of their union.

00:50:55.790 --> 00:50:57.910 align:middle line:84%
So this point is actually
a little more subtle

00:50:57.910 --> 00:50:59.670 align:middle line:84%
than you might
appreciate at this point,

00:50:59.670 --> 00:51:02.480 align:middle line:84%
and I'm going to return
to it at the beginning

00:51:02.480 --> 00:51:04.010 align:middle line:90%
of the next lecture.

00:51:04.010 --> 00:51:07.160 align:middle line:84%
For now, enjoy the
first week of classes

00:51:07.160 --> 00:51:09.380 align:middle line:90%
and have a good weekend.

00:51:09.380 --> 00:51:11.230 align:middle line:90%
Thank you.