WEBVTT

00:00:10.560 --> 00:00:15.220
PATRICK WINSTON: So today we're
gonna talk about a few

00:00:15.220 --> 00:00:19.980
miracles of learning in the
context of the theme that

00:00:19.980 --> 00:00:21.930
we're developing here
in the class.

00:00:24.990 --> 00:00:31.900
We started off with
a discussion

00:00:31.900 --> 00:00:33.260
of some basic methods.

00:00:33.260 --> 00:00:35.610
We talked about nearest
neighbors.

00:00:35.610 --> 00:00:39.250
And we talked about
identification trees.

00:00:39.250 --> 00:00:41.560
And those are kind of basic
things that have been around

00:00:41.560 --> 00:00:42.740
for a long time.

00:00:42.740 --> 00:00:43.690
Still useful.

00:00:43.690 --> 00:00:46.250
Still the right things to do
when you're faced with a

00:00:46.250 --> 00:00:50.840
learning problem and you're not
sure what method to try.

00:00:50.840 --> 00:00:55.870
Then we went on to talk about
some naive biological mimicry.

00:00:55.870 --> 00:01:00.700
We talked about neural nets.

00:01:00.700 --> 00:01:05.650
And we talked about genetic
algorithms.

00:01:05.650 --> 00:01:08.270
And you look at those things and
you think and reflect back

00:01:08.270 --> 00:01:09.440
on what we talked about.

00:01:09.440 --> 00:01:12.010
And you have to say
to yourself, are

00:01:12.010 --> 00:01:14.039
these nugatory ideas?

00:01:14.039 --> 00:01:15.760
Perhaps pistareens?

00:01:15.760 --> 00:01:19.300
Or are they supererogatory
ideas that deserve to be

00:01:19.300 --> 00:01:20.550
center stage?

00:01:24.570 --> 00:01:27.770
Does anybody know what
those words mean?

00:01:27.770 --> 00:01:28.780
A pistareen?

00:01:28.780 --> 00:01:32.000
Well, a pistareen is
a Spanish coin.

00:01:32.000 --> 00:01:32.900
It was so small.

00:01:32.900 --> 00:01:35.130
It was of little worth.

00:01:35.130 --> 00:01:43.460
These ideas like neural nets,
genetic algorithms, I classify

00:01:43.460 --> 00:01:47.820
them as pistareens because
getting them to do something

00:01:47.820 --> 00:01:51.610
is rather like getting a dog
to walk on its hind legs.

00:01:51.610 --> 00:01:54.700
You can make it happen, but they
never do it very well.

00:01:54.700 --> 00:01:57.000
And you have to think it took a
lot of trickery and training

00:01:57.000 --> 00:01:58.250
to make it happen.

00:02:00.795 --> 00:02:08.150
So not too personally
high on those ideas.

00:02:08.150 --> 00:02:10.620
But we teach them to you anyway
because, of course, we

00:02:10.620 --> 00:02:14.250
only editorialize part of time
and part of time we like to

00:02:14.250 --> 00:02:17.310
cover what's in the field.

00:02:17.310 --> 00:02:22.460
Today we're starting a couple
of discussions of mechanisms

00:02:22.460 --> 00:02:26.940
or ideas or things
to know about

00:02:26.940 --> 00:02:27.690
that are quite different.

00:02:27.690 --> 00:02:29.690
Because now we're going to focus
on the problem rather

00:02:29.690 --> 00:02:31.930
than on the mechanism.

00:02:31.930 --> 00:02:34.120
And then a later on we're
going to talk about deep

00:02:34.120 --> 00:02:37.005
theory, FIOS, for
its own sake.

00:02:37.005 --> 00:02:38.560
But this week I want
to talk about

00:02:38.560 --> 00:02:42.090
mechanisms that were devised.

00:02:42.090 --> 00:02:45.760
I want to talk about research
that was done.

00:02:45.760 --> 00:02:47.020
Let me not say mechanisms.

00:02:47.020 --> 00:02:53.530
Let me say research that was
done to attempt an account of

00:02:53.530 --> 00:02:56.040
some of the things that
we humans do well.

00:02:56.040 --> 00:03:00.460
Sometimes without even knowing
that we do it.

00:03:00.460 --> 00:03:03.110
Now Krishna here tells me his
first language was Telugu.

00:03:05.750 --> 00:03:06.790
Telugu.

00:03:06.790 --> 00:03:08.270
I once had another student
whose first

00:03:08.270 --> 00:03:10.160
language was Telugu.

00:03:10.160 --> 00:03:12.100
I said to him, that must
be one of those

00:03:12.100 --> 00:03:14.480
obscure Indian languages.

00:03:14.480 --> 00:03:15.470
And he said, yes.

00:03:15.470 --> 00:03:18.570
It's spoken by 56
million people.

00:03:18.570 --> 00:03:19.756
French is spoken by 52.

00:03:19.756 --> 00:03:22.202
[LAUGHTER]

00:03:22.202 --> 00:03:24.650
PATRICK WINSTON: He's going to
be our experimental subject.

00:03:24.650 --> 00:03:28.040
Krishna, if I pluralize words--
you know what it means

00:03:28.040 --> 00:03:30.120
to pluralize a word.

00:03:30.120 --> 00:03:35.595
So if I say for example, horse,
then if I ask you for

00:03:35.595 --> 00:03:38.220
the plural you'll say horses.

00:03:38.220 --> 00:03:42.500
So if I say dog, what's
the plural?

00:03:42.500 --> 00:03:44.380
STUDENT: Then dogs.

00:03:44.380 --> 00:03:45.222
Or in my language?

00:03:45.222 --> 00:03:45.730
PATRICK WINSTON: No, no, no.

00:03:45.730 --> 00:03:46.680
In English.

00:03:46.680 --> 00:03:47.680
STUDENT: Oh, dogs.

00:03:47.680 --> 00:03:48.977
PATRICK WINSTON: Well,
what about cat?

00:03:48.977 --> 00:03:49.771
STUDENT: Cats.

00:03:49.771 --> 00:03:50.950
PATRICK WINSTON: And
he got it right.

00:03:50.950 --> 00:03:52.250
Isn't that a miracle?

00:03:52.250 --> 00:03:54.570
When did you start
speaking English?

00:03:54.570 --> 00:03:55.560
STUDENT: Second grade.

00:03:55.560 --> 00:03:56.190
PATRICK WINSTON: Second grade.

00:03:56.190 --> 00:03:57.600
But he still got it right.

00:03:57.600 --> 00:04:00.740
But he never learned that he's
actually pluralizing those

00:04:00.740 --> 00:04:02.790
words differently.

00:04:02.790 --> 00:04:05.320
But he is.

00:04:05.320 --> 00:04:08.090
So when you pluralize
dog, what's the

00:04:08.090 --> 00:04:10.600
sound that comes after?

00:04:10.600 --> 00:04:11.770
It's a z sound.

00:04:11.770 --> 00:04:12.870
Zzzzzz.

00:04:12.870 --> 00:04:14.710
Dogzzz.

00:04:14.710 --> 00:04:16.902
If you stick your fingers up
here you can probably feel

00:04:16.902 --> 00:04:18.428
your vocal cords vibrating.

00:04:18.428 --> 00:04:21.040
If you stick a piece of paper in
front of your mouth you'll

00:04:21.040 --> 00:04:23.000
see it vibrate.

00:04:23.000 --> 00:04:26.260
But when you say cats,
the pluralizing sound

00:04:26.260 --> 00:04:28.710
is sss, like that.

00:04:28.710 --> 00:04:29.920
No vocalizing.

00:04:29.920 --> 00:04:32.113
No vibration of the
vocal cords.

00:04:32.113 --> 00:04:35.510
And old Krishna here learned
that rule, as did all of you

00:04:35.510 --> 00:04:38.330
other non-native speakers of
English, effortlessly and

00:04:38.330 --> 00:04:39.190
without noticing it.

00:04:39.190 --> 00:04:40.335
You learned it.

00:04:40.335 --> 00:04:41.650
Buy you always get it right.

00:04:41.650 --> 00:04:44.170
How can that possibly be?

00:04:44.170 --> 00:04:47.250
Well, be the end of hour you'll
know how that might be.

00:04:47.250 --> 00:04:54.060
And you'll experience a case
study in how questions of that

00:04:54.060 --> 00:04:55.909
sort can be approached
with a sort of

00:04:55.909 --> 00:04:57.159
engineering point of view.

00:04:57.159 --> 00:04:59.680
You can say, what if God
were an engineer?

00:04:59.680 --> 00:05:04.730
Or alternatively, what if I were
God and I am an engineer?

00:05:04.730 --> 00:05:07.810
Think about how it might
happen that way.

00:05:07.810 --> 00:05:13.150
So we want to understand how it
might be that the machine

00:05:13.150 --> 00:05:14.700
could learn rules like that.

00:05:14.700 --> 00:05:15.690
Phonological rules.

00:05:15.690 --> 00:05:19.200
Not just that one, but all the
phonological rules you'd

00:05:19.200 --> 00:05:21.540
acquire in a course
on phonology.

00:05:21.540 --> 00:05:27.990
That part of speaking that deals
with those syllabic and

00:05:27.990 --> 00:05:30.740
sub-syllabic sounds.

00:05:30.740 --> 00:05:32.970
The phones of the language.

00:05:32.970 --> 00:05:37.890
So when Yip and Sussman
undertook to solve this

00:05:37.890 --> 00:05:41.950
engineering problem, both being
dedicated engineers, the

00:05:41.950 --> 00:05:44.560
first thing they did was
learn the science.

00:05:44.560 --> 00:05:47.840
So they went to sit at the
foot of Morris Halle, who

00:05:47.840 --> 00:05:51.510
would develop-- was largely
responsible for the

00:05:51.510 --> 00:05:53.280
development theories of

00:05:53.280 --> 00:05:54.940
so-called distinctive features.

00:05:54.940 --> 00:05:57.200
And here's how all that works.

00:05:57.200 --> 00:06:01.030
You start off with a person who
wants to say something.

00:06:03.900 --> 00:06:09.650
And out that person's mouth
comes some sort of acoustic

00:06:09.650 --> 00:06:11.670
pressure wave.

00:06:11.670 --> 00:06:14.780
And if I say, hello, George.

00:06:14.780 --> 00:06:16.560
And you say hello, George.

00:06:16.560 --> 00:06:19.120
Everybody will understand that
we said the same thing.

00:06:19.120 --> 00:06:21.540
But that acoustic waveform won't
look anything alike.

00:06:21.540 --> 00:06:24.790
It'll be very different
for all of us.

00:06:24.790 --> 00:06:29.310
So it's a miracle that words
can be understood.

00:06:29.310 --> 00:06:32.240
In any case, it goes
into an ear.

00:06:32.240 --> 00:06:34.240
And it's processed.

00:06:34.240 --> 00:06:42.325
And out comes a sequence
of distinctive feature.

00:06:51.440 --> 00:06:52.690
Vectors.

00:06:58.120 --> 00:07:04.170
A distinctive feature is a
binary variable like is the

00:07:04.170 --> 00:07:05.880
phone voices or not.

00:07:05.880 --> 00:07:07.830
That is to say, are your
vocal cords vibrating

00:07:07.830 --> 00:07:08.840
when you say it?

00:07:08.840 --> 00:07:11.970
If so, then that's
plus voiced.

00:07:11.970 --> 00:07:14.360
If not, it's minus voiced.

00:07:14.360 --> 00:07:19.440
So according to the original
distinctive feature theory and

00:07:19.440 --> 00:07:22.650
consistent with most of the
theories that have been

00:07:22.650 --> 00:07:25.740
derived since the original one,
there are on the order of

00:07:25.740 --> 00:07:29.300
14 of these distinctive features
that determine which

00:07:29.300 --> 00:07:32.090
phone you're saying.

00:07:32.090 --> 00:07:35.260
So if you say ah, that's
one combination of

00:07:35.260 --> 00:07:36.715
these binary features.

00:07:36.715 --> 00:07:40.659
If you say tuh, that's another
combination of

00:07:40.659 --> 00:07:42.659
these binary features.

00:07:42.659 --> 00:07:45.040
14 of them.

00:07:45.040 --> 00:07:48.909
So how many sounds does that
mean, in principle, there

00:07:48.909 --> 00:07:50.770
could be in a language?

00:07:50.770 --> 00:07:52.165
SEBASTIAN: 2 to the 14th.

00:07:52.165 --> 00:07:58.300
PATRICK WINSTON: And what's
2 the 14th, Sebastian?

00:07:58.300 --> 00:08:01.920
Well, it ought to be about
16,000, don't you think?

00:08:01.920 --> 00:08:03.350
2 to the 10th is 1,000.

00:08:03.350 --> 00:08:05.310
2 the fourth is 16.

00:08:05.310 --> 00:08:09.480
So there are about 16,000
possible combination.

00:08:09.480 --> 00:08:14.100
But no language on Earth has
more than 100 phones.

00:08:14.100 --> 00:08:14.910
That's strange, isn't it?

00:08:14.910 --> 00:08:18.670
Because some of those choices
are probably excluded on

00:08:18.670 --> 00:08:19.390
physical ground.

00:08:19.390 --> 00:08:20.730
But most of them are not.

00:08:20.730 --> 00:08:23.870
So we could have a lot more
phones in our language than we

00:08:23.870 --> 00:08:24.780
actually do.

00:08:24.780 --> 00:08:27.630
English is about 40.

00:08:27.630 --> 00:08:32.070
So the sequence of distinctive
features could be viewed as

00:08:32.070 --> 00:08:41.090
then producing meaning after,
perhaps, a long series of

00:08:41.090 --> 00:08:42.620
operations.

00:08:42.620 --> 00:08:46.730
But in the end, those operations
feedback in here

00:08:46.730 --> 00:08:48.560
because many of the distinctive
features are

00:08:48.560 --> 00:08:50.560
actually hallucinated.

00:08:50.560 --> 00:08:52.320
We think we heard them,
but they're not there.

00:08:52.320 --> 00:08:55.480
Or they're not even in the
acoustic waveform.

00:08:55.480 --> 00:08:57.345
They're there for the
convenience of the phonologist

00:08:57.345 --> 00:09:01.210
who make rules out of them.

00:09:01.210 --> 00:09:10.670
It's remarkable how much of this
feedback there is, and

00:09:10.670 --> 00:09:14.570
even injection from
other modalities.

00:09:14.570 --> 00:09:17.730
Many of you may have heard
about the McGurk Effect.

00:09:17.730 --> 00:09:20.310
Here's who the McGurk
Effect works.

00:09:20.310 --> 00:09:25.270
Look at me while I say ga,
ga, ga, ga, ga, ga.

00:09:25.270 --> 00:09:25.580
OK.

00:09:25.580 --> 00:09:27.120
I said, g-a.

00:09:27.120 --> 00:09:30.050
Now how about ba, ba, ba, ba.

00:09:30.050 --> 00:09:30.460
OK.

00:09:30.460 --> 00:09:35.130
I said ba like a sheep.

00:09:35.130 --> 00:09:40.730
But if I take the sound I make
when I say ba and play it

00:09:40.730 --> 00:09:44.730
while you're taking video of
me saying ga, what do you

00:09:44.730 --> 00:09:46.840
think you hear?

00:09:46.840 --> 00:09:48.000
You don't hear ba.

00:09:48.000 --> 00:09:53.030
Some people report that they
hear a d-a sound like da.

00:09:53.030 --> 00:09:56.200
When I look at it, I can't
make any sense out of it.

00:09:56.200 --> 00:09:58.270
It looks like there's a
disconnection between the

00:09:58.270 --> 00:10:01.470
speech and the video.

00:10:01.470 --> 00:10:03.916
But it does not sound like ba.

00:10:03.916 --> 00:10:07.740
But if I shut my eyes and say
ba, ba, it's absolutely clear

00:10:07.740 --> 00:10:10.350
that it's b-a.

00:10:10.350 --> 00:10:17.280
So what you see has a large
influence on what you hear.

00:10:17.280 --> 00:10:18.600
It's also interesting--

00:10:18.600 --> 00:10:20.690
although a side issue-- it's
also interesting to note that

00:10:20.690 --> 00:10:23.350
it's very difficult pronounced
things correctly if you don't

00:10:23.350 --> 00:10:25.230
see the speaker.

00:10:25.230 --> 00:10:27.360
So many people wonder when they
learn foreign languages

00:10:27.360 --> 00:10:29.600
why they can't speak
like a native.

00:10:29.600 --> 00:10:31.020
And the answer is, they're
not watching the

00:10:31.020 --> 00:10:33.160
mouth of the speaker.

00:10:33.160 --> 00:10:35.665
I was talking to a German friend
once and said, you

00:10:35.665 --> 00:10:39.796
know, I just can't say the
damned umlaut right.

00:10:39.796 --> 00:10:42.470
And he said, oh, the trouble
with you Americans is you

00:10:42.470 --> 00:10:47.080
don't realize that American
cows say moo but

00:10:47.080 --> 00:10:48.530
German cows say muu.

00:10:48.530 --> 00:10:48.910
[LAUGHTER]

00:10:48.910 --> 00:10:51.130
PATRICK WINSTON: And, of
course, I got instantly

00:10:51.130 --> 00:10:53.410
because I could see that the
umlaut sounds are produced

00:10:53.410 --> 00:10:57.290
with protruding lips, which we
don't have any sounds an

00:10:57.290 --> 00:10:59.766
English that require that.

00:10:59.766 --> 00:11:03.180
Ah, but back to what we know
from the phonologists about

00:11:03.180 --> 00:11:04.450
all this stuff.

00:11:04.450 --> 00:11:06.740
If you talk to Morris
Halle, he will tell

00:11:06.740 --> 00:11:09.490
you that over here--

00:11:09.490 --> 00:11:12.100
I like to think of it
as a marionette.

00:11:12.100 --> 00:11:13.860
There are five pieces
of meat down here.

00:11:21.270 --> 00:11:23.330
And the combination of
distinctive features that

00:11:23.330 --> 00:11:27.250
you're trying to utter are
like the control of a

00:11:27.250 --> 00:11:29.320
marionette on those five
pieces of meat.

00:11:29.320 --> 00:11:32.020
So if you want to say an a
sound, the marionette control

00:11:32.020 --> 00:11:36.310
goes into a position that
produces that combination.

00:11:36.310 --> 00:11:36.970
So let's see.

00:11:36.970 --> 00:11:41.110
What does that distinctive
feature sequence look like for

00:11:41.110 --> 00:11:41.920
typical word?

00:11:41.920 --> 00:11:44.020
Well, here's a word.

00:11:44.020 --> 00:11:45.270
A-e-p-l.

00:11:48.250 --> 00:11:50.340
Apples.

00:11:50.340 --> 00:11:56.580
And we can talk about what
distinctive features are

00:11:56.580 --> 00:12:01.380
arrayed in that particular
combination of phones.

00:12:01.380 --> 00:12:03.040
So one of the features
that they like to

00:12:03.040 --> 00:12:05.580
talk about is syllabic.

00:12:05.580 --> 00:12:06.830
Syllabic.

00:12:09.950 --> 00:12:13.240
That roughly means, can that
sound form the sort of core of

00:12:13.240 --> 00:12:14.710
a syllable?

00:12:14.710 --> 00:12:18.930
And the answer is a can,
buy these can't.

00:12:18.930 --> 00:12:22.060
So it's plus, minus,
minus, minus.

00:12:22.060 --> 00:12:29.040
Down here a little ways you'll
run into the voiced feature.

00:12:29.040 --> 00:12:30.735
And for the voiced feature,
well, we can do

00:12:30.735 --> 00:12:32.160
the experiment ourselves.

00:12:32.160 --> 00:12:33.500
Ahh.

00:12:33.500 --> 00:12:35.540
Sounds like it's voices to me.

00:12:35.540 --> 00:12:36.340
Pa.

00:12:36.340 --> 00:12:36.790
No.

00:12:36.790 --> 00:12:37.970
That's not voiced.

00:12:37.970 --> 00:12:38.390
Oo.

00:12:38.390 --> 00:12:39.080
Yep.

00:12:39.080 --> 00:12:40.540
Zzz.

00:12:40.540 --> 00:12:43.180
We already said that
was voiced.

00:12:43.180 --> 00:12:46.940
So that's the combination you
see when you utter apples for

00:12:46.940 --> 00:12:48.670
the voiced feature.

00:12:48.670 --> 00:12:50.540
Then another one is the
continuent one.

00:12:54.190 --> 00:12:57.290
That roughly says is your
vocal apparatus open?

00:12:57.290 --> 00:13:00.370
Is there no obstruction?

00:13:00.370 --> 00:13:04.750
And so ahh plus pa
is constricted.

00:13:04.750 --> 00:13:06.100
Oo, open.

00:13:06.100 --> 00:13:07.960
Zzz, open.

00:13:07.960 --> 00:13:09.960
So that one happens to run right
along with voiced in

00:13:09.960 --> 00:13:12.040
that particular word.

00:13:12.040 --> 00:13:13.640
Oh, and there are
14 altogether.

00:13:13.640 --> 00:13:16.510
But let me just write
down one more.

00:13:16.510 --> 00:13:17.760
The strident one.

00:13:21.980 --> 00:13:25.000
That says, do you use your
tongue to form a

00:13:25.000 --> 00:13:26.800
little jet of air?

00:13:26.800 --> 00:13:30.735
So you don't on aa, pa, oo.

00:13:30.735 --> 00:13:33.220
Buy you do on z.

00:13:33.220 --> 00:13:35.600
So that gets a plus.

00:13:35.600 --> 00:13:39.530
So that's a glimpse through a
soda straw of what it would

00:13:39.530 --> 00:13:43.980
like to represent the word
apples as a set of distinctive

00:13:43.980 --> 00:13:46.350
features all arranged
in a sequence.

00:13:46.350 --> 00:13:49.760
So it's a matrix of features.

00:13:49.760 --> 00:13:52.750
Going down in the columns we
have our distinctive features.

00:13:52.750 --> 00:13:56.250
And going across we have time.

00:13:56.250 --> 00:14:01.230
So as the first thing Sussman
and Yip did in their effort to

00:14:01.230 --> 00:14:04.610
understand how phonological
rules could be learned is to

00:14:04.610 --> 00:14:11.880
design a machine that would
interpret words and sounds and

00:14:11.880 --> 00:14:16.260
things that you see so
as to produce the

00:14:16.260 --> 00:14:18.470
sounds of the language.

00:14:18.470 --> 00:14:21.820
So they imagined the following
kind of machine.

00:14:21.820 --> 00:14:27.320
The machine has some kind of
mystery apparatus over here

00:14:27.320 --> 00:14:31.290
that looks out into the world
and sees what's there.

00:14:31.290 --> 00:14:34.570
So I'm looking out in the world
and I see two apples.

00:14:34.570 --> 00:14:39.090
So what this machine might do
then is, at some point, decide

00:14:39.090 --> 00:14:41.490
that there are two
apples out there.

00:14:41.490 --> 00:14:44.340
Then, thinking in terms of
these guys as computer

00:14:44.340 --> 00:14:50.460
engineers, they think in terms
of a set of registers that

00:14:50.460 --> 00:14:58.700
hold values for concepts like
noun and verb and plural.

00:15:02.500 --> 00:15:05.130
And we've not done anything
with the machine yet.

00:15:05.130 --> 00:15:07.520
We've provided no input.

00:15:07.520 --> 00:15:12.260
So those registers
are all empty.

00:15:12.260 --> 00:15:16.135
Then, up in here, we have
a set of words.

00:15:20.763 --> 00:15:23.470
And they're all kinds
of words.

00:15:23.470 --> 00:15:24.720
Apple is one of them.

00:15:29.290 --> 00:15:34.540
And those words up there know
about how the concept is

00:15:34.540 --> 00:15:37.990
rendered as a sequence of a
phones, that is to say a

00:15:37.990 --> 00:15:41.430
sequence of distinct features.

00:15:41.430 --> 00:15:46.450
Then, over here, most
importantly, they have a set

00:15:46.450 --> 00:15:47.700
of constraints.

00:15:55.660 --> 00:15:58.050
So we'll talk about a particular
constrain, the

00:15:58.050 --> 00:15:59.300
plural constraint.

00:16:01.690 --> 00:16:05.380
Plural constraint number one.

00:16:05.380 --> 00:16:08.040
And it's going to reach around
and connect itself to some

00:16:08.040 --> 00:16:10.570
other parts of the machine.

00:16:10.570 --> 00:16:17.470
Finally, there's a buffer
of phones to be uttered.

00:16:17.470 --> 00:16:20.650
And they're going to flow out
this way to the speaker's

00:16:20.650 --> 00:16:26.770
mouth and get translated into
a acoustic wave form.

00:16:26.770 --> 00:16:29.680
So those are the elements
of the machine.

00:16:29.680 --> 00:16:34.380
Now how are the elements
connected together?

00:16:34.380 --> 00:16:47.460
Well, the words are connected,
of course, into the buffer

00:16:47.460 --> 00:16:51.550
that is used to generate
the sound over

00:16:51.550 --> 00:16:54.360
here on the far left.

00:16:54.360 --> 00:16:58.150
The plural register is
connected to what

00:16:58.150 --> 00:17:00.080
you see in the world.

00:17:00.080 --> 00:17:02.580
What you see in the world is
connected not only to plural

00:17:02.580 --> 00:17:08.530
register, but to all of the
objects in the word

00:17:08.530 --> 00:17:09.780
repertoire.

00:17:12.530 --> 00:17:15.950
This plural constraint here
deserves extra attention

00:17:15.950 --> 00:17:21.868
because it's going to be
desirous of actuating itself

00:17:21.868 --> 00:17:23.839
in the event but the
thing observed in

00:17:23.839 --> 00:17:25.630
the world is plural.

00:17:25.630 --> 00:17:28.520
There are lots of them.

00:17:28.520 --> 00:17:32.285
So it's going to be connected
then to the plural port.

00:17:35.430 --> 00:17:40.050
There's going to be a z sound
port down here connecting to

00:17:40.050 --> 00:17:41.650
that file element
in the buffer.

00:17:46.710 --> 00:17:53.430
And finally, over here is going
to be a plussed voiced

00:17:53.430 --> 00:17:58.370
port, which is going to be
connected to the second

00:17:58.370 --> 00:18:01.270
phoneme in the sequence.

00:18:01.270 --> 00:18:04.130
That's how the machine is
going to be arranged.

00:18:04.130 --> 00:18:07.630
An of course, this is just
one of many constraints.

00:18:07.630 --> 00:18:12.410
But it's a constraint that has
a very peculiar property.

00:18:12.410 --> 00:18:16.410
Information can flow through
it in multiple ways.

00:18:16.410 --> 00:18:18.430
So we think of most programs
as having an

00:18:18.430 --> 00:18:21.430
input and an output.

00:18:21.430 --> 00:18:24.630
But I try to be careful
to draw circles

00:18:24.630 --> 00:18:25.560
here instead of arrows.

00:18:25.560 --> 00:18:29.170
Because these are ports and
information can flow in any

00:18:29.170 --> 00:18:30.890
direction along them.

00:18:30.890 --> 00:18:33.240
What I want to do now is to
show you how this machine

00:18:33.240 --> 00:18:37.540
would react if I suddenly
present it with a pair of

00:18:37.540 --> 00:18:40.570
apples like so.

00:18:40.570 --> 00:18:44.740
So the assumption is that the
vision apparatus comes in and

00:18:44.740 --> 00:18:51.140
produces the notion, the
concept, of two apples.

00:18:51.140 --> 00:18:54.690
So once that has happened--

00:18:54.690 --> 00:18:56.530
that's operation number one--

00:18:59.370 --> 00:19:04.850
then information flows from that
meaning register up here

00:19:04.850 --> 00:19:06.810
to the apple word.

00:19:06.810 --> 00:19:11.700
So that's part of stage
number two.

00:19:11.700 --> 00:19:14.240
Another part of stage number two
is information flows along

00:19:14.240 --> 00:19:21.200
this wire and marks that
as plus plural.

00:19:21.200 --> 00:19:25.290
So operation number one is the
activity of the vision system.

00:19:25.290 --> 00:19:28.770
Activity number two is the flow
of information from that

00:19:28.770 --> 00:19:33.010
vision system into the
word lexicon and

00:19:33.010 --> 00:19:35.090
into this plural register.

00:19:37.690 --> 00:19:38.900
So far so good.

00:19:38.900 --> 00:19:42.560
Here's activity number three.

00:19:42.560 --> 00:19:49.850
This word is also connected
to the registers.

00:19:49.850 --> 00:19:53.900
And information flows along
those wires so as to indicate

00:19:53.900 --> 00:19:56.860
that it's a noun
but not a verb.

00:19:56.860 --> 00:20:00.640
That's part of part
number three.

00:20:00.640 --> 00:20:05.050
At the same time, part number
three, information flows down

00:20:05.050 --> 00:20:11.000
this wire and writes a-p-l
into those are

00:20:11.000 --> 00:20:12.322
elements of the buffer.

00:20:15.940 --> 00:20:20.490
Now this constraint up here,
this box, says, well, I can

00:20:20.490 --> 00:20:22.480
now see some stuff in
that buffer that

00:20:22.480 --> 00:20:24.430
wasn't there before.

00:20:24.430 --> 00:20:28.410
So it says, do I see enough
stuff on my ports to get

00:20:28.410 --> 00:20:33.300
excited about expressing
values on other ports?

00:20:33.300 --> 00:20:33.730
Well, let's see.

00:20:33.730 --> 00:20:34.780
What has it got?

00:20:34.780 --> 00:20:38.890
It's got the elements
in this buffer.

00:20:38.890 --> 00:20:42.270
Also up here in step three
flow the plural thing.

00:20:42.270 --> 00:20:44.400
So it know that the
word is plural.

00:20:44.400 --> 00:20:45.670
So it says, is this voiced?

00:20:48.330 --> 00:20:50.130
P is pa.

00:20:50.130 --> 00:20:51.920
That's not voiced.

00:20:51.920 --> 00:20:53.180
Is this a z sound.

00:20:53.180 --> 00:20:55.410
No, that's not as z sound.

00:20:55.410 --> 00:20:59.930
So it sees what it likes on only
one of its three ports.

00:20:59.930 --> 00:21:01.750
So it says, I'm not going
to do anything.

00:21:01.750 --> 00:21:03.500
I'm [INAUDIBLE].

00:21:03.500 --> 00:21:07.370
I'm not in this particular
combat.

00:21:07.370 --> 00:21:08.440
So far so good.

00:21:08.440 --> 00:21:10.460
What happens next?

00:21:10.460 --> 00:21:15.100
What happens next is that
some time passes.

00:21:15.100 --> 00:21:18.780
And the elements of the buffer
flow to the left toward the

00:21:18.780 --> 00:21:20.820
speaker's mouth.

00:21:20.820 --> 00:21:24.720
So we get an a, p, l.

00:21:24.720 --> 00:21:27.700
Same as we had before,
but shifted over.

00:21:27.700 --> 00:21:28.950
Now what happens?

00:21:31.110 --> 00:21:35.870
Now what happens is that
the l is now in

00:21:35.870 --> 00:21:37.680
the penultimate position.

00:21:37.680 --> 00:21:40.100
So information flows up here.

00:21:40.100 --> 00:21:43.860
Item number four-- oh, I guess
that's item number five.

00:21:43.860 --> 00:21:48.290
Item number four is the leftward
flow of the word.

00:21:48.290 --> 00:21:51.360
So in phase number five, the
p is witnessed by this

00:21:51.360 --> 00:21:52.990
constraint.

00:21:52.990 --> 00:21:54.210
p is--

00:21:54.210 --> 00:21:55.990
sorry, l is witnessed
by this constraint.

00:21:55.990 --> 00:21:57.950
We moved it over one.

00:21:57.950 --> 00:21:58.850
L is lll.

00:21:58.850 --> 00:22:00.570
L is voiced.

00:22:00.570 --> 00:22:03.940
So we have some flow
up here like that.

00:22:03.940 --> 00:22:06.470
That's number five.

00:22:06.470 --> 00:22:10.640
Now we have voiced and
we have plural.

00:22:10.640 --> 00:22:12.940
And we have nothing here.

00:22:12.940 --> 00:22:16.340
So there's a great desire of
this buffer to have something

00:22:16.340 --> 00:22:17.500
written into it.

00:22:17.500 --> 00:22:21.370
So now there's a flow
down in there, of z,

00:22:21.370 --> 00:22:23.170
as item number six.

00:22:23.170 --> 00:22:26.730
So that's how the machine would
work in expressing the

00:22:26.730 --> 00:22:29.580
idea that there are apples
in the field of view.

00:22:35.196 --> 00:22:36.610
Mmm.

00:22:36.610 --> 00:22:37.720
Real apples.

00:22:37.720 --> 00:22:38.970
Not plastic imitations.

00:22:42.870 --> 00:22:46.140
So that's how the
machine works.

00:22:46.140 --> 00:22:47.505
But all those connections
are reversible.

00:22:51.100 --> 00:22:56.545
So if I hear apples then I get
the machine running backwards

00:22:56.545 --> 00:23:00.025
and my visual apparatus
can imagine that there

00:23:00.025 --> 00:23:01.692
are apples out there.

00:23:01.692 --> 00:23:04.030
That's how it works.

00:23:04.030 --> 00:23:07.090
That's just by way of background
the machine that

00:23:07.090 --> 00:23:10.490
they could see it for using the
phonological rules once

00:23:10.490 --> 00:23:11.880
they're learned.

00:23:11.880 --> 00:23:15.400
All the phonological rules
are expressed in these

00:23:15.400 --> 00:23:17.710
constraints.

00:23:17.710 --> 00:23:20.270
But since these constraints are
such that information can

00:23:20.270 --> 00:23:22.990
flow in any direction, they
deserve to be called

00:23:22.990 --> 00:23:24.240
propagators.

00:23:30.290 --> 00:23:33.760
And in the good old days when
everyone took 6.001, they

00:23:33.760 --> 00:23:37.370
learned about propagators as
a kind of architecture for

00:23:37.370 --> 00:23:40.380
building complex systems.

00:23:40.380 --> 00:23:43.060
But in any event, there's
the Sussman-Yip machine.

00:23:43.060 --> 00:23:44.600
And now comes the
big question.

00:23:44.600 --> 00:23:49.010
How do you learn rule
rules like that?

00:23:49.010 --> 00:23:52.120
Well, what we need is we need
some positive examples and

00:23:52.120 --> 00:23:53.370
some negative examples.

00:23:56.910 --> 00:24:00.210
And for the simple classroom
example I've chosen the same

00:24:00.210 --> 00:24:03.780
challenge that I presented
to Krishna.

00:24:03.780 --> 00:24:06.730
We're gonna have
cats and dogs.

00:24:06.730 --> 00:24:10.340
So we're gonna look at the
distinctive features that are

00:24:10.340 --> 00:24:11.620
associated with those words.

00:24:15.220 --> 00:24:17.020
Syllabic.

00:24:17.020 --> 00:24:18.270
Voiced.

00:24:20.050 --> 00:24:21.300
Continuent.

00:24:25.510 --> 00:24:26.760
And strident.

00:24:30.740 --> 00:24:34.130
Just four of the 14 features
that are associated with each

00:24:34.130 --> 00:24:35.940
of the sounds on those words.

00:24:35.940 --> 00:24:37.370
Could you close the
laptop, please?

00:24:40.680 --> 00:24:47.375
Just for the distinctive
features that are arrayed in

00:24:47.375 --> 00:24:51.160
those words by way
of illustration.

00:24:51.160 --> 00:24:52.410
So here we have k-a-t-z.

00:24:56.450 --> 00:24:57.700
Phonetically spelled.

00:25:04.050 --> 00:25:06.810
And if we work that
out, let's see.

00:25:06.810 --> 00:25:08.460
What is syllabic?

00:25:08.460 --> 00:25:08.900
That's not.

00:25:08.900 --> 00:25:09.450
That is.

00:25:09.450 --> 00:25:11.060
That is.

00:25:11.060 --> 00:25:14.090
That's not.

00:25:14.090 --> 00:25:14.610
Voiced?

00:25:14.610 --> 00:25:15.170
Ka.

00:25:15.170 --> 00:25:16.204
Nope.

00:25:16.204 --> 00:25:16.648
Ah.

00:25:16.648 --> 00:25:17.980
Yep.

00:25:17.980 --> 00:25:19.570
T. Nope.

00:25:19.570 --> 00:25:21.540
Z. Yes.

00:25:21.540 --> 00:25:22.660
That can't be right.

00:25:22.660 --> 00:25:25.070
Cats.

00:25:25.070 --> 00:25:26.610
I misspelled it.

00:25:26.610 --> 00:25:27.225
Because cats.

00:25:27.225 --> 00:25:28.310
Sss.

00:25:28.310 --> 00:25:31.590
His a hissing sound but
there's no voicing.

00:25:31.590 --> 00:25:32.910
So that's not as z sound.

00:25:32.910 --> 00:25:35.780
That's an s sound.

00:25:35.780 --> 00:25:36.960
So that's not plus voiced.

00:25:36.960 --> 00:25:38.860
It's minused voiced.

00:25:38.860 --> 00:25:40.980
Continuent.

00:25:40.980 --> 00:25:41.670
Let's see.

00:25:41.670 --> 00:25:43.180
Is my mouth open when I say k?

00:25:43.180 --> 00:25:44.120
No.

00:25:44.120 --> 00:25:44.420
Ah?

00:25:44.420 --> 00:25:45.790
Yes.

00:25:45.790 --> 00:25:46.280
T?

00:25:46.280 --> 00:25:47.660
No.

00:25:47.660 --> 00:25:48.110
S?

00:25:48.110 --> 00:25:49.860
Yes.

00:25:49.860 --> 00:25:50.870
And strident.

00:25:50.870 --> 00:25:53.120
Minus, minus, minus, plus.

00:25:53.120 --> 00:25:56.650
It's only with the s sound that
I have that kind of jet

00:25:56.650 --> 00:25:59.270
forming with my tongue.

00:25:59.270 --> 00:26:00.520
Now we can look at dogs.

00:26:11.950 --> 00:26:16.450
And now we have the z sound
as the pluralization.

00:26:16.450 --> 00:26:18.230
We know that because when
we say it, dogzz.

00:26:18.230 --> 00:26:18.700
Yep.

00:26:18.700 --> 00:26:20.450
There it comes out as a--

00:26:20.450 --> 00:26:23.770
we're only gonna look at the
last two columns because

00:26:23.770 --> 00:26:25.640
they're the only ones that are
going to matter to us.

00:26:25.640 --> 00:26:27.960
So that's plus.

00:26:27.960 --> 00:26:30.906
And that's minus.

00:26:30.906 --> 00:26:32.156
Gu, gu, gu, gu.

00:26:36.330 --> 00:26:37.240
That's plussed.

00:26:37.240 --> 00:26:38.220
And that's plussed.

00:26:38.220 --> 00:26:39.230
They're both voiced.

00:26:39.230 --> 00:26:39.730
Is that right?

00:26:39.730 --> 00:26:41.800
Dogu?

00:26:41.800 --> 00:26:42.700
Gu.

00:26:42.700 --> 00:26:43.170
Gu.

00:26:43.170 --> 00:26:44.430
Is g sound voiced?

00:26:51.430 --> 00:26:53.156
Yeah, I didn't think so.

00:26:53.156 --> 00:26:54.630
G sound is voiced?

00:27:03.420 --> 00:27:04.223
Look-- oh.

00:27:04.223 --> 00:27:08.072
Oh, it is voiced buy it's
not a continuent.

00:27:08.072 --> 00:27:10.400
Just like that.

00:27:10.400 --> 00:27:11.270
Yeah.

00:27:11.270 --> 00:27:12.280
Cat, dogu zz.

00:27:12.280 --> 00:27:12.645
Yeah.

00:27:12.645 --> 00:27:13.760
It is voiced.

00:27:13.760 --> 00:27:16.276
And it has to be for my
example to work out.

00:27:16.276 --> 00:27:20.390
And that's minus, minus,
minus, plus.

00:27:20.390 --> 00:27:22.870
So what we're interested in is,
how come one word gets an

00:27:22.870 --> 00:27:25.020
s sound and how come the other
words gets a z sound?

00:27:28.080 --> 00:27:31.290
Well, it's a pretty sparse
space out there.

00:27:31.290 --> 00:27:34.670
We've already decided that
there are 14,000 possible

00:27:34.670 --> 00:27:37.640
phonemes and there are only
40 in the language.

00:27:37.640 --> 00:27:40.310
So that's one thing
we can consider.

00:27:40.310 --> 00:27:44.800
The other thing that we can
think is that, well, maybe

00:27:44.800 --> 00:27:46.100
this is a logical problem.

00:27:46.100 --> 00:27:47.556
Like the kind of problem
you'd face if you

00:27:47.556 --> 00:27:49.470
were designing a computer.

00:27:49.470 --> 00:27:51.870
And so Sussman and Yip got
stuck for three months

00:27:51.870 --> 00:27:54.250
thinking about the
problem that way.

00:27:54.250 --> 00:27:56.050
Couldn't make any progress
whatsoever.

00:27:56.050 --> 00:27:58.920
And that happens a lot when
you're doing a search.

00:27:58.920 --> 00:28:01.940
You think you've got a way
of approaching it.

00:28:01.940 --> 00:28:03.210
Try to make it work.

00:28:03.210 --> 00:28:05.300
You stay up all night.

00:28:05.300 --> 00:28:06.550
Stay up all night again.

00:28:06.550 --> 00:28:08.310
Still can't make it work.

00:28:08.310 --> 00:28:11.222
Eventually, you abandon ship
and try something else.

00:28:11.222 --> 00:28:14.820
So then they began to say,
well, let's see.

00:28:14.820 --> 00:28:18.880
All we care about is the stuff
before the two ending sounds.

00:28:18.880 --> 00:28:22.580
We care about that part
of the matrix.

00:28:22.580 --> 00:28:25.250
And we care about that
part of the matrix.

00:28:25.250 --> 00:28:29.730
And we can ask, in what ways
are those things different?

00:28:29.730 --> 00:28:31.150
And they're different
all over the place.

00:28:31.150 --> 00:28:33.090
That's why they're
different words.

00:28:33.090 --> 00:28:34.730
We can ask the question a
little bit differently.

00:28:34.730 --> 00:28:38.380
And we can say, what can
we not care about?

00:28:38.380 --> 00:28:41.860
And still retain enough of an
understanding of how the words

00:28:41.860 --> 00:28:46.990
are different so as to put the
proper plural ending on them.

00:28:46.990 --> 00:28:49.075
And they worried about
that for a long time.

00:28:49.075 --> 00:28:50.250
Couldn't find a solution.

00:28:50.250 --> 00:28:52.840
The search space was too big.

00:28:52.840 --> 00:28:57.160
And then they said, maybe what
we ought to do is we ought to

00:28:57.160 --> 00:29:00.550
think about generalizing this
guy here so that we

00:29:00.550 --> 00:29:03.640
don't care about it.

00:29:03.640 --> 00:29:06.080
So now we don't care
about that guy.

00:29:06.080 --> 00:29:09.220
And then he went down through
here saying, well, let's see

00:29:09.220 --> 00:29:11.740
when we have to stop
generalizing.

00:29:11.740 --> 00:29:15.790
Because we've screwed everything
up and we can no

00:29:15.790 --> 00:29:19.200
longer keep the z sound
words separated

00:29:19.200 --> 00:29:22.470
from the s sound words.

00:29:22.470 --> 00:29:24.240
So that eventually distilled
itself down to

00:29:24.240 --> 00:29:25.490
the following algorithm.

00:29:29.710 --> 00:29:37.010
First thing they did was
to collect positive

00:29:37.010 --> 00:29:38.485
and negative examples.

00:29:43.760 --> 00:29:46.870
And there's a positive example
and a negative example.

00:29:46.870 --> 00:29:48.250
That's not enough
to do it right.

00:29:48.250 --> 00:29:51.960
But that's enough to illustrate
the idea.

00:29:51.960 --> 00:29:55.060
So the next thing they did was
something that's extremely

00:29:55.060 --> 00:29:57.930
common in learning anything.

00:29:57.930 --> 00:30:01.950
And that is to pick a positive
example to start from.

00:30:01.950 --> 00:30:05.900
It's actually not a bad idea in
learning anything to start

00:30:05.900 --> 00:30:08.650
with a positive example.

00:30:08.650 --> 00:30:10.480
So they picked a positive
example and they

00:30:10.480 --> 00:30:11.730
called that a seed.

00:30:20.990 --> 00:30:26.070
So in our particular case, cats
is going to be our seed.

00:30:26.070 --> 00:30:29.700
And the question we're going to
ask is, what are the words

00:30:29.700 --> 00:30:34.530
that get pluralized like cat?

00:30:34.530 --> 00:30:37.350
So we've got a positive
and negative example.

00:30:37.350 --> 00:30:38.890
We've picked a seed.

00:30:38.890 --> 00:30:41.385
And now, the next step
is to generalize.

00:30:47.370 --> 00:30:50.360
And what I mean by generalize is
you pick some places in the

00:30:50.360 --> 00:30:54.160
phoneme matrix that you
just don't care about.

00:30:54.160 --> 00:30:56.710
So you may pick a positive
example.

00:30:56.710 --> 00:30:58.100
And you don't care about it.

00:30:58.100 --> 00:31:01.960
So you change it to an asterisk
or, as demonstrated

00:31:01.960 --> 00:31:04.600
in the program I'm about
show you, a ball.

00:31:04.600 --> 00:31:08.820
Or you pick one that's
negative and you

00:31:08.820 --> 00:31:10.240
turn it to a ball.

00:31:10.240 --> 00:31:11.630
Bo.

00:31:11.630 --> 00:31:15.210
So cats, this seed,
becomes a pattern.

00:31:15.210 --> 00:31:18.470
And in order to pluralize the
word this way, you have to

00:31:18.470 --> 00:31:20.100
match all the stuff in here.

00:31:20.100 --> 00:31:22.390
But now what we're going to do
is we're going to gradually

00:31:22.390 --> 00:31:28.860
turn some of those elements into
don't care symbols until

00:31:28.860 --> 00:31:32.910
we get to a point where we've
not cared about so much stuff

00:31:32.910 --> 00:31:34.470
that we think that we
pluralize that one

00:31:34.470 --> 00:31:37.200
with an s sound too.

00:31:37.200 --> 00:31:46.100
So we keep generalizing until
we cover, that is to say we

00:31:46.100 --> 00:31:50.561
admit or match, a negative
example.

00:31:50.561 --> 00:31:53.070
So that's how it works.

00:31:53.070 --> 00:31:54.980
So we generalize like crazy.

00:31:54.980 --> 00:31:58.830
And as soon as we cover a
negative example, we quit.

00:32:01.910 --> 00:32:08.410
Otherwise, we just go back up
here and generalize some more.

00:32:08.410 --> 00:32:12.670
And now we've got to pick a
search technique to decide

00:32:12.670 --> 00:32:14.790
which of these guys to actually
generalize when.

00:32:18.130 --> 00:32:20.940
We could pick one at random.

00:32:20.940 --> 00:32:21.870
And they tried that.

00:32:21.870 --> 00:32:23.641
It didn't work.

00:32:23.641 --> 00:32:26.440
So what they decided is that the
thing that influences the

00:32:26.440 --> 00:32:29.740
pluralization most is the
adjacent phoneme.

00:32:29.740 --> 00:32:32.240
And if that isn't the thing that
solves the problem, it'll

00:32:32.240 --> 00:32:33.760
be the one next to that.

00:32:33.760 --> 00:32:35.837
So in other words, the closer
you are, the more likely you

00:32:35.837 --> 00:32:37.800
are to determine the outcome.

00:32:37.800 --> 00:32:41.770
So these guys over here are
least likely to matter.

00:32:41.770 --> 00:32:43.530
And those are the ones that
are generalized first.

00:32:46.390 --> 00:32:50.390
So if we do that,
what happens?

00:32:50.390 --> 00:32:53.050
Looks like we're going to come
in here and see that there's a

00:32:53.050 --> 00:32:58.025
big difference between the
non-voiced t and the voiced g.

00:32:58.025 --> 00:33:00.130
But that's only a guess because
I've only shown you a

00:33:00.130 --> 00:33:04.900
fraction of the 14 distinctive
features that are involved.

00:33:04.900 --> 00:33:07.952
So I suppose you like to
see a demonstration.

00:33:07.952 --> 00:33:09.202
Yeah.

00:33:25.180 --> 00:33:27.590
So there's our 14 features.

00:33:27.590 --> 00:33:31.350
And that's our seed there,
sitting prominently in the

00:33:31.350 --> 00:33:35.260
display with pluses and minuses
indicating the values

00:33:35.260 --> 00:33:36.860
of the distinctive features
for all three

00:33:36.860 --> 00:33:38.460
of the phones involved.

00:33:38.460 --> 00:33:40.120
That funny left bracket
isn't a mistake.

00:33:40.120 --> 00:33:46.060
That's just one convention for
rendering the ah sound in cat.

00:33:50.510 --> 00:33:53.190
So it's pretty hard to tell from
just that matrix what's

00:33:53.190 --> 00:33:57.400
going to be the determining
feature that separates the

00:33:57.400 --> 00:33:59.630
positive examples from the
negative examples.

00:33:59.630 --> 00:34:01.460
You notice that there
are actually two

00:34:01.460 --> 00:34:02.250
examples down here.

00:34:02.250 --> 00:34:04.460
There's cat and duck.

00:34:04.460 --> 00:34:06.510
Is ducks got an s sound?

00:34:06.510 --> 00:34:06.930
Ducks?

00:34:06.930 --> 00:34:08.650
Yep.

00:34:08.650 --> 00:34:11.420
So dogs and ducks.

00:34:11.420 --> 00:34:14.400
They both get pluralized
with an s sound.

00:34:14.400 --> 00:34:15.449
And then we have
beach doesn't.

00:34:15.449 --> 00:34:17.840
That's beaches.

00:34:17.840 --> 00:34:18.489
Dog.

00:34:18.489 --> 00:34:20.389
We know that's a z.

00:34:20.389 --> 00:34:21.040
Gun.

00:34:21.040 --> 00:34:22.310
Gunz.

00:34:22.310 --> 00:34:25.070
So that's not in the group.

00:34:25.070 --> 00:34:26.900
So we can run this experiment.

00:34:26.900 --> 00:34:27.699
Now here we go.

00:34:27.699 --> 00:34:29.130
We're generalizing like crazy.

00:34:29.130 --> 00:34:30.690
Generalizing, generalizing,
generalizing

00:34:30.690 --> 00:34:33.150
from left to right.

00:34:33.150 --> 00:34:35.810
So nothing in the first
two columns matters.

00:34:35.810 --> 00:34:38.489
Now we get to the t.

00:34:38.489 --> 00:34:39.440
Wow.

00:34:39.440 --> 00:34:40.540
There it is.

00:34:40.540 --> 00:34:43.489
So it looks like you pluralize
with a s sound.

00:34:43.489 --> 00:34:45.780
The sss.

00:34:45.780 --> 00:34:51.830
If, and only if, you're not
voiced and you're not strident

00:34:51.830 --> 00:34:54.940
in the second to the last--

00:34:54.940 --> 00:34:56.330
in the last phone
of the word that

00:34:56.330 --> 00:34:59.400
you're trying to pluralize.

00:34:59.400 --> 00:35:00.950
So that's one phonological
rule that

00:35:00.950 --> 00:35:01.910
the system has learned.

00:35:01.910 --> 00:35:02.330
And guess what?

00:35:02.330 --> 00:35:03.520
It's the same rule
that's found in

00:35:03.520 --> 00:35:05.636
phonological textbooks.

00:35:05.636 --> 00:35:07.225
So now we can try another
experiment.

00:35:13.290 --> 00:35:16.312
So this time we're trying to
deal with dog and gun.

00:35:16.312 --> 00:35:19.630
And our negatives are what was
previously positive plus

00:35:19.630 --> 00:35:23.016
beach, which is still in there
as a negative example.

00:35:23.016 --> 00:35:24.320
So let's see how
that one works.

00:35:32.050 --> 00:35:36.500
Nothing matters except for the
last column, the last phone.

00:35:36.500 --> 00:35:41.390
And now we find out that if the
last sound is voiced, then

00:35:41.390 --> 00:35:46.550
the pluralization gets the z
sound, a voiced determinator.

00:35:46.550 --> 00:35:49.500
And finally, just to
deal with beaches.

00:35:49.500 --> 00:35:51.710
That's beach in it's funny
phonetic spelling.

00:36:04.410 --> 00:36:10.650
So now, if the final sound in
the word is strident, if its

00:36:10.650 --> 00:36:12.430
got this jetty sound--

00:36:12.430 --> 00:36:13.640
beach.

00:36:13.640 --> 00:36:15.530
Beach.

00:36:15.530 --> 00:36:18.840
Then it gets the ea sound.

00:36:18.840 --> 00:36:21.570
So let's go back to experiment
number one.

00:36:21.570 --> 00:36:24.620
Because I want to point out one
small thing about the way

00:36:24.620 --> 00:36:26.180
this works.

00:36:26.180 --> 00:36:28.190
You'll notice that it talks
about coverage and excluded

00:36:28.190 --> 00:36:30.670
down here in the lower
left-hand corner.

00:36:30.670 --> 00:36:33.180
Excluded, well, there are three
negative examples, so

00:36:33.180 --> 00:36:34.630
they better all be excluded.

00:36:34.630 --> 00:36:36.940
You don't want to cover
any of the negatives.

00:36:36.940 --> 00:36:39.220
But it says coverage,
two and two.

00:36:39.220 --> 00:36:42.120
That's because it actually
is doing--

00:36:42.120 --> 00:36:44.530
and now we have the vocabulary
to say it quickly--

00:36:44.530 --> 00:36:47.340
it's doing a beam search
through this space.

00:36:47.340 --> 00:36:48.870
So it's not just doing
a depth first search.

00:36:48.870 --> 00:36:52.840
It's doing a beam search so as
to reduce the possibility of

00:36:52.840 --> 00:36:54.700
overlooking a solution.

00:36:54.700 --> 00:36:56.980
So it says, oh, the coverage.

00:36:56.980 --> 00:37:01.950
Both of the beam search elements
cover both of the

00:37:01.950 --> 00:37:02.950
positive examples.

00:37:02.950 --> 00:37:04.080
And they, in fact, have

00:37:04.080 --> 00:37:06.850
converged to the same solution.

00:37:06.850 --> 00:37:10.920
So that's how the Sussman
and Yip thing worked.

00:37:10.920 --> 00:37:12.470
And then the next question
to ask is, of

00:37:12.470 --> 00:37:16.136
course, why did it work?

00:37:16.136 --> 00:37:20.610
And so the answer,
as articulated

00:37:20.610 --> 00:37:23.270
by Sussman and Yip--

00:37:23.270 --> 00:37:24.816
or rather more by Sussman.

00:37:24.816 --> 00:37:28.290
Or rather more by Yip and a
little bit less by Sussman.

00:37:28.290 --> 00:37:31.840
Yip thinks that it worked
because it's a sparse space.

00:37:31.840 --> 00:37:35.110
And when you have a high
dimensional sparse space, it's

00:37:35.110 --> 00:37:39.660
easy to put a hyperplane into
the space to separate one set

00:37:39.660 --> 00:37:42.100
of examples for another
set of examples.

00:37:42.100 --> 00:37:44.255
So let's consider the
following situation.

00:37:51.390 --> 00:37:57.670
Suppose we have a
one-dimensional situation.

00:37:57.670 --> 00:38:02.320
And we have two white
examples and we

00:38:02.320 --> 00:38:05.670
have two purple examples.

00:38:05.670 --> 00:38:09.700
Well, too bad for us you
can't separate them.

00:38:09.700 --> 00:38:13.150
Now suppose that this is
actually the projection of a

00:38:13.150 --> 00:38:17.280
two-dimensional space that
looks like this.

00:38:17.280 --> 00:38:20.280
Here are the white examples
down here.

00:38:20.280 --> 00:38:24.910
And here are the purple
examples up here.

00:38:24.910 --> 00:38:27.390
Now it's easy to see that you
can separate them with just a

00:38:27.390 --> 00:38:29.810
line that goes across
like that.

00:38:29.810 --> 00:38:34.420
Now let's take this one more
step and suppose that this is

00:38:34.420 --> 00:38:37.326
actually a projection of a
three-dimensional space.

00:38:37.326 --> 00:38:38.610
It looks like this.

00:38:41.610 --> 00:38:43.240
This will be dimension one.

00:38:43.240 --> 00:38:46.630
This'll be two going
back there.

00:38:46.630 --> 00:38:49.900
And this will be
three up here.

00:38:49.900 --> 00:38:53.560
And suppose that the positive
examples are right

00:38:53.560 --> 00:38:54.810
here on this line.

00:39:01.120 --> 00:39:03.980
Let's say this is-- well, we're
gonna draw a little old

00:39:03.980 --> 00:39:05.960
cube like so.

00:39:05.960 --> 00:39:10.420
Those are purple examples
that are up there.

00:39:10.420 --> 00:39:13.150
How many ways are there of
partitioning the space along

00:39:13.150 --> 00:39:13.830
those axes?

00:39:13.830 --> 00:39:16.510
Well, now they're not
even just two.

00:39:16.510 --> 00:39:17.940
They're three.

00:39:17.940 --> 00:39:23.620
So one way to separate the
purple from the white is to

00:39:23.620 --> 00:39:28.290
draw a hyperplane-- or in this
case it's a three dimension,

00:39:28.290 --> 00:39:29.190
so a plane--

00:39:29.190 --> 00:39:33.160
through here on the
number three axis.

00:39:33.160 --> 00:39:36.010
You could also put a plane
in on that axis.

00:39:36.010 --> 00:39:38.100
Or you could do both.

00:39:38.100 --> 00:39:43.458
So in one case your dividing
line would be--

00:39:43.458 --> 00:39:44.210
let's see.

00:39:44.210 --> 00:39:47.240
On the first axis that
would be 1/2.

00:39:47.240 --> 00:39:49.220
And then the don't care.

00:39:49.220 --> 00:39:50.670
Don't care.

00:39:50.670 --> 00:39:53.580
Another solution that
would be don't care.

00:39:53.580 --> 00:39:57.830
And then we divide on the number
2 axis with a plane at

00:39:57.830 --> 00:40:00.430
1/2 and don't care.

00:40:00.430 --> 00:40:07.405
Or we could do it with 1/2,
1/2, and don't care.

00:40:07.405 --> 00:40:10.990
So the higher the dimension of
the space, the easier it is

00:40:10.990 --> 00:40:13.750
sometimes to put in a plane
that separates the data.

00:40:13.750 --> 00:40:17.210
That's why Sussman and Yip think
that we use so little of

00:40:17.210 --> 00:40:18.510
possible phoneme space.

00:40:18.510 --> 00:40:20.910
Because it makes the
thing learnable.

00:40:20.910 --> 00:40:23.920
That's one possibility.

00:40:23.920 --> 00:40:30.090
So one explanation for sparse
space is learnability.

00:40:30.090 --> 00:40:33.790
There's another interesting
possibility, and that is that

00:40:33.790 --> 00:40:37.500
if you have a sparse space, high
dimensional space with 14

00:40:37.500 --> 00:40:42.850
dimensions, and if the 40 points
of your language are

00:40:42.850 --> 00:40:46.260
spread evenly throughout
that space--

00:40:46.260 --> 00:40:47.490
now let me say it
the other way.

00:40:47.490 --> 00:40:51.050
If they are placed at random in
that space, then according

00:40:51.050 --> 00:40:53.190
to the central limit theorem,
then they'll be about equally

00:40:53.190 --> 00:40:55.320
distant from each other.

00:40:55.320 --> 00:40:59.050
So it ensures that the phonemes
are easily separated

00:40:59.050 --> 00:41:01.870
when you speak.

00:41:01.870 --> 00:41:06.480
But if you go to ask a linguist
if that's true, they

00:41:06.480 --> 00:41:07.220
don't know.

00:41:07.220 --> 00:41:08.240
Because they're not looking
at it from a

00:41:08.240 --> 00:41:09.865
computational point of view.

00:41:09.865 --> 00:41:12.840
Well, we can look at it from a
computational point of view.

00:41:12.840 --> 00:41:14.860
So I did that.

00:41:14.860 --> 00:41:16.940
After Sussman and Yip published
their paper.

00:41:16.940 --> 00:41:18.190
And here's the result.

00:41:21.330 --> 00:41:26.690
This is a diagram that shows all
of the phonemes that are

00:41:26.690 --> 00:41:30.170
separated by exactly one
distinctive feature.

00:41:30.170 --> 00:41:32.505
So if you look over in this
corner here, you'll see that

00:41:32.505 --> 00:41:34.640
the constants-- w and x--

00:41:34.640 --> 00:41:38.520
are separated by exactly one
distinctive feature.

00:41:38.520 --> 00:41:43.110
So they're not exactly distant
from each other in the space.

00:41:43.110 --> 00:41:45.020
On the other hand, they are
pretty easy to separate

00:41:45.020 --> 00:41:46.960
relative to the vowels.

00:41:46.960 --> 00:41:49.650
Which are here in this
part of the diagram.

00:41:49.650 --> 00:41:52.230
Which are all tangled up and
the vowels are all close to

00:41:52.230 --> 00:41:52.670
each other.

00:41:52.670 --> 00:41:54.380
So guess what?

00:41:54.380 --> 00:41:57.450
Vowels are much harder to
separate than constants.

00:41:57.450 --> 00:42:00.710
Not surprisingly, because there
are many pairs of them

00:42:00.710 --> 00:42:01.650
that are different.

00:42:01.650 --> 00:42:04.914
And only one distinctive
feature.

00:42:04.914 --> 00:42:05.580
All right.

00:42:05.580 --> 00:42:07.380
So now you back up and
you say, well, gosh.

00:42:07.380 --> 00:42:08.590
That's all been sort
of interesting.

00:42:08.590 --> 00:42:11.180
But what does it teach
us about how to

00:42:11.180 --> 00:42:13.062
do science and stuff?

00:42:13.062 --> 00:42:14.990
And what it teaches us is--

00:42:14.990 --> 00:42:17.320
this is an example.

00:42:17.320 --> 00:42:18.570
Ow.

00:42:24.360 --> 00:42:27.850
This is an example which we can
use to illuminate some of

00:42:27.850 --> 00:42:30.030
thoughts of David Marr, who
I spoke of in a previous

00:42:30.030 --> 00:42:32.882
lecture, connection
with vision.

00:42:32.882 --> 00:42:36.560
But here's Marr's catechism.

00:42:36.560 --> 00:42:38.930
I can't spell very well so I
won't try to respell it.

00:42:38.930 --> 00:42:41.260
But this is Marr's catechism.

00:42:41.260 --> 00:42:43.910
So what Marr said is, when
you're dealing with an AI

00:42:43.910 --> 00:42:45.785
problem, first thing to do is
to specify the problem.

00:42:48.425 --> 00:42:51.380
Gee, that sounds
awfully normal.

00:42:51.380 --> 00:42:57.410
The next thing is to devise
a representation

00:42:57.410 --> 00:42:58.660
suited to the problem.

00:43:02.660 --> 00:43:06.130
The third thing to do,
vocabulary varies, but it's

00:43:06.130 --> 00:43:07.980
something like determine
an approach.

00:43:11.660 --> 00:43:14.395
Sometimes thought
of as a method.

00:43:17.080 --> 00:43:28.048
And then four, pick a mechanism

00:43:28.048 --> 00:43:29.380
or devise an algorithm.

00:43:36.210 --> 00:43:38.710
And, finally, five,
experiment.

00:43:45.840 --> 00:43:48.930
And of course, it never goes
linearly like that.

00:43:48.930 --> 00:43:51.900
You start with the problem and
then you go through a lot of

00:43:51.900 --> 00:43:52.570
loops up here.

00:43:52.570 --> 00:43:54.215
Sometimes even changing
the problem.

00:43:56.740 --> 00:43:58.320
But that's just the scientific
method, right?

00:43:58.320 --> 00:43:59.950
You start with the problem
and you end up with the

00:43:59.950 --> 00:44:01.250
experiment.

00:44:01.250 --> 00:44:05.820
But that's not what people in
AI, over the bulk of its

00:44:05.820 --> 00:44:09.050
existence, have tended to do.

00:44:09.050 --> 00:44:13.300
What they tended to do is to
fall in love with particular

00:44:13.300 --> 00:44:15.170
mechanisms.

00:44:15.170 --> 00:44:16.470
And then they attempt
to apply those

00:44:16.470 --> 00:44:18.690
mechanisms to every problem.

00:44:18.690 --> 00:44:21.680
So you might say, well, gee,
neural nets are so cool.

00:44:21.680 --> 00:44:24.545
I think all of human
intelligence can be explained

00:44:24.545 --> 00:44:27.170
with a suitable neural net.

00:44:27.170 --> 00:44:28.645
That's not the right
way to do it.

00:44:28.645 --> 00:44:30.130
Because that's mechanism envy.

00:44:30.130 --> 00:44:31.275
You fall in love
with mechanism.

00:44:31.275 --> 00:44:34.760
You try to apply it where it
isn't the right thing.

00:44:34.760 --> 00:44:38.350
This is example starting with
the problem and bringing to

00:44:38.350 --> 00:44:40.880
the problem the right
representations, gosh,

00:44:40.880 --> 00:44:42.840
distinctive features.

00:44:42.840 --> 00:44:45.630
Once we've got the right
representation, then the

00:44:45.630 --> 00:44:48.520
constraints emerge, which
enable us to devise an

00:44:48.520 --> 00:44:51.100
approach, write an algorithm,
and do an experiment.

00:44:51.100 --> 00:44:52.830
As they did.

00:44:52.830 --> 00:44:58.560
So this Sussman-Yip thing is an
example of doing AI stuff

00:44:58.560 --> 00:45:02.700
in a way that's congruent with
the Marr's catechism.

00:45:02.700 --> 00:45:04.650
Which I highly recommend.

00:45:04.650 --> 00:45:08.040
They could have come in here and
said, well, we're devotees

00:45:08.040 --> 00:45:10.980
of the idea of neural nets.

00:45:10.980 --> 00:45:14.480
Let's see if we can make a
machine that will properly

00:45:14.480 --> 00:45:17.650
pluralize words using
a neural net.

00:45:17.650 --> 00:45:19.370
That's a loser.

00:45:19.370 --> 00:45:22.140
Because it doesn't match the
problem to the mechanism.

00:45:22.140 --> 00:45:25.450
It tries to force fit the
mechanism into some

00:45:25.450 --> 00:45:28.840
Procrustean bed where
it doesn't

00:45:28.840 --> 00:45:31.820
actually work very well.

00:45:31.820 --> 00:45:34.350
So what this leaves open, of
course, is the question of,

00:45:34.350 --> 00:45:39.140
well, what is a good
representation?

00:45:39.140 --> 00:45:41.910
And here's the other half
Marr's catechism.

00:45:41.910 --> 00:45:45.300
Characteristic number one
is that it makes the

00:45:45.300 --> 00:45:46.550
right things explicit.

00:45:49.540 --> 00:45:51.280
So in this particular
case, it makes

00:45:51.280 --> 00:45:54.750
distinctive features explicit.

00:45:54.750 --> 00:45:59.340
Another thing that Marr was
noted for was stereo vision.

00:45:59.340 --> 00:46:05.830
So in that particular world,
discontinuities in the image,

00:46:05.830 --> 00:46:07.420
when you go across an
edge with the things

00:46:07.420 --> 00:46:10.160
that were made explicit.

00:46:10.160 --> 00:46:12.020
Once you've got to a
representation that makes the

00:46:12.020 --> 00:46:15.850
right things explicit, you can
say, does it also expose

00:46:15.850 --> 00:46:17.100
constraint?

00:46:25.730 --> 00:46:27.320
And if you have a representation
that exposes

00:46:27.320 --> 00:46:28.450
constraint, then you're
off and running.

00:46:28.450 --> 00:46:30.690
Because it's constraint that
you need in order to do the

00:46:30.690 --> 00:46:35.050
processing that leads
to a solution.

00:46:35.050 --> 00:46:36.055
So don't have the right
representation.

00:46:36.055 --> 00:46:38.080
If it doesn't expose
constraints, you're not going

00:46:38.080 --> 00:46:40.930
to be able to make a very
good model out of it.

00:46:40.930 --> 00:46:46.475
And finally, there's a kind
of localness criteria.

00:46:49.730 --> 00:46:52.990
If you have a representation in
which you can see the right

00:46:52.990 --> 00:46:55.780
answer by looking at
descriptions through soda

00:46:55.780 --> 00:46:57.800
straw, that's probably a better
representation than one

00:46:57.800 --> 00:46:59.700
that's all spread out.

00:46:59.700 --> 00:47:01.185
It's true with programs,
right?

00:47:01.185 --> 00:47:03.360
If you can see how they work
by looking through a soda

00:47:03.360 --> 00:47:06.450
straw, you're in much better
situation to understand

00:47:06.450 --> 00:47:08.770
something if you have to look
here and there and on the next

00:47:08.770 --> 00:47:11.560
page and in the next file.

00:47:11.560 --> 00:47:15.510
So all this is basically
common sense.

00:47:15.510 --> 00:47:18.560
But this is kind of common sense
that makes you smarter

00:47:18.560 --> 00:47:20.200
as an engineer and scientist.

00:47:20.200 --> 00:47:23.080
Especially as a scientist
because if you go into a

00:47:23.080 --> 00:47:27.850
problem with mechanism envy,
you're apt to study mechanisms

00:47:27.850 --> 00:47:32.280
in a naive way and never reach
a solution that will be

00:47:32.280 --> 00:47:33.530
satisfactory.