WEBVTT

00:00:00.000 --> 00:00:01.964 align:middle line:90%
[SQUEAKING]

00:00:01.964 --> 00:00:03.916 align:middle line:90%
[RUSTLING]

00:00:03.916 --> 00:00:06.356 align:middle line:90%
[CLICKING]

00:00:06.356 --> 00:00:12.466 align:middle line:90%


00:00:12.466 --> 00:00:18.000 align:middle line:84%
SARA BEERY: OK, so welcome to
lecture two of deep learning.

00:00:18.000 --> 00:00:20.240 align:middle line:84%
Today, we're going to talk
about how you actually

00:00:20.240 --> 00:00:22.760 align:middle line:90%
train a neural network.

00:00:22.760 --> 00:00:25.780 align:middle line:84%
And I imagine that some of this
might be review for some of you

00:00:25.780 --> 00:00:27.880 align:middle line:84%
but, hopefully, the
perspective or the way

00:00:27.880 --> 00:00:31.000 align:middle line:84%
we talk about the process of
training the neural network

00:00:31.000 --> 00:00:33.400 align:middle line:84%
might be a slightly
reformulated from something

00:00:33.400 --> 00:00:34.440 align:middle line:90%
you've seen before.

00:00:34.440 --> 00:00:36.440 align:middle line:84%
I definitely think that
this slide is a nice way

00:00:36.440 --> 00:00:39.140 align:middle line:90%
to characterize the process.

00:00:39.140 --> 00:00:41.960 align:middle line:84%
So today, we're going to
talk about how you actually

00:00:41.960 --> 00:00:44.065 align:middle line:90%
train a neural network.

00:00:44.065 --> 00:00:45.440 align:middle line:84%
First, we're going
to start doing

00:00:45.440 --> 00:00:49.480 align:middle line:84%
a bit of a review of gradient
descent SGD, high-level,

00:00:49.480 --> 00:00:51.690 align:middle line:90%
some optimization ideas.

00:00:51.690 --> 00:00:53.440 align:middle line:84%
Then, we're going to
introduce the concept

00:00:53.440 --> 00:00:55.520 align:middle line:90%
of a computational graph.

00:00:55.520 --> 00:00:58.200 align:middle line:84%
We'll talk about
backpropagation through chains

00:00:58.200 --> 00:01:00.480 align:middle line:84%
and through multi-layer
perceptrons.

00:01:00.480 --> 00:01:04.000 align:middle line:84%
We'll talk about
backpropagation through DAGs.

00:01:04.000 --> 00:01:07.880 align:middle line:84%
And then we will get into
the concept of-- sorry--

00:01:07.880 --> 00:01:10.080 align:middle line:90%
differential programming.

00:01:10.080 --> 00:01:14.783 align:middle line:84%
OK, so just, again, as a
bit of a refresher, when

00:01:14.783 --> 00:01:16.200 align:middle line:84%
we talk about
neural network, it's

00:01:16.200 --> 00:01:21.080 align:middle line:84%
usually something that follows
the rough structure of this,

00:01:21.080 --> 00:01:24.080 align:middle line:84%
where you will have
some input data x.

00:01:24.080 --> 00:01:29.040 align:middle line:84%
And during training time,
you'll have some ground truth

00:01:29.040 --> 00:01:30.140 align:middle line:90%
for that input data.

00:01:30.140 --> 00:01:32.760 align:middle line:84%
This is the thing that you want
the model to learn to predict.

00:01:32.760 --> 00:01:35.380 align:middle line:84%
So in this particular
instance, that is a clownfish.

00:01:35.380 --> 00:01:36.880 align:middle line:84%
You're asking this
model to predict,

00:01:36.880 --> 00:01:39.960 align:middle line:84%
maybe, a visual
category, and that's y.

00:01:39.960 --> 00:01:41.960 align:middle line:90%
And then we have some model.

00:01:41.960 --> 00:01:45.920 align:middle line:84%
This is a chain of
different blocks or layers

00:01:45.920 --> 00:01:47.600 align:middle line:90%
of a neural network.

00:01:47.600 --> 00:01:51.480 align:middle line:84%
Each of those layers will have
parameters associated with them.

00:01:51.480 --> 00:01:53.360 align:middle line:84%
And then once you
send this input

00:01:53.360 --> 00:01:54.880 align:middle line:84%
through all of
those layers, you'll

00:01:54.880 --> 00:01:57.670 align:middle line:84%
calculate a loss based on
the difference between what

00:01:57.670 --> 00:02:00.270 align:middle line:84%
the model predicts
and that ground truth.

00:02:00.270 --> 00:02:02.235 align:middle line:84%
And all of those
parameters that are

00:02:02.235 --> 00:02:04.110 align:middle line:84%
going through the network,
those are actually

00:02:04.110 --> 00:02:05.590 align:middle line:90%
what's getting learned.

00:02:05.590 --> 00:02:07.850 align:middle line:84%
And what when you
say getting learned,

00:02:07.850 --> 00:02:09.789 align:middle line:84%
essentially, we just
mean we wiggle the values

00:02:09.789 --> 00:02:12.750 align:middle line:84%
around in numerical space
until we get something

00:02:12.750 --> 00:02:16.950 align:middle line:84%
that is at least reasonably
close to optimal,

00:02:16.950 --> 00:02:22.870 align:middle line:84%
based on a set of data that
we actually have access to.

00:02:22.870 --> 00:02:23.390 align:middle line:90%
Cool.

00:02:23.390 --> 00:02:25.870 align:middle line:84%
So essentially what
you're trying to find

00:02:25.870 --> 00:02:28.550 align:middle line:84%
is some optimal
set of theta that

00:02:28.550 --> 00:02:34.350 align:middle line:84%
minimizes that loss function
over all of your available data

00:02:34.350 --> 00:02:35.870 align:middle line:90%
set.

00:02:35.870 --> 00:02:38.950 align:middle line:84%
Now, when we talk
about gradient descent,

00:02:38.950 --> 00:02:45.510 align:middle line:84%
here, finding that optimal set
of parameters, what we often

00:02:45.510 --> 00:02:47.830 align:middle line:84%
will talk about is
that that's going

00:02:47.830 --> 00:02:54.070 align:middle line:84%
to be the set of parameters
that is optimal for some cost.

00:02:54.070 --> 00:02:56.550 align:middle line:84%
In this particular
sense, that cost function

00:02:56.550 --> 00:02:59.430 align:middle line:84%
is the sum of the loss over all
the data points in our training

00:02:59.430 --> 00:03:01.717 align:middle line:90%
data set.

00:03:01.717 --> 00:03:04.050 align:middle line:84%
And then how do we actually
fiddle those weights around?

00:03:04.050 --> 00:03:06.070 align:middle line:90%
What's the mechanism?

00:03:06.070 --> 00:03:08.950 align:middle line:84%
That's really just
using optimization

00:03:08.950 --> 00:03:11.870 align:middle line:84%
to try to find those optimal
parameters for that cost

00:03:11.870 --> 00:03:13.230 align:middle line:90%
function.

00:03:13.230 --> 00:03:16.390 align:middle line:84%
So what's the knowledge
that we have about J?

00:03:16.390 --> 00:03:21.110 align:middle line:84%
Well, first, we know that we
can actually evaluate J of theta

00:03:21.110 --> 00:03:22.233 align:middle line:90%
for any given theta.

00:03:22.233 --> 00:03:24.650 align:middle line:84%
So this is something that we
can actually evaluate, right?

00:03:24.650 --> 00:03:27.510 align:middle line:84%
You can calculate the loss over
a training data set or a test

00:03:27.510 --> 00:03:29.310 align:middle line:90%
data set.

00:03:29.310 --> 00:03:34.830 align:middle line:84%
We also know that we can
evaluate the gradient of theta.

00:03:34.830 --> 00:03:37.750 align:middle line:84%
And this is first
order optimization.

00:03:37.750 --> 00:03:40.430 align:middle line:84%
You might take some
approximation of the gradient

00:03:40.430 --> 00:03:42.450 align:middle line:90%
based on a specific point.

00:03:42.450 --> 00:03:44.190 align:middle line:84%
And you can evaluate
that as well.

00:03:44.190 --> 00:03:46.070 align:middle line:84%
And, additionally, we
can also do something

00:03:46.070 --> 00:03:49.110 align:middle line:84%
like evaluating second order
optimization, where you could

00:03:49.110 --> 00:03:50.950 align:middle line:90%
actually calculate the Hessian.

00:03:50.950 --> 00:03:52.390 align:middle line:84%
In practice, this
isn't something

00:03:52.390 --> 00:03:54.580 align:middle line:90%
that we often actually do.

00:03:54.580 --> 00:03:58.100 align:middle line:84%
Usually, we really just
focus on that first and then

00:03:58.100 --> 00:04:00.140 align:middle line:84%
that first order
optimization, where we're just

00:04:00.140 --> 00:04:02.300 align:middle line:84%
looking at that
linear approximation

00:04:02.300 --> 00:04:05.700 align:middle line:84%
to the gradient
at a given point.

00:04:05.700 --> 00:04:07.500 align:middle line:84%
And then what is
gradient descent?

00:04:07.500 --> 00:04:10.340 align:middle line:84%
So if you're trying to
find this theta star--

00:04:10.340 --> 00:04:12.725 align:middle line:90%
that's this optimal value--

00:04:12.725 --> 00:04:14.100 align:middle line:84%
what you're going
to do is you're

00:04:14.100 --> 00:04:16.660 align:middle line:84%
going to start somewhere
on what we call the loss

00:04:16.660 --> 00:04:20.260 align:middle line:84%
landscape or the surface that
represents that cost function.

00:04:20.260 --> 00:04:22.580 align:middle line:84%
And then you're going
to move in the direction

00:04:22.580 --> 00:04:25.860 align:middle line:84%
of the gradient, the
maximal gradient,

00:04:25.860 --> 00:04:28.140 align:middle line:90%
and takes steps over your data.

00:04:28.140 --> 00:04:30.380 align:middle line:84%
And so for each of
these, you'll start

00:04:30.380 --> 00:04:33.540 align:middle line:84%
moving in that downward
direction on that loss landscape

00:04:33.540 --> 00:04:35.452 align:middle line:90%
and eventually find a minimum.

00:04:35.452 --> 00:04:37.660 align:middle line:84%
Now of course, you might
note, if this loss landscape

00:04:37.660 --> 00:04:39.580 align:middle line:84%
is non-convex, if
it's quite complex

00:04:39.580 --> 00:04:41.500 align:middle line:84%
and maybe has
multiple minima, this

00:04:41.500 --> 00:04:43.960 align:middle line:84%
may not be guaranteed to
be the global minimum.

00:04:43.960 --> 00:04:46.020 align:middle line:84%
It might just be
a local minimum.

00:04:46.020 --> 00:04:47.860 align:middle line:84%
And how bad that
local minimum is

00:04:47.860 --> 00:04:52.420 align:middle line:84%
relative to the global
minimum, often, we don't know.

00:04:52.420 --> 00:04:55.500 align:middle line:84%
So if you assume you're
taking a single iteration

00:04:55.500 --> 00:04:59.980 align:middle line:84%
of gradient descent and trying
to find better parameters,

00:04:59.980 --> 00:05:04.620 align:middle line:84%
what we'll do is we will
start with our k-th step

00:05:04.620 --> 00:05:06.820 align:middle line:84%
of those parameters
theta, and then

00:05:06.820 --> 00:05:10.940 align:middle line:84%
we will define some learning
rate, in this case eta,

00:05:10.940 --> 00:05:12.620 align:middle line:84%
and that's going to
say how far we're

00:05:12.620 --> 00:05:15.140 align:middle line:84%
going to step in the
direction of that gradient.

00:05:15.140 --> 00:05:17.700 align:middle line:84%
And then you
calculate, essentially,

00:05:17.700 --> 00:05:23.900 align:middle line:84%
what is the direction of change
that will optimize that loss

00:05:23.900 --> 00:05:25.460 align:middle line:90%
or that cost function.

00:05:25.460 --> 00:05:29.020 align:middle line:84%
So now what you do is you take
the parameters of your model,

00:05:29.020 --> 00:05:32.100 align:middle line:84%
and you move them
in the direction

00:05:32.100 --> 00:05:35.620 align:middle line:84%
of the maximal slope
of your gradient

00:05:35.620 --> 00:05:38.920 align:middle line:84%
with some of learning
rate step value.

00:05:38.920 --> 00:05:42.260 align:middle line:84%
And that gives you the
update to your parameters.

00:05:42.260 --> 00:05:46.700 align:middle line:84%
Now, this might be something
that you could reasonably

00:05:46.700 --> 00:05:50.330 align:middle line:84%
calculate over all of your
existing training data.

00:05:50.330 --> 00:05:54.690 align:middle line:84%
You'd like to, ideally, minimize
the overall loss function,

00:05:54.690 --> 00:05:57.010 align:middle line:84%
which would be some sum
of the individual losses

00:05:57.010 --> 00:06:00.210 align:middle line:84%
over every single example
in your training data.

00:06:00.210 --> 00:06:03.770 align:middle line:84%
But realistically, this is
often not possible to calculate,

00:06:03.770 --> 00:06:07.290 align:middle line:84%
mostly just due to the time,
the computational complexity,

00:06:07.290 --> 00:06:10.610 align:middle line:84%
especially in this day
and age when data sets are

00:06:10.610 --> 00:06:12.530 align:middle line:90%
getting larger and larger.

00:06:12.530 --> 00:06:15.470 align:middle line:84%
So when I started
in machine learning,

00:06:15.470 --> 00:06:18.170 align:middle line:84%
we were lucky if we
had ImageNet scale data

00:06:18.170 --> 00:06:19.890 align:middle line:90%
sets of a million images.

00:06:19.890 --> 00:06:23.010 align:middle line:84%
Now we're dealing with image
data sets in the billions

00:06:23.010 --> 00:06:26.490 align:middle line:84%
and text data sets
even larger than that.

00:06:26.490 --> 00:06:29.890 align:middle line:84%
So instead, in stochastic
gradient descent,

00:06:29.890 --> 00:06:32.490 align:middle line:84%
you assume that
taking the gradient

00:06:32.490 --> 00:06:35.250 align:middle line:84%
over some subset of your data
is a reasonable approximation

00:06:35.250 --> 00:06:38.070 align:middle line:84%
to taking the gradient
over the entire data set.

00:06:38.070 --> 00:06:39.570 align:middle line:84%
So here what you
do is, instead, you

00:06:39.570 --> 00:06:43.030 align:middle line:84%
compute a gradient on a subset
of the data at any given point.

00:06:43.030 --> 00:06:45.050 align:middle line:90%
And we call this a batch.

00:06:45.050 --> 00:06:47.570 align:middle line:84%
So if your batch size is
1, what that would mean

00:06:47.570 --> 00:06:50.130 align:middle line:84%
is you would actually
update your parameters

00:06:50.130 --> 00:06:53.450 align:middle line:84%
after seeing every single
example in the data set.

00:06:53.450 --> 00:06:55.650 align:middle line:84%
You can imagine
that also might be

00:06:55.650 --> 00:06:57.850 align:middle line:84%
suboptimal in terms of the
computational complexity

00:06:57.850 --> 00:06:59.690 align:middle line:90%
of doing those updates.

00:06:59.690 --> 00:07:02.830 align:middle line:84%
If your batch size is N, where
N is the full set of data,

00:07:02.830 --> 00:07:05.170 align:middle line:84%
then this is just
standard gradient descent.

00:07:05.170 --> 00:07:06.670 align:middle line:84%
But of course, it
might be possible.

00:07:06.670 --> 00:07:09.210 align:middle line:84%
This is an easy to fit into
the memory of your GPU,

00:07:09.210 --> 00:07:13.610 align:middle line:84%
or might just be really slow,
computationally intractable.

00:07:13.610 --> 00:07:16.450 align:middle line:84%
So also, if your
gradient direction

00:07:16.450 --> 00:07:20.930 align:middle line:84%
is kind of noisy relative
to averaging over

00:07:20.930 --> 00:07:26.410 align:middle line:84%
all of the examples, which
is standard gradient descent,

00:07:26.410 --> 00:07:27.890 align:middle line:90%
this is pretty normal.

00:07:27.890 --> 00:07:31.770 align:middle line:84%
So taking a batch, you're
not necessarily guaranteed

00:07:31.770 --> 00:07:34.290 align:middle line:84%
that gradient over the
batch is going to match

00:07:34.290 --> 00:07:35.850 align:middle line:90%
the gradient over everything.

00:07:35.850 --> 00:07:38.250 align:middle line:84%
And actually in a lot of
the data that I work with,

00:07:38.250 --> 00:07:41.130 align:middle line:84%
where we have things like
significant imbalance

00:07:41.130 --> 00:07:45.410 align:middle line:84%
over the types of categories
we're interested in predicting,

00:07:45.410 --> 00:07:47.920 align:middle line:84%
this sampling can have a
pretty strong effect in terms

00:07:47.920 --> 00:07:49.560 align:middle line:90%
of how stable that gradient is.

00:07:49.560 --> 00:07:52.560 align:middle line:84%
Because potentially you
have some categories

00:07:52.560 --> 00:07:55.440 align:middle line:84%
that are not seen for
multiple batches at a time,

00:07:55.440 --> 00:07:57.100 align:middle line:90%
and then you see them.

00:07:57.100 --> 00:08:00.000 align:middle line:84%
And at that point, it actually
massively shifts your gradient.

00:08:00.000 --> 00:08:02.560 align:middle line:84%
So this instability is
actually often related

00:08:02.560 --> 00:08:08.880 align:middle line:84%
to how uniformly distributed
your data set is.

00:08:08.880 --> 00:08:12.640 align:middle line:84%
But there are advantages of
doing that stochastic gradient

00:08:12.640 --> 00:08:13.400 align:middle line:90%
descent.

00:08:13.400 --> 00:08:15.240 align:middle line:90%
First, it's faster.

00:08:15.240 --> 00:08:17.920 align:middle line:84%
It's been shown to approximate
the total gradient,

00:08:17.920 --> 00:08:20.120 align:middle line:90%
even with small samples.

00:08:20.120 --> 00:08:23.860 align:middle line:84%
And that noise, it's actually
an implicit regularizer.

00:08:23.860 --> 00:08:27.360 align:middle line:84%
So one thing that can be a
benefit of stochastic gradient

00:08:27.360 --> 00:08:31.240 align:middle line:84%
descent is it can actually
bounce you out of local minima

00:08:31.240 --> 00:08:35.320 align:middle line:84%
and help you reasonably
find closer to what

00:08:35.320 --> 00:08:37.559 align:middle line:90%
might be a global minimum.

00:08:37.559 --> 00:08:39.440 align:middle line:90%
There are disadvantages.

00:08:39.440 --> 00:08:41.659 align:middle line:84%
Like we said, there
can be high variance.

00:08:41.659 --> 00:08:43.140 align:middle line:90%
There can be unstable updates.

00:08:43.140 --> 00:08:45.960 align:middle line:84%
And I've found experimentally
that these updates are

00:08:45.960 --> 00:08:49.320 align:middle line:84%
more unstable the
more unbalanced

00:08:49.320 --> 00:08:51.800 align:middle line:84%
your training data set is
over the set of categories

00:08:51.800 --> 00:08:54.480 align:middle line:84%
that you're interested
to predict on.

00:08:54.480 --> 00:08:55.020 align:middle line:90%
Cool.

00:08:55.020 --> 00:08:57.160 align:middle line:90%
Any questions about SGD?

00:08:57.160 --> 00:08:59.960 align:middle line:84%
How many of you have
seen SGD before?

00:08:59.960 --> 00:09:02.040 align:middle line:90%
OK, that's what I thought.

00:09:02.040 --> 00:09:05.960 align:middle line:84%
So this is, hopefully, a bit
of a refresher for most of you.

00:09:05.960 --> 00:09:10.640 align:middle line:84%
So we also want to introduce
the concept of momentum.

00:09:10.640 --> 00:09:15.080 align:middle line:84%
And this is something that we
might implement within our loss

00:09:15.080 --> 00:09:18.040 align:middle line:90%
or within our gradient descent.

00:09:18.040 --> 00:09:20.600 align:middle line:84%
So, essentially,
the idea of momentum

00:09:20.600 --> 00:09:23.260 align:middle line:84%
is the physical
idea of momentum.

00:09:23.260 --> 00:09:25.780 align:middle line:84%
If you put something
heavy on a slope,

00:09:25.780 --> 00:09:28.720 align:middle line:84%
it will gain speed as
it rolls down the slope.

00:09:28.720 --> 00:09:30.640 align:middle line:84%
And what we do
when we think about

00:09:30.640 --> 00:09:33.120 align:middle line:84%
momentum in gradient
descent is we,

00:09:33.120 --> 00:09:35.360 align:middle line:84%
essentially, are biasing
our gradient steps

00:09:35.360 --> 00:09:38.100 align:middle line:84%
to continue in the direction
of the previous update.

00:09:38.100 --> 00:09:41.600 align:middle line:84%
You're giving it that momentum
based on its past direction

00:09:41.600 --> 00:09:42.417 align:middle line:90%
of movement.

00:09:42.417 --> 00:09:44.750 align:middle line:84%
And so, here, what that looks
like is you're essentially

00:09:44.750 --> 00:09:47.670 align:middle line:84%
just adding in, again,
this parameterized

00:09:47.670 --> 00:09:51.310 align:middle line:84%
by alpha momentum term that is
capturing the direction that you

00:09:51.310 --> 00:09:52.990 align:middle line:90%
moved in the past.

00:09:52.990 --> 00:09:56.313 align:middle line:84%
And this can actually
help or hurt.

00:09:56.313 --> 00:09:58.230 align:middle line:84%
And the strength of that
momentum, that alpha,

00:09:58.230 --> 00:09:59.910 align:middle line:90%
is a hyperparameter.

00:09:59.910 --> 00:10:06.670 align:middle line:84%
A popular example of this type
of momentum in gradient descent

00:10:06.670 --> 00:10:10.110 align:middle line:84%
is Adam, if any of you have
heard of the Adam optimizer.

00:10:10.110 --> 00:10:12.910 align:middle line:84%
So this is something that
is pretty commonly used.

00:10:12.910 --> 00:10:15.470 align:middle line:84%
But how this might
actually help or hurt--

00:10:15.470 --> 00:10:18.550 align:middle line:84%
so here, this is just looking
at optimizing for something

00:10:18.550 --> 00:10:20.390 align:middle line:90%
that is quite sharp.

00:10:20.390 --> 00:10:22.870 align:middle line:84%
And you can see that
if the momentum is 0,

00:10:22.870 --> 00:10:27.150 align:middle line:84%
you'll kind of roll steadily
towards that maximum.

00:10:27.150 --> 00:10:29.710 align:middle line:84%
If you actually have
a momentum of 0.5,

00:10:29.710 --> 00:10:33.910 align:middle line:84%
you're getting to that
minimum much quicker,

00:10:33.910 --> 00:10:37.390 align:middle line:84%
maybe about half the time
the number of training steps.

00:10:37.390 --> 00:10:41.450 align:middle line:84%
But if your momentum is too
high, you start bouncing past,

00:10:41.450 --> 00:10:43.950 align:middle line:84%
and then you have this almost
oscillation and wiggling

00:10:43.950 --> 00:10:46.370 align:middle line:84%
that happens as you're
trying to find that minimum.

00:10:46.370 --> 00:10:50.110 align:middle line:84%
And it ends up taking longer
to optimize your function.

00:10:50.110 --> 00:10:54.070 align:middle line:84%
And there's a really nice
blog post on momentum

00:10:54.070 --> 00:10:58.087 align:middle line:84%
that has some really fun
interactive capacity.

00:10:58.087 --> 00:11:00.170 align:middle line:84%
If you guys are interested,
the link is down here.

00:11:00.170 --> 00:11:01.610 align:middle line:90%
It's from back in 2017.

00:11:01.610 --> 00:11:05.430 align:middle line:84%
But I found this really
helpful for building intuition.

00:11:05.430 --> 00:11:06.150 align:middle line:90%
Cool.

00:11:06.150 --> 00:11:13.390 align:middle line:84%
So talking about loss
landscapes, which of these

00:11:13.390 --> 00:11:16.830 align:middle line:90%
are differentiable?

00:11:16.830 --> 00:11:18.350 align:middle line:90%
So raise of hands.

00:11:18.350 --> 00:11:21.390 align:middle line:84%
Give me one that's
differentiable.

00:11:21.390 --> 00:11:23.290 align:middle line:90%
Yes?

00:11:23.290 --> 00:11:24.550 align:middle line:90%
AUDIENCE: The first one.

00:11:24.550 --> 00:11:25.633 align:middle line:90%
SARA BEERY: The first one.

00:11:25.633 --> 00:11:27.005 align:middle line:90%
Yeah, the top--

00:11:27.005 --> 00:11:27.630 align:middle line:90%
AUDIENCE: Left.

00:11:27.630 --> 00:11:28.670 align:middle line:90%
SARA BEERY: --left, yes.

00:11:28.670 --> 00:11:29.330 align:middle line:90%
Thank you.

00:11:29.330 --> 00:11:32.110 align:middle line:84%
Right and left--
still hard for me.

00:11:32.110 --> 00:11:33.050 align:middle line:90%
Yeah.

00:11:33.050 --> 00:11:34.630 align:middle line:90%
AUDIENCE: The third one?

00:11:34.630 --> 00:11:35.670 align:middle line:90%
SARA BEERY: Yeah.

00:11:35.670 --> 00:11:37.330 align:middle line:84%
The third one,
exactly it's flat,

00:11:37.330 --> 00:11:40.480 align:middle line:90%
but it does have a derivative.

00:11:40.480 --> 00:11:41.360 align:middle line:90%
So, yeah, exactly.

00:11:41.360 --> 00:11:44.140 align:middle line:90%
These two are differentiable.

00:11:44.140 --> 00:11:49.832 align:middle line:84%
So now, which of these have
defined gradients in PyTorch.

00:11:49.832 --> 00:11:51.273 align:middle line:90%
AUDIENCE: All of them.

00:11:51.273 --> 00:11:52.440 align:middle line:90%
SARA BEERY: Raise your hand.

00:11:52.440 --> 00:11:55.340 align:middle line:90%


00:11:55.340 --> 00:11:56.020 align:middle line:90%
Yes?

00:11:56.020 --> 00:11:57.180 align:middle line:84%
AUDIENCE: Probably
all of them, right?

00:11:57.180 --> 00:11:57.847 align:middle line:90%
SARA BEERY: Yes.

00:11:57.847 --> 00:12:00.540 align:middle line:84%
All of them have defined
gradients in PyTorch.

00:12:00.540 --> 00:12:02.620 align:middle line:84%
That's because
PyTorch has autograd.

00:12:02.620 --> 00:12:05.660 align:middle line:84%
It has this nice mechanism
that helps us calculate

00:12:05.660 --> 00:12:07.700 align:middle line:90%
gradients for any function.

00:12:07.700 --> 00:12:10.520 align:middle line:84%
Which of them would be
hard to optimize over?

00:12:10.520 --> 00:12:16.700 align:middle line:90%


00:12:16.700 --> 00:12:17.220 align:middle line:90%
Yeah?

00:12:17.220 --> 00:12:18.580 align:middle line:90%
AUDIENCE: Second one.

00:12:18.580 --> 00:12:19.660 align:middle line:90%
SARA BEERY: Why?

00:12:19.660 --> 00:12:21.655 align:middle line:84%
AUDIENCE: It's got
multiple local minima.

00:12:21.655 --> 00:12:23.280 align:middle line:84%
SARA BEERY: Yep,
multiple local minima.

00:12:23.280 --> 00:12:26.140 align:middle line:84%
So you might end up,
depending on where you start,

00:12:26.140 --> 00:12:27.640 align:middle line:90%
not in the global minimum.

00:12:27.640 --> 00:12:30.900 align:middle line:90%


00:12:30.900 --> 00:12:31.620 align:middle line:90%
Yeah?

00:12:31.620 --> 00:12:33.620 align:middle line:84%
AUDIENCE: The third and the
fourth, since they're both flat.

00:12:33.620 --> 00:12:36.203 align:middle line:84%
SARA BEERY: The third and the
fourth, since they're both flat.

00:12:36.203 --> 00:12:39.200 align:middle line:84%
Yeah, so the third one
is just flat everywhere.

00:12:39.200 --> 00:12:41.800 align:middle line:84%
Essentially, there's no signal
in terms of the gradient.

00:12:41.800 --> 00:12:43.500 align:middle line:84%
The fourth one,
it's flat everywhere

00:12:43.500 --> 00:12:46.960 align:middle line:84%
except for discontinuity, but
anywhere on the actual graph.

00:12:46.960 --> 00:12:49.260 align:middle line:90%
Again, you have no signal.

00:12:49.260 --> 00:12:50.500 align:middle line:90%
Yeah?

00:12:50.500 --> 00:12:54.660 align:middle line:84%
AUDIENCE: The fifth one as well
because the learning rate should

00:12:54.660 --> 00:12:57.540 align:middle line:84%
be very, very small, otherwise
we're just going to diverge.

00:12:57.540 --> 00:13:00.480 align:middle line:84%
SARA BEERY: So this one in
the center of the bottom,

00:13:00.480 --> 00:13:04.160 align:middle line:84%
this is also difficult because
the gradient is very steep.

00:13:04.160 --> 00:13:06.780 align:middle line:84%
It's very, very high as you
get close to the minimum.

00:13:06.780 --> 00:13:09.420 align:middle line:84%
And so the challenge
there is that as you

00:13:09.420 --> 00:13:12.000 align:middle line:84%
get close to that minimum,
the gradient is really large,

00:13:12.000 --> 00:13:14.580 align:middle line:84%
so then it's going to bounce
you somewhere really far away.

00:13:14.580 --> 00:13:15.940 align:middle line:84%
So essentially, it's
really difficult

00:13:15.940 --> 00:13:17.607 align:middle line:84%
to actually ever hit
the minimum because

00:13:17.607 --> 00:13:22.220 align:middle line:84%
of what we call exploding
gradients, yeah, so definitely.

00:13:22.220 --> 00:13:27.960 align:middle line:84%
And one nice intuition around
easy versus hard to optimize,

00:13:27.960 --> 00:13:30.800 align:middle line:84%
you can think about
flowing water.

00:13:30.800 --> 00:13:34.300 align:middle line:84%
So in this example over
here, you would actually,

00:13:34.300 --> 00:13:36.290 align:middle line:84%
no matter where you
started, you would actually

00:13:36.290 --> 00:13:39.650 align:middle line:84%
end up getting to that minimum
because, essentially, you

00:13:39.650 --> 00:13:41.250 align:middle line:90%
can think of water flowing.

00:13:41.250 --> 00:13:44.170 align:middle line:84%
No matter where it came from, it
would still get to that middle.

00:13:44.170 --> 00:13:47.610 align:middle line:84%
Of course, that really only
helps you thinking about,

00:13:47.610 --> 00:13:49.015 align:middle line:90%
is it minimizing?

00:13:49.015 --> 00:13:50.890 align:middle line:84%
So that really is only
filtering out this one

00:13:50.890 --> 00:13:53.890 align:middle line:90%
that has multiple minima.

00:13:53.890 --> 00:14:01.570 align:middle line:84%
So one thing that's true is that
most existing popular neural

00:14:01.570 --> 00:14:06.910 align:middle line:84%
network components will
exhibit certain properties.

00:14:06.910 --> 00:14:09.950 align:middle line:84%
So like we said, these
would be hard to optimize.

00:14:09.950 --> 00:14:12.610 align:middle line:84%
But now I want to go through
what these different landscapes

00:14:12.610 --> 00:14:14.888 align:middle line:84%
would look like, the
pros and cons of them.

00:14:14.888 --> 00:14:16.930 align:middle line:84%
And then we'll get into
some high-level intuition

00:14:16.930 --> 00:14:18.430 align:middle line:84%
and talk about where
the field seems

00:14:18.430 --> 00:14:21.130 align:middle line:84%
to be going in terms
of the types of loss

00:14:21.130 --> 00:14:22.810 align:middle line:90%
that we're calculating.

00:14:22.810 --> 00:14:26.097 align:middle line:84%
So something like this, this is
like maybe the simplest case.

00:14:26.097 --> 00:14:27.930 align:middle line:84%
This is a case that if
you're a theoretician

00:14:27.930 --> 00:14:31.250 align:middle line:84%
is really great because
you can assume convexity.

00:14:31.250 --> 00:14:32.682 align:middle line:90%
You have a single minimum.

00:14:32.682 --> 00:14:34.390 align:middle line:84%
All the gradients
point to it everywhere.

00:14:34.390 --> 00:14:36.330 align:middle line:84%
And the gradient will
gracefully go to 0

00:14:36.330 --> 00:14:38.970 align:middle line:90%
as the minimum is approached.

00:14:38.970 --> 00:14:41.850 align:middle line:84%
Here, we have
discontinuity, but it's well

00:14:41.850 --> 00:14:44.130 align:middle line:84%
defined in terms
of the derivatives

00:14:44.130 --> 00:14:45.910 align:middle line:84%
on both sides of
that discontinuity.

00:14:45.910 --> 00:14:49.330 align:middle line:84%
And this is not a problem
at all for PyTorch.

00:14:49.330 --> 00:14:52.190 align:middle line:84%
Here, even if this is
not completely flat,

00:14:52.190 --> 00:14:54.890 align:middle line:84%
which you can see
here it isn't, it's

00:14:54.890 --> 00:14:56.870 align:middle line:84%
what we call a
vanishing gradient.

00:14:56.870 --> 00:15:01.470 align:middle line:84%
The progress would be really
slow, and noise in your batches,

00:15:01.470 --> 00:15:04.370 align:middle line:84%
for example, might end up
dominating over the signal.

00:15:04.370 --> 00:15:07.250 align:middle line:84%
So this would be very
difficult to optimize.

00:15:07.250 --> 00:15:10.530 align:middle line:84%
Here, this one, just zero
gradient everywhere, basically,

00:15:10.530 --> 00:15:13.330 align:middle line:84%
completely uninformative
as how to make progress,

00:15:13.330 --> 00:15:16.710 align:middle line:84%
and so it's really difficult to
ever hit any low loss region.

00:15:16.710 --> 00:15:17.830 align:middle line:90%
There was a question?

00:15:17.830 --> 00:15:18.330 align:middle line:90%
Yeah?

00:15:18.330 --> 00:15:21.250 align:middle line:84%
AUDIENCE: Can you take a second
to explain the graphs that

00:15:21.250 --> 00:15:22.230 align:middle line:90%
are on the right side.

00:15:22.230 --> 00:15:23.570 align:middle line:90%
SARA BEERY: Oh, yeah, sorry.

00:15:23.570 --> 00:15:25.950 align:middle line:84%
Sometimes I assume something
is intuitive, and it's not.

00:15:25.950 --> 00:15:28.170 align:middle line:90%
So thank you so much for asking.

00:15:28.170 --> 00:15:32.960 align:middle line:84%
So this is for
different parameters.

00:15:32.960 --> 00:15:36.240 align:middle line:84%
All of these are single
versions of parameters.

00:15:36.240 --> 00:15:40.800 align:middle line:84%
So this is now
looking at that plot.

00:15:40.800 --> 00:15:45.720 align:middle line:84%
Now here, this is some optimizer
steps going from 0 to 100,

00:15:45.720 --> 00:15:47.800 align:middle line:84%
so assuming you're
starting from the top.

00:15:47.800 --> 00:15:50.360 align:middle line:84%
And it's basically
stepping through,

00:15:50.360 --> 00:15:52.000 align:middle line:84%
assuming this was
your loss landscape,

00:15:52.000 --> 00:15:55.920 align:middle line:84%
and you had some of
standard learning rate,

00:15:55.920 --> 00:15:59.160 align:middle line:84%
where the model
would actually move

00:15:59.160 --> 00:16:02.180 align:middle line:84%
throughout gradient descent,
stochastic gradient descent.

00:16:02.180 --> 00:16:03.740 align:middle line:84%
So in this case,
there's no signal.

00:16:03.740 --> 00:16:05.320 align:middle line:90%
It just doesn't go anywhere.

00:16:05.320 --> 00:16:07.600 align:middle line:84%
Whereas in the previous
slide or this one,

00:16:07.600 --> 00:16:09.800 align:middle line:84%
for example, you might
start out over here.

00:16:09.800 --> 00:16:12.100 align:middle line:84%
And it might bounce across
and then go back and forth,

00:16:12.100 --> 00:16:15.360 align:middle line:84%
but It will actually,
eventually, find that minimum.

00:16:15.360 --> 00:16:16.380 align:middle line:90%
Thanks for the question.

00:16:16.380 --> 00:16:18.640 align:middle line:90%
It was a good one.

00:16:18.640 --> 00:16:20.880 align:middle line:84%
Yeah, so zero gradient--
this is what we were talking

00:16:20.880 --> 00:16:22.920 align:middle line:90%
about in terms of instability.

00:16:22.920 --> 00:16:25.260 align:middle line:84%
This one has what we call
an exploding gradient.

00:16:25.260 --> 00:16:29.360 align:middle line:84%
The gradient goes to infinity
as the minimizer is approached,

00:16:29.360 --> 00:16:31.120 align:middle line:84%
so you have really
unstable updates,

00:16:31.120 --> 00:16:32.460 align:middle line:90%
and you can really overshoot.

00:16:32.460 --> 00:16:35.063 align:middle line:84%
It can be very difficult to
ever actually hit that minimum.

00:16:35.063 --> 00:16:36.480 align:middle line:84%
And then, here,
when we're talking

00:16:36.480 --> 00:16:38.920 align:middle line:84%
about multiple local
minima, where you initialize

00:16:38.920 --> 00:16:42.240 align:middle line:90%
matters a lot.

00:16:42.240 --> 00:16:45.080 align:middle line:84%
One way that people handle
the fact that many of our lost

00:16:45.080 --> 00:16:47.340 align:middle line:84%
landscapes are not
guaranteed to be convex.

00:16:47.340 --> 00:16:50.040 align:middle line:84%
We're not guaranteed that they
only have a single local minima.

00:16:50.040 --> 00:16:54.640 align:middle line:84%
This is why you might, for
example, try experimentally

00:16:54.640 --> 00:16:56.240 align:middle line:84%
a bunch of different
random seeds,

00:16:56.240 --> 00:16:59.480 align:middle line:84%
and try to then select a model
that ends up looking better.

00:16:59.480 --> 00:17:01.043 align:middle line:84%
And the variance in
your performance,

00:17:01.043 --> 00:17:02.960 align:middle line:84%
based on different random
seeds, will tell you

00:17:02.960 --> 00:17:08.160 align:middle line:84%
something about how unstable
your loss landscape is.

00:17:08.160 --> 00:17:11.020 align:middle line:84%
So there's a couple things
that we wanted to bring up.

00:17:11.020 --> 00:17:11.520 align:middle line:90%
Yeah?

00:17:11.520 --> 00:17:15.002 align:middle line:84%
AUDIENCE: Question about-- you
said that PyTorch [INAUDIBLE].

00:17:15.002 --> 00:17:19.480 align:middle line:90%


00:17:19.480 --> 00:17:22.859 align:middle line:84%
SARA BEERY: Well, PyTorch can
make anything differentiable.

00:17:22.859 --> 00:17:25.599 align:middle line:84%
That doesn't mean it
will be easy to optimize.

00:17:25.599 --> 00:17:27.335 align:middle line:84%
So for example, that
one that's flat,

00:17:27.335 --> 00:17:29.710 align:middle line:84%
it's differentiable in PyTorch,
but that doesn't actually

00:17:29.710 --> 00:17:32.150 align:middle line:84%
mean that it's going to find a
local minimum because it doesn't

00:17:32.150 --> 00:17:32.830 align:middle line:90%
exist.

00:17:32.830 --> 00:17:37.170 align:middle line:84%
So it's computationally possible
to calculate a gradient there,

00:17:37.170 --> 00:17:39.110 align:middle line:84%
but the gradient
will still be 0.

00:17:39.110 --> 00:17:40.762 align:middle line:84%
So essentially,
PyTorch and autograd,

00:17:40.762 --> 00:17:42.470 align:middle line:84%
they mean that even
if you take something

00:17:42.470 --> 00:17:44.690 align:middle line:84%
that would traditionally
be non-differentiable,

00:17:44.690 --> 00:17:47.190 align:middle line:84%
you can still calculate
derivatives in PyTorch.

00:17:47.190 --> 00:17:49.190 align:middle line:84%
So we assume things like
directional derivatives

00:17:49.190 --> 00:17:51.750 align:middle line:84%
around discontinuities,
but that doesn't actually

00:17:51.750 --> 00:17:54.177 align:middle line:84%
mean that every loss
landscape is equal.

00:17:54.177 --> 00:17:56.010 align:middle line:84%
That's kind of what
we're getting into here,

00:17:56.010 --> 00:17:59.550 align:middle line:84%
is these ideas around what
makes a good versus a bad loss

00:17:59.550 --> 00:18:00.290 align:middle line:90%
landscape.

00:18:00.290 --> 00:18:01.910 align:middle line:90%
Does that make sense?

00:18:01.910 --> 00:18:02.710 align:middle line:90%
Cool.

00:18:02.710 --> 00:18:03.210 align:middle line:90%
OK.

00:18:03.210 --> 00:18:04.835 align:middle line:84%
So one thing that we
wanted to bring up

00:18:04.835 --> 00:18:07.790 align:middle line:84%
was this idea of an
evolution strategy.

00:18:07.790 --> 00:18:09.990 align:middle line:90%
So these are gradient like.

00:18:09.990 --> 00:18:12.590 align:middle line:84%
They find locally
loss-minimizing directions

00:18:12.590 --> 00:18:14.350 align:middle line:90%
in parameter space.

00:18:14.350 --> 00:18:15.910 align:middle line:84%
And the way that
they do that is they

00:18:15.910 --> 00:18:19.790 align:middle line:84%
sample small perturbations
around your existing parameter

00:18:19.790 --> 00:18:22.110 align:middle line:84%
values and then move
towards the perturbations

00:18:22.110 --> 00:18:24.790 align:middle line:90%
that achieve lower loss.

00:18:24.790 --> 00:18:29.630 align:middle line:84%
So for example, you could take
some perturbation or error

00:18:29.630 --> 00:18:33.670 align:middle line:90%
that's normally distributed.

00:18:33.670 --> 00:18:37.870 align:middle line:84%
And then when you're actually
calculating your values,

00:18:37.870 --> 00:18:39.750 align:middle line:84%
you calculate them
of the loss function

00:18:39.750 --> 00:18:44.750 align:middle line:84%
plus some amount of perturbation
in that normal space.

00:18:44.750 --> 00:18:46.430 align:middle line:84%
And then when you're
doing your updates,

00:18:46.430 --> 00:18:51.710 align:middle line:84%
now you have this almost
robustness that comes in.

00:18:51.710 --> 00:18:56.350 align:middle line:84%
And it will actually select
successfully minimize, even

00:18:56.350 --> 00:18:58.430 align:middle line:90%
something that looks like this.

00:18:58.430 --> 00:19:00.830 align:middle line:84%
So here, even if
we started up here

00:19:00.830 --> 00:19:06.830 align:middle line:84%
in this non-minimal
situation, because we're

00:19:06.830 --> 00:19:08.490 align:middle line:84%
perturbing the values
of the function,

00:19:08.490 --> 00:19:11.270 align:middle line:84%
it might bounce us over
into a better value,

00:19:11.270 --> 00:19:13.330 align:middle line:84%
just by sampling
this cost function,

00:19:13.330 --> 00:19:15.350 align:middle line:90%
these different things.

00:19:15.350 --> 00:19:18.030 align:middle line:84%
So this is something,
these evolution strategies

00:19:18.030 --> 00:19:20.750 align:middle line:84%
is something that sometimes
people will take into account.

00:19:20.750 --> 00:19:22.350 align:middle line:84%
And then another
important one is

00:19:22.350 --> 00:19:24.710 align:middle line:90%
this idea of gradient clipping.

00:19:24.710 --> 00:19:26.620 align:middle line:84%
And this is a
mechanism that helps us

00:19:26.620 --> 00:19:30.520 align:middle line:84%
with, basically, peakiness or
pointiness in our loss function.

00:19:30.520 --> 00:19:32.820 align:middle line:84%
So we'll successfully
minimize this function.

00:19:32.820 --> 00:19:34.680 align:middle line:90%
And so, here, it's quite simple.

00:19:34.680 --> 00:19:38.140 align:middle line:84%
If the gradient gets too
high above some large value,

00:19:38.140 --> 00:19:40.080 align:middle line:84%
you just scale them
back to that value.

00:19:40.080 --> 00:19:42.660 align:middle line:84%
So you essentially just
put a clipping filter

00:19:42.660 --> 00:19:44.380 align:middle line:90%
on top of your gradient.

00:19:44.380 --> 00:19:48.085 align:middle line:84%
And this is a useful
and commonly used hack.

00:19:48.085 --> 00:19:49.460 align:middle line:84%
So here, essentially,
what you do

00:19:49.460 --> 00:19:53.120 align:middle line:84%
is you just calculate the
change in your parameters.

00:19:53.120 --> 00:19:57.100 align:middle line:84%
And then you move in the
direction of that maximal change

00:19:57.100 --> 00:20:00.120 align:middle line:84%
but clipped to
some maximum value.

00:20:00.120 --> 00:20:03.700 align:middle line:84%
So you won't let your model
move too far in any direction.

00:20:03.700 --> 00:20:07.100 align:middle line:84%
And this, then, helps us with
not oscillating back and forth

00:20:07.100 --> 00:20:09.100 align:middle line:90%
across something.

00:20:09.100 --> 00:20:13.220 align:middle line:84%
So what then is important
in a loss function?

00:20:13.220 --> 00:20:16.700 align:middle line:84%
So a few things that are
consistently showing up--

00:20:16.700 --> 00:20:18.380 align:middle line:84%
one-- though this
is, of course, not

00:20:18.380 --> 00:20:21.740 align:middle line:84%
required-- everywhere
continuous, another one,

00:20:21.740 --> 00:20:23.240 align:middle line:90%
everywhere differentiable.

00:20:23.240 --> 00:20:26.260 align:middle line:84%
And the third one would
be everywhere smooth.

00:20:26.260 --> 00:20:32.260 align:middle line:84%
So one of the most commonly
used non-linearities is ReLU.

00:20:32.260 --> 00:20:37.060 align:middle line:84%
And this is something that
is everywhere continuous,

00:20:37.060 --> 00:20:39.120 align:middle line:84%
almost everywhere
differentiable,

00:20:39.120 --> 00:20:41.140 align:middle line:90%
but it's not everywhere smooth.

00:20:41.140 --> 00:20:47.460 align:middle line:90%
We have this kink in it.

00:20:47.460 --> 00:20:53.220 align:middle line:84%
Now, recently, there's
been a lot of work

00:20:53.220 --> 00:20:56.540 align:middle line:84%
that will instead use
something like a ReLU, which is

00:20:56.540 --> 00:21:00.820 align:middle line:90%
a Gaussian error linear unit.

00:21:00.820 --> 00:21:04.540 align:middle line:84%
And this one, essentially,
just takes the parameters times

00:21:04.540 --> 00:21:06.207 align:middle line:84%
phi of the
parameters, where this

00:21:06.207 --> 00:21:08.540 align:middle line:84%
is the cumulative distribution
function for the Gaussian

00:21:08.540 --> 00:21:09.520 align:middle line:90%
distribution.

00:21:09.520 --> 00:21:12.540 align:middle line:84%
It looks something like
this in two dimensions.

00:21:12.540 --> 00:21:14.440 align:middle line:84%
And this one is
everywhere continuous,

00:21:14.440 --> 00:21:17.347 align:middle line:84%
everywhere differentiable,
and everywhere smooth.

00:21:17.347 --> 00:21:19.180 align:middle line:84%
The important thing I
want to point out here

00:21:19.180 --> 00:21:22.930 align:middle line:84%
is that we can agree that
these three things might

00:21:22.930 --> 00:21:24.650 align:middle line:90%
make optimization easier.

00:21:24.650 --> 00:21:28.170 align:middle line:84%
And even if we don't
precisely know experimentally

00:21:28.170 --> 00:21:29.770 align:middle line:84%
or theoretically,
which properties

00:21:29.770 --> 00:21:33.190 align:middle line:84%
are needed for training
neural networks,

00:21:33.190 --> 00:21:35.370 align:middle line:84%
trends do seem to be
moving towards functions

00:21:35.370 --> 00:21:39.330 align:middle line:84%
which satisfy all three,
so things like this GeLU.

00:21:39.330 --> 00:21:42.010 align:middle line:84%
Any questions about
optimization and SGD?

00:21:42.010 --> 00:21:42.810 align:middle line:90%
Yeah?

00:21:42.810 --> 00:21:46.970 align:middle line:84%
AUDIENCE: What was the
[INAUDIBLE], use that for again?

00:21:46.970 --> 00:21:49.685 align:middle line:84%
SARA BEERY: So
evolution strategies?

00:21:49.685 --> 00:21:50.310 align:middle line:90%
AUDIENCE: Yeah.

00:21:50.310 --> 00:21:53.330 align:middle line:84%
SARA BEERY: Essentially,
instead of always moving

00:21:53.330 --> 00:21:54.930 align:middle line:84%
in the direction
of the gradient,

00:21:54.930 --> 00:22:00.150 align:middle line:84%
you just randomly sample around
your current parameter values,

00:22:00.150 --> 00:22:03.930 align:middle line:84%
and then you move in the
random direction that maximizes

00:22:03.930 --> 00:22:06.770 align:middle line:90%
or that minimizes your cost.

00:22:06.770 --> 00:22:09.770 align:middle line:90%
AUDIENCE: [INAUDIBLE]

00:22:09.770 --> 00:22:12.210 align:middle line:84%
SARA BEERY: Essentially,
it's helping you maybe handle

00:22:12.210 --> 00:22:15.250 align:middle line:84%
things like vanishing
gradients or some

00:22:15.250 --> 00:22:18.570 align:middle line:84%
of these parts of
your loss landscape

00:22:18.570 --> 00:22:21.250 align:middle line:84%
that might be not
actually functional

00:22:21.250 --> 00:22:23.410 align:middle line:90%
to optimize super efficiently.

00:22:23.410 --> 00:22:26.210 align:middle line:84%
And I think you can
almost think of it--

00:22:26.210 --> 00:22:31.137 align:middle line:84%
one intuition, It's almost like
regularization in your loss,

00:22:31.137 --> 00:22:31.970 align:middle line:90%
if that makes sense.

00:22:31.970 --> 00:22:33.690 align:middle line:84%
AUDIENCE: Also,
follow up, how common

00:22:33.690 --> 00:22:37.850 align:middle line:84%
is it that people use
everything together?

00:22:37.850 --> 00:22:43.450 align:middle line:84%
SARA BEERY: Yeah,
I mean, I think

00:22:43.450 --> 00:22:48.050 align:middle line:84%
everything's a bit of a
big, soup pot these days.

00:22:48.050 --> 00:22:50.890 align:middle line:84%
There's a lot of these
different experimental tricks.

00:22:50.890 --> 00:22:53.810 align:middle line:84%
Some of them, even if
they seem quite different,

00:22:53.810 --> 00:22:56.010 align:middle line:84%
almost seem like they
achieve the same ends, when

00:22:56.010 --> 00:22:59.730 align:middle line:84%
you talk about the
experimental performance.

00:22:59.730 --> 00:23:03.630 align:middle line:84%
I would say that a lot of
the really top-performing,

00:23:03.630 --> 00:23:07.450 align:middle line:84%
state-of-the-art models on
different types of benchmarks

00:23:07.450 --> 00:23:10.210 align:middle line:84%
are often ones that incorporate
a lot of these tricks,

00:23:10.210 --> 00:23:16.570 align:middle line:84%
different hacks and mechanisms
to try to get a little better

00:23:16.570 --> 00:23:19.680 align:middle line:84%
at finding some maybe
slightly better optimum,

00:23:19.680 --> 00:23:22.680 align:middle line:84%
but I don't think we actually
have a really great recipe

00:23:22.680 --> 00:23:26.400 align:middle line:84%
for given a certain data set,
given a certain architecture,

00:23:26.400 --> 00:23:30.520 align:middle line:84%
what is the right way to
optimize that architecture?

00:23:30.520 --> 00:23:31.340 align:middle line:90%
Yeah?

00:23:31.340 --> 00:23:32.840 align:middle line:84%
AUDIENCE: Can you
please explain how

00:23:32.840 --> 00:23:38.520 align:middle line:84%
ReLU and GeLU relates to
this question of [INAUDIBLE]?

00:23:38.520 --> 00:23:40.760 align:middle line:84%
SARA BEERY: Yeah,
so, again, I'm not

00:23:40.760 --> 00:23:44.240 align:middle line:84%
I'm not saying here that
GeLU is much better than ReLU

00:23:44.240 --> 00:23:47.200 align:middle line:84%
all the time because that's
not necessarily true.

00:23:47.200 --> 00:23:51.520 align:middle line:84%
It's just pointing out that the
trend in these different loss

00:23:51.520 --> 00:23:53.440 align:middle line:84%
terms for neural
network training

00:23:53.440 --> 00:23:57.920 align:middle line:84%
does at least seem to be moving
away from functions that don't

00:23:57.920 --> 00:23:59.680 align:middle line:90%
satisfy all three of these.

00:23:59.680 --> 00:24:02.000 align:middle line:84%
But again, we don't
necessarily know,

00:24:02.000 --> 00:24:03.960 align:middle line:84%
or we haven't decided
as a community

00:24:03.960 --> 00:24:09.000 align:middle line:84%
or even gotten close to
some theoretical guarantee

00:24:09.000 --> 00:24:14.280 align:middle line:84%
as a community as to what
is the optimal mechanism.

00:24:14.280 --> 00:24:15.600 align:middle line:90%
Yeah?

00:24:15.600 --> 00:24:17.960 align:middle line:84%
AUDIENCE: Are these
loss functions?

00:24:17.960 --> 00:24:19.468 align:middle line:90%
Are these activation functions?

00:24:19.468 --> 00:24:21.260 align:middle line:84%
SARA BEERY: These are
activation functions.

00:24:21.260 --> 00:24:21.860 align:middle line:90%
Yeah, yeah.

00:24:21.860 --> 00:24:23.443 align:middle line:84%
AUDIENCE: Always
positive [INAUDIBLE]?

00:24:23.443 --> 00:24:25.020 align:middle line:90%
SARA BEERY: Hmm?

00:24:25.020 --> 00:24:25.520 align:middle line:90%
Sorry?

00:24:25.520 --> 00:24:29.200 align:middle line:84%
AUDIENCE: [INAUDIBLE] loss
functions also not be negative?

00:24:29.200 --> 00:24:35.720 align:middle line:84%
SARA BEERY: Oh, well,
it depends on the way

00:24:35.720 --> 00:24:36.780 align:middle line:90%
you've set up your loss.

00:24:36.780 --> 00:24:39.880 align:middle line:84%
You definitely can
have loss functions

00:24:39.880 --> 00:24:42.733 align:middle line:90%
that will have negative values.

00:24:42.733 --> 00:24:43.900 align:middle line:90%
I mean, that depends, right?

00:24:43.900 --> 00:24:45.640 align:middle line:84%
And I guess what
you're saying is,

00:24:45.640 --> 00:24:47.500 align:middle line:84%
if you're trying
to minimize a cost,

00:24:47.500 --> 00:24:49.845 align:middle line:84%
you want the gradients to
move in the direction of that,

00:24:49.845 --> 00:24:52.220 align:middle line:84%
but that's not guaranteed if
you can't guarantee the loss

00:24:52.220 --> 00:24:53.680 align:middle line:90%
landscape is convex everywhere.

00:24:53.680 --> 00:24:57.400 align:middle line:84%
There will be parts where
it goes up and down.

00:24:57.400 --> 00:24:59.173 align:middle line:90%
Yeah?

00:24:59.173 --> 00:25:00.840 align:middle line:84%
AUDIENCE: Is there
any issue with having

00:25:00.840 --> 00:25:03.120 align:middle line:90%
a non-monotonic function?

00:25:03.120 --> 00:25:07.120 align:middle line:84%
SARA BEERY: A non-monotonic
activation function or loss

00:25:07.120 --> 00:25:07.740 align:middle line:90%
function?

00:25:07.740 --> 00:25:11.160 align:middle line:90%
AUDIENCE: Activation function.

00:25:11.160 --> 00:25:12.680 align:middle line:84%
SARA BEERY: I mean,
there definitely

00:25:12.680 --> 00:25:15.350 align:middle line:84%
exist, in theory,
activation functions that

00:25:15.350 --> 00:25:18.510 align:middle line:84%
are non-monotonic,
but I don't think

00:25:18.510 --> 00:25:20.010 align:middle line:90%
that they're often widely used.

00:25:20.010 --> 00:25:21.492 align:middle line:90%
I mean, it seems to make sense.

00:25:21.492 --> 00:25:22.950 align:middle line:84%
AUDIENCE: [INAUDIBLE]
seems common.

00:25:22.950 --> 00:25:24.117 align:middle line:90%
SARA BEERY: It's not, right?

00:25:24.117 --> 00:25:27.452 align:middle line:84%
Yeah, it has this part where
it dips down right here.

00:25:27.452 --> 00:25:28.910 align:middle line:84%
AUDIENCE: Yeah,
that's what I mean.

00:25:28.910 --> 00:25:30.230 align:middle line:90%
SARA BEERY: Yeah.

00:25:30.230 --> 00:25:32.310 align:middle line:90%
But then this is the question.

00:25:32.310 --> 00:25:35.150 align:middle line:84%
Is it better to be smooth
or to be monotonic?

00:25:35.150 --> 00:25:38.830 align:middle line:84%
And that's where I think the
jury's a little bit still out.

00:25:38.830 --> 00:25:40.890 align:middle line:84%
But, yeah, I guess we
could put monotonic here,

00:25:40.890 --> 00:25:43.350 align:middle line:84%
but I feel like these three
are the ones, at least

00:25:43.350 --> 00:25:47.270 align:middle line:84%
from my perspective, that seem
to be captured most consistently

00:25:47.270 --> 00:25:51.470 align:middle line:84%
across the activation functions
that people tend to use

00:25:51.470 --> 00:25:53.030 align:middle line:90%
or are moving towards.

00:25:53.030 --> 00:25:53.630 align:middle line:90%
Yeah?

00:25:53.630 --> 00:25:56.650 align:middle line:84%
AUDIENCE: [INAUDIBLE]
for the perturbation,

00:25:56.650 --> 00:25:59.650 align:middle line:84%
can you talk about it from a
numerical mathematical equation?

00:25:59.650 --> 00:26:01.190 align:middle line:84%
What are the strength
of levers that

00:26:01.190 --> 00:26:05.090 align:middle line:84%
affects that, making the
magnitude bigger or smaller,

00:26:05.090 --> 00:26:06.690 align:middle line:84%
and also from the
application sense?

00:26:06.690 --> 00:26:08.887 align:middle line:84%
If you're looking at
a data set, what--

00:26:08.887 --> 00:26:09.970 align:middle line:90%
SARA BEERY: You mean this?

00:26:09.970 --> 00:26:10.790 align:middle line:90%
AUDIENCE: Yeah.

00:26:10.790 --> 00:26:15.510 align:middle line:84%
SARA BEERY: I mean, so
definitely those parameters

00:26:15.510 --> 00:26:18.790 align:middle line:84%
are something that you would
need to experimentally find

00:26:18.790 --> 00:26:21.930 align:middle line:84%
optimals for, for any
given data set or problem.

00:26:21.930 --> 00:26:25.692 align:middle line:84%
It's not like there's
some universal ideal here.

00:26:25.692 --> 00:26:27.150 align:middle line:84%
And I think it will
probably depend

00:26:27.150 --> 00:26:31.270 align:middle line:84%
a lot on which data set, which
architecture, for example.

00:26:31.270 --> 00:26:35.190 align:middle line:84%
So I don't have any big
rules of thumb for you there.

00:26:35.190 --> 00:26:40.790 align:middle line:84%
Though, I would say one
generally good strategy

00:26:40.790 --> 00:26:42.270 align:middle line:84%
is you take a
paper that used it,

00:26:42.270 --> 00:26:43.890 align:middle line:84%
and you look at the
values they used,

00:26:43.890 --> 00:26:46.430 align:middle line:84%
and you start
there, just usually

00:26:46.430 --> 00:26:49.070 align:middle line:84%
because that they've
already done it, right?

00:26:49.070 --> 00:26:50.590 align:middle line:84%
They've already
done it-- defined

00:26:50.590 --> 00:26:52.450 align:middle line:90%
their optimal parameters.

00:26:52.450 --> 00:26:54.990 align:middle line:84%
And sometimes those can be
really far off for a new data

00:26:54.990 --> 00:26:57.190 align:middle line:84%
set, but it's often
not a bad strategy

00:26:57.190 --> 00:26:59.750 align:middle line:84%
to at least start with
something that has trained

00:26:59.750 --> 00:27:03.150 align:middle line:90%
for someone in the past.

00:27:03.150 --> 00:27:05.630 align:middle line:90%
OK, I'm going to move on.

00:27:05.630 --> 00:27:07.670 align:middle line:90%
Great questions though, guys.

00:27:07.670 --> 00:27:09.630 align:middle line:84%
It's much more fun
for a professor

00:27:09.630 --> 00:27:13.000 align:middle line:84%
when people are engaged, so
it's great to hear from you.

00:27:13.000 --> 00:27:14.886 align:middle line:90%
[LAUGHS]

00:27:14.886 --> 00:27:16.740 align:middle line:90%


00:27:16.740 --> 00:27:17.260 align:middle line:90%
Cool.

00:27:17.260 --> 00:27:18.677 align:middle line:84%
So I'm going to
move on to talking

00:27:18.677 --> 00:27:20.540 align:middle line:90%
about computation graphs.

00:27:20.540 --> 00:27:23.660 align:middle line:84%
So we're going to move
to thinking of this

00:27:23.660 --> 00:27:26.600 align:middle line:84%
all from this perspective
of differential programming.

00:27:26.600 --> 00:27:31.660 align:middle line:84%
Essentially, here, we're
talking about computation graph.

00:27:31.660 --> 00:27:34.500 align:middle line:84%
Here, this is something like a
graph of different functional

00:27:34.500 --> 00:27:36.115 align:middle line:90%
transformations.

00:27:36.115 --> 00:27:37.740 align:middle line:84%
Each of those functional
transformation

00:27:37.740 --> 00:27:39.420 align:middle line:90%
is a node in the graph.

00:27:39.420 --> 00:27:41.580 align:middle line:84%
And then when you string
all of those together,

00:27:41.580 --> 00:27:45.020 align:middle line:84%
you can perform some
useful computation.

00:27:45.020 --> 00:27:49.600 align:middle line:84%
So intuitively, maybe this
is some sort of decision tree

00:27:49.600 --> 00:27:52.900 align:middle line:84%
that's a very simple
version of a graph

00:27:52.900 --> 00:27:54.660 align:middle line:90%
of functional transformations.

00:27:54.660 --> 00:27:56.940 align:middle line:84%
So maybe I'm saying,
all right, I'm

00:27:56.940 --> 00:27:59.860 align:middle line:84%
going to try to make
it to the classroom.

00:27:59.860 --> 00:28:04.740 align:middle line:84%
So I'm going to go
into a building.

00:28:04.740 --> 00:28:07.620 align:middle line:84%
And then maybe on
the right, you've

00:28:07.620 --> 00:28:09.643 align:middle line:84%
gone into building
45, which is correct.

00:28:09.643 --> 00:28:11.560 align:middle line:84%
And on the left, you've
gone into building 32,

00:28:11.560 --> 00:28:13.060 align:middle line:90%
which is incorrect.

00:28:13.060 --> 00:28:15.620 align:middle line:84%
And then there's these
different decisions

00:28:15.620 --> 00:28:18.860 align:middle line:84%
you might make based on
the values of your input

00:28:18.860 --> 00:28:19.940 align:middle line:90%
at any given time.

00:28:19.940 --> 00:28:21.980 align:middle line:84%
So if you did go
into building 45,

00:28:21.980 --> 00:28:26.580 align:middle line:84%
then you could maybe take the
elevator or take the stairs.

00:28:26.580 --> 00:28:29.220 align:middle line:84%
And that decision over
what action you'd take

00:28:29.220 --> 00:28:31.483 align:middle line:84%
could be a really, really
simple heuristic version

00:28:31.483 --> 00:28:33.900 align:middle line:84%
of a functional transformation
because you're transforming

00:28:33.900 --> 00:28:36.740 align:middle line:90%
your inputs to some output.

00:28:36.740 --> 00:28:42.580 align:middle line:84%
Of course, what we
actually see here

00:28:42.580 --> 00:28:45.420 align:middle line:84%
when we think about machine
learning systems is each

00:28:45.420 --> 00:28:47.940 align:middle line:84%
of these maybe
blocks are something

00:28:47.940 --> 00:28:50.860 align:middle line:84%
like a layer in a
neural network or even

00:28:50.860 --> 00:28:53.340 align:middle line:90%
an entire neural network.

00:28:53.340 --> 00:28:57.820 align:middle line:84%
And there's nothing that
says what the size of these

00:28:57.820 --> 00:29:01.420 align:middle line:84%
are, any of restrictions when
you're thinking about it really

00:29:01.420 --> 00:29:05.460 align:middle line:84%
from this high-level perspective
of a computational graph.

00:29:05.460 --> 00:29:08.910 align:middle line:84%
So deep learning
primarily deals with DAGs.

00:29:08.910 --> 00:29:10.950 align:middle line:84%
These are directed
acyclic graph.

00:29:10.950 --> 00:29:12.690 align:middle line:84%
So directed means
the information only

00:29:12.690 --> 00:29:15.290 align:middle line:84%
goes one direction
on any given edge.

00:29:15.290 --> 00:29:19.450 align:middle line:84%
Acyclic means you don't
have loops or cycles.

00:29:19.450 --> 00:29:22.330 align:middle line:84%
And we also assume that every
single one of these nodes

00:29:22.330 --> 00:29:23.850 align:middle line:90%
is differentiable.

00:29:23.850 --> 00:29:25.650 align:middle line:84%
Though we did just say
that using PyTorch,

00:29:25.650 --> 00:29:28.230 align:middle line:84%
you can calculate a derivative
for almost anything,

00:29:28.230 --> 00:29:31.850 align:middle line:84%
so that becomes less
of a hard constraint.

00:29:31.850 --> 00:29:36.970 align:middle line:84%
So if we now think about
this from the perspective

00:29:36.970 --> 00:29:40.090 align:middle line:84%
of, for example, a really
simple multilayer perceptron,

00:29:40.090 --> 00:29:43.170 align:middle line:84%
like we talked
about on Thursday.

00:29:43.170 --> 00:29:45.290 align:middle line:84%
Here, you can think
of that now, again,

00:29:45.290 --> 00:29:47.230 align:middle line:84%
as one of these
directed acyclic graph.

00:29:47.230 --> 00:29:48.610 align:middle line:90%
You have your input x.

00:29:48.610 --> 00:29:50.950 align:middle line:84%
You send that through
this linear layer,

00:29:50.950 --> 00:29:53.250 align:middle line:90%
which is one node in the graph.

00:29:53.250 --> 00:29:55.370 align:middle line:84%
Then that gives you
some hidden unit z.

00:29:55.370 --> 00:29:57.750 align:middle line:84%
You send that through an
activation function or ReLU

00:29:57.750 --> 00:29:59.890 align:middle line:84%
that's, again, some
computation that takes an input

00:29:59.890 --> 00:30:02.970 align:middle line:84%
and changes it to
create an output.

00:30:02.970 --> 00:30:05.845 align:middle line:84%
Now, another hidden
layer h, hidden state h.

00:30:05.845 --> 00:30:08.470 align:middle line:84%
And then you take that and run
it through another linear layer.

00:30:08.470 --> 00:30:10.730 align:middle line:84%
And that gives
you your output y.

00:30:10.730 --> 00:30:12.730 align:middle line:90%
So even an MLP--

00:30:12.730 --> 00:30:16.210 align:middle line:84%
easy to represent as
a computation graph.

00:30:16.210 --> 00:30:20.450 align:middle line:84%
So now let's talk about what a
forward pass actually looks like

00:30:20.450 --> 00:30:24.810 align:middle line:84%
and what's required to
calculate a forward pass.

00:30:24.810 --> 00:30:28.230 align:middle line:84%
So in the simplest
sense, for any model,

00:30:28.230 --> 00:30:32.030 align:middle line:84%
for any set of parameters, for
anything within that function,

00:30:32.030 --> 00:30:34.050 align:middle line:90%
you take in some input value.

00:30:34.050 --> 00:30:35.610 align:middle line:84%
You take in some
theta, which are

00:30:35.610 --> 00:30:39.930 align:middle line:84%
the parameters of that
component or of function.

00:30:39.930 --> 00:30:44.410 align:middle line:84%
And then you map through that
functional transformation,

00:30:44.410 --> 00:30:46.850 align:middle line:84%
the inputs, the
parameters, and you get out

00:30:46.850 --> 00:30:48.530 align:middle line:90%
the outputs of that layer.

00:30:48.530 --> 00:30:54.090 align:middle line:84%
So that's a forward pass for any
type of computational function.

00:30:54.090 --> 00:30:56.330 align:middle line:84%
Now, you can have
multiple layers, right?

00:30:56.330 --> 00:30:59.570 align:middle line:84%
This computation graph, for
example, could represent an MLP.

00:30:59.570 --> 00:31:03.280 align:middle line:84%
Now you have your inputs
and your parameters

00:31:03.280 --> 00:31:04.787 align:middle line:90%
for that first component.

00:31:04.787 --> 00:31:05.620 align:middle line:90%
Then, you get those.

00:31:05.620 --> 00:31:06.520 align:middle line:90%
That's an output.

00:31:06.520 --> 00:31:08.540 align:middle line:84%
That becomes the input
of the next component,

00:31:08.540 --> 00:31:10.240 align:middle line:90%
et cetera, et cetera.

00:31:10.240 --> 00:31:15.360 align:middle line:84%
And then, at the end, maybe you
have something like this loss

00:31:15.360 --> 00:31:17.380 align:middle line:84%
that you use to
calculate your cost.

00:31:17.380 --> 00:31:23.440 align:middle line:84%
So then when you're thinking
about learning, now here,

00:31:23.440 --> 00:31:25.580 align:middle line:84%
in order to calculate
this learning,

00:31:25.580 --> 00:31:27.280 align:middle line:84%
that's where we
think about computing

00:31:27.280 --> 00:31:30.120 align:middle line:84%
the gradients of the
cost with respect

00:31:30.120 --> 00:31:33.680 align:middle line:84%
to all of those model
parameters that happen all

00:31:33.680 --> 00:31:35.100 align:middle line:90%
the way through the network.

00:31:35.100 --> 00:31:37.960 align:middle line:84%
And so by design,
every single layer

00:31:37.960 --> 00:31:41.000 align:middle line:84%
will be differentiable
with respect to its inputs

00:31:41.000 --> 00:31:43.420 align:middle line:84%
and, therefore, this is
actually possible to compute.

00:31:43.420 --> 00:31:46.560 align:middle line:84%
It is possible to compute
that gradient with respect

00:31:46.560 --> 00:31:48.020 align:middle line:84%
to all of these
model parameters.

00:31:48.020 --> 00:31:49.680 align:middle line:84%
So you can figure
out how you need

00:31:49.680 --> 00:31:55.640 align:middle line:84%
to update those model parameters
to move to a more optimal cost.

00:31:55.640 --> 00:31:56.440 align:middle line:90%
OK.

00:31:56.440 --> 00:32:00.320 align:middle line:84%
So a bit of an aside,
just in case any of you

00:32:00.320 --> 00:32:05.360 align:middle line:84%
are not brushed up on
your matrix calculus

00:32:05.360 --> 00:32:07.495 align:middle line:90%
because this will show up a lot.

00:32:07.495 --> 00:32:09.120 align:middle line:84%
We're going to be
doing a lot of matrix

00:32:09.120 --> 00:32:10.622 align:middle line:90%
multiplications in this class.

00:32:10.622 --> 00:32:12.580 align:middle line:84%
Hopefully, this is super
review for all of you,

00:32:12.580 --> 00:32:14.360 align:middle line:90%
but just a reminder.

00:32:14.360 --> 00:32:18.240 align:middle line:84%
So we're going to assume that
our inputs are, for example,

00:32:18.240 --> 00:32:20.960 align:middle line:90%
column vectors of size n by 1.

00:32:20.960 --> 00:32:25.920 align:middle line:84%
Now if we define a function on
that vector, y equals f of x.

00:32:25.920 --> 00:32:29.600 align:middle line:84%
If y is a scalar,
then the derivative

00:32:29.600 --> 00:32:33.560 align:middle line:84%
of that output with
respect to x, the input,

00:32:33.560 --> 00:32:37.640 align:middle line:84%
would be a horizontal
vector, where

00:32:37.640 --> 00:32:39.680 align:middle line:84%
for each element of
that vector, you're

00:32:39.680 --> 00:32:42.760 align:middle line:84%
getting the partial
derivative of y

00:32:42.760 --> 00:32:45.960 align:middle line:84%
with respect to that first
element, or second element,

00:32:45.960 --> 00:32:47.200 align:middle line:90%
or third element of x.

00:32:47.200 --> 00:32:53.200 align:middle line:84%
And then that is a row
vector of size 1 by n.

00:32:53.200 --> 00:32:57.120 align:middle line:84%
If y is a vector
of n by 1, then we

00:32:57.120 --> 00:33:01.870 align:middle line:84%
assume the Jacobian formulation
of the partial derivative of y

00:33:01.870 --> 00:33:05.750 align:middle line:84%
with respect to x, where
now you, essentially, end up

00:33:05.750 --> 00:33:08.270 align:middle line:84%
with something that's a
matrix of size m by n,

00:33:08.270 --> 00:33:11.950 align:middle line:84%
so m rows, which is the
size of y, by n columns

00:33:11.950 --> 00:33:12.990 align:middle line:90%
is the size of x.

00:33:12.990 --> 00:33:15.430 align:middle line:84%
And each element
of that matrix will

00:33:15.430 --> 00:33:18.150 align:middle line:84%
be the partial derivative
of that element of y

00:33:18.150 --> 00:33:21.430 align:middle line:90%
with that element of x,

00:33:21.430 --> 00:33:25.530 align:middle line:84%
Now if y is a scalar
and x is a matrix--

00:33:25.530 --> 00:33:29.510 align:middle line:84%
so, for example, an
image of size n by m--

00:33:29.510 --> 00:33:33.470 align:middle line:84%
then the partial derivative of
y with respect to that matrix

00:33:33.470 --> 00:33:36.350 align:middle line:84%
x will be a matrix
of size m by n,

00:33:36.350 --> 00:33:39.830 align:middle line:84%
where now here, each
component corresponds

00:33:39.830 --> 00:33:43.590 align:middle line:90%
to that relevant component of x.

00:33:43.590 --> 00:33:45.710 align:middle line:84%
And note that by
taking that derivative,

00:33:45.710 --> 00:33:48.550 align:middle line:90%
we flipped the rows and columns.

00:33:48.550 --> 00:33:50.870 align:middle line:90%
So it's been transposed.

00:33:50.870 --> 00:33:54.037 align:middle line:84%
And according to
Wikipedia, the three types

00:33:54.037 --> 00:33:55.870 align:middle line:84%
of derivatives that
have not been considered

00:33:55.870 --> 00:33:59.250 align:middle line:84%
are those involving vectors by
matrices, matrices by vectors,

00:33:59.250 --> 00:34:01.350 align:middle line:90%
and matrices by matrices.

00:34:01.350 --> 00:34:03.310 align:middle line:84%
Notation is not
widely agreed upon.

00:34:03.310 --> 00:34:06.630 align:middle line:84%
We're now going to see
any of in this class.

00:34:06.630 --> 00:34:09.830 align:middle line:90%
Cool, so chain rule.

00:34:09.830 --> 00:34:13.710 align:middle line:84%
For the function h
of x is f of g of x,

00:34:13.710 --> 00:34:15.630 align:middle line:90%
how do you take the derivative?

00:34:15.630 --> 00:34:18.150 align:middle line:90%
Who remembers the chain rule?

00:34:18.150 --> 00:34:19.590 align:middle line:90%
Yeah?

00:34:19.590 --> 00:34:24.630 align:middle line:84%
AUDIENCE: x times the
derivative of g of x plus

00:34:24.630 --> 00:34:27.462 align:middle line:90%
f prime of x times g of x.

00:34:27.462 --> 00:34:28.670 align:middle line:90%
You're basically [INAUDIBLE].

00:34:28.670 --> 00:34:30.983 align:middle line:90%
SARA BEERY: OK.

00:34:30.983 --> 00:34:33.150 align:middle line:90%
AUDIENCE: [INAUDIBLE]

00:34:33.150 --> 00:34:34.810 align:middle line:84%
SARA BEERY: I don't
think that's right.

00:34:34.810 --> 00:34:38.010 align:middle line:90%
[LAUGHS] It's all right, man.

00:34:38.010 --> 00:34:38.730 align:middle line:90%
It's all right.

00:34:38.730 --> 00:34:39.230 align:middle line:90%
[LAUGHS]

00:34:39.230 --> 00:34:42.030 align:middle line:90%
AUDIENCE: [INAUDIBLE]

00:34:42.030 --> 00:34:42.710 align:middle line:90%
SARA BEERY: OK.

00:34:42.710 --> 00:34:45.150 align:middle line:90%
Go for it.

00:34:45.150 --> 00:34:49.750 align:middle line:84%
AUDIENCE: g x prime
with the f prime of gx.

00:34:49.750 --> 00:34:50.670 align:middle line:90%
SARA BEERY: Yes.

00:34:50.670 --> 00:34:53.510 align:middle line:84%
So the derivative
of x f with respect

00:34:53.510 --> 00:34:58.500 align:middle line:84%
to g of x times the derivative
of g with respect to x, yeah.

00:34:58.500 --> 00:34:59.700 align:middle line:90%
Good job, dude.

00:34:59.700 --> 00:35:00.710 align:middle line:90%
Way to go for it.

00:35:00.710 --> 00:35:02.460 align:middle line:84%
No one's going to judge
anyone for getting

00:35:02.460 --> 00:35:03.960 align:middle line:90%
something wrong in this class.

00:35:03.960 --> 00:35:07.460 align:middle line:90%
[LAUGHS, APPLAUSE]

00:35:07.460 --> 00:35:09.780 align:middle line:84%
All right, so then, actually,
if we write this out

00:35:09.780 --> 00:35:13.380 align:middle line:84%
as z is f of u and
u is g of x, then

00:35:13.380 --> 00:35:15.940 align:middle line:90%
you can write it out this way.

00:35:15.940 --> 00:35:17.580 align:middle line:84%
So essentially, you
can write this out

00:35:17.580 --> 00:35:20.740 align:middle line:84%
as the multiplication of
two partial derivatives.

00:35:20.740 --> 00:35:26.980 align:middle line:84%
So that means that p would
be the length of vector u.

00:35:26.980 --> 00:35:33.100 align:middle line:84%
And m would be-- wait,
what the fuck is p?

00:35:33.100 --> 00:35:34.200 align:middle line:90%
What am I talking about?

00:35:34.200 --> 00:35:38.740 align:middle line:90%


00:35:38.740 --> 00:35:40.860 align:middle line:84%
All right Essentially,
what we want to get at

00:35:40.860 --> 00:35:43.020 align:middle line:84%
is what size are
each of these things?

00:35:43.020 --> 00:35:46.500 align:middle line:84%
So who can tell
me what the shape

00:35:46.500 --> 00:35:50.680 align:middle line:84%
of the partial of z with respect
to x would be when x equals a?

00:35:50.680 --> 00:35:55.500 align:middle line:90%


00:35:55.500 --> 00:35:59.060 align:middle line:84%
I can give you one
freebie, if it's helpful.

00:35:59.060 --> 00:36:07.940 align:middle line:84%
So if we say the size of you
is p, the size of z is m,

00:36:07.940 --> 00:36:13.540 align:middle line:84%
and the size of x is n, then
the partial of z with respect

00:36:13.540 --> 00:36:16.460 align:middle line:90%
to x would be an M by N matrix.

00:36:16.460 --> 00:36:21.680 align:middle line:84%
So what size would the partial
of z with respect to u be?

00:36:21.680 --> 00:36:26.160 align:middle line:90%


00:36:26.160 --> 00:36:26.660 align:middle line:90%
Yeah?

00:36:26.660 --> 00:36:27.368 align:middle line:90%
AUDIENCE: m by u.

00:36:27.368 --> 00:36:29.940 align:middle line:90%


00:36:29.940 --> 00:36:31.420 align:middle line:90%
SARA BEERY: m by p.

00:36:31.420 --> 00:36:33.280 align:middle line:84%
This is really
confusingly written.

00:36:33.280 --> 00:36:34.780 align:middle line:84%
I'm going to rewrite
this next year.

00:36:34.780 --> 00:36:38.380 align:middle line:84%
p is the length of vector
u equals the size of u

00:36:38.380 --> 00:36:39.980 align:middle line:90%
is confusing.

00:36:39.980 --> 00:36:41.200 align:middle line:90%
But, yes, m by p.

00:36:41.200 --> 00:36:44.604 align:middle line:90%
And then what's the last one?

00:36:44.604 --> 00:36:45.816 align:middle line:90%
AUDIENCE: p by n.

00:36:45.816 --> 00:36:48.200 align:middle line:84%
SARA BEERY: p by
n, exactly Cool.

00:36:48.200 --> 00:36:53.090 align:middle line:84%
So therefore, if you said the
size of z is 1, the size of u

00:36:53.090 --> 00:36:57.090 align:middle line:84%
is 2, the size of x is 4, then
essentially what you're saying

00:36:57.090 --> 00:37:00.890 align:middle line:84%
is this derivative would
be, essentially, equal

00:37:00.890 --> 00:37:06.690 align:middle line:84%
to this product of two matrices,
though one is really a vector.

00:37:06.690 --> 00:37:07.930 align:middle line:90%
All right.

00:37:07.930 --> 00:37:09.530 align:middle line:84%
We're bringing
this up because we

00:37:09.530 --> 00:37:14.010 align:middle line:84%
will see a lot of matrices
times matrices in the next bit.

00:37:14.010 --> 00:37:17.970 align:middle line:84%
So that's the end of our
little bit of a vector calculus

00:37:17.970 --> 00:37:20.730 align:middle line:90%
or matrix calculus reminder.

00:37:20.730 --> 00:37:23.450 align:middle line:90%
So let's talk about backprop.

00:37:23.450 --> 00:37:25.230 align:middle line:84%
How does any of this
become possible?

00:37:25.230 --> 00:37:28.730 align:middle line:84%
How do we calculate the
update in those parameters?

00:37:28.730 --> 00:37:30.690 align:middle line:84%
Essentially, what
we need to calculate

00:37:30.690 --> 00:37:34.170 align:middle line:84%
is the partial derivative
of j with respect

00:37:34.170 --> 00:37:41.050 align:middle line:84%
to, for example, those initial
layers in the model, theta 1.

00:37:41.050 --> 00:37:47.010 align:middle line:84%
So by the chain rule, you can
essentially factor it out.

00:37:47.010 --> 00:37:51.330 align:middle line:84%
And you can map that
back through all

00:37:51.330 --> 00:37:55.850 align:middle line:84%
of these different layers,
so the inputs and outputs

00:37:55.850 --> 00:37:59.090 align:middle line:84%
of x as it goes through
all these different layers.

00:37:59.090 --> 00:38:02.490 align:middle line:84%
And if you need to calculate
the partial derivative of j

00:38:02.490 --> 00:38:06.570 align:middle line:84%
with respect to that
second layer's parameters,

00:38:06.570 --> 00:38:10.650 align:middle line:84%
you'll note that it's
the same essentially.

00:38:10.650 --> 00:38:13.847 align:middle line:84%
But the thing that's different
is, here, these last terms

00:38:13.847 --> 00:38:14.430 align:middle line:90%
are different.

00:38:14.430 --> 00:38:17.570 align:middle line:84%
But everything for
those two are shared.

00:38:17.570 --> 00:38:20.010 align:middle line:84%
And this is really
the trick, right?

00:38:20.010 --> 00:38:22.850 align:middle line:84%
We could separately compute
all of the derivatives using

00:38:22.850 --> 00:38:25.650 align:middle line:84%
the chain rule, but because
these terms in the gray box

00:38:25.650 --> 00:38:28.130 align:middle line:84%
are shared, we only need
to compute them once.

00:38:28.130 --> 00:38:30.770 align:middle line:84%
So back propagation is a
pretty simple algorithm

00:38:30.770 --> 00:38:34.410 align:middle line:84%
for propagating shared terms
through the computation graph.

00:38:34.410 --> 00:38:35.900 align:middle line:84%
It's basically an
efficiency trick,

00:38:35.900 --> 00:38:37.650 align:middle line:84%
but it's one that makes
it computationally

00:38:37.650 --> 00:38:41.250 align:middle line:84%
practical for very large models
to be able to actually calculate

00:38:41.250 --> 00:38:43.130 align:middle line:90%
these gradients.

00:38:43.130 --> 00:38:46.970 align:middle line:84%
So if this is a forward pass,
we're sending data forward

00:38:46.970 --> 00:38:49.640 align:middle line:84%
through the network, and
then computing the outputs

00:38:49.640 --> 00:38:52.320 align:middle line:84%
and calculating
the overall loss.

00:38:52.320 --> 00:38:54.960 align:middle line:84%
Then a backwards
pass look like this,

00:38:54.960 --> 00:38:58.960 align:middle line:84%
which is that we send our error
signal, the gradients, backwards

00:38:58.960 --> 00:39:01.460 align:middle line:84%
through the network
from the outputs,

00:39:01.460 --> 00:39:05.560 align:middle line:84%
and the loss back to the
inputs and the parameters.

00:39:05.560 --> 00:39:07.320 align:middle line:90%
Does this make sense?

00:39:07.320 --> 00:39:08.800 align:middle line:90%
Cool.

00:39:08.800 --> 00:39:12.280 align:middle line:84%
So what is a backwards pass
look like for a generic layer?

00:39:12.280 --> 00:39:14.600 align:middle line:84%
We're going to keep track
of two kinds of arrays

00:39:14.600 --> 00:39:16.320 align:middle line:90%
of partial derivatives here.

00:39:16.320 --> 00:39:20.840 align:middle line:84%
So first, L is the gradient of
the layer outputs with regard

00:39:20.840 --> 00:39:21.972 align:middle line:90%
to the layer inputs.

00:39:21.972 --> 00:39:23.180 align:middle line:90%
This will be a matrix, right?

00:39:23.180 --> 00:39:26.560 align:middle line:84%
This is, what way does
each of these parameters

00:39:26.560 --> 00:39:30.800 align:middle line:84%
need to change to move
in the maximal direction

00:39:30.800 --> 00:39:32.980 align:middle line:90%
of optimization?

00:39:32.980 --> 00:39:39.640 align:middle line:90%


00:39:39.640 --> 00:39:41.480 align:middle line:84%
So here, this is
essentially a matrix

00:39:41.480 --> 00:39:46.640 align:middle line:84%
that's just telling
you which would

00:39:46.640 --> 00:39:50.240 align:middle line:90%
these parameters optimally move.

00:39:50.240 --> 00:39:52.660 align:middle line:84%
And then second, we're
going to keep track of g.

00:39:52.660 --> 00:39:55.240 align:middle line:84%
And this is the gradient
of the cost with respect

00:39:55.240 --> 00:39:56.260 align:middle line:90%
to the activations.

00:39:56.260 --> 00:39:58.280 align:middle line:90%
And this will be a row vector.

00:39:58.280 --> 00:40:03.160 align:middle line:84%
So this is basically
saying, yeah, essentially,

00:40:03.160 --> 00:40:04.640 align:middle line:84%
what's the partial
derivative of j

00:40:04.640 --> 00:40:08.040 align:middle line:84%
with respect to x, with
respect to that input?

00:40:08.040 --> 00:40:10.040 align:middle line:84%
So now, if we want to
update the parameters,

00:40:10.040 --> 00:40:14.080 align:middle line:84%
that's actually really easy
if we have both L and g.

00:40:14.080 --> 00:40:16.740 align:middle line:84%
Essentially, it's just
a multiplication, right?

00:40:16.740 --> 00:40:19.800 align:middle line:84%
You have g out, and you
multiply that by L of theta.

00:40:19.800 --> 00:40:21.680 align:middle line:84%
And that matrix
multiplication actually

00:40:21.680 --> 00:40:23.000 align:middle line:90%
just updates your layer.

00:40:23.000 --> 00:40:25.440 align:middle line:84%
That tells you
which way to move.

00:40:25.440 --> 00:40:27.440 align:middle line:84%
And so now you take
your parameters,

00:40:27.440 --> 00:40:29.240 align:middle line:84%
and you change them
by your learning

00:40:29.240 --> 00:40:33.320 align:middle line:90%
rate multiplied by that update.

00:40:33.320 --> 00:40:36.520 align:middle line:84%
So how do we actually get
L and g for each layer?

00:40:36.520 --> 00:40:39.160 align:middle line:84%
So L comes from that derivative
function of the layer, which

00:40:39.160 --> 00:40:40.540 align:middle line:90%
we assume is provided.

00:40:40.540 --> 00:40:43.420 align:middle line:84%
We assume for any given layer,
that it is differentiable,

00:40:43.420 --> 00:40:46.440 align:middle line:84%
and that we have a function that
lets us take that derivative.

00:40:46.440 --> 00:40:50.160 align:middle line:84%
And then g is something that
can be computed iteratively

00:40:50.160 --> 00:40:51.340 align:middle line:90%
via recurrence.

00:40:51.340 --> 00:40:55.440 align:middle line:84%
So the input
gradient will always

00:40:55.440 --> 00:41:01.543 align:middle line:84%
be equal to the output gradient
times that L for that layer.

00:41:01.543 --> 00:41:03.960 align:middle line:84%
So this is, essentially, the
back propagation of the error

00:41:03.960 --> 00:41:07.000 align:middle line:90%
signal back through the model.

00:41:07.000 --> 00:41:09.480 align:middle line:84%
And then all of
this machinery is

00:41:09.480 --> 00:41:12.420 align:middle line:84%
used to then compute
parameter update directions,

00:41:12.420 --> 00:41:15.520 align:middle line:90%
so which way do we need to go?

00:41:15.520 --> 00:41:18.720 align:middle line:84%
So the full algorithm--
forward then backward.

00:41:18.720 --> 00:41:20.880 align:middle line:90%
First, we're moving forward.

00:41:20.880 --> 00:41:22.880 align:middle line:90%
Then we go backward.

00:41:22.880 --> 00:41:26.185 align:middle line:90%
And then we update and repeat.

00:41:26.185 --> 00:41:28.560 align:middle line:84%
And this is, essentially, how
you train a neural network,

00:41:28.560 --> 00:41:29.060 align:middle line:90%
right?

00:41:29.060 --> 00:41:31.300 align:middle line:84%
You do a forward
pass, calculate loss.

00:41:31.300 --> 00:41:33.480 align:middle line:84%
You take that loss,
propagate it backward,

00:41:33.480 --> 00:41:36.120 align:middle line:84%
update all your parameters,
and do it again,

00:41:36.120 --> 00:41:41.320 align:middle line:84%
and again, and again, until
you found some likelihood that

00:41:41.320 --> 00:41:44.990 align:middle line:84%
makes you think that your
model is finished training.

00:41:44.990 --> 00:41:47.870 align:middle line:84%
Often, we'll do this-- we'll
talk more about mechanisms

00:41:47.870 --> 00:41:49.870 align:middle line:84%
for this later, but it's
often using something

00:41:49.870 --> 00:41:52.510 align:middle line:84%
like a validation
set and looking

00:41:52.510 --> 00:41:59.750 align:middle line:84%
for some sort of plateauing of
change on that validation set.

00:41:59.750 --> 00:42:04.590 align:middle line:84%
So if we're doing
back propagation,

00:42:04.590 --> 00:42:08.070 align:middle line:84%
now what we've got here,
in the next set of slides,

00:42:08.070 --> 00:42:14.750 align:middle line:84%
are essentially cheat sheets
for a lot of this stuff.

00:42:14.750 --> 00:42:17.290 align:middle line:84%
It's really just intended
to be a pretty nice resource

00:42:17.290 --> 00:42:19.870 align:middle line:84%
if you need to refer back and
think about how you might need

00:42:19.870 --> 00:42:21.590 align:middle line:90%
to implement some of these.

00:42:21.590 --> 00:42:27.070 align:middle line:84%
So now a forward pass
through some hidden layer.

00:42:27.070 --> 00:42:30.350 align:middle line:84%
So now we have the layer
before and the layer after.

00:42:30.350 --> 00:42:31.550 align:middle line:90%
We take this input.

00:42:31.550 --> 00:42:35.750 align:middle line:84%
We run it through the function,
and then we have this output.

00:42:35.750 --> 00:42:38.230 align:middle line:84%
Now when we're doing
that backwards pass,

00:42:38.230 --> 00:42:41.670 align:middle line:84%
we're taking the derivative
of with respect to the output.

00:42:41.670 --> 00:42:45.630 align:middle line:84%
We take the partial derivative
with respect to the input.

00:42:45.630 --> 00:42:49.970 align:middle line:84%
And then we also
are calculating how

00:42:49.970 --> 00:42:52.390 align:middle line:84%
we're taking the parameters
in, and then we're

00:42:52.390 --> 00:42:55.430 align:middle line:84%
calculating the change to
those parameters going out.

00:42:55.430 --> 00:43:00.390 align:middle line:84%
So this means that layer L
for this explicit layer--

00:43:00.390 --> 00:43:04.250 align:middle line:84%
so all the inputs are going
in with the full lines.

00:43:04.250 --> 00:43:08.350 align:middle line:84%
And the outputs are always
in the dotted lines here.

00:43:08.350 --> 00:43:13.950 align:middle line:84%
So the inputs are then the
previous layers output,

00:43:13.950 --> 00:43:21.230 align:middle line:84%
the change in the parameters
with respect to the next layer,

00:43:21.230 --> 00:43:23.630 align:middle line:90%
and the current parameters.

00:43:23.630 --> 00:43:26.470 align:middle line:84%
And then the three
outputs of every layer

00:43:26.470 --> 00:43:30.922 align:middle line:84%
are going to be the output
running through the function,

00:43:30.922 --> 00:43:32.630 align:middle line:84%
the change in the
parameters with respect

00:43:32.630 --> 00:43:39.210 align:middle line:84%
to the previous layer, and the
actual update to the parameters.

00:43:39.210 --> 00:43:40.740 align:middle line:84%
So given those
inputs, we just need

00:43:40.740 --> 00:43:43.460 align:middle line:90%
to evaluate those three things.

00:43:43.460 --> 00:43:47.300 align:middle line:84%
And this actually
makes training happen.

00:43:47.300 --> 00:43:50.220 align:middle line:84%
So now if we take a
summary-- forward pass--

00:43:50.220 --> 00:43:52.420 align:middle line:84%
we'll compute the outputs
for all the layers.

00:43:52.420 --> 00:43:55.100 align:middle line:84%
Backward pass-- we compute the
loss derivatives iteratively

00:43:55.100 --> 00:43:56.085 align:middle line:90%
from top to bottom.

00:43:56.085 --> 00:43:58.460 align:middle line:84%
And then the parameter updates--
we compute the gradients

00:43:58.460 --> 00:43:59.627 align:middle line:90%
with respect to the weights.

00:43:59.627 --> 00:44:02.820 align:middle line:84%
And then we update
all of those weights.

00:44:02.820 --> 00:44:07.200 align:middle line:84%
So now what we actually do
is we do this over batches.

00:44:07.200 --> 00:44:10.042 align:middle line:84%
We're not doing this for a
single data point at a time.

00:44:10.042 --> 00:44:11.500 align:middle line:84%
And technically,
what that means is

00:44:11.500 --> 00:44:14.460 align:middle line:84%
we want to minimize the
average cost over lots

00:44:14.460 --> 00:44:17.000 align:middle line:84%
of data points, the
entire size of a batch.

00:44:17.000 --> 00:44:21.980 align:middle line:84%
And in this day and age,
with very large memory GPUs,

00:44:21.980 --> 00:44:26.100 align:middle line:84%
that batch size could be
thousands, thousands of data

00:44:26.100 --> 00:44:28.180 align:middle line:90%
points at once.

00:44:28.180 --> 00:44:30.900 align:middle line:84%
So then the gradient
of the total cost

00:44:30.900 --> 00:44:33.860 align:middle line:84%
is just the average of all
the gradients of all the

00:44:33.860 --> 00:44:34.900 align:middle line:90%
per data point costs.

00:44:34.900 --> 00:44:37.320 align:middle line:84%
Because when you
take a derivative,

00:44:37.320 --> 00:44:41.940 align:middle line:84%
it can move inside the sum,
so pretty straightforward.

00:44:41.940 --> 00:44:44.733 align:middle line:84%
So now if we actually have
this as a linear layer,

00:44:44.733 --> 00:44:46.400 align:middle line:84%
let's talk through
what that looks like.

00:44:46.400 --> 00:44:49.940 align:middle line:84%
Because it turns out everything
gets really nice and simple.

00:44:49.940 --> 00:44:53.160 align:middle line:84%
So first, forward
propagation is quite simple.

00:44:53.160 --> 00:44:57.580 align:middle line:84%
If we have a linear layer,
the weights of that layer

00:44:57.580 --> 00:44:59.500 align:middle line:90%
are just a matrix.

00:44:59.500 --> 00:45:01.960 align:middle line:84%
So now if you want to
calculate the output.

00:45:01.960 --> 00:45:04.220 align:middle line:84%
You just take that
matrix and multiply it

00:45:04.220 --> 00:45:08.680 align:middle line:84%
by our vector of inputs,
so pretty straightforward.

00:45:08.680 --> 00:45:12.620 align:middle line:84%
If you want to calculate
backpropagation to that input,

00:45:12.620 --> 00:45:16.580 align:middle line:84%
this is essentially
the input gradient

00:45:16.580 --> 00:45:18.940 align:middle line:84%
is going to be the output
gradient multiplied

00:45:18.940 --> 00:45:21.740 align:middle line:84%
by the update to those
parameters with respect

00:45:21.740 --> 00:45:22.740 align:middle line:90%
to the input.

00:45:22.740 --> 00:45:26.200 align:middle line:84%
So essentially, that just
looks like this, right?

00:45:26.200 --> 00:45:27.820 align:middle line:84%
This is just that
matrix that we've

00:45:27.820 --> 00:45:30.620 align:middle line:84%
defined that represents the
direction that you would want

00:45:30.620 --> 00:45:33.205 align:middle line:90%
to move in for those weights.

00:45:33.205 --> 00:45:34.580 align:middle line:84%
And then we look
at the component

00:45:34.580 --> 00:45:36.792 align:middle line:84%
of the output with respect
to the j-th component

00:45:36.792 --> 00:45:39.250 align:middle line:84%
of the input, the i-th component
of the output with respect

00:45:39.250 --> 00:45:41.570 align:middle line:84%
to the j-th component
of the input.

00:45:41.570 --> 00:45:46.930 align:middle line:90%
It turns out that, that is Wij.

00:45:46.930 --> 00:45:51.050 align:middle line:84%
And so that actually
means that this is really

00:45:51.050 --> 00:45:52.930 align:middle line:90%
just the weights again.

00:45:52.930 --> 00:45:57.430 align:middle line:84%
So this is really nice
and simple, right?

00:45:57.430 --> 00:46:00.330 align:middle line:90%


00:46:00.330 --> 00:46:03.010 align:middle line:84%
In the opposite direction,
you multiply the gradient

00:46:03.010 --> 00:46:08.570 align:middle line:84%
by the weights to get
the gradient coming in.

00:46:08.570 --> 00:46:12.050 align:middle line:84%
Cool, so your forward
and backward pass

00:46:12.050 --> 00:46:14.850 align:middle line:84%
are just multiplying
by your weight matrix,

00:46:14.850 --> 00:46:16.810 align:middle line:90%
just in a different order.

00:46:16.810 --> 00:46:18.890 align:middle line:90%
Pretty nice.

00:46:18.890 --> 00:46:23.290 align:middle line:84%
And finally, now
let's see how we

00:46:23.290 --> 00:46:27.350 align:middle line:84%
use those outputs to compute
the weights update equation.

00:46:27.350 --> 00:46:30.810 align:middle line:84%
So how do we actually do that
backprop through to the weights?

00:46:30.810 --> 00:46:35.590 align:middle line:90%
So that looks like this, right?

00:46:35.590 --> 00:46:38.290 align:middle line:84%
We're trying to take the
partial derivative of the cost

00:46:38.290 --> 00:46:40.210 align:middle line:90%
with respect to those weights.

00:46:40.210 --> 00:46:43.050 align:middle line:84%
And so that's going to be the
output gradient multiplied

00:46:43.050 --> 00:46:44.670 align:middle line:90%
by this partial derivative.

00:46:44.670 --> 00:46:47.690 align:middle line:90%


00:46:47.690 --> 00:46:50.650 align:middle line:84%
If we look at the
parameter Wij and look

00:46:50.650 --> 00:46:53.970 align:middle line:84%
at how that parameter
changes the cost,

00:46:53.970 --> 00:46:57.490 align:middle line:84%
only the Ith component of the
output is going to change.

00:46:57.490 --> 00:47:00.810 align:middle line:84%
So this means that,
essentially, you

00:47:00.810 --> 00:47:08.330 align:middle line:84%
can break this down through this
simple almost decomposition.

00:47:08.330 --> 00:47:13.690 align:middle line:84%
And you end up getting this
because of what we just

00:47:13.690 --> 00:47:19.170 align:middle line:84%
pointed out, so essentially the
partial derivative of the output

00:47:19.170 --> 00:47:25.770 align:middle line:84%
at that i-th component with
respect to the weights matrix.

00:47:25.770 --> 00:47:29.130 align:middle line:84%
The corresponding element
ij is actually just going

00:47:29.130 --> 00:47:31.770 align:middle line:90%
to be equal to the input of j.

00:47:31.770 --> 00:47:36.120 align:middle line:90%
And so that means the following.

00:47:36.120 --> 00:47:38.520 align:middle line:84%
Essentially, when you want
to calculate the update

00:47:38.520 --> 00:47:40.560 align:middle line:84%
to the weights,
all you're doing is

00:47:40.560 --> 00:47:44.960 align:middle line:84%
you're multiplying the
input to the function

00:47:44.960 --> 00:47:47.040 align:middle line:90%
by the output gradient.

00:47:47.040 --> 00:47:50.520 align:middle line:84%
Again, it's a really
nice simplification

00:47:50.520 --> 00:47:54.200 align:middle line:84%
that comes out just because
everything is linear.

00:47:54.200 --> 00:47:56.880 align:middle line:84%
So then when we actually
want to update the weights,

00:47:56.880 --> 00:48:02.680 align:middle line:84%
you're literally just taking
eta, and transposing this,

00:48:02.680 --> 00:48:08.400 align:middle line:84%
and multiplying it or
adding it in to the weights.

00:48:08.400 --> 00:48:12.640 align:middle line:84%
Any questions about of very
nice, simple reformulation

00:48:12.640 --> 00:48:15.480 align:middle line:84%
that gives us essentially a
bunch of matrix multiplications?

00:48:15.480 --> 00:48:16.980 align:middle line:90%
Yes?

00:48:16.980 --> 00:48:19.885 align:middle line:90%


00:48:19.885 --> 00:48:21.260 align:middle line:84%
AUDIENCE: Do we
have enough time?

00:48:21.260 --> 00:48:23.040 align:middle line:84%
Can you work through
a small example

00:48:23.040 --> 00:48:26.480 align:middle line:84%
with the computational graph
on the board or something?

00:48:26.480 --> 00:48:28.200 align:middle line:84%
SARA BEERY: So at the
end of the lecture,

00:48:28.200 --> 00:48:30.700 align:middle line:84%
I actually do have
an example of that.

00:48:30.700 --> 00:48:33.880 align:middle line:84%
And so either we'll
get to it at the end--

00:48:33.880 --> 00:48:35.320 align:middle line:90%
I'm a bit worried about time--

00:48:35.320 --> 00:48:36.633 align:middle line:90%
or it's in the slides.

00:48:36.633 --> 00:48:38.300 align:middle line:84%
And so you can actually
work through it.

00:48:38.300 --> 00:48:40.560 align:middle line:84%
And there are the
solutions at the end.

00:48:40.560 --> 00:48:41.880 align:middle line:90%
Sounds good?

00:48:41.880 --> 00:48:42.840 align:middle line:90%
Cool.

00:48:42.840 --> 00:48:47.280 align:middle line:84%
All right, so cheat
sheet for a linear layer.

00:48:47.280 --> 00:48:51.440 align:middle line:84%
The output is equal to the
weights times the input.

00:48:51.440 --> 00:48:53.480 align:middle line:84%
The gradient in is
equal to the output

00:48:53.480 --> 00:48:55.320 align:middle line:90%
gradient times the weights.

00:48:55.320 --> 00:48:58.280 align:middle line:84%
And the update to the
weights is equal to the input

00:48:58.280 --> 00:49:01.120 align:middle line:90%
times the output gradient.

00:49:01.120 --> 00:49:05.653 align:middle line:84%
And then you actually do that
update based on adding in that,

00:49:05.653 --> 00:49:07.320 align:middle line:84%
basically, direction
you want to move in

00:49:07.320 --> 00:49:10.600 align:middle line:84%
multiplied by your
learning rate.

00:49:10.600 --> 00:49:13.492 align:middle line:84%
So now let's look at a
whole multilayer perceptron.

00:49:13.492 --> 00:49:15.700 align:middle line:84%
What happens if we put all
these operations together?

00:49:15.700 --> 00:49:18.920 align:middle line:84%
So first, you have
some linear layer.

00:49:18.920 --> 00:49:21.660 align:middle line:84%
Then, you'll have some ReLU,
some activation function.

00:49:21.660 --> 00:49:24.080 align:middle line:84%
In particular, here,
we're talking about ReLU.

00:49:24.080 --> 00:49:26.180 align:middle line:84%
Then, you have
another linear layer.

00:49:26.180 --> 00:49:29.190 align:middle line:84%
So now this will be
weights one, weights two.

00:49:29.190 --> 00:49:30.750 align:middle line:84%
And then you
calculate some loss,

00:49:30.750 --> 00:49:34.070 align:middle line:84%
based on the difference between
your output and the ground

00:49:34.070 --> 00:49:35.390 align:middle line:90%
truth.

00:49:35.390 --> 00:49:38.750 align:middle line:84%
So now to do a backwards
pass, you essentially

00:49:38.750 --> 00:49:42.110 align:middle line:84%
assume a slight
change in convention.

00:49:42.110 --> 00:49:44.550 align:middle line:84%
This will clarify
this nice connection

00:49:44.550 --> 00:49:46.830 align:middle line:84%
between forward and
backward directions.

00:49:46.830 --> 00:49:49.077 align:middle line:84%
So instead of representing
gradients as row vectors,

00:49:49.077 --> 00:49:50.910 align:middle line:84%
we're going to transpose
them and treat them

00:49:50.910 --> 00:49:52.870 align:middle line:90%
as column vectors.

00:49:52.870 --> 00:49:54.390 align:middle line:84%
And then that
backwards operation

00:49:54.390 --> 00:49:57.230 align:middle line:84%
for the transposed vectors
will follow from that matrix

00:49:57.230 --> 00:50:01.430 align:middle line:84%
identity, that AB transpose is
B transpose A transpose, if you

00:50:01.430 --> 00:50:05.070 align:middle line:90%
remember basic matrix math.

00:50:05.070 --> 00:50:07.950 align:middle line:84%
So essentially,
this then reveals

00:50:07.950 --> 00:50:11.390 align:middle line:84%
this interesting connection
between forward and backward.

00:50:11.390 --> 00:50:14.550 align:middle line:84%
So backward for a linear
layer is the same operation

00:50:14.550 --> 00:50:17.230 align:middle line:84%
as forward but with
the weights transposed.

00:50:17.230 --> 00:50:23.030 align:middle line:84%
So here, now you're
going backward, you get,

00:50:23.030 --> 00:50:26.990 align:middle line:84%
again, weights transposed
times the gradient.

00:50:26.990 --> 00:50:29.510 align:middle line:84%
And then this is an
interesting thing.

00:50:29.510 --> 00:50:31.750 align:middle line:84%
A ReLU wouldn't be a ReLU
on the backwards pass

00:50:31.750 --> 00:50:36.710 align:middle line:84%
because you don't want to
re pass them through ReLU.

00:50:36.710 --> 00:50:38.550 align:middle line:84%
Essentially, what you
need to turn it into

00:50:38.550 --> 00:50:42.030 align:middle line:90%
is roughly like a gating matrix.

00:50:42.030 --> 00:50:44.550 align:middle line:84%
This will be parameterized
by the function

00:50:44.550 --> 00:50:45.990 align:middle line:84%
by the function
of the activations

00:50:45.990 --> 00:50:47.790 align:middle line:90%
from that forward pass.

00:50:47.790 --> 00:50:50.550 align:middle line:84%
So this would be like a,
b, and c, the activations

00:50:50.550 --> 00:50:53.190 align:middle line:84%
that you had sending those
weights through the ReLU

00:50:53.190 --> 00:50:54.470 align:middle line:90%
on the forward pass.

00:50:54.470 --> 00:50:56.430 align:middle line:84%
What this is doing
is making sure

00:50:56.430 --> 00:50:58.430 align:middle line:84%
that you're not
passing gradients

00:50:58.430 --> 00:51:00.550 align:middle line:84%
for the components
of the weights that

00:51:00.550 --> 00:51:10.490 align:middle line:84%
corresponded to the parts that
were in that 0 part of the ReLU.

00:51:10.490 --> 00:51:12.830 align:middle line:84%
You don't want to send
gradients for parts

00:51:12.830 --> 00:51:14.470 align:middle line:90%
that should be masked out.

00:51:14.470 --> 00:51:15.170 align:middle line:90%
Yeah?

00:51:15.170 --> 00:51:18.190 align:middle line:84%
AUDIENCE: Why is that
matrix diagonalized?

00:51:18.190 --> 00:51:22.310 align:middle line:84%
SARA BEERY: That's just so that
it's operating on each element

00:51:22.310 --> 00:51:25.150 align:middle line:90%
independently.

00:51:25.150 --> 00:51:26.500 align:middle line:90%
Cool.

00:51:26.500 --> 00:51:31.840 align:middle line:84%
And then, now, you can mask
the values of the gradients,

00:51:31.840 --> 00:51:34.180 align:middle line:84%
make sure you're not
sending gradients

00:51:34.180 --> 00:51:36.320 align:middle line:84%
for things that are
going into that 0,

00:51:36.320 --> 00:51:38.420 align:middle line:90%
that negative part of the ReLU.

00:51:38.420 --> 00:51:40.940 align:middle line:84%
And then, essentially,
what this means

00:51:40.940 --> 00:51:46.940 align:middle line:84%
is that backprop is
still a linear model.

00:51:46.940 --> 00:51:50.300 align:middle line:84%
Because even that ReLU, this
is still a linear operation

00:51:50.300 --> 00:51:52.220 align:middle line:90%
that we've taken.

00:51:52.220 --> 00:51:55.180 align:middle line:90%
And there's a cool intuition.

00:51:55.180 --> 00:51:58.090 align:middle line:84%
Phil told this to me,
and I really liked it.

00:51:58.090 --> 00:51:59.840 align:middle line:84%
If you're thinking
about a loss landscape,

00:51:59.840 --> 00:52:02.060 align:middle line:84%
no matter how complex
that loss landscape is,

00:52:02.060 --> 00:52:05.440 align:middle line:84%
if we're taking a first order
approximation of the derivative,

00:52:05.440 --> 00:52:07.140 align:middle line:84%
essentially, what
we're doing is we're

00:52:07.140 --> 00:52:11.498 align:middle line:84%
fitting a plane to this
complex curved loss landscape,

00:52:11.498 --> 00:52:13.540 align:middle line:84%
and then we're moving the
direction of the plane.

00:52:13.540 --> 00:52:15.580 align:middle line:84%
So by definition,
it must be linear.

00:52:15.580 --> 00:52:17.780 align:middle line:84%
We're moving on
a planar surface,

00:52:17.780 --> 00:52:21.420 align:middle line:84%
even if the actual underlying
loss is not planar.

00:52:21.420 --> 00:52:23.580 align:middle line:90%
Make sense?

00:52:23.580 --> 00:52:24.960 align:middle line:90%
Cool.

00:52:24.960 --> 00:52:29.300 align:middle line:84%
So this means that forward and
backward for a linear model is

00:52:29.300 --> 00:52:35.180 align:middle line:84%
actually all just one really
big, flipped out linear model.

00:52:35.180 --> 00:52:41.180 align:middle line:84%
You can treat the entire thing
as a single iteration, as just

00:52:41.180 --> 00:52:42.740 align:middle line:90%
one big linear model.

00:52:42.740 --> 00:52:45.660 align:middle line:84%
And it can be thought of as a
single forward pass, in a way,

00:52:45.660 --> 00:52:48.100 align:middle line:84%
even though the actual
had a forward pass is only

00:52:48.100 --> 00:52:51.140 align:middle line:84%
the first half, and the
backward passes the second half.

00:52:51.140 --> 00:52:52.573 align:middle line:90%
Yeah?

00:52:52.573 --> 00:52:54.240 align:middle line:84%
AUDIENCE: Comparing
forward and backward

00:52:54.240 --> 00:52:57.600 align:middle line:84%
pass, you said that to
compute the backward pass,

00:52:57.600 --> 00:53:00.480 align:middle line:84%
you need to have the
x in at each layer.

00:53:00.480 --> 00:53:05.680 align:middle line:84%
So I think that you need to have
in memory for all the layers

00:53:05.680 --> 00:53:07.820 align:middle line:84%
what is the value
of the embedding.

00:53:07.820 --> 00:53:10.260 align:middle line:84%
But as in the forward
pass, you don't need it.

00:53:10.260 --> 00:53:13.620 align:middle line:84%
You only need it
for the-- you can

00:53:13.620 --> 00:53:18.060 align:middle line:84%
discard the previous
embedding, for instance,

00:53:18.060 --> 00:53:20.980 align:middle line:90%
with your [INAUDIBLE].

00:53:20.980 --> 00:53:22.957 align:middle line:84%
Once it has passed
through the first layer,

00:53:22.957 --> 00:53:25.290 align:middle line:84%
then you can just erase the
[INAUDIBLE] from your memory

00:53:25.290 --> 00:53:26.730 align:middle line:90%
as you're only going forward.

00:53:26.730 --> 00:53:31.530 align:middle line:84%
Whereas in the backward pass,
you need to save all of them.

00:53:31.530 --> 00:53:32.330 align:middle line:90%
SARA BEERY: Yes.

00:53:32.330 --> 00:53:36.250 align:middle line:84%
So you do need to save the
intermediate representation so

00:53:36.250 --> 00:53:38.990 align:middle line:84%
that you can efficiently
compute back propagation.

00:53:38.990 --> 00:53:39.550 align:middle line:90%
That's true.

00:53:39.550 --> 00:53:42.175 align:middle line:84%
AUDIENCE: So there is not real
symmetry [INAUDIBLE] the forward

00:53:42.175 --> 00:53:43.290 align:middle line:90%
and backward passes.

00:53:43.290 --> 00:53:44.130 align:middle line:90%
SARA BEERY: No.

00:53:44.130 --> 00:53:46.338 align:middle line:84%
But but, I mean, you can
actually even see that here.

00:53:46.338 --> 00:53:51.330 align:middle line:84%
There's this asymmetry in where
the information is flowing.

00:53:51.330 --> 00:53:52.730 align:middle line:90%
Does that make sense?

00:53:52.730 --> 00:53:54.410 align:middle line:90%
AUDIENCE: Yes.

00:53:54.410 --> 00:53:55.410 align:middle line:90%
SARA BEERY: Yeah?

00:53:55.410 --> 00:53:58.730 align:middle line:84%
AUDIENCE: Is there a specific
reason why we don't see

00:53:58.730 --> 00:54:00.330 align:middle line:90%
[INAUDIBLE]?

00:54:00.330 --> 00:54:02.930 align:middle line:84%
SARA BEERY: Oh, that was
just for simplicity here.

00:54:02.930 --> 00:54:04.930 align:middle line:84%
So, yeah, in
practice, there would

00:54:04.930 --> 00:54:06.670 align:middle line:84%
be bias terms floating
around everywhere,

00:54:06.670 --> 00:54:10.050 align:middle line:84%
but we just wanted to
capture the intuition.

00:54:10.050 --> 00:54:11.770 align:middle line:90%
Yeah?

00:54:11.770 --> 00:54:22.580 align:middle line:84%
AUDIENCE: When we're going
back through our ReLU layer,

00:54:22.580 --> 00:54:26.650 align:middle line:84%
do we need to know
the pre-activations--

00:54:26.650 --> 00:54:28.180 align:middle line:90%
we need to have that in memory?

00:54:28.180 --> 00:54:29.930 align:middle line:84%
SARA BEERY: You basically
need to save off

00:54:29.930 --> 00:54:32.010 align:middle line:84%
what the value of
the activations

00:54:32.010 --> 00:54:34.140 align:middle line:84%
were for each component
of your input.

00:54:34.140 --> 00:54:35.890 align:middle line:84%
And you keep those in
memory, and then you

00:54:35.890 --> 00:54:39.890 align:middle line:84%
build that sparse matrix
that lets you actually

00:54:39.890 --> 00:54:42.510 align:middle line:90%
gate those going back through.

00:54:42.510 --> 00:54:43.010 align:middle line:90%
Yeah.

00:54:43.010 --> 00:54:43.870 align:middle line:90%
Cool.

00:54:43.870 --> 00:54:44.370 align:middle line:90%
OK.

00:54:44.370 --> 00:54:49.170 align:middle line:84%
So let's move on to
talking about DAGs.

00:54:49.170 --> 00:54:52.250 align:middle line:84%
So what if your
DAG isn't a chain?

00:54:52.250 --> 00:54:54.250 align:middle line:84%
And actually, this
is pretty common.

00:54:54.250 --> 00:55:00.170 align:middle line:84%
We often have connections within
our machine-learning models

00:55:00.170 --> 00:55:04.030 align:middle line:84%
that are not just
chains, for example,

00:55:04.030 --> 00:55:06.130 align:middle line:84%
sharing weights, or
splitting weights,

00:55:06.130 --> 00:55:09.970 align:middle line:84%
or having things come
together, be concatenated

00:55:09.970 --> 00:55:16.250 align:middle line:84%
from different heads or being
split into different objectives.

00:55:16.250 --> 00:55:19.320 align:middle line:84%
So let's talk about what
this actually looks like.

00:55:19.320 --> 00:55:21.520 align:middle line:84%
So there's actually
only two operations

00:55:21.520 --> 00:55:23.840 align:middle line:84%
that you need to make
everything we just

00:55:23.840 --> 00:55:26.960 align:middle line:84%
talked about work for any
arbitrary directed-acyclic

00:55:26.960 --> 00:55:27.500 align:middle line:90%
graph.

00:55:27.500 --> 00:55:29.520 align:middle line:84%
And those are merging
and branching.

00:55:29.520 --> 00:55:32.697 align:middle line:84%
So assuming you have
some merge where, here--

00:55:32.697 --> 00:55:34.780 align:middle line:84%
and that merge could be a
lot of different things.

00:55:34.780 --> 00:55:35.655 align:middle line:90%
It could be addition.

00:55:35.655 --> 00:55:38.080 align:middle line:84%
It could be concatenation,
what have you.

00:55:38.080 --> 00:55:42.320 align:middle line:84%
You have some components,
some functional components

00:55:42.320 --> 00:55:45.000 align:middle line:84%
that you're merging
together to move forward

00:55:45.000 --> 00:55:48.350 align:middle line:84%
through your
directed-acyclic graph.

00:55:48.350 --> 00:55:50.600 align:middle line:84%
Now, if you actually need
to propagate a gradient back

00:55:50.600 --> 00:55:53.280 align:middle line:84%
through that,
essentially, you just

00:55:53.280 --> 00:55:55.760 align:middle line:84%
need to make sure that you
track the gradient with respect

00:55:55.760 --> 00:55:57.340 align:middle line:90%
to the correct input variables.

00:55:57.340 --> 00:56:01.960 align:middle line:84%
So if, here, you just only
pass the gradient that

00:56:01.960 --> 00:56:07.280 align:middle line:84%
corresponds to the x superscript
a, and on the other path,

00:56:07.280 --> 00:56:10.600 align:middle line:84%
you only pass the gradient that
corresponds to the components

00:56:10.600 --> 00:56:13.400 align:middle line:90%
of that other part of the path.

00:56:13.400 --> 00:56:17.440 align:middle line:84%
Similarly, if you're
branching, so now you're

00:56:17.440 --> 00:56:22.000 align:middle line:84%
splitting up your information
in some way or even just

00:56:22.000 --> 00:56:23.100 align:middle line:90%
duplicating it.

00:56:23.100 --> 00:56:25.580 align:middle line:84%
Now, if you're propagating a
gradient back through that,

00:56:25.580 --> 00:56:27.280 align:middle line:90%
it's as simple as a sum.

00:56:27.280 --> 00:56:30.240 align:middle line:84%
You just take the gradients
from both, and you sum them,

00:56:30.240 --> 00:56:33.480 align:middle line:84%
and you send that back
through that input.

00:56:33.480 --> 00:56:34.560 align:middle line:90%
Cool.

00:56:34.560 --> 00:56:38.440 align:middle line:84%
So there's a much more detailed
derivation in the lecture notes,

00:56:38.440 --> 00:56:40.520 align:middle line:90%
if you're interested.

00:56:40.520 --> 00:56:42.080 align:middle line:84%
But so now let's
talk about what this

00:56:42.080 --> 00:56:43.800 align:middle line:90%
means for parameter sharing.

00:56:43.800 --> 00:56:46.140 align:middle line:90%
So say you have some chain.

00:56:46.140 --> 00:56:48.640 align:middle line:84%
And now you have
the parameters are

00:56:48.640 --> 00:56:52.400 align:middle line:84%
separate for each
component of that chain.

00:56:52.400 --> 00:56:55.160 align:middle line:84%
You might instead
have something where

00:56:55.160 --> 00:56:56.680 align:middle line:84%
those parameters
are actually shared

00:56:56.680 --> 00:56:58.480 align:middle line:84%
across those
different components.

00:56:58.480 --> 00:57:00.200 align:middle line:84%
Maybe you're reusing
parameters that

00:57:00.200 --> 00:57:02.480 align:middle line:84%
were pre-trained on
ImageNet or something,

00:57:02.480 --> 00:57:05.920 align:middle line:84%
and you're using them in
multiple parts of your network.

00:57:05.920 --> 00:57:07.760 align:middle line:84%
There, again, you
can actually just

00:57:07.760 --> 00:57:10.120 align:middle line:84%
think about flipping
this on its side.

00:57:10.120 --> 00:57:11.767 align:middle line:90%
All that is a branch, right?

00:57:11.767 --> 00:57:13.600 align:middle line:84%
So now, if you're passing
your gradient back

00:57:13.600 --> 00:57:15.870 align:middle line:84%
through that branch,
you just need

00:57:15.870 --> 00:57:18.390 align:middle line:84%
to be summing the gradients
that are coming in

00:57:18.390 --> 00:57:21.830 align:middle line:84%
from these different
branch dimensions.

00:57:21.830 --> 00:57:24.790 align:middle line:84%
So parameter sharing,
some ingredients.

00:57:24.790 --> 00:57:25.290 align:middle line:90%
All right.

00:57:25.290 --> 00:57:27.430 align:middle line:84%
So towards the end
of the lecture, now

00:57:27.430 --> 00:57:30.790 align:middle line:84%
I want to move on and talk
about this kind of meta concept

00:57:30.790 --> 00:57:32.590 align:middle line:84%
of what differential
programming is

00:57:32.590 --> 00:57:35.590 align:middle line:84%
and maybe why neural network
training is, essentially,

00:57:35.590 --> 00:57:38.070 align:middle line:90%
just differential programming.

00:57:38.070 --> 00:57:40.310 align:middle line:84%
So when we think
about deep learning,

00:57:40.310 --> 00:57:42.150 align:middle line:84%
generally, we think
about it this way.

00:57:42.150 --> 00:57:46.230 align:middle line:84%
You have some
network, again, a DAG.

00:57:46.230 --> 00:57:48.950 align:middle line:84%
And then you define these
forward and backward passes

00:57:48.950 --> 00:57:50.710 align:middle line:90%
through that network.

00:57:50.710 --> 00:57:52.930 align:middle line:84%
Differential programming
takes the same perspective,

00:57:52.930 --> 00:57:58.350 align:middle line:84%
but now we basically just are
actually programming that model.

00:57:58.350 --> 00:58:02.670 align:middle line:84%
And this is really what all
of these different existing

00:58:02.670 --> 00:58:03.230 align:middle line:90%
languages--

00:58:03.230 --> 00:58:08.470 align:middle line:84%
PyTorch, TensorFlow, if
anyone still uses it, JAX--

00:58:08.470 --> 00:58:12.110 align:middle line:84%
these are all basically
libraries that

00:58:12.110 --> 00:58:16.550 align:middle line:84%
enable us to really efficiently
do differential programming.

00:58:16.550 --> 00:58:19.510 align:middle line:84%
So deep networks are popular
for a few different reasons.

00:58:19.510 --> 00:58:21.350 align:middle line:90%
First, they're easy to optimize.

00:58:21.350 --> 00:58:22.910 align:middle line:90%
Everything is differentiable.

00:58:22.910 --> 00:58:25.190 align:middle line:90%
Second, they're compositional.

00:58:25.190 --> 00:58:28.390 align:middle line:84%
It's essentially like
block-based programming.

00:58:28.390 --> 00:58:31.390 align:middle line:84%
And so differential
programming surfaced

00:58:31.390 --> 00:58:35.443 align:middle line:84%
as an emerging term for general
models with those properties.

00:58:35.443 --> 00:58:38.110 align:middle line:84%
These tweets are a bit old now,
but I still think they're funny.

00:58:38.110 --> 00:58:41.590 align:middle line:84%
Yann LeCun said, "deep learning
has outlived its usefulness

00:58:41.590 --> 00:58:44.430 align:middle line:84%
as a buzz phrase, deep
learning, est mort.

00:58:44.430 --> 00:58:46.470 align:middle line:90%
Vive differential programming."

00:58:46.470 --> 00:58:49.270 align:middle line:84%
Essentially, even
years ago, this

00:58:49.270 --> 00:58:51.510 align:middle line:84%
was like maybe this is
the new buzzword for what

00:58:51.510 --> 00:58:53.230 align:middle line:90%
we're talking about here.

00:58:53.230 --> 00:58:56.190 align:middle line:84%
Tom Ditterich, who is
significantly less splashy

00:58:56.190 --> 00:58:59.230 align:middle line:84%
in his tweets, said that,
"deep learning is essentially

00:58:59.230 --> 00:59:02.330 align:middle line:84%
a new style of programming,
differential programming.

00:59:02.330 --> 00:59:03.750 align:middle line:84%
And the field is
trying to work up

00:59:03.750 --> 00:59:07.022 align:middle line:84%
what the reusable constructs
in this style are.

00:59:07.022 --> 00:59:08.730 align:middle line:84%
And there are some
that we actually know.

00:59:08.730 --> 00:59:11.140 align:middle line:84%
Some of these blocks
have been well defined,

00:59:11.140 --> 00:59:14.780 align:middle line:84%
things like convolution,
pooling, LSTMs, GANs, VAEs,

00:59:14.780 --> 00:59:18.860 align:middle line:84%
memory units, routing
units, et cetera."

00:59:18.860 --> 00:59:23.640 align:middle line:84%
So this is taking
that perspective.

00:59:23.640 --> 00:59:26.872 align:middle line:84%
This is a paper from 2017 called
Neural Module Networks, where

00:59:26.872 --> 00:59:28.580 align:middle line:84%
essentially, here,
you can see that there

00:59:28.580 --> 00:59:31.360 align:middle line:84%
are these different
components of the model.

00:59:31.360 --> 00:59:35.260 align:middle line:84%
And some of them might be
architectural, and some of them

00:59:35.260 --> 00:59:38.740 align:middle line:84%
might have parameters
that are learned or not.

00:59:38.740 --> 00:59:42.420 align:middle line:84%
So you're sending in something
like, where is the dog?

00:59:42.420 --> 00:59:44.367 align:middle line:84%
And then maybe that
gets routed to a parser,

00:59:44.367 --> 00:59:46.200 align:middle line:84%
and that parser might
not be learned at all.

00:59:46.200 --> 00:59:48.480 align:middle line:84%
It might just be
a standard parser.

00:59:48.480 --> 00:59:52.060 align:middle line:84%
So that component might not have
any differential parameters that

00:59:52.060 --> 00:59:54.220 align:middle line:90%
are actually being learned.

00:59:54.220 --> 00:59:56.300 align:middle line:84%
But then, potentially,
there's a CNN

00:59:56.300 --> 00:59:58.340 align:middle line:84%
down here that's
parsing the objects that

00:59:58.340 --> 00:59:59.740 align:middle line:90%
are seen in an image.

00:59:59.740 --> 01:00:02.300 align:middle line:84%
And that might
actually be something

01:00:02.300 --> 01:00:04.300 align:middle line:84%
that you have updates
being propagated to

01:00:04.300 --> 01:00:06.540 align:middle line:90%
from your training data set.

01:00:06.540 --> 01:00:09.640 align:middle line:84%
And there's an interesting
blog post from Andrej Karpathy

01:00:09.640 --> 01:00:13.580 align:middle line:84%
that talks about this as
what he calls Software 2.0.

01:00:13.580 --> 01:00:16.100 align:middle line:84%
Essentially, here, he's saying
that Software 1.0, which

01:00:16.100 --> 01:00:18.940 align:middle line:84%
is the software we had
originally, where everything

01:00:18.940 --> 01:00:22.540 align:middle line:84%
is defined, that might be
just that little, red dot.

01:00:22.540 --> 01:00:24.700 align:middle line:84%
Software 2.0 is
actually defining

01:00:24.700 --> 01:00:30.000 align:middle line:84%
a space of possible
software systems,

01:00:30.000 --> 01:00:32.380 align:middle line:84%
and then you're actually
optimizing within that space

01:00:32.380 --> 01:00:38.260 align:middle line:84%
to build the system that you
need to solve your problem.

01:00:38.260 --> 01:00:42.380 align:middle line:84%
So here, instead of thinking
about this DAG, where

01:00:42.380 --> 01:00:46.340 align:middle line:84%
every component is
actually differentiable

01:00:46.340 --> 01:00:48.360 align:middle line:84%
or every component
is differentiable--

01:00:48.360 --> 01:00:50.120 align:middle line:84%
where every
component is learned,

01:00:50.120 --> 01:00:52.900 align:middle line:84%
instead, you might have a bunch
of the components of your system

01:00:52.900 --> 01:00:55.300 align:middle line:84%
be explicitly
programmed by a human.

01:00:55.300 --> 01:00:59.660 align:middle line:84%
A really obvious example of this
is we often explicitly program

01:00:59.660 --> 01:01:03.100 align:middle line:84%
the normalization values that we
want to use for natural images

01:01:03.100 --> 01:01:05.620 align:middle line:84%
directly into the
pre-processing of our data.

01:01:05.620 --> 01:01:07.490 align:middle line:84%
This is something
that we just define.

01:01:07.490 --> 01:01:10.250 align:middle line:84%
We don't learn what those
normalization values should be.

01:01:10.250 --> 01:01:13.090 align:middle line:84%
And then there are
these other components

01:01:13.090 --> 01:01:18.970 align:middle line:84%
where this part is actually
programmed by back propagation,

01:01:18.970 --> 01:01:21.330 align:middle line:84%
so, for example, programmed
by tuning the behavior

01:01:21.330 --> 01:01:23.770 align:middle line:90%
to match the training examples.

01:01:23.770 --> 01:01:25.850 align:middle line:84%
And the cool thing
is because we assume

01:01:25.850 --> 01:01:29.050 align:middle line:84%
that every piece of this
system is differentiable,

01:01:29.050 --> 01:01:32.690 align:middle line:84%
you can actually optimize
any node or any edge

01:01:32.690 --> 01:01:35.650 align:middle line:90%
with regard to any scalar cost.

01:01:35.650 --> 01:01:40.810 align:middle line:84%
So here, you could think about
calculating how the cost would

01:01:40.810 --> 01:01:46.010 align:middle line:84%
change when the weights of the
specific yellow function change.

01:01:46.010 --> 01:01:49.570 align:middle line:84%
You could also
think about looking

01:01:49.570 --> 01:01:52.910 align:middle line:84%
at how the cost changes when
the input data changes, right?

01:01:52.910 --> 01:01:54.730 align:middle line:84%
You don't actually
need to specifically do

01:01:54.730 --> 01:01:57.190 align:middle line:90%
this for weights.

01:01:57.190 --> 01:01:58.010 align:middle line:90%
Yeah?

01:01:58.010 --> 01:02:02.890 align:middle line:84%
AUDIENCE: So is this how
we create a metric that

01:02:02.890 --> 01:02:06.570 align:middle line:84%
can determine what might be
good for a human to train versus

01:02:06.570 --> 01:02:07.670 align:middle line:90%
the computer to train?

01:02:07.670 --> 01:02:10.050 align:middle line:84%
Or is there another
qualitative or empirical metric

01:02:10.050 --> 01:02:11.770 align:middle line:90%
that we might use?

01:02:11.770 --> 01:02:14.170 align:middle line:90%
SARA BEERY: So a big question.

01:02:14.170 --> 01:02:17.330 align:middle line:84%
So he was asking if
this is something

01:02:17.330 --> 01:02:21.010 align:middle line:84%
that helps us build metrics
around what maybe we might want

01:02:21.010 --> 01:02:26.130 align:middle line:84%
to optimize via data versus what
we want to just define directly.

01:02:26.130 --> 01:02:29.970 align:middle line:84%
I think, experimentally, we
often find those things out

01:02:29.970 --> 01:02:31.930 align:middle line:90%
for different applications.

01:02:31.930 --> 01:02:35.250 align:middle line:84%
What parts does it
actually good to--

01:02:35.250 --> 01:02:39.190 align:middle line:84%
essentially, anytime a human is
programming part of this system,

01:02:39.190 --> 01:02:40.790 align:middle line:90%
it's constraining the system.

01:02:40.790 --> 01:02:42.430 align:middle line:84%
And that constraint
can be useful,

01:02:42.430 --> 01:02:43.830 align:middle line:90%
but it could also be unuseful.

01:02:43.830 --> 01:02:46.930 align:middle line:84%
I think about when I
started in machine learning,

01:02:46.930 --> 01:02:47.890 align:middle line:90%
it was built around--

01:02:47.890 --> 01:02:49.390 align:middle line:84%
I'm going to date
myself right now--

01:02:49.390 --> 01:02:51.950 align:middle line:84%
but we did a lot of what we
called feature engineering,

01:02:51.950 --> 01:02:53.690 align:middle line:84%
which is essentially
like us building

01:02:53.690 --> 01:02:56.650 align:middle line:84%
a bunch of explicit constraints
into how the data might be

01:02:56.650 --> 01:03:00.690 align:middle line:84%
processed to then go through a
very, very simple neural network

01:03:00.690 --> 01:03:04.120 align:middle line:84%
or even something more
simple than that, something

01:03:04.120 --> 01:03:06.240 align:middle line:90%
like an SDM.

01:03:06.240 --> 01:03:08.720 align:middle line:84%
And those bottlenecks,
it turned out,

01:03:08.720 --> 01:03:10.420 align:middle line:90%
were not necessarily optimal.

01:03:10.420 --> 01:03:11.932 align:middle line:90%
We were defining this.

01:03:11.932 --> 01:03:13.640 align:middle line:84%
We were programming
it, but it turned out

01:03:13.640 --> 01:03:18.580 align:middle line:84%
our best ideas were not as good
as just a much more complex,

01:03:18.580 --> 01:03:21.185 align:middle line:84%
larger models that were
learned somewhat end-to-end.

01:03:21.185 --> 01:03:22.060 align:middle line:90%
Does that make sense?

01:03:22.060 --> 01:03:23.540 align:middle line:90%
So I think the jury's still out.

01:03:23.540 --> 01:03:25.290 align:middle line:84%
When you talk about
metrics, that's really

01:03:25.290 --> 01:03:27.400 align:middle line:90%
where that red square comes in.

01:03:27.400 --> 01:03:29.680 align:middle line:84%
You define what that
cost function is.

01:03:29.680 --> 01:03:34.240 align:middle line:84%
So you have to figure out how
you are going to define the loss

01:03:34.240 --> 01:03:36.260 align:middle line:84%
and what that then looks
like for your cost.

01:03:36.260 --> 01:03:40.860 align:middle line:84%
And that definition will
affect what your optimal are.

01:03:40.860 --> 01:03:41.360 align:middle line:90%
Cool.

01:03:41.360 --> 01:03:45.360 align:middle line:84%
So let's talk a little bit
about optimizing parameters

01:03:45.360 --> 01:03:47.960 align:middle line:90%
versus optimizing inputs.

01:03:47.960 --> 01:03:50.765 align:middle line:84%
So this is the standard
approach, right?

01:03:50.765 --> 01:03:52.640 align:middle line:84%
This is looking at how
much the total cost is

01:03:52.640 --> 01:03:55.380 align:middle line:84%
increased or decreased by
changing the parameters.

01:03:55.380 --> 01:03:58.280 align:middle line:84%
This is all the
different parameters

01:03:58.280 --> 01:04:00.000 align:middle line:90%
throughout the network.

01:04:00.000 --> 01:04:02.600 align:middle line:84%
This is what we
commonly think of,

01:04:02.600 --> 01:04:08.520 align:middle line:84%
but you can also think about
optimizing the input for a fixed

01:04:08.520 --> 01:04:11.440 align:middle line:90%
set of parameters.

01:04:11.440 --> 01:04:15.200 align:middle line:84%
So here, this could be
something how much the chameleon

01:04:15.200 --> 01:04:20.360 align:middle line:84%
score is increased or decreased
by changing the input pixels.

01:04:20.360 --> 01:04:26.040 align:middle line:84%
And so here, now, if you
think about this from a way

01:04:26.040 --> 01:04:28.960 align:middle line:84%
where you're optimizing maybe
class logits before the softmax

01:04:28.960 --> 01:04:32.400 align:middle line:84%
or optimizing the class
probabilities after the softmax,

01:04:32.400 --> 01:04:36.480 align:middle line:84%
the maybe easiest way to
increase the probability softmax

01:04:36.480 --> 01:04:40.160 align:middle line:84%
given to a class is often to
make the alternatives unlikely,

01:04:40.160 --> 01:04:43.040 align:middle line:84%
rather than to make the
class of interest likely.

01:04:43.040 --> 01:04:45.460 align:middle line:84%
Whereas if you optimize
pre softmax logits,

01:04:45.460 --> 01:04:47.720 align:middle line:84%
this tends to actually
be a bit more stable.

01:04:47.720 --> 01:04:49.680 align:middle line:84%
So what we find is if
you want to actually make

01:04:49.680 --> 01:04:52.960 align:middle line:84%
an image that maximizes
the cat output neuron--

01:04:52.960 --> 01:04:57.040 align:middle line:84%
again, we are doing
this maybe pre softmax--

01:04:57.040 --> 01:04:58.580 align:middle line:84%
it might look
something like this.

01:04:58.580 --> 01:05:01.350 align:middle line:84%
So this is an
interesting, again,

01:05:01.350 --> 01:05:04.870 align:middle line:84%
blog post back from 2017,
but on feature visualization.

01:05:04.870 --> 01:05:08.230 align:middle line:84%
So this is telling you what
a given trained model thinks

01:05:08.230 --> 01:05:11.910 align:middle line:90%
is most cat like.

01:05:11.910 --> 01:05:13.950 align:middle line:84%
And this is something
you can learn.

01:05:13.950 --> 01:05:17.390 align:middle line:84%
Or you can actually
also do things

01:05:17.390 --> 01:05:19.790 align:middle line:84%
that say, for example, make
an image that maximizes

01:05:19.790 --> 01:05:23.710 align:middle line:84%
the value of a neuron,
j, on a layer, L.

01:05:23.710 --> 01:05:27.990 align:middle line:84%
So using this as a mechanism to
probe what the model is paying

01:05:27.990 --> 01:05:29.110 align:middle line:90%
attention to.

01:05:29.110 --> 01:05:31.950 align:middle line:84%
So that might look
something like this.

01:05:31.950 --> 01:05:34.150 align:middle line:84%
This is saying that
this particular neuron

01:05:34.150 --> 01:05:37.870 align:middle line:84%
and this particular layer is
looking for stuff that looks

01:05:37.870 --> 01:05:40.630 align:middle line:90%
like this in a hand-wavy way.

01:05:40.630 --> 01:05:43.985 align:middle line:84%
And this is the idea behind,
for example, DeepDream.

01:05:43.985 --> 01:05:46.610 align:middle line:84%
I don't know if any of you saw
this, a couple of years old now.

01:05:46.610 --> 01:05:49.870 align:middle line:84%
But I think these are
psychedelic and beautiful.

01:05:49.870 --> 01:05:56.790 align:middle line:84%
And this is looking at what
these models, I guess, dream of.

01:05:56.790 --> 01:06:00.605 align:middle line:84%
And this is an idea
of what's to come.

01:06:00.605 --> 01:06:02.230 align:middle line:84%
But if you're
comfortable with the idea

01:06:02.230 --> 01:06:05.670 align:middle line:84%
that you can optimize anything
with respect to anything,

01:06:05.670 --> 01:06:09.070 align:middle line:84%
here, we're looking
at maximizing clip

01:06:09.070 --> 01:06:12.910 align:middle line:84%
as a model that tries
to make descriptions

01:06:12.910 --> 01:06:15.270 align:middle line:84%
and images that
correspond to them

01:06:15.270 --> 01:06:17.990 align:middle line:84%
close together in a
joint embedding space.

01:06:17.990 --> 01:06:20.710 align:middle line:84%
So you want to build a text
encoder and an image encoder

01:06:20.710 --> 01:06:23.370 align:middle line:84%
where things that are
similar, semantically,

01:06:23.370 --> 01:06:25.350 align:middle line:84%
from text to images,
are quite close

01:06:25.350 --> 01:06:28.270 align:middle line:84%
together in the learned
embedding space, that projection

01:06:28.270 --> 01:06:30.230 align:middle line:84%
that you're projecting
each of them into.

01:06:30.230 --> 01:06:32.670 align:middle line:84%
Now, if you have
something like CLIP,

01:06:32.670 --> 01:06:39.670 align:middle line:84%
then you can, for example,
tie your CLIP model together,

01:06:39.670 --> 01:06:42.270 align:middle line:84%
the image encoder and the text
encoder from your CLIP model,

01:06:42.270 --> 01:06:45.510 align:middle line:84%
together with a GAN, something
like a generative adversarial

01:06:45.510 --> 01:06:46.470 align:middle line:90%
network.

01:06:46.470 --> 01:06:48.910 align:middle line:84%
And then you can
optimize the embedding

01:06:48.910 --> 01:06:52.510 align:middle line:84%
or that input parameter that
goes into that image generator.

01:06:52.510 --> 01:06:56.580 align:middle line:84%
And you can figure
out, for a given input,

01:06:56.580 --> 01:06:59.780 align:middle line:84%
how to optimize this so that
your output of your image

01:06:59.780 --> 01:07:03.620 align:middle line:84%
generator is the thing
that's closest in CLIP space

01:07:03.620 --> 01:07:07.260 align:middle line:90%
to any given prompt.

01:07:07.260 --> 01:07:10.620 align:middle line:84%
And if you actually look at
what that learns over time,

01:07:10.620 --> 01:07:13.500 align:middle line:84%
you can get stuff
that's fascinating.

01:07:13.500 --> 01:07:17.500 align:middle line:84%
So this is a way to actually
generate an image from text.

01:07:17.500 --> 01:07:20.360 align:middle line:84%
And here, again, all these
trapezoids are neural networks.

01:07:20.360 --> 01:07:21.520 align:middle line:90%
You can plug them together.

01:07:21.520 --> 01:07:23.500 align:middle line:84%
You can take the components
trained in one way

01:07:23.500 --> 01:07:25.340 align:middle line:90%
and use them in another way.

01:07:25.340 --> 01:07:27.740 align:middle line:84%
Really, the idea is you
can optimize modules

01:07:27.740 --> 01:07:30.020 align:middle line:84%
with respect to all
the other modules,

01:07:30.020 --> 01:07:33.500 align:middle line:90%
and the world's your oyster.

01:07:33.500 --> 01:07:34.980 align:middle line:90%
Cool.

01:07:34.980 --> 01:07:36.900 align:middle line:90%
So what we covered today--

01:07:36.900 --> 01:07:39.220 align:middle line:84%
review of gradient
descent and SGD,

01:07:39.220 --> 01:07:41.880 align:middle line:84%
computation graphs,
backprop through chains,

01:07:41.880 --> 01:07:45.820 align:middle line:84%
backprop through MLPs,
backprop through DAGs,

01:07:45.820 --> 01:07:47.980 align:middle line:90%
and differential programming.

01:07:47.980 --> 01:07:49.080 align:middle line:90%
Any questions?

01:07:49.080 --> 01:07:54.420 align:middle line:90%


01:07:54.420 --> 01:07:55.443 align:middle line:90%
Yes?

01:07:55.443 --> 01:07:57.360 align:middle line:84%
AUDIENCE: I have a
question about [INAUDIBLE].

01:07:57.360 --> 01:08:02.540 align:middle line:90%


01:08:02.540 --> 01:08:05.620 align:middle line:84%
So first, why do you have to
train the network [INAUDIBLE]?

01:08:05.620 --> 01:08:12.260 align:middle line:90%


01:08:12.260 --> 01:08:14.260 align:middle line:84%
SARA BEERY: So essentially
what we're doing here

01:08:14.260 --> 01:08:15.720 align:middle line:90%
is you had CLIP.

01:08:15.720 --> 01:08:17.899 align:middle line:84%
You have this text encoder
and this image encoder.

01:08:17.899 --> 01:08:20.995 align:middle line:84%
And they were already trained,
and now they're fixed, right?

01:08:20.995 --> 01:08:23.620 align:middle line:84%
So the idea is, here, the text
and the image encoder, actually,

01:08:23.620 --> 01:08:25.700 align:middle line:84%
we're not backpropping
through them at all.

01:08:25.700 --> 01:08:27.220 align:middle line:90%
They're now defined.

01:08:27.220 --> 01:08:28.740 align:middle line:90%
They're not being learned.

01:08:28.740 --> 01:08:31.660 align:middle line:84%
Similarly, you've separately
trained an image generation

01:08:31.660 --> 01:08:33.899 align:middle line:84%
model that takes in
any given embedding

01:08:33.899 --> 01:08:36.819 align:middle line:84%
and generates an image
from that embedding.

01:08:36.819 --> 01:08:39.260 align:middle line:90%
Now that one is also fixed.

01:08:39.260 --> 01:08:41.517 align:middle line:84%
So now you have a lot of
parameters in a model,

01:08:41.517 --> 01:08:42.600 align:middle line:90%
but all of them are fixed.

01:08:42.600 --> 01:08:44.180 align:middle line:84%
And the only thing
you're optimizing

01:08:44.180 --> 01:08:50.540 align:middle line:84%
is this hidden parameter
that's going to be the input

01:08:50.540 --> 01:08:55.490 align:middle line:84%
to the image generator, such
that this is as is maximal.

01:08:55.490 --> 01:08:59.250 align:middle line:84%
So you're trying to
find the sine of z value

01:08:59.250 --> 01:09:02.830 align:middle line:84%
that maximizes this with
respect to that input text.

01:09:02.830 --> 01:09:08.580 align:middle line:90%
AUDIENCE: So

01:09:08.580 --> 01:09:09.330 align:middle line:90%
SARA BEERY: Sorry?

01:09:09.330 --> 01:09:15.050 align:middle line:84%
AUDIENCE: [INAUDIBLE]
optimization [INAUDIBLE]?

01:09:15.050 --> 01:09:17.370 align:middle line:84%
SARA BEERY: Yeah, so all the
parameters here are fixed.

01:09:17.370 --> 01:09:20.689 align:middle line:84%
But you are going
to need to calculate

01:09:20.689 --> 01:09:23.370 align:middle line:84%
how to backprop the signal
back through these models

01:09:23.370 --> 01:09:26.250 align:middle line:90%
to get to that one.

01:09:26.250 --> 01:09:27.609 align:middle line:90%
Yeah?

01:09:27.609 --> 01:09:30.820 align:middle line:84%
AUDIENCE: Can you explain more
about embedding [INAUDIBLE]?

01:09:30.820 --> 01:09:32.090 align:middle line:90%
SARA BEERY: Yeah, absolutely.

01:09:32.090 --> 01:09:33.970 align:middle line:90%
So when I say embedding--

01:09:33.970 --> 01:09:37.529 align:middle line:84%
I might also say
representation--

01:09:37.529 --> 01:09:40.729 align:middle line:84%
these are all
terminology that we often

01:09:40.729 --> 01:09:43.770 align:middle line:84%
will throw around in the
machine-learning community.

01:09:43.770 --> 01:09:47.170 align:middle line:84%
But usually it means
something like this.

01:09:47.170 --> 01:09:51.290 align:middle line:90%
So you have some neural network.

01:09:51.290 --> 01:09:53.529 align:middle line:84%
And we'll call them
an encoder sometimes,

01:09:53.529 --> 01:09:55.250 align:middle line:90%
again, just terminology.

01:09:55.250 --> 01:09:59.050 align:middle line:84%
But this neural network is
essentially building a mapping

01:09:59.050 --> 01:10:02.150 align:middle line:84%
from a high-dimensional
input or a complex input,

01:10:02.150 --> 01:10:06.050 align:middle line:84%
something like an image, to a
low-dimensional representation

01:10:06.050 --> 01:10:10.210 align:middle line:84%
or a low-dimensional
embedding of that input data.

01:10:10.210 --> 01:10:15.270 align:middle line:84%
And so that might be a vector
of length 2048, for example,

01:10:15.270 --> 01:10:16.190 align:middle line:90%
or 1024.

01:10:16.190 --> 01:10:20.650 align:middle line:84%
Those are both common
embedding sizes.

01:10:20.650 --> 01:10:22.430 align:middle line:84%
But really, when we
talk about embedding,

01:10:22.430 --> 01:10:24.410 align:middle line:84%
it's essentially just saying
we've trained the neural network

01:10:24.410 --> 01:10:25.230 align:middle line:90%
already.

01:10:25.230 --> 01:10:29.450 align:middle line:84%
Now we're using it
to give us a low rank

01:10:29.450 --> 01:10:31.650 align:middle line:90%
representation of our data.

01:10:31.650 --> 01:10:34.170 align:middle line:90%
Does that make sense?

01:10:34.170 --> 01:10:34.860 align:middle line:90%
Yes?

01:10:34.860 --> 01:10:36.610 align:middle line:84%
AUDIENCE: I notice
that some of the slides

01:10:36.610 --> 01:10:39.170 align:middle line:84%
have used loss and
the others cost.

01:10:39.170 --> 01:10:40.585 align:middle line:90%
Are they interchangeable?

01:10:40.585 --> 01:10:41.210 align:middle line:90%
SARA BEERY: No.

01:10:41.210 --> 01:10:43.530 align:middle line:84%
This is me being a
little fast and loose.

01:10:43.530 --> 01:10:45.770 align:middle line:84%
The way we defined it
is that the loss is

01:10:45.770 --> 01:10:48.630 align:middle line:84%
like a function over
all of the data points.

01:10:48.630 --> 01:10:53.600 align:middle line:84%
And then the cost might
be the overall calculation

01:10:53.600 --> 01:10:56.020 align:middle line:84%
of that loss over, maybe, all
of your data or something.

01:10:56.020 --> 01:10:59.880 align:middle line:84%
But I think it's
a little semantic.

01:10:59.880 --> 01:11:01.937 align:middle line:90%
Yeah.

01:11:01.937 --> 01:11:03.520 align:middle line:84%
It's not the end of
the world to think

01:11:03.520 --> 01:11:05.100 align:middle line:84%
of them somewhat
interchangeably.

01:11:05.100 --> 01:11:08.600 align:middle line:84%
Essentially, both are you're
defining some mechanism

01:11:08.600 --> 01:11:12.880 align:middle line:90%
for evaluating optimality.

01:11:12.880 --> 01:11:13.703 align:middle line:90%
Yeah?

01:11:13.703 --> 01:11:15.120 align:middle line:84%
AUDIENCE: How easy
is it to switch

01:11:15.120 --> 01:11:18.560 align:middle line:84%
between optimizing parameters
versus optimizing inputs?

01:11:18.560 --> 01:11:22.500 align:middle line:84%
SARA BEERY: So lucky for us,
things like PyTorch exist.

01:11:22.500 --> 01:11:24.520 align:middle line:84%
And so actually this
is just a matter

01:11:24.520 --> 01:11:28.360 align:middle line:84%
of defining what you
want to update, basically

01:11:28.360 --> 01:11:30.317 align:middle line:84%
where you want to
freeze your gradients,

01:11:30.317 --> 01:11:32.400 align:middle line:84%
and where you don't want
to freeze your gradients.

01:11:32.400 --> 01:11:34.220 align:middle line:84%
And that's not
actually so hard to do.

01:11:34.220 --> 01:11:34.762 align:middle line:90%
AUDIENCE: OK.

01:11:34.762 --> 01:11:37.760 align:middle line:84%
If we were to do that switch,
could we use the prior work,

01:11:37.760 --> 01:11:42.102 align:middle line:84%
convert it in a way that we
can use the prior [INAUDIBLE]?

01:11:42.102 --> 01:11:43.060 align:middle line:90%
SARA BEERY: Yeah, yeah.

01:11:43.060 --> 01:11:45.700 align:middle line:84%
So this is where that
modularity comes in.

01:11:45.700 --> 01:11:46.700 align:middle line:90%
Sorry, guys.

01:11:46.700 --> 01:11:48.740 align:middle line:90%
If just if you could be quiet.

01:11:48.740 --> 01:11:50.740 align:middle line:84%
So this is where the
modularity comes in, right?

01:11:50.740 --> 01:11:59.040 align:middle line:84%
So we talked about here, you
can have these models where

01:11:59.040 --> 01:12:01.000 align:middle line:84%
some parts are
programmed by a human

01:12:01.000 --> 01:12:03.700 align:middle line:84%
and some parts are actually
being programmed by backprop.

01:12:03.700 --> 01:12:05.360 align:middle line:84%
But you could also think
of the ones that are quote,

01:12:05.360 --> 01:12:06.610 align:middle line:90%
unquote programmed by a human.

01:12:06.610 --> 01:12:08.320 align:middle line:84%
Those might have just
actually previously

01:12:08.320 --> 01:12:09.540 align:middle line:90%
been programmed by backprop.

01:12:09.540 --> 01:12:12.480 align:middle line:84%
And now you're taking those
weights and plugging them in.

01:12:12.480 --> 01:12:13.540 align:middle line:90%
Yeah?

01:12:13.540 --> 01:12:17.252 align:middle line:90%
AUDIENCE: [INAUDIBLE]

01:12:17.252 --> 01:12:17.960 align:middle line:90%
SARA BEERY: Yeah.

01:12:17.960 --> 01:12:21.700 align:middle line:84%
AUDIENCE: That was
through [INAUDIBLE]?

01:12:21.700 --> 01:12:24.240 align:middle line:84%
SARA BEERY: Sorry, if everyone
could just keep it quiet so

01:12:24.240 --> 01:12:25.490 align:middle line:90%
that I can hear the questions.

01:12:25.490 --> 01:12:26.300 align:middle line:90%
Thank you so much.

01:12:26.300 --> 01:12:26.800 align:middle line:90%
Yeah?

01:12:26.800 --> 01:12:29.132 align:middle line:84%
AUDIENCE: [INAUDIBLE]
attention block [INAUDIBLE]?

01:12:29.132 --> 01:12:31.045 align:middle line:90%


01:12:31.045 --> 01:12:32.920 align:middle line:84%
SARA BEERY: Is that what
I would call a what?

01:12:32.920 --> 01:12:33.960 align:middle line:90%
AUDIENCE: Attention block.

01:12:33.960 --> 01:12:35.252 align:middle line:90%
SARA BEERY: An attention block.

01:12:35.252 --> 01:12:43.220 align:middle line:84%
Oh, so this is a
cosine similarity.

01:12:43.220 --> 01:12:46.050 align:middle line:84%
And cosine similarity is
a component of attention,

01:12:46.050 --> 01:12:48.270 align:middle line:84%
but it doesn't actually
make up all of what we

01:12:48.270 --> 01:12:49.590 align:middle line:90%
would call an attention block.

01:12:49.590 --> 01:12:52.990 align:middle line:84%
So an attention block
generally defined

01:12:52.990 --> 01:12:54.750 align:middle line:84%
for a transformer or
something, will also

01:12:54.750 --> 01:12:58.070 align:middle line:84%
include the projections into
the shared space where you're

01:12:58.070 --> 01:13:00.775 align:middle line:84%
going to take that similarity,
as well as a projection out

01:13:00.775 --> 01:13:02.150 align:middle line:84%
into the space
where you actually

01:13:02.150 --> 01:13:06.230 align:middle line:84%
want to use the information
about what's valid.

01:13:06.230 --> 01:13:07.390 align:middle line:90%
Yeah?

01:13:07.390 --> 01:13:11.110 align:middle line:84%
AUDIENCE: I noticed
that [INAUDIBLE].

01:13:11.110 --> 01:13:13.598 align:middle line:84%
SARA BEERY: It'll be by the
end of the day today, sorry.

01:13:13.598 --> 01:13:15.890 align:middle line:84%
We were working on it on our
meeting right before this.

01:13:15.890 --> 01:13:17.950 align:middle line:84%
And I thought it might
be up, but not quite yet.

01:13:17.950 --> 01:13:19.970 align:middle line:90%
Yes?

01:13:19.970 --> 01:13:20.970 align:middle line:90%
Did you have a question?

01:13:20.970 --> 01:13:21.470 align:middle line:90%
No?

01:13:21.470 --> 01:13:22.010 align:middle line:90%
OK.

01:13:22.010 --> 01:13:22.510 align:middle line:90%
Yes?

01:13:22.510 --> 01:13:25.830 align:middle line:90%
AUDIENCE: [INAUDIBLE]

01:13:25.830 --> 01:13:29.318 align:middle line:90%


01:13:29.318 --> 01:13:31.610 align:middle line:84%
SARA BEERY: You're going to
need to be a little louder.

01:13:31.610 --> 01:13:34.310 align:middle line:90%
Sorry.

01:13:34.310 --> 01:13:36.530 align:middle line:84%
If we can just be
quiet as we leave.

01:13:36.530 --> 01:13:38.070 align:middle line:90%
Thank you so much.

01:13:38.070 --> 01:13:38.950 align:middle line:90%
All right.

01:13:38.950 --> 01:13:43.950 align:middle line:84%
AUDIENCE: So what's the
difference between the backdrop

01:13:43.950 --> 01:13:45.787 align:middle line:90%
for MLP and DAG?

01:13:45.787 --> 01:13:47.870 align:middle line:84%
Because they are both
computational graphs, right?

01:13:47.870 --> 01:13:49.870 align:middle line:84%
SARA BEERY: Yeah, so
there's no difference.

01:13:49.870 --> 01:13:52.510 align:middle line:84%
And this is where
the only two things

01:13:52.510 --> 01:13:57.550 align:middle line:84%
you need to take all the stuff
we talked about from backprop

01:13:57.550 --> 01:14:00.250 align:middle line:84%
through MLPs and translate it
to backprop through any DAG

01:14:00.250 --> 01:14:00.770 align:middle line:90%
are these.

01:14:00.770 --> 01:14:03.310 align:middle line:84%
You need a merging operation
and a branching operation.

01:14:03.310 --> 01:14:04.710 align:middle line:84%
So now, essentially,
they're just

01:14:04.710 --> 01:14:06.170 align:middle line:84%
changed with merging
and branching.

01:14:06.170 --> 01:14:07.650 align:middle line:90%
And that represents any DAG.

01:14:07.650 --> 01:14:08.350 align:middle line:90%
AUDIENCE: OK.

01:14:08.350 --> 01:14:09.570 align:middle line:90%
SARA BEERY: Yeah Yes?

01:14:09.570 --> 01:14:11.153 align:middle line:84%
AUDIENCE: This is
just a different way

01:14:11.153 --> 01:14:14.108 align:middle line:90%
of looking at backpropagation.

01:14:14.108 --> 01:14:16.150 align:middle line:84%
SARA BEERY: Yeah, so
essentially all we're saying

01:14:16.150 --> 01:14:18.710 align:middle line:84%
is that all the really
detailed examples

01:14:18.710 --> 01:14:20.270 align:middle line:84%
we showed of
backpropagation, those

01:14:20.270 --> 01:14:24.870 align:middle line:84%
were all for a really
simple chain-based DAG.

01:14:24.870 --> 01:14:27.670 align:middle line:84%
It's just you have a component,
and it goes, and a component,

01:14:27.670 --> 01:14:30.187 align:middle line:84%
and it goes, so it
was all a chain.

01:14:30.187 --> 01:14:32.270 align:middle line:84%
And now what we're saying
is that you can actually

01:14:32.270 --> 01:14:34.150 align:middle line:84%
take everything
we talked about it

01:14:34.150 --> 01:14:38.910 align:middle line:84%
and apply it to any
structure of a DAG.

01:14:38.910 --> 01:14:42.420 align:middle line:84%
Because now it's just chains,
where sometimes they merge,

01:14:42.420 --> 01:14:43.940 align:middle line:90%
and sometimes they branch.

01:14:43.940 --> 01:14:45.660 align:middle line:84%
But there's really
simple operations

01:14:45.660 --> 01:14:47.900 align:middle line:84%
that we can use to pass the
gradients through merges

01:14:47.900 --> 01:14:49.060 align:middle line:90%
and branches.

01:14:49.060 --> 01:14:52.780 align:middle line:84%
And so that's how you expand
that definition of backprop

01:14:52.780 --> 01:14:54.680 align:middle line:90%
to any DAG.

01:14:54.680 --> 01:14:55.180 align:middle line:90%
Yeah?

01:14:55.180 --> 01:14:59.220 align:middle line:90%
AUDIENCE: [INAUDIBLE]

01:14:59.220 --> 01:15:00.220 align:middle line:90%
SARA BEERY: Sorry, what?

01:15:00.220 --> 01:15:03.060 align:middle line:90%
AUDIENCE: [INAUDIBLE]

01:15:03.060 --> 01:15:05.890 align:middle line:90%


01:15:05.890 --> 01:15:06.640 align:middle line:90%
SARA BEERY: Sorry.

01:15:06.640 --> 01:15:08.400 align:middle line:84%
I'm just I'm having
trouble hearing you.

01:15:08.400 --> 01:15:11.083 align:middle line:84%
AUDIENCE: What is the
[INAUDIBLE] for the gradients?

01:15:11.083 --> 01:15:11.875 align:middle line:90%
SARA BEERY: Mm-hmm.

01:15:11.875 --> 01:15:13.172 align:middle line:90%
AUDIENCE: [INAUDIBLE]

01:15:13.172 --> 01:15:13.880 align:middle line:90%
SARA BEERY: Yeah.

01:15:13.880 --> 01:15:15.372 align:middle line:90%
AUDIENCE: [INAUDIBLE].

01:15:15.372 --> 01:15:19.340 align:middle line:90%


01:15:19.340 --> 01:15:20.393 align:middle line:90%
Why is it a sum?

01:15:20.393 --> 01:15:21.560 align:middle line:90%
SARA BEERY: Why is it a sum?

01:15:21.560 --> 01:15:24.130 align:middle line:90%


01:15:24.130 --> 01:15:30.620 align:middle line:84%
Because the change
in this, as it's

01:15:30.620 --> 01:15:33.500 align:middle line:84%
coming in from that
and that, you're

01:15:33.500 --> 01:15:35.240 align:middle line:90%
able to just aggregate it.

01:15:35.240 --> 01:15:43.420 align:middle line:84%
And the reason for that
is, essentially, I'm

01:15:43.420 --> 01:15:45.980 align:middle line:84%
not going to talk this through
really simply right now.

01:15:45.980 --> 01:15:49.260 align:middle line:84%
Come to office hours,
and we can talk about it.

01:15:49.260 --> 01:15:50.260 align:middle line:90%
Yes?

01:15:50.260 --> 01:15:53.380 align:middle line:84%
AUDIENCE: So I know that
PyTorch builds its computation

01:15:53.380 --> 01:15:54.773 align:middle line:90%
graph dynamically.

01:15:54.773 --> 01:15:56.940 align:middle line:84%
For instance, like imagine
that in your forward pass

01:15:56.940 --> 01:15:58.500 align:middle line:84%
you have a conditional,
which is like

01:15:58.500 --> 01:16:01.500 align:middle line:84%
if the magnitude of the
gradient exceeds whatever,

01:16:01.500 --> 01:16:04.417 align:middle line:84%
then I want to apply
a CLIP operation.

01:16:04.417 --> 01:16:06.500 align:middle line:84%
I have a slightly different
use case for research.

01:16:06.500 --> 01:16:09.060 align:middle line:84%
Do you know how
it handles if you

01:16:09.060 --> 01:16:12.480 align:middle line:84%
apply like a non-torch
operation to a tensor?

01:16:12.480 --> 01:16:15.180 align:middle line:84%
If you apply some
of UDF which is just

01:16:15.180 --> 01:16:19.140 align:middle line:84%
written using normal
Python arguments,

01:16:19.140 --> 01:16:23.500 align:middle line:84%
does it basically just suddenly
say in the computation graph,

01:16:23.500 --> 01:16:26.800 align:middle line:84%
even if I use torch operators
before that node in the graph,

01:16:26.800 --> 01:16:28.640 align:middle line:90%
it is now not optimizable.

01:16:28.640 --> 01:16:31.100 align:middle line:84%
And I'm just going to treat
that basically as the way

01:16:31.100 --> 01:16:33.340 align:middle line:90%
I would treat a data input?

01:16:33.340 --> 01:16:36.000 align:middle line:84%
SARA BEERY: So there might
be hacks around this.

01:16:36.000 --> 01:16:38.810 align:middle line:84%
But generally, if
you want to have

01:16:38.810 --> 01:16:41.910 align:middle line:84%
any operation within a neural
network defined in PyTorch,

01:16:41.910 --> 01:16:45.650 align:middle line:84%
you have to define it as a torch
operation, which is basically

01:16:45.650 --> 01:16:49.050 align:middle line:84%
you are required to define a
gradient for that operation.

01:16:49.050 --> 01:16:51.750 align:middle line:84%
So you have to assume that
everything's differentiable.

01:16:51.750 --> 01:16:53.208 align:middle line:84%
And this is where
things get nasty.

01:16:53.208 --> 01:16:55.417 align:middle line:84%
Because if something doesn't
have a simple gradient--

01:16:55.417 --> 01:16:57.570 align:middle line:84%
so there's a lot of built
in stuff in PyTorch that

01:16:57.570 --> 01:16:59.570 align:middle line:84%
handles what those
approximations to those

01:16:59.570 --> 01:17:00.870 align:middle line:90%
gradients would look like.

01:17:00.870 --> 01:17:03.578 align:middle line:84%
But if you have something that's
really complicated, for example,

01:17:03.578 --> 01:17:06.110 align:middle line:84%
like an occupancy model of
species, which is something

01:17:06.110 --> 01:17:07.610 align:middle line:84%
that we work with,
then you actually

01:17:07.610 --> 01:17:11.030 align:middle line:84%
might have to define that
as a torch operation.

01:17:11.030 --> 01:17:11.890 align:middle line:90%
It also depends.

01:17:11.890 --> 01:17:16.050 align:middle line:84%
Sometimes there are choices
you can make based on

01:17:16.050 --> 01:17:18.190 align:middle line:84%
whether or not that makes
sense in the structure.

01:17:18.190 --> 01:17:20.328 align:middle line:84%
You might not actually
want to learn all the way

01:17:20.328 --> 01:17:22.370 align:middle line:84%
through that thing, if it
becomes computationally

01:17:22.370 --> 01:17:23.953 align:middle line:84%
intractable to
calculate the gradient.

01:17:23.953 --> 01:17:27.930 align:middle line:84%
AUDIENCE: So I guess that's
my question is, will PyTorch--

01:17:27.930 --> 01:17:32.250 align:middle line:84%
say you didn't define the
gradient for that operation

01:17:32.250 --> 01:17:34.970 align:middle line:84%
or how to compute
it, can PyTorch

01:17:34.970 --> 01:17:37.970 align:middle line:84%
still optimize the
rest of the model

01:17:37.970 --> 01:17:42.450 align:middle line:84%
and just treat what comes in
from the output of your UDF

01:17:42.450 --> 01:17:47.210 align:middle line:84%
as basically being like,
you had input some x sub i,

01:17:47.210 --> 01:17:50.690 align:middle line:90%
x of basically a data vector.

01:17:50.690 --> 01:17:53.487 align:middle line:84%
SARA BEERY: I might just
define that separately.

01:17:53.487 --> 01:17:55.570 align:middle line:84%
I might define that as
part of your pre-processing

01:17:55.570 --> 01:17:56.310 align:middle line:90%
for your data.

01:17:56.310 --> 01:17:59.730 align:middle line:84%
So the pre-processing
steps don't necessarily

01:17:59.730 --> 01:18:02.790 align:middle line:84%
need to be differentiable,
things like data augmentation.

01:18:02.790 --> 01:18:06.370 align:middle line:84%
So you could put that UDF in
your pre-processing, your data

01:18:06.370 --> 01:18:07.490 align:middle line:90%
loader.

01:18:07.490 --> 01:18:09.505 align:middle line:84%
And then whatever
comes out of it

01:18:09.505 --> 01:18:12.130 align:middle line:84%
would then be the data that you
were going to optimize towards.

01:18:12.130 --> 01:18:15.570 align:middle line:84%
That's probably the
simplest thing to do.

01:18:15.570 --> 01:18:18.290 align:middle line:84%
But again,
computational complexity

01:18:18.290 --> 01:18:20.430 align:middle line:84%
can get really nasty
with some of this stuff,

01:18:20.430 --> 01:18:21.712 align:middle line:90%
so maybe that's not too bad.

01:18:21.712 --> 01:18:23.170 align:middle line:84%
But I've definitely
run into issues

01:18:23.170 --> 01:18:28.050 align:middle line:84%
where I've tried to build simple
statistical models into machine

01:18:28.050 --> 01:18:31.510 align:middle line:84%
learning and with varying
levels of success.

01:18:31.510 --> 01:18:35.100 align:middle line:84%
And often it's just like, OK,
it's technically possible,

01:18:35.100 --> 01:18:37.760 align:middle line:84%
but it's intractable
to actually learn.

01:18:37.760 --> 01:18:38.760 align:middle line:90%
Yeah?

01:18:38.760 --> 01:18:42.240 align:middle line:84%
AUDIENCE: On this example,
for the superscript a and b,

01:18:42.240 --> 01:18:44.960 align:middle line:90%
what do they really represent?

01:18:44.960 --> 01:18:48.000 align:middle line:84%
SARA BEERY: Oh, it's intended
it's intended to be--

01:18:48.000 --> 01:18:50.600 align:middle line:90%
it's intended to be open ended.

01:18:50.600 --> 01:18:53.220 align:middle line:84%
So they could literally just
be a duplicate of copy of both.

01:18:53.220 --> 01:18:57.400 align:middle line:84%
Or it could be some splitting
of the embedding vector,

01:18:57.400 --> 01:18:59.160 align:middle line:90%
or et cetera, et cetera.

01:18:59.160 --> 01:19:01.600 align:middle line:90%
So it's any branching operation.

01:19:01.600 --> 01:19:02.400 align:middle line:90%
Yeah.

01:19:02.400 --> 01:19:05.200 align:middle line:84%
Any differentiable
branching operation.

01:19:05.200 --> 01:19:06.453 align:middle line:90%
OK, one last question.

01:19:06.453 --> 01:19:08.120 align:middle line:84%
AUDIENCE: I have a
[INAUDIBLE] question.

01:19:08.120 --> 01:19:09.500 align:middle line:90%
SARA BEERY: Yeah.

01:19:09.500 --> 01:19:12.683 align:middle line:84%
AUDIENCE: The problem set-- was
that supposed to be [INAUDIBLE]?

01:19:12.683 --> 01:19:14.100 align:middle line:84%
SARA BEERY: It
will be there soon.

01:19:14.100 --> 01:19:16.070 align:middle line:84%
I thought it was up
already, but it's not.

01:19:16.070 --> 01:19:16.820 align:middle line:90%
AUDIENCE: 3:00 PM.

01:19:16.820 --> 01:19:17.420 align:middle line:90%
SARA BEERY: 3:00 PM.

01:19:17.420 --> 01:19:17.920 align:middle line:90%
AUDIENCE: OK.

01:19:17.920 --> 01:19:18.420 align:middle line:90%
[LAUGHS]

01:19:18.420 --> 01:19:18.863 align:middle line:90%


01:19:18.863 --> 01:19:19.780 align:middle line:90%
SARA BEERY: All right.

01:19:19.780 --> 01:19:21.000 align:middle line:90%
Thank you, guys, so much.

01:19:21.000 --> 01:19:22.910 align:middle line:90%
See you soon.

01:19:22.910 --> 01:19:34.000 align:middle line:90%