WEBVTT

00:00:00.000 --> 00:00:01.961 align:middle line:90%
[SQUEAKING]

00:00:01.961 --> 00:00:03.909 align:middle line:90%
[RUSTLING]

00:00:03.909 --> 00:00:07.318 align:middle line:90%
[CLICKING]

00:00:07.318 --> 00:00:09.812 align:middle line:90%


00:00:09.812 --> 00:00:11.520 align:middle line:84%
SPEAKER: Thanks for
having me here today,

00:00:11.520 --> 00:00:17.440 align:middle line:84%
and also thanks for
accepting my request

00:00:17.440 --> 00:00:20.520 align:middle line:90%
for presenting on this topic.

00:00:20.520 --> 00:00:24.640 align:middle line:84%
So just for some context,
shamelessly reintroducing

00:00:24.640 --> 00:00:28.970 align:middle line:84%
myself, I work as a
research engineer at Hugging

00:00:28.970 --> 00:00:35.000 align:middle line:84%
Face I try to focus on different
facets of diffusion models

00:00:35.000 --> 00:00:39.040 align:middle line:90%
for image and video generation.

00:00:39.040 --> 00:00:40.920 align:middle line:84%
Part of my time
at Hugging Face is

00:00:40.920 --> 00:00:45.420 align:middle line:84%
focused on maintaining and
growing the diffusers library,

00:00:45.420 --> 00:00:48.640 align:middle line:84%
but this talk is not going to
be about the diffusers library,

00:00:48.640 --> 00:00:50.600 align:middle line:84%
and the rest of my
time at Hugging Face

00:00:50.600 --> 00:00:54.040 align:middle line:84%
is also spent around
researching diffusion models

00:00:54.040 --> 00:00:56.640 align:middle line:84%
from many different
perspectives.

00:00:56.640 --> 00:01:01.060 align:middle line:84%
And I'm going to be covering
some of that in this talk.

00:01:01.060 --> 00:01:03.200 align:middle line:84%
As this talk
suggests, it's going

00:01:03.200 --> 00:01:08.400 align:middle line:84%
to be about optimizing
the full stack for image

00:01:08.400 --> 00:01:10.760 align:middle line:90%
and video generation models.

00:01:10.760 --> 00:01:15.800 align:middle line:84%
In particular, we will
see how the optimization

00:01:15.800 --> 00:01:19.880 align:middle line:84%
aspects of these models are
a little nontrivial and also

00:01:19.880 --> 00:01:25.480 align:middle line:84%
kind of atypical in
nature, and also some

00:01:25.480 --> 00:01:28.400 align:middle line:84%
perspectives into how
the whole optimization

00:01:28.400 --> 00:01:34.200 align:middle line:84%
landscape can transcend
beyond latency and speed.

00:01:34.200 --> 00:01:38.320 align:middle line:90%
So with that, I'll start.

00:01:38.320 --> 00:01:43.560 align:middle line:84%
I'll try to also give you a
very hand-wavy introduction

00:01:43.560 --> 00:01:47.040 align:middle line:84%
to diffusion models
because diffusion model

00:01:47.040 --> 00:01:50.400 align:middle line:84%
itself will take another
lecture to cover.

00:01:50.400 --> 00:01:52.960 align:middle line:84%
So I'll try to give
you just enough context

00:01:52.960 --> 00:01:56.360 align:middle line:84%
so that we can step through
the rest of the discussion

00:01:56.360 --> 00:02:01.200 align:middle line:84%
for this talk, and hopefully,
by the end of this talk,

00:02:01.200 --> 00:02:05.360 align:middle line:84%
there will be time for some
Q&A. In case there isn't, you

00:02:05.360 --> 00:02:07.700 align:middle line:84%
can always reach
out to me via email,

00:02:07.700 --> 00:02:10.020 align:middle line:84%
and we can always
discuss things offline.

00:02:10.020 --> 00:02:12.560 align:middle line:90%


00:02:12.560 --> 00:02:16.560 align:middle line:84%
I want to get the excitement
rolling by giving you

00:02:16.560 --> 00:02:22.520 align:middle line:84%
some outputs from some of the
recent text-to-image models,

00:02:22.520 --> 00:02:23.980 align:middle line:90%
but this is from PixArt-Alpha.

00:02:23.980 --> 00:02:26.680 align:middle line:84%
This is from back in
the days in 2024--

00:02:26.680 --> 00:02:29.600 align:middle line:90%
I think 2023, not even 2024.

00:02:29.600 --> 00:02:35.200 align:middle line:84%
So this is quite old, but you
can see the quality is not bad.

00:02:35.200 --> 00:02:41.780 align:middle line:84%
This is from DALL-E 3, OpenAI,
and this is from 2024 September

00:02:41.780 --> 00:02:42.780 align:middle line:90%
if I remember correctly.

00:02:42.780 --> 00:02:44.000 align:middle line:90%
This is from Flux.

00:02:44.000 --> 00:02:48.200 align:middle line:84%
And this example is by
far my most favorite

00:02:48.200 --> 00:02:52.920 align:middle line:84%
because even imagining a
tiny astronaut, let alone

00:02:52.920 --> 00:02:57.420 align:middle line:84%
trying to imagine it hatching
from an egg, is wild.

00:02:57.420 --> 00:02:58.500 align:middle line:90%
It's simply wild.

00:02:58.500 --> 00:03:01.600 align:middle line:84%
But somehow, these models
are expressive enough

00:03:01.600 --> 00:03:04.340 align:middle line:84%
to be able to come up with
something as real as this.

00:03:04.340 --> 00:03:07.440 align:middle line:90%


00:03:07.440 --> 00:03:10.880 align:middle line:84%
There's an ongoing emergence
around text-to-video models

00:03:10.880 --> 00:03:15.800 align:middle line:84%
and the whole moniker
around world models as well.

00:03:15.800 --> 00:03:17.680 align:middle line:90%
This is one of those.

00:03:17.680 --> 00:03:21.000 align:middle line:84%
And then we have
got this cool cat.

00:03:21.000 --> 00:03:25.240 align:middle line:84%
And then you have got very
cinematic light landscape

00:03:25.240 --> 00:03:26.280 align:middle line:90%
frames.

00:03:26.280 --> 00:03:29.960 align:middle line:84%
All of these are open
models except for DALL-E 3

00:03:29.960 --> 00:03:34.520 align:middle line:84%
that I've shown in
the previous slide.

00:03:34.520 --> 00:03:40.120 align:middle line:84%
As a cautionary note, I'm going
to be interchangeably using

00:03:40.120 --> 00:03:43.680 align:middle line:84%
diffusion and flow in this
talk, which also means

00:03:43.680 --> 00:03:46.560 align:middle line:84%
that the points
and the discussions

00:03:46.560 --> 00:03:52.760 align:middle line:84%
that we'll do in this talk, they
will apply to both diffusion

00:03:52.760 --> 00:03:54.960 align:middle line:90%
and flow models.

00:03:54.960 --> 00:03:58.960 align:middle line:84%
Flow model-- I like
to explain them

00:03:58.960 --> 00:04:03.120 align:middle line:84%
is a generalization
of diffusion models.

00:04:03.120 --> 00:04:05.080 align:middle line:84%
But they are not
exactly the same.

00:04:05.080 --> 00:04:08.800 align:middle line:84%
But for the purposes
of this talk,

00:04:08.800 --> 00:04:12.680 align:middle line:90%
I'll interchangeably use them.

00:04:12.680 --> 00:04:15.400 align:middle line:84%
Now let me try to give
you some context as to how

00:04:15.400 --> 00:04:17.160 align:middle line:90%
diffusion models work.

00:04:17.160 --> 00:04:20.200 align:middle line:84%
I like to think of diffusion
models as the following.

00:04:20.200 --> 00:04:25.200 align:middle line:84%
What happens when we start
from a random noise drawn

00:04:25.200 --> 00:04:26.560 align:middle line:90%
from a Gaussian?

00:04:26.560 --> 00:04:32.260 align:middle line:84%
And we then try to slowly
denoise it over a period of time

00:04:32.260 --> 00:04:34.260 align:middle line:84%
unless we get
something realistic.

00:04:34.260 --> 00:04:35.400 align:middle line:90%
It can be an image.

00:04:35.400 --> 00:04:36.780 align:middle line:90%
It can be a video.

00:04:36.780 --> 00:04:40.160 align:middle line:84%
Or it can be even a
piece of audio as well.

00:04:40.160 --> 00:04:46.080 align:middle line:84%
But the crucial point
is we are starting

00:04:46.080 --> 00:04:48.880 align:middle line:84%
from a pure random
noise, and then we

00:04:48.880 --> 00:04:52.600 align:middle line:84%
are iterating over a period
of time, denoising it,

00:04:52.600 --> 00:04:56.760 align:middle line:84%
so that it becomes a clean and
realistic representation of what

00:04:56.760 --> 00:05:01.760 align:middle line:84%
we want, so it can be considered
as some of an iterative

00:05:01.760 --> 00:05:03.600 align:middle line:90%
denoising process.

00:05:03.600 --> 00:05:06.060 align:middle line:84%
And we can also condition
this denoising process.

00:05:06.060 --> 00:05:09.360 align:middle line:84%
When we try to condition
with text descriptions,

00:05:09.360 --> 00:05:13.080 align:middle line:84%
we can do tasks
like text-to-image

00:05:13.080 --> 00:05:14.460 align:middle line:90%
like we are seeing here.

00:05:14.460 --> 00:05:18.320 align:middle line:90%


00:05:18.320 --> 00:05:20.880 align:middle line:84%
Now the thing is,
diffusion models

00:05:20.880 --> 00:05:26.760 align:middle line:84%
can have different variants,
like when we are operating

00:05:26.760 --> 00:05:32.360 align:middle line:84%
directly on the pixel
space, that family of model

00:05:32.360 --> 00:05:36.380 align:middle line:84%
is typically referred to as
pixel space diffusion model.

00:05:36.380 --> 00:05:38.120 align:middle line:84%
But pixel space
diffusion is pretty

00:05:38.120 --> 00:05:43.720 align:middle line:84%
intensive from both memory
and compute standpoints, which

00:05:43.720 --> 00:05:47.960 align:middle line:84%
is exactly why operating
on the latent space

00:05:47.960 --> 00:05:51.130 align:middle line:84%
is almost the de facto
for diffusion models.

00:05:51.130 --> 00:05:54.510 align:middle line:84%
And in this diagram, we
can get a sense of how

00:05:54.510 --> 00:05:56.470 align:middle line:90%
latent space diffusion works.

00:05:56.470 --> 00:06:01.150 align:middle line:84%
Now we need to have some
of an encoder and a decoder

00:06:01.150 --> 00:06:05.870 align:middle line:84%
in order to be able to encode
the pixel representation

00:06:05.870 --> 00:06:08.590 align:middle line:84%
of an image into its
latent representation.

00:06:08.590 --> 00:06:12.190 align:middle line:84%
And then when we are
converting that latent space

00:06:12.190 --> 00:06:15.710 align:middle line:84%
into the pixel space, we
need to have some decoder.

00:06:15.710 --> 00:06:19.770 align:middle line:84%
Or typically, for the case of
latent space diffusion models,

00:06:19.770 --> 00:06:25.310 align:middle line:84%
we usually use a VAE, which
has an encoder variant as well

00:06:25.310 --> 00:06:27.230 align:middle line:90%
as a decoder variant.

00:06:27.230 --> 00:06:32.870 align:middle line:84%
So yeah, just note that
this talk is mostly

00:06:32.870 --> 00:06:35.910 align:middle line:84%
going to be about latent
space diffusion models.

00:06:35.910 --> 00:06:38.890 align:middle line:84%
But the approaches are
fairly general enough.

00:06:38.890 --> 00:06:40.770 align:middle line:90%
They will also apply equally.

00:06:40.770 --> 00:06:44.350 align:middle line:84%
They should apply equally to
pixel space diffusion models

00:06:44.350 --> 00:06:46.950 align:middle line:90%
as well.

00:06:46.950 --> 00:06:51.590 align:middle line:84%
Now, let's try to also see
how different components

00:06:51.590 --> 00:06:56.110 align:middle line:84%
within a diffusion model are
connected with one another

00:06:56.110 --> 00:06:59.430 align:middle line:84%
because unlike language models
or visual language models,

00:06:59.430 --> 00:07:03.510 align:middle line:84%
diffusion, any modern state
of that diffusion model

00:07:03.510 --> 00:07:05.110 align:middle line:90%
is not a single model.

00:07:05.110 --> 00:07:08.790 align:middle line:84%
That's a very important
distinction to be aware of.

00:07:08.790 --> 00:07:13.030 align:middle line:84%
Now let's say we want to
take this text prompt,

00:07:13.030 --> 00:07:15.830 align:middle line:84%
and we want to generate
a video out of it.

00:07:15.830 --> 00:07:19.150 align:middle line:84%
Now let's see what are the
typical components that

00:07:19.150 --> 00:07:23.190 align:middle line:84%
are involved and how
they are connected.

00:07:23.190 --> 00:07:26.070 align:middle line:84%
Now, to be able to have
salient representations

00:07:26.070 --> 00:07:28.230 align:middle line:84%
of the textual
description, we need

00:07:28.230 --> 00:07:32.550 align:middle line:84%
to have text encoder,
which is fairly intuitive.

00:07:32.550 --> 00:07:35.870 align:middle line:84%
And then once we have
the text embeddings out,

00:07:35.870 --> 00:07:40.390 align:middle line:84%
we start the-- this is the
inference workflow, by the way.

00:07:40.390 --> 00:07:42.430 align:middle line:84%
Now, with the text
embeddings out,

00:07:42.430 --> 00:07:48.470 align:middle line:84%
we start off with a pure random
noise as I was mentioning.

00:07:48.470 --> 00:07:52.270 align:middle line:84%
These are our noisy latents as
we are operating on the latent

00:07:52.270 --> 00:07:55.750 align:middle line:84%
space, and we also
need to have access

00:07:55.750 --> 00:07:58.530 align:middle line:84%
to a component we
term as scheduler,

00:07:58.530 --> 00:08:00.550 align:middle line:84%
which is a nonparametric
component, which

00:08:00.550 --> 00:08:03.110 align:middle line:84%
takes care of--
which basically makes

00:08:03.110 --> 00:08:08.190 align:middle line:84%
the model aware of the position
it is in the entire denoising

00:08:08.190 --> 00:08:10.910 align:middle line:90%
trajectory.

00:08:10.910 --> 00:08:15.270 align:middle line:84%
And this is the beast that
we will try to optimize.

00:08:15.270 --> 00:08:17.030 align:middle line:90%
But more on that later.

00:08:17.030 --> 00:08:20.110 align:middle line:84%
It can be a typical
unit-like architecture

00:08:20.110 --> 00:08:23.350 align:middle line:84%
or a transformer-like
architecture.

00:08:23.350 --> 00:08:28.070 align:middle line:84%
And broadly, this is referred to
as the diffusion network, which

00:08:28.070 --> 00:08:31.390 align:middle line:84%
is conditioned on the time
step, like the position

00:08:31.390 --> 00:08:34.950 align:middle line:84%
in the denoising
iterations that it is in,

00:08:34.950 --> 00:08:36.650 align:middle line:84%
the text embeddings
in this case,

00:08:36.650 --> 00:08:38.750 align:middle line:90%
and also the noisy latents.

00:08:38.750 --> 00:08:43.750 align:middle line:84%
And then we invoke it over a
period of time, hence the loop.

00:08:43.750 --> 00:08:46.950 align:middle line:84%
Now, once we get our refined
legends out the diffusion

00:08:46.950 --> 00:08:50.110 align:middle line:84%
network, it's passed
off to a decoder,

00:08:50.110 --> 00:08:52.150 align:middle line:90%
and we finally get our frames.

00:08:52.150 --> 00:08:55.630 align:middle line:84%
In case of an image, it will
just be a single frame or just

00:08:55.630 --> 00:08:58.990 align:middle line:90%
the single final image.

00:08:58.990 --> 00:09:03.750 align:middle line:84%
Now, now that we have some
very basic understanding

00:09:03.750 --> 00:09:06.110 align:middle line:84%
of the different components
that are involved

00:09:06.110 --> 00:09:10.490 align:middle line:84%
in a standard text-to-image or
text-to-video diffusion model,

00:09:10.490 --> 00:09:13.930 align:middle line:84%
and how those components are
connected with one another,

00:09:13.930 --> 00:09:17.710 align:middle line:84%
how they draw the chronology in
the whole inference workflow.

00:09:17.710 --> 00:09:21.510 align:middle line:84%
I'll try to now
motivate why we might

00:09:21.510 --> 00:09:25.510 align:middle line:84%
have to optimize them to
be able to make anything

00:09:25.510 --> 00:09:28.950 align:middle line:90%
useful out of them.

00:09:28.950 --> 00:09:30.910 align:middle line:84%
Now let's try to
look at the memory

00:09:30.910 --> 00:09:33.750 align:middle line:84%
footprint of the individual
model-level components

00:09:33.750 --> 00:09:38.430 align:middle line:84%
in a fairly state of the art
and recent model called Flux.

00:09:38.430 --> 00:09:42.110 align:middle line:84%
It uses two text
encoders, which is also

00:09:42.110 --> 00:09:45.270 align:middle line:84%
kind of common in the literature
of text-to-image diffusion

00:09:45.270 --> 00:09:48.150 align:middle line:90%
models.

00:09:48.150 --> 00:09:51.170 align:middle line:84%
For Flux, it uses
two text encoders.

00:09:51.170 --> 00:09:55.750 align:middle line:84%
It has got a T5-XXL, and
it has got a clip large.

00:09:55.750 --> 00:10:01.110 align:middle line:84%
If you are interested in knowing
why we need different text

00:10:01.110 --> 00:10:03.530 align:middle line:84%
encoders and why that
might be beneficial,

00:10:03.530 --> 00:10:05.110 align:middle line:90%
we can talk about it later.

00:10:05.110 --> 00:10:07.650 align:middle line:84%
But in the interest of
time, I'll just keep going.

00:10:07.650 --> 00:10:11.397 align:middle line:84%
But that's an interesting
question to ask.

00:10:11.397 --> 00:10:13.230 align:middle line:84%
And then you have got
the transformer, which

00:10:13.230 --> 00:10:15.110 align:middle line:90%
is like the diffusion network.

00:10:15.110 --> 00:10:18.790 align:middle line:84%
And as we can see, it has
got the highest footprint.

00:10:18.790 --> 00:10:21.630 align:middle line:84%
And you have got
the decoder, which

00:10:21.630 --> 00:10:24.430 align:middle line:84%
will be responsible for
taking the defined latents out

00:10:24.430 --> 00:10:26.750 align:middle line:84%
of the transformer and
then decoding it back

00:10:26.750 --> 00:10:28.690 align:middle line:90%
to the pixel space.

00:10:28.690 --> 00:10:31.430 align:middle line:90%


00:10:31.430 --> 00:10:33.910 align:middle line:84%
And as I was mentioning,
the transformer

00:10:33.910 --> 00:10:39.030 align:middle line:84%
or the diffusion network here is
the most compute intensive unit.

00:10:39.030 --> 00:10:42.470 align:middle line:84%
And it's compute bound,
unlike language models.

00:10:42.470 --> 00:10:48.030 align:middle line:84%
Hence, consequently, the
optimization literature

00:10:48.030 --> 00:10:51.470 align:middle line:84%
from the language
modeling world might not

00:10:51.470 --> 00:10:57.110 align:middle line:84%
carry over to the diffusion
world very gracefully.

00:10:57.110 --> 00:11:00.910 align:middle line:84%
Now, even when using
the brain float16, which

00:11:00.910 --> 00:11:03.670 align:middle line:84%
is very common in
the diffusion world,

00:11:03.670 --> 00:11:08.470 align:middle line:84%
it takes about 34 gigs
to generate a 1024

00:11:08.470 --> 00:11:11.610 align:middle line:84%
by 1024 resolution
image, which is standard,

00:11:11.610 --> 00:11:17.310 align:middle line:84%
and it takes about 7 seconds on
a fairly beefy GPU such as H100,

00:11:17.310 --> 00:11:18.990 align:middle line:90%
without any optimizations.

00:11:18.990 --> 00:11:23.910 align:middle line:84%
So I hope you we can see
how these numbers are

00:11:23.910 --> 00:11:27.310 align:middle line:84%
kind of staggeringly
high because if we have

00:11:27.310 --> 00:11:31.630 align:middle line:84%
to wait for about 7 seconds
to generate one single image,

00:11:31.630 --> 00:11:34.310 align:middle line:90%
it's not good.

00:11:34.310 --> 00:11:36.930 align:middle line:84%
And then videos can be
even worse, like a fight.

00:11:36.930 --> 00:11:42.190 align:middle line:84%
So just to give you some
perspective, a 5 second, 16 FPS,

00:11:42.190 --> 00:11:48.670 align:middle line:84%
720p video can take about 30
minutes to generate fairly

00:11:48.670 --> 00:11:53.910 align:middle line:84%
state-of-the-art open
video generation model.

00:11:53.910 --> 00:11:55.930 align:middle line:90%
Now these numbers are heavy.

00:11:55.930 --> 00:11:59.250 align:middle line:84%
These numbers should
concern us in case

00:11:59.250 --> 00:12:01.870 align:middle line:84%
we are interested in
operationalizing some

00:12:01.870 --> 00:12:04.830 align:middle line:84%
of these models because
long-running tasks

00:12:04.830 --> 00:12:07.330 align:middle line:84%
are bad for user-facing
applications.

00:12:07.330 --> 00:12:12.710 align:middle line:84%
It hinders interactions because
if the interactivity aspect

00:12:12.710 --> 00:12:17.250 align:middle line:84%
of creative applications
is hindered,

00:12:17.250 --> 00:12:21.390 align:middle line:84%
it's not good for your users,
and it's also power inefficient.

00:12:21.390 --> 00:12:25.310 align:middle line:84%
And it also slows
down the overall rate

00:12:25.310 --> 00:12:27.710 align:middle line:84%
at which you would want
to iterate and make

00:12:27.710 --> 00:12:29.330 align:middle line:84%
some improvements
to these models.

00:12:29.330 --> 00:12:33.150 align:middle line:84%
So it's kind of a bad experience
for both the developers

00:12:33.150 --> 00:12:37.470 align:middle line:84%
as well as the users
of these models.

00:12:37.470 --> 00:12:43.910 align:middle line:84%
Now, as I was hinting at
the iterative transformer.

00:12:43.910 --> 00:12:46.570 align:middle line:84%
It's at the root
of all evil here,

00:12:46.570 --> 00:12:49.190 align:middle line:84%
basically, because
it's both compute bound

00:12:49.190 --> 00:12:55.590 align:middle line:84%
and same time very memory hungry
now and a natural question

00:12:55.590 --> 00:12:58.710 align:middle line:84%
because the transformer
takes the most

00:12:58.710 --> 00:13:01.910 align:middle line:84%
amount of time in the
whole generation workflow.

00:13:01.910 --> 00:13:04.750 align:middle line:84%
A natural question
here to ask would be

00:13:04.750 --> 00:13:07.230 align:middle line:90%
do we just optimize for speed?

00:13:07.230 --> 00:13:10.430 align:middle line:84%
But what happens
when you also try

00:13:10.430 --> 00:13:14.790 align:middle line:84%
to motivate the
entire optimization

00:13:14.790 --> 00:13:17.470 align:middle line:84%
landscape with an
application, the application

00:13:17.470 --> 00:13:21.310 align:middle line:84%
in which this diffusion
network is going to get used?

00:13:21.310 --> 00:13:26.410 align:middle line:84%
If we consider things
from that perspective.

00:13:26.410 --> 00:13:28.830 align:middle line:84%
There will be other
factors that will influence

00:13:28.830 --> 00:13:32.310 align:middle line:84%
the whole process, like the use
case in which this network is

00:13:32.310 --> 00:13:34.930 align:middle line:84%
used, the level of
user interaction,

00:13:34.930 --> 00:13:37.230 align:middle line:84%
the rapidity of the
user interaction,

00:13:37.230 --> 00:13:39.410 align:middle line:84%
the kind of throughput
we will have to meet,

00:13:39.410 --> 00:13:43.770 align:middle line:84%
and also the deployment
hardware that's available to us.

00:13:43.770 --> 00:13:46.830 align:middle line:90%


00:13:46.830 --> 00:13:50.510 align:middle line:84%
So I think I was
able to motivate

00:13:50.510 --> 00:13:56.710 align:middle line:84%
why you would want to work
with an optimized version

00:13:56.710 --> 00:14:01.750 align:middle line:84%
of the diffusion network, and
also why speed might not just

00:14:01.750 --> 00:14:07.390 align:middle line:84%
be the only factor that we want
to care about in this case.

00:14:07.390 --> 00:14:12.030 align:middle line:84%
Hence, I am going to talk
a little more about how

00:14:12.030 --> 00:14:16.270 align:middle line:84%
optimization becomes a
factor that transcends well

00:14:16.270 --> 00:14:17.450 align:middle line:90%
beyond just speed.

00:14:17.450 --> 00:14:20.510 align:middle line:90%


00:14:20.510 --> 00:14:22.670 align:middle line:84%
Effectively and
essentially, I like

00:14:22.670 --> 00:14:25.110 align:middle line:84%
to think of the
whole optimization

00:14:25.110 --> 00:14:31.990 align:middle line:84%
landscape as a few guiding
principles, when should,

00:14:31.990 --> 00:14:34.430 align:middle line:84%
when should the
optimization take place?

00:14:34.430 --> 00:14:38.630 align:middle line:84%
And in this case, the motivation
basically comes from the app

00:14:38.630 --> 00:14:42.130 align:middle line:84%
where should we
perform optimization.

00:14:42.130 --> 00:14:43.630 align:middle line:84%
And in this case,
it's roughly going

00:14:43.630 --> 00:14:47.130 align:middle line:84%
to be the diffusion network,
the transformer model.

00:14:47.130 --> 00:14:50.110 align:middle line:84%
That's the most
compute-intensive element

00:14:50.110 --> 00:14:51.910 align:middle line:90%
in our workflow.

00:14:51.910 --> 00:14:54.490 align:middle line:84%
And do we know what
we want to optimize?

00:14:54.490 --> 00:14:56.930 align:middle line:84%
Do we want to
optimize throughput?

00:14:56.930 --> 00:14:58.670 align:middle line:90%
Do we want to optimize memory?

00:14:58.670 --> 00:15:02.430 align:middle line:90%
Or do we want to optimize both?

00:15:02.430 --> 00:15:03.890 align:middle line:90%
And how should we optimize?

00:15:03.890 --> 00:15:08.230 align:middle line:84%
This is where this talk is
going to be focusing on.

00:15:08.230 --> 00:15:10.710 align:middle line:84%
So yeah, if we are
going to be discussing

00:15:10.710 --> 00:15:13.510 align:middle line:84%
a couple of different
recipes and approaches

00:15:13.510 --> 00:15:17.430 align:middle line:84%
as to how we can optimize
the transformer network

00:15:17.430 --> 00:15:22.270 align:middle line:84%
and also how we
can go beyond that.

00:15:22.270 --> 00:15:28.830 align:middle line:84%
Now I think there needs to
be more motivation and more

00:15:28.830 --> 00:15:33.660 align:middle line:84%
awareness to be put when it
comes to the kind of hardware

00:15:33.660 --> 00:15:37.060 align:middle line:84%
that we have access to
especially when we are deploying

00:15:37.060 --> 00:15:38.820 align:middle line:90%
these models.

00:15:38.820 --> 00:15:43.340 align:middle line:84%
Now in this picture, we can see
how the throughput immediately

00:15:43.340 --> 00:15:48.500 align:middle line:84%
improves when we try to use
shapes that are particularly

00:15:48.500 --> 00:15:50.460 align:middle line:90%
optimized for a given hardware.

00:15:50.460 --> 00:15:52.840 align:middle line:84%
Now, on the extreme
right-hand side,

00:15:52.840 --> 00:15:58.020 align:middle line:84%
the bar shows, it's
basically configured

00:15:58.020 --> 00:16:01.140 align:middle line:84%
with the most optimized
shapes for a given hardware,

00:16:01.140 --> 00:16:04.020 align:middle line:84%
and it gives us the most
amount of throughput.

00:16:04.020 --> 00:16:08.280 align:middle line:84%
And the number of total model
parameters is not changing.

00:16:08.280 --> 00:16:09.120 align:middle line:90%
That's not changing.

00:16:09.120 --> 00:16:13.460 align:middle line:84%
We are just using the shapes
within the model, maybe,

00:16:13.460 --> 00:16:20.260 align:middle line:84%
the number of attention heads
or maybe the hidden dimension

00:16:20.260 --> 00:16:22.420 align:middle line:84%
of the transformer
blocks and so on.

00:16:22.420 --> 00:16:24.500 align:middle line:84%
And it shows how
just changing that

00:16:24.500 --> 00:16:28.620 align:middle line:84%
and how optimizing that with
respect to the given hardware

00:16:28.620 --> 00:16:32.500 align:middle line:84%
can immediately render
some positive gains

00:16:32.500 --> 00:16:34.820 align:middle line:90%
in terms of throughput.

00:16:34.820 --> 00:16:38.540 align:middle line:90%
So that's what I was mentioning.

00:16:38.540 --> 00:16:41.580 align:middle line:84%
Nontrivial gains can
come from kernels

00:16:41.580 --> 00:16:44.660 align:middle line:84%
that are hyper specialized
for a given hardware,

00:16:44.660 --> 00:16:49.620 align:middle line:84%
and also model architectures
where the shapes of the model

00:16:49.620 --> 00:16:55.140 align:middle line:84%
are kind of optimized and have
been made efficient with respect

00:16:55.140 --> 00:16:56.340 align:middle line:90%
to the hardware.

00:16:56.340 --> 00:17:00.020 align:middle line:84%
It's going to be
operationalized.

00:17:00.020 --> 00:17:05.099 align:middle line:84%
Now there's also this
aspect of efficiency

00:17:05.099 --> 00:17:07.660 align:middle line:84%
that's going to keep
coming up because we

00:17:07.660 --> 00:17:11.220 align:middle line:84%
want to make our model
as efficient as possible.

00:17:11.220 --> 00:17:16.220 align:middle line:84%
But efficiency is also often a
term that has so many misnomers.

00:17:16.220 --> 00:17:20.380 align:middle line:84%
The popular belief is smaller
models are almost always faster

00:17:20.380 --> 00:17:21.560 align:middle line:90%
and more efficient.

00:17:21.560 --> 00:17:24.940 align:middle line:90%
The reality is no, not always.

00:17:24.940 --> 00:17:29.100 align:middle line:84%
And I have this very popular
figure from this paper called

00:17:29.100 --> 00:17:33.260 align:middle line:84%
"the efficiency misnomer" as
the title of my slide goes.

00:17:33.260 --> 00:17:38.300 align:middle line:84%
It basically shows how models
with higher number of parameters

00:17:38.300 --> 00:17:42.340 align:middle line:84%
may not have the
highest number of flops,

00:17:42.340 --> 00:17:48.940 align:middle line:84%
and also it may not
have a lower throughput.

00:17:48.940 --> 00:17:51.840 align:middle line:84%
Now, in this case, it's
basically reversed.

00:17:51.840 --> 00:17:55.500 align:middle line:84%
That is, the models with a
lower number of parameters

00:17:55.500 --> 00:17:58.260 align:middle line:84%
has more flops and
has a lower amount

00:17:58.260 --> 00:18:04.140 align:middle line:84%
of throughput than their
apparently heavier counterparts.

00:18:04.140 --> 00:18:07.620 align:middle line:84%
So that kind of drives
this point home.

00:18:07.620 --> 00:18:11.260 align:middle line:84%
That efficiency may not
always be related to models

00:18:11.260 --> 00:18:13.540 align:middle line:90%
being smaller.

00:18:13.540 --> 00:18:16.080 align:middle line:84%
There are more angles
to take a look at it.

00:18:16.080 --> 00:18:19.220 align:middle line:90%


00:18:19.220 --> 00:18:22.980 align:middle line:84%
Now the big elephant
is also transformer.

00:18:22.980 --> 00:18:26.500 align:middle line:84%
I'm going to keep coming back
to this transformer thing.

00:18:26.500 --> 00:18:31.500 align:middle line:84%
But also, when it comes to high
resolution generation of images

00:18:31.500 --> 00:18:39.020 align:middle line:84%
and videos, the effects of
the compute-bounded regime

00:18:39.020 --> 00:18:43.460 align:middle line:84%
of our transformer can become
really, really evident.

00:18:43.460 --> 00:18:46.620 align:middle line:84%
First, the inputs are
of high dimensional.

00:18:46.620 --> 00:18:51.260 align:middle line:84%
For example, for 4K
image generation,

00:18:51.260 --> 00:18:55.700 align:middle line:84%
even if we were to operate with
a compression factor of 8x,

00:18:55.700 --> 00:18:58.540 align:middle line:84%
we have got this huge
dimensionality problem.

00:18:58.540 --> 00:18:59.920 align:middle line:90%
We have got batch size.

00:18:59.920 --> 00:19:02.900 align:middle line:84%
Then we have nominating
channels, which

00:19:02.900 --> 00:19:05.460 align:middle line:90%
can be 64, 128, and so on.

00:19:05.460 --> 00:19:09.340 align:middle line:84%
And then you have got your
latent width and latent--

00:19:09.340 --> 00:19:11.880 align:middle line:84%
you have got your latent
height and latent width,

00:19:11.880 --> 00:19:14.080 align:middle line:90%
which is also like 512 and 512.

00:19:14.080 --> 00:19:19.020 align:middle line:84%
Now doing attention at this
high-dimensional space can

00:19:19.020 --> 00:19:21.620 align:middle line:90%
immediately restrict--

00:19:21.620 --> 00:19:27.020 align:middle line:84%
can immediately impose memory
implications as well as

00:19:27.020 --> 00:19:28.980 align:middle line:90%
speed implications.

00:19:28.980 --> 00:19:32.780 align:middle line:84%
Two very popular ways to deal
with this high-dimensionality

00:19:32.780 --> 00:19:36.460 align:middle line:84%
problem, in order to achieve
some efficiency gains,

00:19:36.460 --> 00:19:41.700 align:middle line:84%
is we increase the compression
factor, like from 8x,

00:19:41.700 --> 00:19:45.780 align:middle line:84%
we compress even higher,
maybe 16x or 24x,

00:19:45.780 --> 00:19:52.040 align:middle line:84%
and we try to recover the lost
information through other means.

00:19:52.040 --> 00:19:55.120 align:middle line:84%
There are several works
that have explored this,

00:19:55.120 --> 00:19:58.300 align:middle line:84%
but I just wanted to give
you a high-level overview

00:19:58.300 --> 00:20:00.420 align:middle line:84%
of the typical
things that are done

00:20:00.420 --> 00:20:04.740 align:middle line:84%
when we have this problem of
high-dimensional representation

00:20:04.740 --> 00:20:05.780 align:middle line:90%
spaces.

00:20:05.780 --> 00:20:11.960 align:middle line:84%
Now when we-- just to elaborate
a bit further on this point,

00:20:11.960 --> 00:20:14.960 align:middle line:84%
when we try to compress
higher and higher,

00:20:14.960 --> 00:20:18.460 align:middle line:84%
we also end up losing a lot of
information redundancy, which

00:20:18.460 --> 00:20:21.260 align:middle line:84%
might be actually
necessary study in order

00:20:21.260 --> 00:20:25.620 align:middle line:84%
to model the expressivity
into the overall workflow.

00:20:25.620 --> 00:20:31.580 align:middle line:84%
And if we do miss out on that,
our quality can take a huge hit.

00:20:31.580 --> 00:20:35.420 align:middle line:84%
Now, that's also why
it's essential to try

00:20:35.420 --> 00:20:38.620 align:middle line:84%
to recover the
information loss incurred

00:20:38.620 --> 00:20:44.220 align:middle line:84%
during this heavy compression,
so some kinds of other means.

00:20:44.220 --> 00:20:46.780 align:middle line:84%
And as a motivating
example, I want

00:20:46.780 --> 00:20:52.160 align:middle line:84%
to show you this picture
from this very popular

00:20:52.160 --> 00:20:55.500 align:middle line:84%
high resolution image
synthesis model called SANA,

00:20:55.500 --> 00:20:59.500 align:middle line:84%
which also happens to operate
on a highly compressed latent

00:20:59.500 --> 00:21:00.180 align:middle line:90%
space.

00:21:00.180 --> 00:21:05.020 align:middle line:84%
And as we can see, using
extreme compression

00:21:05.020 --> 00:21:11.340 align:middle line:84%
can have significant
gains on the speed.

00:21:11.340 --> 00:21:15.900 align:middle line:84%
So in this case, we can see a
25x reduction in the generation

00:21:15.900 --> 00:21:19.920 align:middle line:84%
latency when operating
with 4K images.

00:21:19.920 --> 00:21:24.060 align:middle line:90%


00:21:24.060 --> 00:21:31.340 align:middle line:84%
Now, the good thing
about trying to do

00:21:31.340 --> 00:21:34.180 align:middle line:84%
this more from a first
principles approach

00:21:34.180 --> 00:21:38.100 align:middle line:84%
is these approaches
should be complemented

00:21:38.100 --> 00:21:39.620 align:middle line:84%
with the kinds of
techniques that we

00:21:39.620 --> 00:21:42.320 align:middle line:84%
have for latency
optimization, for example,

00:21:42.320 --> 00:21:45.220 align:middle line:84%
maybe an optimized
kernel or maybe

00:21:45.220 --> 00:21:47.160 align:middle line:84%
using things like
FlashAttention.

00:21:47.160 --> 00:21:50.460 align:middle line:90%


00:21:50.460 --> 00:21:52.300 align:middle line:84%
So at this point
in time, I would

00:21:52.300 --> 00:21:58.220 align:middle line:84%
expect us to have a good
understanding of how

00:21:58.220 --> 00:22:00.940 align:middle line:84%
we can optimize the
model architecture

00:22:00.940 --> 00:22:05.220 align:middle line:84%
and also have a
model architecture

00:22:05.220 --> 00:22:10.060 align:middle line:84%
that respects the hardware
that it will get deployed to.

00:22:10.060 --> 00:22:12.580 align:middle line:84%
Now, ultimately,
since we are also

00:22:12.580 --> 00:22:17.020 align:middle line:84%
focusing on the applications
in which these models are going

00:22:17.020 --> 00:22:20.780 align:middle line:84%
to be used, I think
it also makes sense

00:22:20.780 --> 00:22:23.740 align:middle line:84%
to try to think about
how we can also optimize

00:22:23.740 --> 00:22:25.740 align:middle line:90%
a little bit for the use case.

00:22:25.740 --> 00:22:28.620 align:middle line:90%
So let's see how.

00:22:28.620 --> 00:22:33.180 align:middle line:84%
So we know that my
model is topping

00:22:33.180 --> 00:22:35.880 align:middle line:84%
all the standard
generation benchmarks.

00:22:35.880 --> 00:22:38.020 align:middle line:90%
That's all well and good.

00:22:38.020 --> 00:22:40.860 align:middle line:84%
But the use cases
in which the model

00:22:40.860 --> 00:22:44.020 align:middle line:84%
is going to get deployed
to, they might be different.

00:22:44.020 --> 00:22:46.220 align:middle line:84%
For example, maybe
the model would

00:22:46.220 --> 00:22:48.500 align:middle line:90%
be needed for photorealism.

00:22:48.500 --> 00:22:53.220 align:middle line:84%
So in that case, the users would
demand for specific attributes

00:22:53.220 --> 00:22:55.620 align:middle line:90%
to be better than others.

00:22:55.620 --> 00:22:59.500 align:middle line:84%
And then maybe the model
will be used in the context

00:22:59.500 --> 00:23:00.940 align:middle line:90%
of interactive generation.

00:23:00.940 --> 00:23:05.620 align:middle line:84%
In those cases, we would rather
have real-time interaction

00:23:05.620 --> 00:23:07.120 align:middle line:90%
and lowest possible latency.

00:23:07.120 --> 00:23:11.780 align:middle line:84%
And maybe quality
doesn't matter that much.

00:23:11.780 --> 00:23:17.340 align:middle line:84%
So in these cases, we will
have to finetune whatever model

00:23:17.340 --> 00:23:19.420 align:middle line:90%
that we have for specific needs.

00:23:19.420 --> 00:23:25.500 align:middle line:84%
And that entire process should
be driven by the use cases.

00:23:25.500 --> 00:23:30.460 align:middle line:84%
Now one such way to do
that, one such way to inject

00:23:30.460 --> 00:23:32.380 align:middle line:84%
some form of use
case awareness would

00:23:32.380 --> 00:23:35.900 align:middle line:84%
be to do preference alignment,
where the model is finetuned

00:23:35.900 --> 00:23:38.020 align:middle line:90%
to output what's preferred.

00:23:38.020 --> 00:23:41.380 align:middle line:84%
In this case, you have got
a text prompt, a cyberpunk

00:23:41.380 --> 00:23:44.220 align:middle line:84%
cat with a neon
sign that says Sana,

00:23:44.220 --> 00:23:48.060 align:middle line:84%
and you have got two
images, two candidates.

00:23:48.060 --> 00:23:51.540 align:middle line:84%
Now, the model will be
trained on this triplet,

00:23:51.540 --> 00:23:55.060 align:middle line:84%
and it will be basically
trained to output what

00:23:55.060 --> 00:23:57.200 align:middle line:90%
the users typically prefer.

00:23:57.200 --> 00:23:59.820 align:middle line:90%


00:23:59.820 --> 00:24:06.840 align:middle line:84%
We might also want to extend the
way users prompt these models.

00:24:06.840 --> 00:24:10.700 align:middle line:84%
So far, we have been seeing
text-to-image models, which

00:24:10.700 --> 00:24:16.060 align:middle line:84%
is basically users provide
natural text description of what

00:24:16.060 --> 00:24:18.460 align:middle line:84%
they want to see in
the final outputs.

00:24:18.460 --> 00:24:19.920 align:middle line:90%
Maybe that's not sufficient.

00:24:19.920 --> 00:24:23.220 align:middle line:84%
Maybe I also want
the output image

00:24:23.220 --> 00:24:26.140 align:middle line:90%
to follow a particular pose.

00:24:26.140 --> 00:24:30.500 align:middle line:84%
Maybe I want to condition the
model with more structural forms

00:24:30.500 --> 00:24:33.460 align:middle line:84%
of inputs such as pose,
segmentation, map,

00:24:33.460 --> 00:24:35.300 align:middle line:90%
Canny map for example.

00:24:35.300 --> 00:24:38.900 align:middle line:84%
And in this case, I'm
basically prompting the model

00:24:38.900 --> 00:24:43.260 align:middle line:84%
with a certain pose as well
as a textual description.

00:24:43.260 --> 00:24:46.460 align:middle line:84%
And as you can see,
it works, basically.

00:24:46.460 --> 00:24:49.740 align:middle line:84%
So maybe the textual
prompt was not enough.

00:24:49.740 --> 00:24:52.500 align:middle line:84%
And hence there
was a need to allow

00:24:52.500 --> 00:24:55.980 align:middle line:84%
the users to have more
control over what they can

00:24:55.980 --> 00:25:00.060 align:middle line:90%
provide to the model as inputs.

00:25:00.060 --> 00:25:02.060 align:middle line:84%
If training is not
possible because training

00:25:02.060 --> 00:25:05.560 align:middle line:84%
is an expensive gig, even with
parameter efficient finetuning,

00:25:05.560 --> 00:25:09.260 align:middle line:84%
the results might be
far from expected.

00:25:09.260 --> 00:25:11.970 align:middle line:84%
In those cases, we can
also incorporate things

00:25:11.970 --> 00:25:16.130 align:middle line:84%
like inference time
scaling, where we might want

00:25:16.130 --> 00:25:18.090 align:middle line:90%
to search for better noise.

00:25:18.090 --> 00:25:23.090 align:middle line:84%
As we saw a couple slides
back, when performing inference

00:25:23.090 --> 00:25:27.330 align:middle line:84%
with diffusion models, we try
to start from a random noise,

00:25:27.330 --> 00:25:29.450 align:middle line:90%
and then we try to denoise it.

00:25:29.450 --> 00:25:31.310 align:middle line:84%
We could search
for better noise,

00:25:31.310 --> 00:25:35.090 align:middle line:84%
like the initial noise
could be better because not

00:25:35.090 --> 00:25:37.730 align:middle line:90%
all noises are equal.

00:25:37.730 --> 00:25:42.610 align:middle line:84%
They are not going to lead
to same quality outputs.

00:25:42.610 --> 00:25:46.570 align:middle line:84%
So we might want to
search for better noise,

00:25:46.570 --> 00:25:49.090 align:middle line:84%
or maybe we want to search
for a better prompt.

00:25:49.090 --> 00:25:54.170 align:middle line:84%
And it's also possible to
search for both, actually.

00:25:54.170 --> 00:26:00.970 align:middle line:84%
Now this gives us a depiction
of what I wanted to convey.

00:26:00.970 --> 00:26:05.370 align:middle line:84%
So you basically generate an
image from a language prompt.

00:26:05.370 --> 00:26:08.310 align:middle line:84%
And then you
compute some metric.

00:26:08.310 --> 00:26:11.010 align:middle line:84%
In this case, it's
the ClipScore.

00:26:11.010 --> 00:26:16.390 align:middle line:84%
And then if the metric
is OK, then all good.

00:26:16.390 --> 00:26:18.010 align:middle line:90%
Our job is done.

00:26:18.010 --> 00:26:20.810 align:middle line:84%
If the metric is
not OK, maybe we

00:26:20.810 --> 00:26:24.810 align:middle line:84%
ask some language model to
come up with a better prompt.

00:26:24.810 --> 00:26:34.530 align:middle line:84%
And then we enter into this
interactive scaling framework.

00:26:34.530 --> 00:26:39.230 align:middle line:84%
Now Clip Score here I used clip
score as a motivating example.

00:26:39.230 --> 00:26:44.090 align:middle line:84%
But this metric should
be use case-specific.

00:26:44.090 --> 00:26:47.210 align:middle line:84%
So this is the last
section of my talk, where

00:26:47.210 --> 00:26:49.570 align:middle line:84%
I want to give you
a flavor of some

00:26:49.570 --> 00:26:52.970 align:middle line:84%
of the more advanced
optimization techniques

00:26:52.970 --> 00:26:55.650 align:middle line:90%
that we can incorporate.

00:26:55.650 --> 00:26:58.610 align:middle line:84%
So knowledge distillation
is already very popular.

00:26:58.610 --> 00:27:01.530 align:middle line:84%
We have had many
papers in the community

00:27:01.530 --> 00:27:04.490 align:middle line:90%
that talk about this technique.

00:27:04.490 --> 00:27:08.330 align:middle line:84%
So the basic idea
is we have a larger

00:27:08.330 --> 00:27:15.210 align:middle line:84%
model that is probably memory
intensive and also slower and so

00:27:15.210 --> 00:27:15.970 align:middle line:90%
on.

00:27:15.970 --> 00:27:19.530 align:middle line:84%
And we want to distill that
large model into a smaller

00:27:19.530 --> 00:27:22.090 align:middle line:84%
one, which is
often known as some

00:27:22.090 --> 00:27:27.690 align:middle line:84%
of a compressed representation
of that larger model.

00:27:27.690 --> 00:27:29.750 align:middle line:84%
And in the case of
diffusion models,

00:27:29.750 --> 00:27:32.050 align:middle line:84%
we have got large
number of iterations.

00:27:32.050 --> 00:27:34.890 align:middle line:84%
And we can typically
reduce the number

00:27:34.890 --> 00:27:39.410 align:middle line:84%
of steps or iterations needed to
generate something reasonable.

00:27:39.410 --> 00:27:42.170 align:middle line:84%
And in the literature
of diffusion models,

00:27:42.170 --> 00:27:45.770 align:middle line:84%
it's known as time
step distillation.

00:27:45.770 --> 00:27:50.890 align:middle line:84%
And also, one advantage of
working with diffusion models

00:27:50.890 --> 00:27:54.650 align:middle line:84%
is that we can combine
architectural compression

00:27:54.650 --> 00:27:58.210 align:middle line:84%
with time step
distillation and come

00:27:58.210 --> 00:28:05.610 align:middle line:84%
to be able to benefit from
the best of both worlds.

00:28:05.610 --> 00:28:09.370 align:middle line:84%
So at a glance, if
I were to give you

00:28:09.370 --> 00:28:14.490 align:middle line:84%
an overview of the things
that I find to be beneficial

00:28:14.490 --> 00:28:17.290 align:middle line:84%
when trying to approach
the optimization

00:28:17.290 --> 00:28:21.490 align:middle line:84%
landscape for diffusion
or flow models is

00:28:21.490 --> 00:28:24.530 align:middle line:84%
we should incorporate
hardware awareness

00:28:24.530 --> 00:28:26.450 align:middle line:84%
when designing the
model architecture

00:28:26.450 --> 00:28:30.850 align:middle line:84%
and also shapes
that are optimized

00:28:30.850 --> 00:28:35.050 align:middle line:84%
for a given hardware for
potentially better efficiency.

00:28:35.050 --> 00:28:38.970 align:middle line:84%
And it also can be very
beneficial to make the model

00:28:38.970 --> 00:28:42.130 align:middle line:84%
architecture flexible,
to be able to extrapolate

00:28:42.130 --> 00:28:44.970 align:middle line:84%
to use cases like high
resolution image and video

00:28:44.970 --> 00:28:47.610 align:middle line:90%
synthesis.

00:28:47.610 --> 00:28:51.610 align:middle line:84%
And we can complement
the benefits

00:28:51.610 --> 00:28:54.170 align:middle line:84%
that come from
model architectures

00:28:54.170 --> 00:28:59.010 align:middle line:84%
with latency optimization
techniques such as better

00:28:59.010 --> 00:29:04.190 align:middle line:84%
kernels, FlashAttention,
and those techniques.

00:29:04.190 --> 00:29:08.610 align:middle line:84%
And if the use cases demand
for it, we should post train.

00:29:08.610 --> 00:29:12.750 align:middle line:84%
And if latency
requirements, demand for it,

00:29:12.750 --> 00:29:16.010 align:middle line:84%
we might want to
just distill as well.

00:29:16.010 --> 00:29:19.970 align:middle line:90%
So I think that's about it.

00:29:19.970 --> 00:29:23.130 align:middle line:84%
That went quicker
than I was expecting.

00:29:23.130 --> 00:29:26.790 align:middle line:84%
And most of the
content of this talk

00:29:26.790 --> 00:29:29.610 align:middle line:84%
I have written in
this blog post.

00:29:29.610 --> 00:29:33.690 align:middle line:84%
So yeah, that's
about it from my end.

00:29:33.690 --> 00:29:37.410 align:middle line:84%
So if there's any questions,
I'm happy to take them.

00:29:37.410 --> 00:29:38.650 align:middle line:90%
So yeah.

00:29:38.650 --> 00:29:40.920 align:middle line:84%
HOST: Awesome,
thank you so much.

00:29:40.920 --> 00:30:02.000 align:middle line:90%