WEBVTT

00:00:00.000 --> 00:00:01.492 align:middle line:90%


00:00:01.492 --> 00:00:07.940 align:middle line:84%
[SQUEAKING]
[RUSTLING] [CLICKING]

00:00:07.940 --> 00:00:11.220 align:middle line:90%


00:00:11.220 --> 00:00:14.720 align:middle line:84%
SPEAKER: OK, today is the
last of the formal lectures,

00:00:14.720 --> 00:00:20.520 align:middle line:84%
and I'm very pleased that
Samuel Bruce is here to present.

00:00:20.520 --> 00:00:23.160 align:middle line:84%
Sam and I have been
working together,

00:00:23.160 --> 00:00:27.840 align:middle line:84%
and hopefully, you'll see this
is the culmination and synthesis

00:00:27.840 --> 00:00:31.800 align:middle line:84%
of bringing computer science
and economics together

00:00:31.800 --> 00:00:33.640 align:middle line:90%
on the same page.

00:00:33.640 --> 00:00:36.260 align:middle line:90%
Thanks, Sam.

00:00:36.260 --> 00:00:37.260 align:middle line:90%
SAMUEL BRUCE: Thank you.

00:00:37.260 --> 00:00:38.340 align:middle line:90%
Yeah, I'm Sam.

00:00:38.340 --> 00:00:41.040 align:middle line:84%
So I'm currently
a graduate student

00:00:41.040 --> 00:00:43.560 align:middle line:84%
at MIT in computer
science and been

00:00:43.560 --> 00:00:47.280 align:middle line:84%
working with Professor Townsend
for the last year or so

00:00:47.280 --> 00:00:49.280 align:middle line:84%
and really trying to take
the things that we've

00:00:49.280 --> 00:00:53.400 align:middle line:84%
learned in this class and other
computer science and economics

00:00:53.400 --> 00:00:56.480 align:middle line:84%
concepts and bring them together
into practical applications that

00:00:56.480 --> 00:00:59.410 align:middle line:84%
can be deployed as good
examples for infrastructure

00:00:59.410 --> 00:01:01.570 align:middle line:90%
to be built in the future.

00:01:01.570 --> 00:01:04.569 align:middle line:84%
So today, we're going to
talk about some of the things

00:01:04.569 --> 00:01:08.570 align:middle line:84%
that we've been looking into--
specifically market mechanisms

00:01:08.570 --> 00:01:10.810 align:middle line:90%
and alternative equilibria.

00:01:10.810 --> 00:01:13.850 align:middle line:84%
So we're going to be using
algorithmic game theory

00:01:13.850 --> 00:01:17.690 align:middle line:84%
and coordinated equilibria,
which we'll mention,

00:01:17.690 --> 00:01:21.450 align:middle line:84%
to find computationally feasible
algorithms for mechanisms

00:01:21.450 --> 00:01:23.270 align:middle line:90%
we can put in real code.

00:01:23.270 --> 00:01:25.850 align:middle line:90%


00:01:25.850 --> 00:01:28.730 align:middle line:84%
So just a roadmap of what
we're going to go through.

00:01:28.730 --> 00:01:32.010 align:middle line:84%
We're going to start with
Dubey's limit order market

00:01:32.010 --> 00:01:35.410 align:middle line:84%
mechanism, which is a very
specific mechanism that we're

00:01:35.410 --> 00:01:38.970 align:middle line:84%
trying to implement on
the blockchain in code.

00:01:38.970 --> 00:01:41.530 align:middle line:84%
We're going to talk about the
complexity of computing Nash

00:01:41.530 --> 00:01:44.930 align:middle line:84%
equilibria and how it relates
to that market mechanism.

00:01:44.930 --> 00:01:47.570 align:middle line:84%
We're going to talk about
correlated and coarse correlated

00:01:47.570 --> 00:01:51.170 align:middle line:84%
equilibria, which are alternate
equilibria to Nash that

00:01:51.170 --> 00:01:55.050 align:middle line:84%
might be more tractable,
easy to compute.

00:01:55.050 --> 00:01:56.810 align:middle line:84%
We're going to talk
about why it matters,

00:01:56.810 --> 00:02:00.730 align:middle line:84%
what the constraints on
blockchain algorithms are,

00:02:00.730 --> 00:02:03.850 align:middle line:84%
and tractability concerns
that arise when trying to find

00:02:03.850 --> 00:02:06.038 align:middle line:90%
equilibria of mechanisms.

00:02:06.038 --> 00:02:08.330 align:middle line:84%
And finally, we're going to
talk about specific machine

00:02:08.330 --> 00:02:11.710 align:middle line:84%
learning algorithms, which can
be used to improve this process,

00:02:11.710 --> 00:02:14.130 align:middle line:90%
make it more efficient.

00:02:14.130 --> 00:02:17.170 align:middle line:84%
So we're going to jump right
into Dubey's limit order

00:02:17.170 --> 00:02:18.110 align:middle line:90%
mechanism.

00:02:18.110 --> 00:02:24.230 align:middle line:84%
And this is a specific game
from Dubey from a while ago,

00:02:24.230 --> 00:02:28.410 align:middle line:84%
basically detailing a limit
order market that in equilibria

00:02:28.410 --> 00:02:30.170 align:middle line:90%
has very nice properties.

00:02:30.170 --> 00:02:33.170 align:middle line:84%
And just the notation
of the market here--

00:02:33.170 --> 00:02:39.290 align:middle line:84%
we have n players
notated by i and k goods.

00:02:39.290 --> 00:02:40.950 align:middle line:90%
Each player has an endowment.

00:02:40.950 --> 00:02:43.930 align:middle line:84%
We'll denote that with a
and a utility function.

00:02:43.930 --> 00:02:45.570 align:middle line:84%
And in each time
period, a player

00:02:45.570 --> 00:02:47.810 align:middle line:90%
gets to submit a strategy.

00:02:47.810 --> 00:02:51.450 align:middle line:84%
And what the strategy looks like
is these four quantities here--

00:02:51.450 --> 00:02:54.490 align:middle line:90%
p, q, p tilde, and q tilde.

00:02:54.490 --> 00:03:03.980 align:middle line:84%
And what those denote-- denotes
is I will buy q units of good j

00:03:03.980 --> 00:03:06.300 align:middle line:90%
if the price is p or less.

00:03:06.300 --> 00:03:09.700 align:middle line:84%
Or conversely, I
will sell q units

00:03:09.700 --> 00:03:11.780 align:middle line:90%
if the price is p tilde or more.

00:03:11.780 --> 00:03:15.100 align:middle line:84%
So in one strategy
vector, each player

00:03:15.100 --> 00:03:19.020 align:middle line:84%
I is denoting what they
would do as a limit order

00:03:19.020 --> 00:03:22.900 align:middle line:84%
over all n goods-- or
all k goods, sorry.

00:03:22.900 --> 00:03:27.060 align:middle line:84%
And they're allowed to both
say that they will buy and sell

00:03:27.060 --> 00:03:29.940 align:middle line:84%
all goods, even the same
ones, at whatever price

00:03:29.940 --> 00:03:31.460 align:middle line:90%
they so choose.

00:03:31.460 --> 00:03:34.020 align:middle line:84%
So a lot of flexibility
in the strategy space.

00:03:34.020 --> 00:03:37.260 align:middle line:84%
They're allowed to submit a
bid of these four vectors.

00:03:37.260 --> 00:03:39.220 align:middle line:84%
And they can choose
any combination of what

00:03:39.220 --> 00:03:42.940 align:middle line:84%
they would like to buy,
sell, and what prices.

00:03:42.940 --> 00:03:45.340 align:middle line:84%
The only constraint
on their strategy

00:03:45.340 --> 00:03:49.780 align:middle line:84%
is that they may not sell goods
that are not in their endowment.

00:03:49.780 --> 00:03:52.020 align:middle line:84%
They cannot go over their
endowment set for selling

00:03:52.020 --> 00:03:54.020 align:middle line:90%
in each time period.

00:03:54.020 --> 00:03:56.700 align:middle line:84%
So the full strategy
space is the combination

00:03:56.700 --> 00:03:58.120 align:middle line:84%
of all of these
player strategies.

00:03:58.120 --> 00:04:01.300 align:middle line:90%
We'll denote that as S here.

00:04:01.300 --> 00:04:06.500 align:middle line:84%
And in outcome consists
of x where each x is just

00:04:06.500 --> 00:04:09.860 align:middle line:84%
the bundle of goods that
each agent buys or sells

00:04:09.860 --> 00:04:11.340 align:middle line:84%
in the time period--
so the trades

00:04:11.340 --> 00:04:12.840 align:middle line:84%
in the limit order
that are actually

00:04:12.840 --> 00:04:15.660 align:middle line:90%
executed each time period.

00:04:15.660 --> 00:04:19.019 align:middle line:84%
We denote the entire game
as this gamma function

00:04:19.019 --> 00:04:22.660 align:middle line:84%
here, where E is our market,
as we've described it above,

00:04:22.660 --> 00:04:25.500 align:middle line:84%
and also this weight
lambda, where lambda

00:04:25.500 --> 00:04:29.700 align:middle line:84%
is a penalty to players who
at the end of the period

00:04:29.700 --> 00:04:31.200 align:middle line:90%
have negative credit.

00:04:31.200 --> 00:04:33.740 align:middle line:84%
So notice when I said the
constraints on the strategy

00:04:33.740 --> 00:04:37.380 align:middle line:84%
functions, there was no
constraint that a player cannot

00:04:37.380 --> 00:04:40.740 align:middle line:84%
buy goods he does not
have the money for.

00:04:40.740 --> 00:04:44.940 align:middle line:84%
If a player is endowed
with goods and money,

00:04:44.940 --> 00:04:47.320 align:middle line:84%
they may sell
goods and buy goods

00:04:47.320 --> 00:04:50.060 align:middle line:84%
as long as they don't sell
goods they don't have.

00:04:50.060 --> 00:04:54.370 align:middle line:84%
And at the end of the period, if
they after their trade settle,

00:04:54.370 --> 00:04:56.870 align:middle line:84%
have negative
credit or the goods

00:04:56.870 --> 00:04:59.630 align:middle line:84%
that they buy outweigh
the goods they

00:04:59.630 --> 00:05:03.590 align:middle line:84%
sell plus their endowed
money, they incur a penalty.

00:05:03.590 --> 00:05:08.430 align:middle line:84%
So it's allowed, but they are
penalized weighted by lambda.

00:05:08.430 --> 00:05:11.535 align:middle line:84%
So the actual mechanism
implementation

00:05:11.535 --> 00:05:12.910 align:middle line:84%
for each of the
goods we're going

00:05:12.910 --> 00:05:15.090 align:middle line:84%
to establish a
separate trading post.

00:05:15.090 --> 00:05:17.830 align:middle line:84%
So each good essentially
has its own market

00:05:17.830 --> 00:05:21.950 align:middle line:84%
where the specific orders
for each good will filter.

00:05:21.950 --> 00:05:25.910 align:middle line:84%
The players submit their own
strategy in each time period.

00:05:25.910 --> 00:05:28.870 align:middle line:84%
And for each trading
post, all of the bids

00:05:28.870 --> 00:05:31.670 align:middle line:84%
that the players
submit are accrued,

00:05:31.670 --> 00:05:35.870 align:middle line:84%
and you match the highest
buyer with the lowest seller.

00:05:35.870 --> 00:05:38.150 align:middle line:84%
So intuitively right--
if the highest buyer

00:05:38.150 --> 00:05:40.030 align:middle line:84%
is less than the
lowest seller, no trade

00:05:40.030 --> 00:05:41.562 align:middle line:84%
will occur because
the seller won't

00:05:41.562 --> 00:05:43.270 align:middle line:84%
trade for a price that
low, and the buyer

00:05:43.270 --> 00:05:45.150 align:middle line:84%
won't trade for a
price that high.

00:05:45.150 --> 00:05:47.110 align:middle line:84%
But if there is some
intersection where

00:05:47.110 --> 00:05:49.910 align:middle line:84%
some buyer is willing to
pay and some seller is

00:05:49.910 --> 00:05:52.360 align:middle line:84%
willing to sell for
less, that trade

00:05:52.360 --> 00:05:54.680 align:middle line:84%
will execute at
the buyer's amount.

00:05:54.680 --> 00:05:56.020 align:middle line:90%
And it'll match up.

00:05:56.020 --> 00:06:01.580 align:middle line:84%
So if the first buyer only
wants to buy two units,

00:06:01.580 --> 00:06:03.760 align:middle line:84%
but the first seller
wants to sell 4,

00:06:03.760 --> 00:06:05.860 align:middle line:84%
the first buyer will
buy those two units,

00:06:05.860 --> 00:06:08.520 align:middle line:84%
and then the mechanism will
move to the second-highest buyer

00:06:08.520 --> 00:06:12.060 align:middle line:84%
to execute the rest of the
seller's order and vice versa.

00:06:12.060 --> 00:06:14.120 align:middle line:84%
If the seller
wants to sell less,

00:06:14.120 --> 00:06:17.360 align:middle line:84%
then it will execute in
full to that first buyer,

00:06:17.360 --> 00:06:20.320 align:middle line:84%
and then we'll move
on to the next seller.

00:06:20.320 --> 00:06:23.500 align:middle line:84%
So as I mentioned
on the last slide,

00:06:23.500 --> 00:06:26.400 align:middle line:84%
each player can
borrow at no interest

00:06:26.400 --> 00:06:27.920 align:middle line:90%
to fund their purchases.

00:06:27.920 --> 00:06:32.000 align:middle line:84%
So in each period, they can make
whatever trades they'd like.

00:06:32.000 --> 00:06:34.240 align:middle line:84%
However, if at the
end of the period

00:06:34.240 --> 00:06:37.520 align:middle line:84%
they have a net negative credit
between the goods they purchased

00:06:37.520 --> 00:06:40.240 align:middle line:84%
and the goods they sold
and their endowed money,

00:06:40.240 --> 00:06:44.040 align:middle line:84%
they incur a weighted
cost by lambda based off

00:06:44.040 --> 00:06:47.600 align:middle line:84%
of the amount of negative
credit they have.

00:06:47.600 --> 00:06:52.440 align:middle line:84%
So here we're going to define
the equilibria of the mechanism.

00:06:52.440 --> 00:06:55.000 align:middle line:90%
So we have our strategy set.

00:06:55.000 --> 00:06:58.800 align:middle line:84%
And this is a specific strategy
for one time period, little s.

00:06:58.800 --> 00:07:00.720 align:middle line:84%
So all the players
submit their bids.

00:07:00.720 --> 00:07:04.000 align:middle line:84%
We have a payoff function, which
is just the utility functions

00:07:04.000 --> 00:07:06.360 align:middle line:90%
for each player, capital pi.

00:07:06.360 --> 00:07:10.400 align:middle line:84%
And we denote the
mechanism to be efficient

00:07:10.400 --> 00:07:15.200 align:middle line:84%
if there is no subset of
players, subset m, who

00:07:15.200 --> 00:07:18.600 align:middle line:84%
can play an alternate
strategy and find

00:07:18.600 --> 00:07:23.100 align:middle line:84%
a Pareto-dominant outcome
for that time period.

00:07:23.100 --> 00:07:28.420 align:middle line:84%
So here, we have our payout
function across all players.

00:07:28.420 --> 00:07:32.400 align:middle line:84%
If the payout of the
strategy given E-- so E

00:07:32.400 --> 00:07:33.540 align:middle line:90%
is the deviation here.

00:07:33.540 --> 00:07:35.760 align:middle line:84%
So this is the
payout with deviation

00:07:35.760 --> 00:07:37.960 align:middle line:84%
is greater than the
payout without deviation

00:07:37.960 --> 00:07:39.440 align:middle line:90%
for all players.

00:07:39.440 --> 00:07:42.840 align:middle line:84%
And for some player, the payout
with deviation is greater.

00:07:42.840 --> 00:07:46.800 align:middle line:84%
It's a Pareto-dominating
bundle, and that is not

00:07:46.800 --> 00:07:49.310 align:middle line:90%
the equilibrium in this set.

00:07:49.310 --> 00:07:50.510 align:middle line:90%
It's not efficient here.

00:07:50.510 --> 00:07:52.250 align:middle line:90%
That's how we define it.

00:07:52.250 --> 00:07:54.810 align:middle line:84%
So we can have a coalition
of arbitrary size

00:07:54.810 --> 00:07:58.330 align:middle line:84%
all the way up to every
subset of the n players.

00:07:58.330 --> 00:08:02.090 align:middle line:84%
We call a non-cooperative
equilibrium of the mechanism

00:08:02.090 --> 00:08:05.390 align:middle line:84%
if it is efficient
individually for each player.

00:08:05.390 --> 00:08:08.810 align:middle line:84%
So that's very similar to the
idea of a Nash equilibrium.

00:08:08.810 --> 00:08:11.550 align:middle line:84%
If, given the strategies
of other players,

00:08:11.550 --> 00:08:15.790 align:middle line:84%
no individual can deviate and
find a Pareto dominating bundle,

00:08:15.790 --> 00:08:18.210 align:middle line:84%
they can't make someone
or themselves better

00:08:18.210 --> 00:08:20.410 align:middle line:84%
without making
anyone else worse.

00:08:20.410 --> 00:08:23.850 align:middle line:84%
We call that a non-cooperative
equilibrium, and that

00:08:23.850 --> 00:08:26.870 align:middle line:84%
will be very important for
the goal of this mechanism,

00:08:26.870 --> 00:08:31.290 align:middle line:84%
because very good properties
arise in this equilibrium.

00:08:31.290 --> 00:08:37.130 align:middle line:84%
So further, a slightly stronger
statement, if it's efficient,

00:08:37.130 --> 00:08:41.470 align:middle line:84%
this holds for any arbitrary
coalition of any size,

00:08:41.470 --> 00:08:43.929 align:middle line:90%
up to any subset of n players.

00:08:43.929 --> 00:08:46.930 align:middle line:84%
Then we call it a strong
non-cooperative equilibria.

00:08:46.930 --> 00:08:49.610 align:middle line:84%
So the equilibria
is strong if there

00:08:49.610 --> 00:08:52.490 align:middle line:84%
is no group of players
that can form a coalition

00:08:52.490 --> 00:08:54.750 align:middle line:84%
and create an
alternative strategy,

00:08:54.750 --> 00:08:58.130 align:middle line:84%
that Pareto dominates
the original one.

00:08:58.130 --> 00:09:00.690 align:middle line:84%
A couple more
terminology points here--

00:09:00.690 --> 00:09:03.510 align:middle line:84%
we call the non-cooperative
equilibria active

00:09:03.510 --> 00:09:06.650 align:middle line:84%
if each trading post has
at least two active buyers

00:09:06.650 --> 00:09:07.710 align:middle line:90%
and active sellers.

00:09:07.710 --> 00:09:10.530 align:middle line:84%
So there's trade happening
at each trading post,

00:09:10.530 --> 00:09:14.250 align:middle line:84%
and we call it tight if all
active buyers and sellers are

00:09:14.250 --> 00:09:18.090 align:middle line:90%
quoting the same price.

00:09:18.090 --> 00:09:20.030 align:middle line:84%
And feel free to stop
me with any questions.

00:09:20.030 --> 00:09:21.113 align:middle line:90%
There's a lot of notation.

00:09:21.113 --> 00:09:21.690 align:middle line:90%
Yes.

00:09:21.690 --> 00:09:25.850 align:middle line:84%
AUDIENCE: So the strategies
can say for one price,

00:09:25.850 --> 00:09:28.210 align:middle line:84%
I want to buy three goods,
and for another price,

00:09:28.210 --> 00:09:29.670 align:middle line:90%
I want to buy five goods.

00:09:29.670 --> 00:09:31.733 align:middle line:84%
It's just one price
and one quantity.

00:09:31.733 --> 00:09:32.650 align:middle line:90%
SAMUEL BRUCE: Correct.

00:09:32.650 --> 00:09:34.410 align:middle line:90%
Yes.

00:09:34.410 --> 00:09:37.330 align:middle line:84%
So now we look at
the properties that

00:09:37.330 --> 00:09:39.710 align:middle line:84%
arise in this
non-cooperative equilibria.

00:09:39.710 --> 00:09:41.490 align:middle line:84%
And I'm going to be
using NE throughout

00:09:41.490 --> 00:09:43.150 align:middle line:90%
for non-cooperative equilibria.

00:09:43.150 --> 00:09:44.520 align:middle line:90%
It's not Nash equilibria.

00:09:44.520 --> 00:09:48.340 align:middle line:84%
It's similar, but again
slightly stronger.

00:09:48.340 --> 00:09:52.260 align:middle line:84%
For any market and any
negative credit penalty,

00:09:52.260 --> 00:09:57.100 align:middle line:84%
the active non-cooperative
equilibria are all competitive.

00:09:57.100 --> 00:10:00.760 align:middle line:84%
The tight active non-cooperative
equilibria are also competitive.

00:10:00.760 --> 00:10:06.060 align:middle line:84%
So all of these game outcomes
where coalitions can't do better

00:10:06.060 --> 00:10:08.460 align:middle line:84%
by deviating, they
actually coincide perfectly

00:10:08.460 --> 00:10:11.280 align:middle line:84%
with the competitive
equilibria of the market game.

00:10:11.280 --> 00:10:13.540 align:middle line:90%
So it's very favorable.

00:10:13.540 --> 00:10:16.820 align:middle line:84%
And every tight active
non-cooperative equilibria

00:10:16.820 --> 00:10:19.980 align:middle line:84%
is strong, meaning
that if it's tight

00:10:19.980 --> 00:10:23.180 align:middle line:84%
and everyone is quoting the same
prices for the active traders,

00:10:23.180 --> 00:10:25.100 align:middle line:84%
then there are no
coalitions that

00:10:25.100 --> 00:10:30.780 align:middle line:84%
exist that could change a
Pareto-dominant strategy.

00:10:30.780 --> 00:10:34.660 align:middle line:84%
So there are actually
exists an application

00:10:34.660 --> 00:10:37.340 align:middle line:84%
on the blockchain
now that implements

00:10:37.340 --> 00:10:40.560 align:middle line:84%
a lot of Dubey's mechanism
pretty faithfully.

00:10:40.560 --> 00:10:42.260 align:middle line:90%
It's called SPEEDEX.

00:10:42.260 --> 00:10:45.930 align:middle line:84%
It was developed,
published in 2023.

00:10:45.930 --> 00:10:49.350 align:middle line:84%
It's a decentralized
limit order exchange built

00:10:49.350 --> 00:10:51.830 align:middle line:90%
on a blockchain technology.

00:10:51.830 --> 00:10:54.430 align:middle line:84%
So it's fully
operational, and it exists

00:10:54.430 --> 00:10:56.830 align:middle line:90%
in a very similar operation.

00:10:56.830 --> 00:11:02.070 align:middle line:84%
Users submit bids, limit
orders to the mechanism.

00:11:02.070 --> 00:11:03.642 align:middle line:90%
In this case, it's an exchange.

00:11:03.642 --> 00:11:05.350 align:middle line:84%
So your bid would look
something like I'm

00:11:05.350 --> 00:11:09.030 align:middle line:84%
willing to trade
$100 for 110 euros,

00:11:09.030 --> 00:11:13.910 align:middle line:84%
or some currency exchange is
what the current application is.

00:11:13.910 --> 00:11:16.750 align:middle line:84%
And then in each block
on the blockchain,

00:11:16.750 --> 00:11:20.430 align:middle line:84%
the mechanism can calculate
the market-clearing prices

00:11:20.430 --> 00:11:23.550 align:middle line:84%
for each good separately
at those trading posts

00:11:23.550 --> 00:11:26.862 align:middle line:84%
and executes all outstanding
orders at that price.

00:11:26.862 --> 00:11:29.070 align:middle line:84%
So similarly, it's going to
group up the limit orders

00:11:29.070 --> 00:11:35.767 align:middle line:84%
by separate markets
for each good.

00:11:35.767 --> 00:11:37.350 align:middle line:84%
It's going to find
the clearing price,

00:11:37.350 --> 00:11:40.030 align:middle line:84%
and it's going to execute so
that all trade that can happen

00:11:40.030 --> 00:11:41.510 align:middle line:90%
will happen.

00:11:41.510 --> 00:11:44.790 align:middle line:84%
So one difference that's
very important to note

00:11:44.790 --> 00:11:47.870 align:middle line:84%
is here, our
mechanism is executing

00:11:47.870 --> 00:11:50.850 align:middle line:84%
all trades at the
clearing price--

00:11:50.850 --> 00:11:54.270 align:middle line:84%
so the exact intersection price
where the buyers and sellers

00:11:54.270 --> 00:11:56.590 align:middle line:90%
want to sell the same amount.

00:11:56.590 --> 00:11:59.430 align:middle line:84%
In Dubey's mechanism,
remember, all buyers

00:11:59.430 --> 00:12:01.250 align:middle line:90%
sold at the price they quote.

00:12:01.250 --> 00:12:04.090 align:middle line:84%
They did not operate at
the intersection price.

00:12:04.090 --> 00:12:07.110 align:middle line:84%
It was at the buyer's
price for each limit order.

00:12:07.110 --> 00:12:10.690 align:middle line:84%
And that may seem unfavorable
in Dubey's mechanism,

00:12:10.690 --> 00:12:13.730 align:middle line:84%
because it gives a lot
of power to the sellers,

00:12:13.730 --> 00:12:16.590 align:middle line:84%
because they extract
all of the willingness

00:12:16.590 --> 00:12:18.670 align:middle line:90%
to pay from the buyers.

00:12:18.670 --> 00:12:22.367 align:middle line:84%
When their orders are higher
than some intersection price,

00:12:22.367 --> 00:12:23.950 align:middle line:84%
the seller actually
is able to extract

00:12:23.950 --> 00:12:27.670 align:middle line:84%
the value in between their sell
price and the buyer's buy price.

00:12:27.670 --> 00:12:30.710 align:middle line:84%
However, Dubey
proves in his paper

00:12:30.710 --> 00:12:33.670 align:middle line:84%
that executing at the
intersection price

00:12:33.670 --> 00:12:36.570 align:middle line:84%
breaks down some of the
assumptions of the paper,

00:12:36.570 --> 00:12:40.120 align:middle line:84%
and we can no longer guarantee
that the equilibria of the game

00:12:40.120 --> 00:12:41.760 align:middle line:90%
are competitive.

00:12:41.760 --> 00:12:44.440 align:middle line:90%
So that is a big problem.

00:12:44.440 --> 00:12:47.540 align:middle line:84%
Obviously, if the game
equilibria is not competitive,

00:12:47.540 --> 00:12:49.920 align:middle line:84%
it's not a great model for
how agents are going to act,

00:12:49.920 --> 00:12:51.800 align:middle line:90%
and goods are going to clear.

00:12:51.800 --> 00:12:56.200 align:middle line:84%
So SPEEDEX operates mostly
in the scope of the paper

00:12:56.200 --> 00:12:58.360 align:middle line:90%
without this big change.

00:12:58.360 --> 00:12:59.840 align:middle line:84%
And the algorithm
behind the scenes

00:12:59.840 --> 00:13:02.960 align:middle line:84%
is using it's a tatonnement
iterative process.

00:13:02.960 --> 00:13:04.900 align:middle line:90%
That's not extremely important.

00:13:04.900 --> 00:13:07.340 align:middle line:84%
It's just good to know that
it's an iterative strategy.

00:13:07.340 --> 00:13:11.880 align:middle line:84%
So essentially, at each block,
it groups all of the limit

00:13:11.880 --> 00:13:14.800 align:middle line:84%
orders for each
good, and then it's

00:13:14.800 --> 00:13:17.040 align:middle line:84%
going to iteratively
try and quote

00:13:17.040 --> 00:13:20.480 align:middle line:84%
prices that will clear
the market in an algorithm

00:13:20.480 --> 00:13:23.040 align:middle line:84%
until it finally finds
the one that clears

00:13:23.040 --> 00:13:26.240 align:middle line:90%
the buyers and the sellers.

00:13:26.240 --> 00:13:27.720 align:middle line:84%
Some of the outcomes
of SPEEDEX--

00:13:27.720 --> 00:13:29.400 align:middle line:90%
that it does really well.

00:13:29.400 --> 00:13:32.360 align:middle line:84%
It has a great
runtime per block.

00:13:32.360 --> 00:13:35.720 align:middle line:84%
Its runtime is of the order
of the number of assets

00:13:35.720 --> 00:13:38.530 align:middle line:84%
squared times the
logarithm of the number

00:13:38.530 --> 00:13:40.450 align:middle line:90%
of outstanding orders.

00:13:40.450 --> 00:13:42.130 align:middle line:84%
And the logarithm
part is very important

00:13:42.130 --> 00:13:45.530 align:middle line:84%
because you can imagine
in most exchange markets

00:13:45.530 --> 00:13:48.970 align:middle line:84%
or limit order markets, we're
expecting a relatively small

00:13:48.970 --> 00:13:50.570 align:middle line:90%
fixed number of goods.

00:13:50.570 --> 00:13:53.570 align:middle line:84%
However, there may be
many, many, many orders

00:13:53.570 --> 00:13:55.030 align:middle line:90%
from many, many agents.

00:13:55.030 --> 00:13:57.450 align:middle line:84%
So an algorithm that
scales logarithmically

00:13:57.450 --> 00:13:59.570 align:middle line:84%
in those number
of orders is going

00:13:59.570 --> 00:14:02.530 align:middle line:84%
to be very favorable
for execution.

00:14:02.530 --> 00:14:05.130 align:middle line:84%
The clearing prices
are also consistent.

00:14:05.130 --> 00:14:07.250 align:middle line:84%
Remember I said, it's
an exchange specifically

00:14:07.250 --> 00:14:08.730 align:middle line:90%
between currencies.

00:14:08.730 --> 00:14:12.010 align:middle line:84%
So if you have multiple
currencies here--

00:14:12.010 --> 00:14:15.410 align:middle line:84%
A, B, C-- there's no
arbitrage opportunities

00:14:15.410 --> 00:14:17.010 align:middle line:84%
between the
different currencies.

00:14:17.010 --> 00:14:18.970 align:middle line:84%
The ratio of prices
between A and B

00:14:18.970 --> 00:14:22.290 align:middle line:84%
times the ratio between B
and C is the same as A and C.

00:14:22.290 --> 00:14:25.850 align:middle line:84%
So there's no opportunity
to submit different exchange

00:14:25.850 --> 00:14:31.690 align:middle line:84%
bids on different goods
and extract arbitrage.

00:14:31.690 --> 00:14:37.490 align:middle line:84%
And another good thing about
having these prices consistent

00:14:37.490 --> 00:14:41.890 align:middle line:84%
is that all transactions in one
block occur at the same prices.

00:14:41.890 --> 00:14:45.530 align:middle line:84%
Now, that may seem
intuitive, but it's actually

00:14:45.530 --> 00:14:48.810 align:middle line:84%
not how most previous
decentralized exchanges

00:14:48.810 --> 00:14:50.050 align:middle line:90%
are implemented.

00:14:50.050 --> 00:14:53.490 align:middle line:84%
Most decentralized exchanges
will operate on each order

00:14:53.490 --> 00:14:56.170 align:middle line:84%
as it comes in,
ordered in the block,

00:14:56.170 --> 00:14:58.390 align:middle line:84%
and it opens up to
front-running attacks.

00:14:58.390 --> 00:15:01.550 align:middle line:84%
And what a front-running
attack is is essentially,

00:15:01.550 --> 00:15:05.010 align:middle line:84%
you can imagine some node
in the blockchain who

00:15:05.010 --> 00:15:09.010 align:middle line:84%
has a lot of computing power
and therefore validates

00:15:09.010 --> 00:15:10.950 align:middle line:90%
and proposes blocks very often.

00:15:10.950 --> 00:15:14.110 align:middle line:84%
Remember how often
a block is-- sorry,

00:15:14.110 --> 00:15:17.170 align:middle line:84%
how often a node
successfully proposes

00:15:17.170 --> 00:15:20.250 align:middle line:84%
a block is proportional
to their computing power.

00:15:20.250 --> 00:15:22.370 align:middle line:84%
So you may have a
specific node who

00:15:22.370 --> 00:15:24.770 align:middle line:84%
can spy on the
transactions that are being

00:15:24.770 --> 00:15:26.110 align:middle line:90%
submitted in the next block.

00:15:26.110 --> 00:15:27.930 align:middle line:84%
They see them as
they get submitted,

00:15:27.930 --> 00:15:31.210 align:middle line:84%
and they can add in a
transaction of their own, T

00:15:31.210 --> 00:15:32.890 align:middle line:90%
prime.

00:15:32.890 --> 00:15:35.720 align:middle line:84%
And if they can reorder the
transactions, which they can,

00:15:35.720 --> 00:15:38.020 align:middle line:84%
because generally, in a
block, all transactions

00:15:38.020 --> 00:15:40.500 align:middle line:90%
can be ordered in any way.

00:15:40.500 --> 00:15:42.900 align:middle line:84%
With old decentralized
exchanges,

00:15:42.900 --> 00:15:45.380 align:middle line:84%
they can actually
extract extra rent

00:15:45.380 --> 00:15:47.960 align:middle line:84%
from the transaction
they spied on.

00:15:47.960 --> 00:15:50.020 align:middle line:84%
You can imagine
some transaction T

00:15:50.020 --> 00:15:54.300 align:middle line:84%
wants to buy this currency at
this price as a limit order.

00:15:54.300 --> 00:15:56.520 align:middle line:84%
It's going to execute
at this price.

00:15:56.520 --> 00:15:59.100 align:middle line:84%
But if you submit some
transaction that executes just

00:15:59.100 --> 00:16:01.300 align:middle line:84%
before it, you
can actually raise

00:16:01.300 --> 00:16:04.177 align:middle line:84%
the price they pay a little bit
and take a small amount of money

00:16:04.177 --> 00:16:04.760 align:middle line:90%
in the middle.

00:16:04.760 --> 00:16:07.220 align:middle line:84%
And this is called a
front-running attack,

00:16:07.220 --> 00:16:10.260 align:middle line:84%
and it's an example of
miner-extractable value.

00:16:10.260 --> 00:16:12.820 align:middle line:84%
So miner-extractable
value here just refers

00:16:12.820 --> 00:16:16.420 align:middle line:84%
to any node who's
mining in Bitcoin

00:16:16.420 --> 00:16:18.460 align:middle line:84%
or other proof-of-work
blockchains that

00:16:18.460 --> 00:16:22.040 align:middle line:84%
has a lot of computing power,
has a lot of processing power,

00:16:22.040 --> 00:16:24.500 align:middle line:84%
can reorder these transactions
to benefit themselves

00:16:24.500 --> 00:16:29.580 align:middle line:84%
and extract extra surplus
from other transactions.

00:16:29.580 --> 00:16:33.120 align:middle line:84%
And SPEEDEX eliminates that
opportunity, which is very good.

00:16:33.120 --> 00:16:35.780 align:middle line:90%


00:16:35.780 --> 00:16:37.820 align:middle line:84%
AUDIENCE: But the
Dubey-- is it true

00:16:37.820 --> 00:16:44.563 align:middle line:84%
that every non-cooperative
equilibrium has no arbitrage?

00:16:44.563 --> 00:16:45.480 align:middle line:90%
There couldn't be one.

00:16:45.480 --> 00:16:47.440 align:middle line:90%
That has arbitrage.

00:16:47.440 --> 00:16:51.540 align:middle line:84%
SAMUEL BRUCE: So in the
specific example of Dubey,

00:16:51.540 --> 00:16:53.780 align:middle line:90%
it's not an exchange market.

00:16:53.780 --> 00:16:57.202 align:middle line:84%
And again, the prices
are competitive,

00:16:57.202 --> 00:16:58.660 align:middle line:84%
which means that
no, there wouldn't

00:16:58.660 --> 00:17:01.980 align:middle line:90%
be arbitrage opportunities.

00:17:01.980 --> 00:17:07.099 align:middle line:84%
So looking at SPEEDEX and
looking at Dubey's mechanism,

00:17:07.099 --> 00:17:11.420 align:middle line:84%
we raise some concerns on
the actual tractability

00:17:11.420 --> 00:17:13.859 align:middle line:90%
of finding equilibria.

00:17:13.859 --> 00:17:17.300 align:middle line:84%
So right, we've talked about
the non-cooperative equilibria,

00:17:17.300 --> 00:17:20.180 align:middle line:84%
where there's no individual
agents or coalitions who

00:17:20.180 --> 00:17:23.000 align:middle line:84%
can deviate and find
Pareto-dominant strategies.

00:17:23.000 --> 00:17:26.380 align:middle line:84%
And it's very nice
in equilibria.

00:17:26.380 --> 00:17:32.710 align:middle line:84%
However, nowhere in either
SPEEDEX or Dubey's proofs

00:17:32.710 --> 00:17:34.710 align:middle line:84%
do they mention
how it can converge

00:17:34.710 --> 00:17:37.350 align:middle line:84%
on our non-cooperative
equilibria.

00:17:37.350 --> 00:17:40.310 align:middle line:84%
And in general, finding
non-cooperative or Nash

00:17:40.310 --> 00:17:42.950 align:middle line:84%
equilibria, because
they're similar concepts,

00:17:42.950 --> 00:17:45.430 align:middle line:84%
requires each agent
to rationally predict

00:17:45.430 --> 00:17:47.230 align:middle line:84%
the strategies of
the other players

00:17:47.230 --> 00:17:50.570 align:middle line:84%
and submit their optimal
bid vector accordingly.

00:17:50.570 --> 00:17:53.310 align:middle line:84%
If you think back to
our reward function,

00:17:53.310 --> 00:17:55.043 align:middle line:84%
you're looking at your
deviation compared

00:17:55.043 --> 00:17:56.210 align:middle line:90%
to the strategies of others.

00:17:56.210 --> 00:17:59.830 align:middle line:84%
So for you to accurately compute
and think about as an agent

00:17:59.830 --> 00:18:02.390 align:middle line:84%
what your strategy
should have to be,

00:18:02.390 --> 00:18:05.150 align:middle line:84%
you able to estimate the
strategies of others.

00:18:05.150 --> 00:18:07.550 align:middle line:84%
And when we're talking
about almost continuously

00:18:07.550 --> 00:18:08.610 align:middle line:90%
valued prices--

00:18:08.610 --> 00:18:11.830 align:middle line:84%
I put almost in parentheses
here because if we're

00:18:11.830 --> 00:18:13.750 align:middle line:84%
talking about computing
these, it's always

00:18:13.750 --> 00:18:14.970 align:middle line:90%
going to be finite valued.

00:18:14.970 --> 00:18:17.550 align:middle line:84%
But let's say a very,
very large number

00:18:17.550 --> 00:18:20.730 align:middle line:84%
of possible prices, a very
large number of goods,

00:18:20.730 --> 00:18:22.950 align:middle line:84%
a very large number
of quantities,

00:18:22.950 --> 00:18:26.390 align:middle line:84%
the strategy space for each
agent, which is the four vectors

00:18:26.390 --> 00:18:29.480 align:middle line:84%
all over those things,
it gets extremely large.

00:18:29.480 --> 00:18:31.220 align:middle line:84%
And it's very hard
to search over.

00:18:31.220 --> 00:18:34.640 align:middle line:84%
And it's very hard
for a rational agent

00:18:34.640 --> 00:18:38.640 align:middle line:84%
to reason well about what other
players are going to submit

00:18:38.640 --> 00:18:42.040 align:middle line:84%
and what they ought to
submit, making it concerning

00:18:42.040 --> 00:18:43.800 align:middle line:84%
whether or not we
would actually find

00:18:43.800 --> 00:18:47.440 align:middle line:84%
the non-cooperative
equilibria of our game.

00:18:47.440 --> 00:18:51.680 align:middle line:84%
And even when we iterate many
times over the mechanism,

00:18:51.680 --> 00:18:55.360 align:middle line:84%
it may be completely intractable
to just find ourselves

00:18:55.360 --> 00:18:56.980 align:middle line:90%
in a non-cooperative equilibria.

00:18:56.980 --> 00:19:00.880 align:middle line:84%
It may be very hard for agents
playing in our mechanism

00:19:00.880 --> 00:19:05.080 align:middle line:84%
to find themselves to our
equilibria on their own.

00:19:05.080 --> 00:19:10.080 align:middle line:84%
So we're going to formalize
that a bit more first on what

00:19:10.080 --> 00:19:12.740 align:middle line:84%
computing a Nash equilibria
actually looks like.

00:19:12.740 --> 00:19:14.648 align:middle line:84%
And we're going to
talk about Nash here

00:19:14.648 --> 00:19:15.940 align:middle line:90%
because it's very well-studied.

00:19:15.940 --> 00:19:19.840 align:middle line:84%
The non-cooperative equilibria,
because of the coalitions,

00:19:19.840 --> 00:19:21.460 align:middle line:90%
is at least as hard as Nash.

00:19:21.460 --> 00:19:25.280 align:middle line:84%
So anything that holds here for
Nash will hold for that as well.

00:19:25.280 --> 00:19:28.460 align:middle line:84%
Nash equilibria to compute
them in algorithms--

00:19:28.460 --> 00:19:30.500 align:middle line:90%
it's PPAD-complete.

00:19:30.500 --> 00:19:33.620 align:middle line:84%
And we'll describe what that
means exactly in a minute.

00:19:33.620 --> 00:19:36.080 align:middle line:84%
But importantly,
all known algorithms

00:19:36.080 --> 00:19:37.890 align:middle line:90%
are exponential algorithms.

00:19:37.890 --> 00:19:39.640 align:middle line:84%
So there are no
polynomial time algorithms

00:19:39.640 --> 00:19:43.200 align:middle line:84%
that can compute Nash equilibria
in the numbers of players

00:19:43.200 --> 00:19:45.520 align:middle line:90%
and strategies.

00:19:45.520 --> 00:19:50.200 align:middle line:84%
Competitive equilibria are also
in the same complexity class,

00:19:50.200 --> 00:19:52.080 align:middle line:84%
PPAD-complete for general
exchange economies,

00:19:52.080 --> 00:19:53.640 align:middle line:84%
so we can't get
around it by finding

00:19:53.640 --> 00:19:55.920 align:middle line:84%
the competitive equilibria
and working backwards

00:19:55.920 --> 00:19:57.220 align:middle line:90%
to see if it's non-cooperative.

00:19:57.220 --> 00:19:59.280 align:middle line:90%
That won't work either.

00:19:59.280 --> 00:20:02.680 align:middle line:84%
So again, an agent faced with
many goods and continuous prices

00:20:02.680 --> 00:20:05.560 align:middle line:84%
has a very complex optimization
problem that computers

00:20:05.560 --> 00:20:07.680 align:middle line:90%
take exponential time to solve.

00:20:07.680 --> 00:20:11.320 align:middle line:84%
So convergence is very
dubious, and it's very hard

00:20:11.320 --> 00:20:13.640 align:middle line:84%
to find a situation in
which a rational agent will

00:20:13.640 --> 00:20:16.520 align:middle line:84%
be able to make this
optimization themselves.

00:20:16.520 --> 00:20:18.023 align:middle line:90%
AUDIENCE: Do we know?

00:20:18.023 --> 00:20:19.440 align:middle line:84%
Are there a lot
of Nash equilibria

00:20:19.440 --> 00:20:23.680 align:middle line:90%
that are not non-cooperative?

00:20:23.680 --> 00:20:26.330 align:middle line:84%
Because every
non-cooperative equilibrium

00:20:26.330 --> 00:20:28.570 align:middle line:84%
is a Nash equilibrium
with the extra requirement

00:20:28.570 --> 00:20:32.810 align:middle line:84%
that when you deviate, there
is no one that's worse off.

00:20:32.810 --> 00:20:36.130 align:middle line:84%
So then are there lots
of Nash equilibria

00:20:36.130 --> 00:20:38.555 align:middle line:90%
that would be non-cooperative?

00:20:38.555 --> 00:20:39.930 align:middle line:84%
Because I mean,
presumably, those

00:20:39.930 --> 00:20:41.750 align:middle line:84%
could arise to for
whatever reason.

00:20:41.750 --> 00:20:45.810 align:middle line:84%
SAMUEL BRUCE: Yes, so
from the definitions,

00:20:45.810 --> 00:20:51.450 align:middle line:84%
all Nash equilibria are
non-competitive and vice versa--

00:20:51.450 --> 00:20:53.890 align:middle line:84%
sorry, yeah,
non-cooperative, sorry.

00:20:53.890 --> 00:20:57.330 align:middle line:84%
They're the same in the
instance that the deviation

00:20:57.330 --> 00:20:59.570 align:middle line:90%
player is a single player.

00:20:59.570 --> 00:21:02.530 align:middle line:84%
The coalition,
which we talk about

00:21:02.530 --> 00:21:05.130 align:middle line:84%
to get to our strong
non-cooperative equilibria,

00:21:05.130 --> 00:21:07.410 align:middle line:90%
is the added constraint.

00:21:07.410 --> 00:21:11.010 align:middle line:84%
And that would be a specific
subset of Nash equilibria.

00:21:11.010 --> 00:21:13.930 align:middle line:84%
AUDIENCE: But even if
the deviating coalition

00:21:13.930 --> 00:21:17.265 align:middle line:84%
is a single player, I mean,
in a Nash equilibrium,

00:21:17.265 --> 00:21:19.890 align:middle line:84%
a single player can deviate, and
it could make one other player

00:21:19.890 --> 00:21:20.750 align:middle line:90%
worse off.

00:21:20.750 --> 00:21:25.650 align:middle line:90%
So it's not necessarily right.

00:21:25.650 --> 00:21:29.370 align:middle line:84%
So it doesn't mean-- like it
doesn't have to be the case that

00:21:29.370 --> 00:21:30.188 align:middle line:90%
it is--

00:21:30.188 --> 00:21:31.230 align:middle line:90%
what were you calling it?

00:21:31.230 --> 00:21:32.230 align:middle line:90%
What are we calling it?

00:21:32.230 --> 00:21:34.770 align:middle line:90%


00:21:34.770 --> 00:21:37.090 align:middle line:90%
Efficient or Pareto-improving.

00:21:37.090 --> 00:21:45.050 align:middle line:84%
AUDIENCE: The Bayes' theorem is
that the selfish individual Nash

00:21:45.050 --> 00:21:47.650 align:middle line:84%
equilibrium without
the extra criteria

00:21:47.650 --> 00:21:53.370 align:middle line:84%
for the game he
specifies will be

00:21:53.370 --> 00:21:56.063 align:middle line:84%
the Walrasian
competitive outcome.

00:21:56.063 --> 00:21:56.730 align:middle line:90%
AUDIENCE: I see.

00:21:56.730 --> 00:21:57.230 align:middle line:90%
OK.

00:21:57.230 --> 00:21:59.010 align:middle line:84%
So that is a result
for this game then.

00:21:59.010 --> 00:22:01.330 align:middle line:84%
SAMUEL BRUCE: Yes, it's
specific to the mechanism.

00:22:01.330 --> 00:22:02.930 align:middle line:84%
AUDIENCE: The corollary
is competitive

00:22:02.930 --> 00:22:05.810 align:middle line:84%
equilibria are in
the core so they

00:22:05.810 --> 00:22:08.190 align:middle line:90%
can't be blocked by any subset.

00:22:08.190 --> 00:22:11.970 align:middle line:84%
So the spirit at least
of the extra condition

00:22:11.970 --> 00:22:16.490 align:middle line:84%
is satisfied once we get
to the Walrasian outcome.

00:22:16.490 --> 00:22:17.390 align:middle line:90%
AUDIENCE: OK.

00:22:17.390 --> 00:22:19.610 align:middle line:84%
Does it have this
result because it

00:22:19.610 --> 00:22:21.570 align:middle line:84%
assumes that they
can only submit

00:22:21.570 --> 00:22:23.742 align:middle line:90%
one price and one quantity?

00:22:23.742 --> 00:22:24.700 align:middle line:90%
AUDIENCE: I don't know.

00:22:24.700 --> 00:22:26.200 align:middle line:90%
SAMUEL BRUCE: Likely.

00:22:26.200 --> 00:22:28.820 align:middle line:90%


00:22:28.820 --> 00:22:30.760 align:middle line:90%
So just some definitions here.

00:22:30.760 --> 00:22:34.620 align:middle line:84%
Before we go into the
formal definition for PPAD.

00:22:34.620 --> 00:22:38.260 align:middle line:84%
General complexity
theory NP is the class

00:22:38.260 --> 00:22:43.020 align:middle line:84%
of problems that have
polynomial time verifiers.

00:22:43.020 --> 00:22:44.920 align:middle line:84%
And what that
means, essentially,

00:22:44.920 --> 00:22:49.140 align:middle line:84%
is that given the problem
and a potential solution,

00:22:49.140 --> 00:22:51.940 align:middle line:84%
you can find in polynomial
time whether or not

00:22:51.940 --> 00:22:53.500 align:middle line:90%
the solution is correct.

00:22:53.500 --> 00:22:57.100 align:middle line:84%
So it may take exponential
time to search and find

00:22:57.100 --> 00:22:58.242 align:middle line:90%
a correct solution.

00:22:58.242 --> 00:22:59.700 align:middle line:84%
But given a potential
solution, you

00:22:59.700 --> 00:23:01.880 align:middle line:84%
can say yes or no
in polynomial time.

00:23:01.880 --> 00:23:04.060 align:middle line:90%
That's the class NP.

00:23:04.060 --> 00:23:07.420 align:middle line:84%
A hard problem in
a complexity class

00:23:07.420 --> 00:23:09.340 align:middle line:84%
denotes any problem
that is known

00:23:09.340 --> 00:23:12.100 align:middle line:84%
to be at least as
difficult to solve

00:23:12.100 --> 00:23:15.620 align:middle line:84%
as any other problem in
the complexity class.

00:23:15.620 --> 00:23:19.060 align:middle line:84%
And then further,
a complete problem

00:23:19.060 --> 00:23:22.910 align:middle line:84%
is a problem in the class
where any problem can

00:23:22.910 --> 00:23:25.030 align:middle line:90%
be mapped to that problem.

00:23:25.030 --> 00:23:26.590 align:middle line:84%
And what that
essentially means is

00:23:26.590 --> 00:23:29.650 align:middle line:84%
if you have a complete
problem in a class,

00:23:29.650 --> 00:23:34.990 align:middle line:84%
you can take any other potential
problem, and in polynomial time,

00:23:34.990 --> 00:23:36.990 align:middle line:84%
a solution to the
complete problem

00:23:36.990 --> 00:23:40.550 align:middle line:84%
can be turned into a solution
to the other problem.

00:23:40.550 --> 00:23:42.610 align:middle line:84%
So a little bit of
nomenclature there.

00:23:42.610 --> 00:23:45.630 align:middle line:84%
It's worth noting all
complete problems are hard.

00:23:45.630 --> 00:23:51.030 align:middle line:84%
So for a problem to be a map to
any other problem in the class,

00:23:51.030 --> 00:23:54.070 align:middle line:84%
it must be at least as hard as
any other problem in the class,

00:23:54.070 --> 00:23:56.750 align:middle line:84%
and there must be
some algorithm that

00:23:56.750 --> 00:23:59.670 align:middle line:84%
allows it to be converted
quickly from a solution of 1

00:23:59.670 --> 00:24:02.350 align:middle line:90%
to a solution of another.

00:24:02.350 --> 00:24:04.290 align:middle line:84%
And a lot of times
in complexity theory,

00:24:04.290 --> 00:24:06.553 align:middle line:84%
we're looking for these
complete problems.

00:24:06.553 --> 00:24:08.470 align:middle line:84%
Because if you can prove
a problem is complete

00:24:08.470 --> 00:24:11.830 align:middle line:84%
for a complexity class, you
know it cannot be simplified any

00:24:11.830 --> 00:24:14.370 align:middle line:84%
more, and it cannot possibly
be more complicated.

00:24:14.370 --> 00:24:16.870 align:middle line:84%
So it's a very tight bound
on exactly how difficult

00:24:16.870 --> 00:24:18.776 align:middle line:84%
the problem is to
solve, because you

00:24:18.776 --> 00:24:22.670 align:middle line:84%
know it exists as a full
member of that class,

00:24:22.670 --> 00:24:25.570 align:middle line:84%
and we're going to
get to it in a second.

00:24:25.570 --> 00:24:27.070 align:middle line:84%
But Nash equilibria
are proven to be

00:24:27.070 --> 00:24:28.650 align:middle line:90%
complete for this class, PPAD.

00:24:28.650 --> 00:24:30.910 align:middle line:84%
And that's what we're
going to talk about here.

00:24:30.910 --> 00:24:34.030 align:middle line:84%
Long, long time ago, Nash
proved that every game

00:24:34.030 --> 00:24:37.130 align:middle line:84%
has a Nash equilibrium, whether
in pure or mixed strategies.

00:24:37.130 --> 00:24:40.830 align:middle line:84%
There is guaranteed to be a
Nash equilibria of every game,

00:24:40.830 --> 00:24:44.790 align:middle line:84%
and simple classes of
games, such as two-player,

00:24:44.790 --> 00:24:47.710 align:middle line:84%
zero-sum games and
some graphical games,

00:24:47.710 --> 00:24:49.970 align:middle line:84%
they have polynomial
algorithms to find them.

00:24:49.970 --> 00:24:53.450 align:middle line:84%
And there's minimax theorems
that you can look into.

00:24:53.450 --> 00:24:55.930 align:middle line:84%
And these are in very
specific situations,

00:24:55.930 --> 00:24:59.070 align:middle line:84%
great, efficient algorithms
for solving Nash equilibria.

00:24:59.070 --> 00:25:01.210 align:middle line:84%
But in the case
of general games--

00:25:01.210 --> 00:25:05.750 align:middle line:84%
so n players each having their
own strategy set and arbitrary

00:25:05.750 --> 00:25:09.230 align:middle line:84%
utility functions that maps
from submitted strategies

00:25:09.230 --> 00:25:11.390 align:middle line:90%
to utilities--

00:25:11.390 --> 00:25:12.890 align:middle line:90%
it's much more complicated.

00:25:12.890 --> 00:25:14.910 align:middle line:84%
So they started
looking into this

00:25:14.910 --> 00:25:16.910 align:middle line:90%
a long-- maybe 30 years ago.

00:25:16.910 --> 00:25:21.020 align:middle line:84%
And for this idea, they
created a new complexity class,

00:25:21.020 --> 00:25:24.367 align:middle line:84%
which is called TFNP
or NP total functions.

00:25:24.367 --> 00:25:26.200 align:middle line:84%
And this is the class
of all search problems

00:25:26.200 --> 00:25:27.940 align:middle line:84%
that are guaranteed
to have a solution.

00:25:27.940 --> 00:25:30.440 align:middle line:84%
So remember, Nash fits in there
because we know every Nash

00:25:30.440 --> 00:25:31.380 align:middle line:90%
equilibrium.

00:25:31.380 --> 00:25:33.600 align:middle line:84%
Every game has a
Nash equilibrium.

00:25:33.600 --> 00:25:35.160 align:middle line:84%
So we create this
complexity class

00:25:35.160 --> 00:25:38.400 align:middle line:84%
where it's search problems, but
there's a guaranteed solution.

00:25:38.400 --> 00:25:41.360 align:middle line:84%
However, this is
a semantic class,

00:25:41.360 --> 00:25:44.960 align:middle line:84%
which just means that there
are no complete problems

00:25:44.960 --> 00:25:45.700 align:middle line:90%
for the class.

00:25:45.700 --> 00:25:47.840 align:middle line:84%
It's too general
of a description.

00:25:47.840 --> 00:25:50.040 align:middle line:84%
Games-- search games with
a guaranteed solution

00:25:50.040 --> 00:25:54.080 align:middle line:84%
is too general to really fit
one problem to define the class.

00:25:54.080 --> 00:25:56.860 align:middle line:84%
So we break it up into
some more specific classes.

00:25:56.860 --> 00:26:00.040 align:middle line:84%
And one of those
classes, PPAD, which

00:26:00.040 --> 00:26:04.160 align:middle line:84%
stands for polynomial
parity for directed

00:26:04.160 --> 00:26:08.560 align:middle line:84%
graphs-- polynomial parity
argument in directed graphs.

00:26:08.560 --> 00:26:13.520 align:middle line:84%
And it's defined by the complete
problem in any directed graph

00:26:13.520 --> 00:26:15.180 align:middle line:90%
with one unbalanced node.

00:26:15.180 --> 00:26:18.860 align:middle line:84%
So you can picture some
arbitrary directed graph,

00:26:18.860 --> 00:26:20.660 align:middle line:84%
and there's one node
that's unbalanced,

00:26:20.660 --> 00:26:22.800 align:middle line:84%
meaning that they're
in degree to the node

00:26:22.800 --> 00:26:24.640 align:middle line:84%
is different than
their outdegree.

00:26:24.640 --> 00:26:27.640 align:middle line:84%
There's guaranteed to be
another node in that graph that

00:26:27.640 --> 00:26:29.360 align:middle line:90%
is also unbalanced.

00:26:29.360 --> 00:26:32.200 align:middle line:84%
You can imagine if some
node has two in and one

00:26:32.200 --> 00:26:34.940 align:middle line:84%
out that can't
come from nowhere.

00:26:34.940 --> 00:26:36.680 align:middle line:84%
There must also be
a node somewhere

00:26:36.680 --> 00:26:40.960 align:middle line:84%
that has, say, one in and two
outs or some other combination

00:26:40.960 --> 00:26:43.820 align:middle line:84%
of nodes that allows
us to get unbalanced.

00:26:43.820 --> 00:26:46.480 align:middle line:84%
So you can't have an
unbalanced node by itself.

00:26:46.480 --> 00:26:49.520 align:middle line:84%
This is defined as
the class PPAD finding

00:26:49.520 --> 00:26:52.400 align:middle line:90%
that other node given one.

00:26:52.400 --> 00:26:56.560 align:middle line:84%
And Nash equilibria was proven
to be a member of this class

00:26:56.560 --> 00:27:00.440 align:middle line:84%
by Daskalakis, Goldberg,
and Papadimitriou.

00:27:00.440 --> 00:27:02.440 align:middle line:84%
Now, I won't take you
through the proof of this

00:27:02.440 --> 00:27:05.520 align:middle line:84%
because it's very long
and very complicated.

00:27:05.520 --> 00:27:07.680 align:middle line:84%
But it's just the
intuition here is

00:27:07.680 --> 00:27:12.440 align:middle line:84%
what matters that we have
found like algorithms for Nash,

00:27:12.440 --> 00:27:15.750 align:middle line:84%
and they are specifically
very hard and complete,

00:27:15.750 --> 00:27:18.730 align:middle line:84%
meaning that no simpler
algorithm can exist.

00:27:18.730 --> 00:27:21.090 align:middle line:84%
And it will always be
exponential algorithms

00:27:21.090 --> 00:27:26.610 align:middle line:84%
to find Nash equilibria
under most cases.

00:27:26.610 --> 00:27:31.490 align:middle line:84%
So if we know Nash equilibria
are extremely hard to compute

00:27:31.490 --> 00:27:34.050 align:middle line:84%
and we still want to find
the solutions to our game

00:27:34.050 --> 00:27:38.090 align:middle line:84%
and find areas where agents
will be able to reasonably find

00:27:38.090 --> 00:27:41.130 align:middle line:84%
equilibria, we can start
looking at alternative types

00:27:41.130 --> 00:27:43.755 align:middle line:84%
of equilibria because
there's more than just Nash.

00:27:43.755 --> 00:27:45.130 align:middle line:84%
Specifically,
we're going to talk

00:27:45.130 --> 00:27:47.890 align:middle line:84%
about equilibrium concepts
that require coordination

00:27:47.890 --> 00:27:49.110 align:middle line:90%
among the players.

00:27:49.110 --> 00:27:50.930 align:middle line:84%
Remember, in Nash
equilibria, each player

00:27:50.930 --> 00:27:52.890 align:middle line:90%
is acting on their own.

00:27:52.890 --> 00:27:58.490 align:middle line:84%
However, if we add a
coordinator to the game,

00:27:58.490 --> 00:27:59.950 align:middle line:90%
it might make things simpler.

00:27:59.950 --> 00:28:02.530 align:middle line:84%
So instead of each player
computing their optimal strategy

00:28:02.530 --> 00:28:05.010 align:middle line:84%
profile, we're going to
add a central coordinator

00:28:05.010 --> 00:28:07.530 align:middle line:84%
to the game who
chooses the strategies

00:28:07.530 --> 00:28:09.530 align:middle line:90%
of each player for them.

00:28:09.530 --> 00:28:13.020 align:middle line:84%
So the coordinator keeps a
joint probability distribution

00:28:13.020 --> 00:28:15.380 align:middle line:90%
across all player strategies--

00:28:15.380 --> 00:28:17.220 align:middle line:90%
call that p.

00:28:17.220 --> 00:28:19.380 align:middle line:84%
And for each
iteration of the game,

00:28:19.380 --> 00:28:21.180 align:middle line:84%
the coordinator is
just going to draw

00:28:21.180 --> 00:28:24.260 align:middle line:84%
from that distribution at
random, select some strategy

00:28:24.260 --> 00:28:26.460 align:middle line:84%
profile for all
the players, and is

00:28:26.460 --> 00:28:30.340 align:middle line:84%
going to tell each player
their selected strategy.

00:28:30.340 --> 00:28:34.180 align:middle line:84%
Each player on their own knows
the underlying distribution p,

00:28:34.180 --> 00:28:37.540 align:middle line:84%
and they know their
selected strategy.

00:28:37.540 --> 00:28:40.588 align:middle line:84%
If I'm the coordinator, I
could tell you, like, oh, I'm

00:28:40.588 --> 00:28:41.880 align:middle line:90%
drawing from this distribution.

00:28:41.880 --> 00:28:43.780 align:middle line:84%
Here's what I'm
telling you to do.

00:28:43.780 --> 00:28:47.300 align:middle line:84%
I'm not going to tell you what
I told the other players to do.

00:28:47.300 --> 00:28:49.620 align:middle line:84%
So you only learn
your own strategy.

00:28:49.620 --> 00:28:53.080 align:middle line:84%
However, by knowing the
underlying distribution,

00:28:53.080 --> 00:28:56.620 align:middle line:84%
you can then reason about
what other players might

00:28:56.620 --> 00:28:59.220 align:middle line:90%
have been signaled to play.

00:28:59.220 --> 00:29:01.940 align:middle line:84%
Under our coordinator
mechanism, we

00:29:01.940 --> 00:29:04.260 align:middle line:84%
will call it a
correlated equilibrium

00:29:04.260 --> 00:29:07.220 align:middle line:84%
when each player, knowing
their signaled strategy

00:29:07.220 --> 00:29:12.140 align:middle line:84%
and the underlying distribution,
has no incentive to not play

00:29:12.140 --> 00:29:13.980 align:middle line:90%
the strategy they were given.

00:29:13.980 --> 00:29:18.140 align:middle line:84%
So in other words,
in our notation here,

00:29:18.140 --> 00:29:21.220 align:middle line:84%
over all possible strategies
of the other players,

00:29:21.220 --> 00:29:25.500 align:middle line:84%
S negative I here, your utility
of playing your selected

00:29:25.500 --> 00:29:28.580 align:middle line:84%
strategy given they're
playing that strategy

00:29:28.580 --> 00:29:30.820 align:middle line:84%
times the probability
of that happening

00:29:30.820 --> 00:29:35.780 align:middle line:84%
must be greater than some
fixed deviation S prime

00:29:35.780 --> 00:29:38.460 align:middle line:84%
across again all other
possible strategies

00:29:38.460 --> 00:29:41.840 align:middle line:84%
of each player weighted by the
probability of that happening.

00:29:41.840 --> 00:29:43.700 align:middle line:84%
So each player must
individually be

00:29:43.700 --> 00:29:46.500 align:middle line:84%
incentive compatible
to play their strategy.

00:29:46.500 --> 00:29:49.780 align:middle line:84%
When that's the case, we'll call
it a correlated equilibrium.

00:29:49.780 --> 00:29:54.140 align:middle line:84%
Now each player here faces
the number of strategies

00:29:54.140 --> 00:29:55.440 align:middle line:90%
squared constraints.

00:29:55.440 --> 00:29:58.140 align:middle line:84%
And you can see that
here because you

00:29:58.140 --> 00:30:01.180 align:middle line:84%
could get any signal that's
in your strategy set.

00:30:01.180 --> 00:30:02.680 align:middle line:84%
So that's your
number of strategies.

00:30:02.680 --> 00:30:04.420 align:middle line:84%
And for each of
those signals, it

00:30:04.420 --> 00:30:08.540 align:middle line:84%
must be at least as good
as any pure deviation.

00:30:08.540 --> 00:30:12.070 align:middle line:84%
So for each strategy,
we have again the number

00:30:12.070 --> 00:30:14.070 align:middle line:84%
of strategy conditions--
so number of strategies

00:30:14.070 --> 00:30:17.270 align:middle line:84%
squared for our total
constraints on each player.

00:30:17.270 --> 00:30:20.470 align:middle line:84%
We call it correlated
because knowing your strategy

00:30:20.470 --> 00:30:22.510 align:middle line:84%
gives you information on
the possible strategies

00:30:22.510 --> 00:30:23.710 align:middle line:90%
of other players.

00:30:23.710 --> 00:30:26.390 align:middle line:84%
So with the coordinator, you're
able to correlate players

00:30:26.390 --> 00:30:29.750 align:middle line:84%
together so that the
joint distribution favors

00:30:29.750 --> 00:30:32.470 align:middle line:84%
playing certain actions
together as a group

00:30:32.470 --> 00:30:36.870 align:middle line:84%
amongst multiple players, even
when individually they may not

00:30:36.870 --> 00:30:39.630 align:middle line:90%
all come to that conclusion.

00:30:39.630 --> 00:30:42.902 align:middle line:84%
So this is an example
here, and we'll

00:30:42.902 --> 00:30:45.110 align:middle line:84%
keep bringing this one back
up throughout the lecture

00:30:45.110 --> 00:30:47.090 align:middle line:84%
to demonstrate
different equilibria.

00:30:47.090 --> 00:30:49.430 align:middle line:90%
But we have player one here.

00:30:49.430 --> 00:30:51.170 align:middle line:84%
Their utility matrix
is on the left.

00:30:51.170 --> 00:30:55.070 align:middle line:84%
And they choose between top and
bottom as their two strategies.

00:30:55.070 --> 00:30:57.770 align:middle line:84%
Player 2 on the right chooses
between left and right,

00:30:57.770 --> 00:30:59.470 align:middle line:84%
and this is their
utility matrix.

00:30:59.470 --> 00:31:02.210 align:middle line:84%
And then D here just
denotes the distribution.

00:31:02.210 --> 00:31:05.230 align:middle line:84%
So each p is the
probability that--

00:31:05.230 --> 00:31:07.350 align:middle line:84%
the joint probability
of that outcome.

00:31:07.350 --> 00:31:10.390 align:middle line:84%
So PTL is the probability
that player 1 plays t,

00:31:10.390 --> 00:31:13.150 align:middle line:84%
and player 2 plays L
and so on and so forth.

00:31:13.150 --> 00:31:14.550 align:middle line:84%
So this is what
those constraints

00:31:14.550 --> 00:31:18.510 align:middle line:84%
would look like for each
player in this situation.

00:31:18.510 --> 00:31:23.510 align:middle line:84%
So player one in
this row here, this

00:31:23.510 --> 00:31:28.550 align:middle line:84%
is the strategy of playing their
signal when their signal to play

00:31:28.550 --> 00:31:29.570 align:middle line:90%
the top row.

00:31:29.570 --> 00:31:33.350 align:middle line:84%
So weighted average
of what is probably

00:31:33.350 --> 00:31:36.470 align:middle line:84%
going to happen if they play the
top row-- so their utility when

00:31:36.470 --> 00:31:38.327 align:middle line:84%
they play top row
and the other player

00:31:38.327 --> 00:31:40.910 align:middle line:84%
plays left times the probability
the other player was signaled

00:31:40.910 --> 00:31:43.070 align:middle line:84%
to play left plus
their probability

00:31:43.070 --> 00:31:45.910 align:middle line:84%
when the other player plays
right times the probability

00:31:45.910 --> 00:31:47.390 align:middle line:84%
that they were
told to play right.

00:31:47.390 --> 00:31:49.450 align:middle line:84%
And we compare it
to any deviation.

00:31:49.450 --> 00:31:53.150 align:middle line:84%
So the deviation here for player
one, they were told to play top.

00:31:53.150 --> 00:31:55.390 align:middle line:84%
They could deviate
to play bottom.

00:31:55.390 --> 00:31:57.090 align:middle line:84%
And now the probabilities
are the same.

00:31:57.090 --> 00:31:59.010 align:middle line:84%
Because you were still
told to play top.

00:31:59.010 --> 00:32:00.510 align:middle line:84%
So it's still the
same probabilities

00:32:00.510 --> 00:32:03.027 align:middle line:84%
that the other player was told
to play left versus right.

00:32:03.027 --> 00:32:05.110 align:middle line:84%
But your utilities would
change because now you're

00:32:05.110 --> 00:32:06.140 align:middle line:90%
playing bottom.

00:32:06.140 --> 00:32:10.147 align:middle line:84%
So if the coordinator said
you play top, they play left,

00:32:10.147 --> 00:32:10.980 align:middle line:90%
but you play bottom.

00:32:10.980 --> 00:32:12.220 align:middle line:90%
You get this utility instead.

00:32:12.220 --> 00:32:15.680 align:middle line:84%
That's why it's here, and
then likewise with the right.

00:32:15.680 --> 00:32:18.280 align:middle line:84%
We do the same thing
for player 1 being

00:32:18.280 --> 00:32:20.840 align:middle line:84%
told to play the bottom
strategy, player 2 being

00:32:20.840 --> 00:32:24.040 align:middle line:84%
told to play the left strategy
and the right strategy.

00:32:24.040 --> 00:32:27.080 align:middle line:84%
And if all these
constraints hold,

00:32:27.080 --> 00:32:33.760 align:middle line:84%
then the distribution D is
a coordinated equilibrium.

00:32:33.760 --> 00:32:36.200 align:middle line:84%
Moving from
correlated equilibria,

00:32:36.200 --> 00:32:38.700 align:middle line:84%
we're going to talk next about
coarse correlated equilibria,

00:32:38.700 --> 00:32:42.320 align:middle line:84%
which is an even weaker
version of equilibria.

00:32:42.320 --> 00:32:44.040 align:middle line:84%
So we're in the
same setting where

00:32:44.040 --> 00:32:45.840 align:middle line:84%
we have our coordinator
with a distribution

00:32:45.840 --> 00:32:47.920 align:middle line:90%
drawing strategies for players.

00:32:47.920 --> 00:32:50.720 align:middle line:84%
However, now the
players do not get

00:32:50.720 --> 00:32:55.000 align:middle line:84%
to learn their selected
strategy before committing

00:32:55.000 --> 00:32:56.280 align:middle line:90%
to the mechanism.

00:32:56.280 --> 00:32:58.440 align:middle line:84%
So instead of having
the information of what

00:32:58.440 --> 00:33:02.080 align:middle line:84%
you were told to play
and using that to guess

00:33:02.080 --> 00:33:04.850 align:middle line:84%
at what other players might
have been told to play,

00:33:04.850 --> 00:33:07.070 align:middle line:84%
now it's just the
case that you must,

00:33:07.070 --> 00:33:11.690 align:middle line:84%
in expectation over all
possible strategies,

00:33:11.690 --> 00:33:14.530 align:middle line:84%
be better off than
some fixed deviation

00:33:14.530 --> 00:33:17.090 align:middle line:90%
for all possible strategies.

00:33:17.090 --> 00:33:18.930 align:middle line:84%
So prior to learning
their strategy,

00:33:18.930 --> 00:33:21.370 align:middle line:84%
each player i must be
incentive compatible.

00:33:21.370 --> 00:33:23.850 align:middle line:84%
And walking through
the notation here,

00:33:23.850 --> 00:33:28.210 align:middle line:84%
this term on the right,
this sum, is just the sum.

00:33:28.210 --> 00:33:31.730 align:middle line:84%
It's the total probability
of S negative i

00:33:31.730 --> 00:33:34.510 align:middle line:84%
because it's summing over
all of your strategies

00:33:34.510 --> 00:33:36.730 align:middle line:84%
the probability of you
getting that strategy

00:33:36.730 --> 00:33:38.210 align:middle line:90%
and then getting that one.

00:33:38.210 --> 00:33:41.330 align:middle line:84%
So this is the total probability
of the other players playing

00:33:41.330 --> 00:33:42.370 align:middle line:90%
S negative i.

00:33:42.370 --> 00:33:45.970 align:middle line:84%
We multiply that times the
utility for the deviation,

00:33:45.970 --> 00:33:48.330 align:middle line:84%
and we sum over all
possible other strategies

00:33:48.330 --> 00:33:51.430 align:middle line:84%
to get the expected
utility of any deviation.

00:33:51.430 --> 00:33:53.770 align:middle line:90%
That's what this right term is.

00:33:53.770 --> 00:33:56.550 align:middle line:84%
If the expectation of
playing the distribution,

00:33:56.550 --> 00:34:00.930 align:middle line:84%
as you're told, is better or
equal to some fixed deviation,

00:34:00.930 --> 00:34:04.170 align:middle line:84%
we'll call it a coarse
correlated equilibria.

00:34:04.170 --> 00:34:07.690 align:middle line:84%
And here the number of
constraints is actually lesser.

00:34:07.690 --> 00:34:11.610 align:middle line:84%
So because it's not conditioning
on your own strategies,

00:34:11.610 --> 00:34:14.929 align:middle line:84%
there's only O of number
of strategies constraints

00:34:14.929 --> 00:34:17.230 align:middle line:84%
because it's just your
total expectation.

00:34:17.230 --> 00:34:21.010 align:middle line:84%
Compare it to all
possible fixed deviations.

00:34:21.010 --> 00:34:23.730 align:middle line:84%
So we'll go back to
our example here.

00:34:23.730 --> 00:34:26.810 align:middle line:84%
We have the same utility
matrices and distribution.

00:34:26.810 --> 00:34:28.730 align:middle line:84%
And our constraints
are similar, but they

00:34:28.730 --> 00:34:29.790 align:middle line:90%
look slightly different.

00:34:29.790 --> 00:34:32.090 align:middle line:84%
So the top row for
each player here

00:34:32.090 --> 00:34:34.810 align:middle line:84%
is just denoting
their expected--

00:34:34.810 --> 00:34:36.570 align:middle line:84%
the expected utility
that they receive

00:34:36.570 --> 00:34:38.170 align:middle line:90%
from playing the distribution.

00:34:38.170 --> 00:34:42.010 align:middle line:84%
So intuitively for player 1,
it's three times the probability

00:34:42.010 --> 00:34:44.030 align:middle line:84%
that this box is chosen
plus 1, that one,

00:34:44.030 --> 00:34:45.250 align:middle line:90%
and so on and so forth.

00:34:45.250 --> 00:34:46.350 align:middle line:90%
Player 2, it's similar.

00:34:46.350 --> 00:34:48.630 align:middle line:84%
It's just their utility
matrix times the probability.

00:34:48.630 --> 00:34:50.130 align:middle line:90%
That's where they end up.

00:34:50.130 --> 00:34:53.690 align:middle line:84%
And in each case, for
each player that expected

00:34:53.690 --> 00:34:57.730 align:middle line:84%
utility needs to be
greater than any deviation

00:34:57.730 --> 00:34:58.710 align:middle line:90%
to a fixed strategy.

00:34:58.710 --> 00:35:01.820 align:middle line:84%
So for here that
looks like three--

00:35:01.820 --> 00:35:05.660 align:middle line:84%
if they deviate to the top
row, which is this one here,

00:35:05.660 --> 00:35:09.020 align:middle line:84%
you have the probability that
any left is signaled for player

00:35:09.020 --> 00:35:15.520 align:middle line:84%
2 times you play the top fixed
plus any deviation-- or sorry,

00:35:15.520 --> 00:35:17.500 align:middle line:84%
any time the right
player is told

00:35:17.500 --> 00:35:22.660 align:middle line:84%
to play right times your
top deviation utility.

00:35:22.660 --> 00:35:27.060 align:middle line:84%
So again, constraints to make
sure that you want to play as

00:35:27.060 --> 00:35:27.680 align:middle line:90%
told.

00:35:27.680 --> 00:35:30.220 align:middle line:84%
In this case, it's
just in expectation.

00:35:30.220 --> 00:35:33.900 align:middle line:84%
For each player, we
have these constraints,

00:35:33.900 --> 00:35:39.140 align:middle line:84%
and if they all hold, then it's
a coarse correlated equilibria.

00:35:39.140 --> 00:35:41.540 align:middle line:84%
Some important things
about correlated and coarse

00:35:41.540 --> 00:35:42.980 align:middle line:90%
correlated equilibria.

00:35:42.980 --> 00:35:47.660 align:middle line:84%
First, all Nash
equilibria are correlated.

00:35:47.660 --> 00:35:49.960 align:middle line:84%
They're just a subset of
correlated equilibria.

00:35:49.960 --> 00:35:52.180 align:middle line:84%
And likewise, all
correlated are a subset

00:35:52.180 --> 00:35:53.360 align:middle line:90%
of the coarse correlated.

00:35:53.360 --> 00:35:55.880 align:middle line:84%
And you can think of that
in terms of the constraints.

00:35:55.880 --> 00:35:57.540 align:middle line:90%
They get weaker and weaker.

00:35:57.540 --> 00:36:01.060 align:middle line:84%
So anything that fits the
constraints of a Nash equilibria

00:36:01.060 --> 00:36:02.940 align:middle line:84%
is going to automatically
fit the constraints

00:36:02.940 --> 00:36:04.580 align:middle line:84%
of a correlated
and similarly for

00:36:04.580 --> 00:36:06.500 align:middle line:90%
coarse correlated equilibria.

00:36:06.500 --> 00:36:08.940 align:middle line:84%
Nash equilibrium is
actually a special case

00:36:08.940 --> 00:36:12.340 align:middle line:84%
of correlated equilibria
where the probabilities

00:36:12.340 --> 00:36:16.460 align:middle line:90%
for each player are independent.

00:36:16.460 --> 00:36:18.820 align:middle line:84%
So you could imagine
the coordinator

00:36:18.820 --> 00:36:20.980 align:middle line:84%
who has his underlying
probability distribution,

00:36:20.980 --> 00:36:23.740 align:middle line:84%
just builds it in such a
way that the probability

00:36:23.740 --> 00:36:26.620 align:middle line:84%
of any joint
distribution is just

00:36:26.620 --> 00:36:30.060 align:middle line:84%
the product of the probability
for each individual player

00:36:30.060 --> 00:36:31.420 align:middle line:90%
playing that strategy.

00:36:31.420 --> 00:36:36.460 align:middle line:84%
If there's no correlation
between the player strategies,

00:36:36.460 --> 00:36:40.780 align:middle line:84%
then it's a Nash equilibrium,
but still in the same space.

00:36:40.780 --> 00:36:42.340 align:middle line:84%
The difference
between the correlated

00:36:42.340 --> 00:36:43.798 align:middle line:84%
and the coarse
correlated again can

00:36:43.798 --> 00:36:47.660 align:middle line:84%
be thought of as a difference in
commitment and when you commit.

00:36:47.660 --> 00:36:49.190 align:middle line:84%
In the coarse
correlated equilibria,

00:36:49.190 --> 00:36:50.940 align:middle line:84%
it requires you to
commit to the mechanism

00:36:50.940 --> 00:36:53.060 align:middle line:84%
before you even
learn your strategy.

00:36:53.060 --> 00:36:55.700 align:middle line:84%
So you must think in expectation
the mechanism will be better

00:36:55.700 --> 00:36:58.470 align:middle line:84%
for you, but you don't know what
you're being told to play yet.

00:36:58.470 --> 00:37:01.430 align:middle line:84%
In the correlated equilibria,
you get to learn your own signal

00:37:01.430 --> 00:37:03.028 align:middle line:90%
and then make your decision.

00:37:03.028 --> 00:37:05.070 align:middle line:84%
So that's why it's a little
stricter, because you

00:37:05.070 --> 00:37:07.530 align:middle line:84%
must be able to learn your
signal and still think,

00:37:07.530 --> 00:37:09.830 align:middle line:90%
yeah, I want to play this.

00:37:09.830 --> 00:37:13.150 align:middle line:84%
Because the coordinator
gives you that strategy

00:37:13.150 --> 00:37:14.550 align:middle line:84%
and now you're
simply just trying

00:37:14.550 --> 00:37:18.190 align:middle line:84%
to find incentive-compatible
allocations,

00:37:18.190 --> 00:37:20.950 align:middle line:84%
it greatly reduces the
complexity for each player's

00:37:20.950 --> 00:37:22.170 align:middle line:90%
optimization problem.

00:37:22.170 --> 00:37:24.850 align:middle line:84%
Now instead of trying to predict
the strategies of others,

00:37:24.850 --> 00:37:26.990 align:middle line:84%
they know this
underlying distribution.

00:37:26.990 --> 00:37:28.790 align:middle line:84%
They can assume that
other players are

00:37:28.790 --> 00:37:31.430 align:middle line:84%
going to play the strategy
that they were told,

00:37:31.430 --> 00:37:33.990 align:middle line:84%
and they just have to form
these linear constraints

00:37:33.990 --> 00:37:37.190 align:middle line:84%
about when it's beneficial for
them to play their strategy

00:37:37.190 --> 00:37:38.790 align:middle line:90%
and when it's not.

00:37:38.790 --> 00:37:41.790 align:middle line:84%
So it really is
helpful when we're

00:37:41.790 --> 00:37:43.407 align:middle line:84%
talking about our
algorithm design

00:37:43.407 --> 00:37:45.990 align:middle line:84%
because we're no longer talking
about our exponential strategy

00:37:45.990 --> 00:37:50.190 align:middle line:84%
space, but just a very
linear space for each player.

00:37:50.190 --> 00:37:53.610 align:middle line:84%
One thing is it does require
trust in that coordinator.

00:37:53.610 --> 00:37:56.520 align:middle line:84%
You could imagine some
nefarious coordinator

00:37:56.520 --> 00:37:58.960 align:middle line:84%
that tells you their
distribution is p,

00:37:58.960 --> 00:38:01.240 align:middle line:84%
but it's really some
alternate distribution q.

00:38:01.240 --> 00:38:03.200 align:middle line:84%
And q is really favorable
to other players

00:38:03.200 --> 00:38:05.320 align:middle line:84%
and really favorable
against you.

00:38:05.320 --> 00:38:07.340 align:middle line:84%
You would not want to
play in that mechanism.

00:38:07.340 --> 00:38:09.720 align:middle line:84%
So you must believe
that you're receiving

00:38:09.720 --> 00:38:12.560 align:middle line:84%
fair draws from the distribution
that you're playing.

00:38:12.560 --> 00:38:15.440 align:middle line:84%
And our blockchain
implementations actually

00:38:15.440 --> 00:38:17.380 align:middle line:84%
kind of inherently
solve this issue for us

00:38:17.380 --> 00:38:20.800 align:middle line:84%
because smart contract code
deployed on the blockchain

00:38:20.800 --> 00:38:23.200 align:middle line:90%
is visible to all participants.

00:38:23.200 --> 00:38:26.120 align:middle line:84%
And you can actually see
by inspecting the code

00:38:26.120 --> 00:38:28.580 align:middle line:84%
on the contracts what
the distribution is,

00:38:28.580 --> 00:38:30.880 align:middle line:84%
making sure that they're
giving you fair draws.

00:38:30.880 --> 00:38:34.320 align:middle line:84%
And since every player gets
to learn the distribution,

00:38:34.320 --> 00:38:36.360 align:middle line:84%
it's revealing no
more information

00:38:36.360 --> 00:38:39.960 align:middle line:84%
to verify that the
mechanism is true.

00:38:39.960 --> 00:38:42.080 align:middle line:84%
Importantly, all
those constraints

00:38:42.080 --> 00:38:45.020 align:middle line:84%
for both the correlated and the
cross correlated equilibria,

00:38:45.020 --> 00:38:49.920 align:middle line:84%
they're linear combinations
of the strategy probabilities.

00:38:49.920 --> 00:38:52.400 align:middle line:84%
When our constraints
are all linear

00:38:52.400 --> 00:38:54.800 align:middle line:84%
in terms of our
variables, we can easily

00:38:54.800 --> 00:38:57.280 align:middle line:84%
find the correct
distributions that

00:38:57.280 --> 00:39:00.220 align:middle line:84%
solve our constraints with
linear programming methods.

00:39:00.220 --> 00:39:02.280 align:middle line:84%
These are polynomial
time algorithms

00:39:02.280 --> 00:39:05.000 align:middle line:84%
that allow us to put
in our constraints

00:39:05.000 --> 00:39:07.120 align:middle line:84%
and get out a solution,
a distribution that

00:39:07.120 --> 00:39:09.120 align:middle line:90%
will work for us.

00:39:09.120 --> 00:39:11.120 align:middle line:84%
Because we kind of did
the math earlier but

00:39:11.120 --> 00:39:13.680 align:middle line:84%
if we have n players
k strategies,

00:39:13.680 --> 00:39:16.480 align:middle line:84%
you're going to have O
of NK squared constraints

00:39:16.480 --> 00:39:19.200 align:middle line:84%
for correlated equilibria
or NK for coarse correlated

00:39:19.200 --> 00:39:20.238 align:middle line:90%
equilibria.

00:39:20.238 --> 00:39:22.280 align:middle line:84%
You have the added
constraints that probabilities

00:39:22.280 --> 00:39:25.960 align:middle line:84%
must sum to 1, that they
must be greater than zero.

00:39:25.960 --> 00:39:28.800 align:middle line:84%
But this is a very
reasonable number

00:39:28.800 --> 00:39:31.100 align:middle line:84%
of constraints for a
linear programming problem.

00:39:31.100 --> 00:39:34.080 align:middle line:84%
This is very easily handled
by many libraries that exist

00:39:34.080 --> 00:39:36.200 align:middle line:90%
and many algorithms.

00:39:36.200 --> 00:39:40.320 align:middle line:84%
And additionally, one great
aspect about linear programming

00:39:40.320 --> 00:39:43.000 align:middle line:84%
is at no extra
cost, you're allowed

00:39:43.000 --> 00:39:46.260 align:middle line:84%
to add in an optimization
to the problem.

00:39:46.260 --> 00:39:50.720 align:middle line:84%
So instead of just finding some
arbitrary distribution that

00:39:50.720 --> 00:39:53.770 align:middle line:84%
meets all of our incentive
compatibility constraints,

00:39:53.770 --> 00:39:56.330 align:middle line:84%
we can actually
maximize over something

00:39:56.330 --> 00:39:58.650 align:middle line:84%
in the set of
distributions that work.

00:39:58.650 --> 00:40:03.370 align:middle line:84%
So a classic one is the
Pareto planner's problem.

00:40:03.370 --> 00:40:05.570 align:middle line:84%
Given weights lambda
for each player,

00:40:05.570 --> 00:40:07.490 align:middle line:84%
you can maximize
the expected utility

00:40:07.490 --> 00:40:09.810 align:middle line:84%
for each player in a
lambda-weighted sum.

00:40:09.810 --> 00:40:12.530 align:middle line:84%
That's a very
typical maximization

00:40:12.530 --> 00:40:14.630 align:middle line:84%
you'd want to do as
some coordinator,

00:40:14.630 --> 00:40:17.250 align:middle line:84%
to try and maximize the
lambda-weighted sums

00:40:17.250 --> 00:40:18.710 align:middle line:90%
of utilities of players.

00:40:18.710 --> 00:40:21.190 align:middle line:84%
And you can do this in linear
programming at no extra cost.

00:40:21.190 --> 00:40:25.370 align:middle line:84%
It's baked into the
algorithm already.

00:40:25.370 --> 00:40:28.890 align:middle line:84%
This is an example of what that
linear program might look like.

00:40:28.890 --> 00:40:33.910 align:middle line:84%
So again, here we have our
maximization problem now.

00:40:33.910 --> 00:40:36.430 align:middle line:84%
And we're just assuming that
the lambda weights are one.

00:40:36.430 --> 00:40:38.970 align:middle line:84%
So this is just you can see
some of the player's utilities

00:40:38.970 --> 00:40:41.910 align:middle line:84%
for top left times the
probability of that

00:40:41.910 --> 00:40:44.370 align:middle line:84%
and so on and so forth for
all the other probabilities.

00:40:44.370 --> 00:40:46.970 align:middle line:84%
We can maximize the expected
utility for each player

00:40:46.970 --> 00:40:50.740 align:middle line:84%
by maximizing this sum, subject
to our constraints here.

00:40:50.740 --> 00:40:54.080 align:middle line:84%
So these are just our correlated
equilibria constraints.

00:40:54.080 --> 00:40:56.260 align:middle line:84%
I reorganized them a
little bit to make it clear

00:40:56.260 --> 00:40:58.980 align:middle line:84%
that it's very simple
to express them just

00:40:58.980 --> 00:41:02.520 align:middle line:84%
in terms of our probabilities
and greater than zero.

00:41:02.520 --> 00:41:05.460 align:middle line:84%
So it's easy to add these
to some matrix format

00:41:05.460 --> 00:41:07.560 align:middle line:90%
or some constraint algorithm.

00:41:07.560 --> 00:41:10.160 align:middle line:84%
And just plug this in to
our linear program solver,

00:41:10.160 --> 00:41:12.420 align:middle line:84%
and we'll get a
distribution that's

00:41:12.420 --> 00:41:15.460 align:middle line:84%
correlated or coarse correlated
and maximizes the expected

00:41:15.460 --> 00:41:19.502 align:middle line:84%
utility for the players--
so very promising.

00:41:19.502 --> 00:41:21.460 align:middle line:84%
We've been talking a lot
about these correlated

00:41:21.460 --> 00:41:22.880 align:middle line:84%
and coarse correlated
equilibria.

00:41:22.880 --> 00:41:23.540 align:middle line:90%
Why?

00:41:23.540 --> 00:41:25.540 align:middle line:84%
This polynomial time
algorithm is really good

00:41:25.540 --> 00:41:27.180 align:middle line:90%
compared to the Nash.

00:41:27.180 --> 00:41:30.580 align:middle line:84%
And here's a little bit
more on why that matters.

00:41:30.580 --> 00:41:33.020 align:middle line:84%
I said there are Nash
equilibria algorithms that

00:41:33.020 --> 00:41:35.820 align:middle line:84%
are in that class PPAD
that are exponential.

00:41:35.820 --> 00:41:38.180 align:middle line:84%
But if those algorithms
exist, why do we really

00:41:38.180 --> 00:41:39.687 align:middle line:90%
care about how long they take?

00:41:39.687 --> 00:41:42.020 align:middle line:84%
You could imagine just running
it for a really long time

00:41:42.020 --> 00:41:44.480 align:middle line:84%
and eventually getting a
solution to the Nash problem.

00:41:44.480 --> 00:41:47.380 align:middle line:84%
And Nash are tighter
equilibria, and it'll

00:41:47.380 --> 00:41:50.780 align:middle line:84%
be better probably for the
market if we can find that.

00:41:50.780 --> 00:41:53.900 align:middle line:84%
However, on our
blockchain implementation,

00:41:53.900 --> 00:41:58.380 align:middle line:84%
those exponential algorithms
are very, very hard to use

00:41:58.380 --> 00:42:02.020 align:middle line:84%
and pretty much are incompatible
with most architectures.

00:42:02.020 --> 00:42:04.518 align:middle line:84%
And the reason why is we'll
take Ethereum, for example,

00:42:04.518 --> 00:42:06.060 align:middle line:84%
because we've been
talking about that

00:42:06.060 --> 00:42:08.740 align:middle line:90%
in terms of smart contracts.

00:42:08.740 --> 00:42:13.220 align:middle line:84%
When a new block is
committed, each node

00:42:13.220 --> 00:42:16.720 align:middle line:84%
individually runs all of the
transactions on that node.

00:42:16.720 --> 00:42:18.340 align:middle line:84%
And for smart
contracts, that looks

00:42:18.340 --> 00:42:21.203 align:middle line:84%
like all the smart contract
updates, the function

00:42:21.203 --> 00:42:22.620 align:middle line:84%
calls and all the
smart contracts,

00:42:22.620 --> 00:42:24.560 align:middle line:84%
are run individually
on every node,

00:42:24.560 --> 00:42:26.620 align:middle line:84%
across every node
in the blockchain.

00:42:26.620 --> 00:42:29.060 align:middle line:84%
And to ensure that
this can be done

00:42:29.060 --> 00:42:31.740 align:middle line:84%
in a reasonable amount of
time so you can add blocks

00:42:31.740 --> 00:42:35.460 align:middle line:84%
fairly frequently, there
are very strict limits

00:42:35.460 --> 00:42:38.140 align:middle line:84%
on the complexity
allowed by each contract

00:42:38.140 --> 00:42:40.180 align:middle line:90%
and by each function call.

00:42:40.180 --> 00:42:45.060 align:middle line:84%
Ethereum, for example, will not
let you program any contracts

00:42:45.060 --> 00:42:50.270 align:middle line:84%
in which you cannot determine
how long it will take at compile

00:42:50.270 --> 00:42:50.950 align:middle line:90%
time.

00:42:50.950 --> 00:42:54.390 align:middle line:84%
So you can't pass in a variable
that determines the length

00:42:54.390 --> 00:42:56.270 align:middle line:84%
of a loop, or you
can't have "while"

00:42:56.270 --> 00:43:00.030 align:middle line:84%
loops where there's not
a guaranteed ending.

00:43:00.030 --> 00:43:03.070 align:middle line:84%
Anything in programming that has
a non-deterministic ending, you

00:43:03.070 --> 00:43:04.830 align:middle line:84%
can't do in Ethereum,
and anything

00:43:04.830 --> 00:43:09.150 align:middle line:84%
that's more than even low-factor
polynomial time complexity

00:43:09.150 --> 00:43:15.150 align:middle line:84%
gets very, very, very expensive
even for known fixed bounds.

00:43:15.150 --> 00:43:17.490 align:middle line:84%
And this is true on
most blockchains.

00:43:17.490 --> 00:43:19.790 align:middle line:84%
It differs between
them, but most of them

00:43:19.790 --> 00:43:23.230 align:middle line:84%
require that code is written
so that polynomial execution

00:43:23.230 --> 00:43:25.750 align:middle line:84%
time is guaranteed
regardless of what input

00:43:25.750 --> 00:43:27.510 align:middle line:90%
is passed to the contract.

00:43:27.510 --> 00:43:30.435 align:middle line:84%
So we need these
polynomial time algorithms,

00:43:30.435 --> 00:43:32.310 align:middle line:84%
and we need them to be
fairly simple in order

00:43:32.310 --> 00:43:35.470 align:middle line:90%
for our mechanism to work.

00:43:35.470 --> 00:43:37.930 align:middle line:84%
We have our new
equilibrium concept,

00:43:37.930 --> 00:43:41.870 align:middle line:84%
and we want to see if we can use
it to help the players induce

00:43:41.870 --> 00:43:43.690 align:middle line:90%
equilibria in Dubey's game.

00:43:43.690 --> 00:43:45.470 align:middle line:84%
That's our end goal
here, is to find

00:43:45.470 --> 00:43:48.030 align:middle line:84%
our correlated or coarse
correlated equilibria,

00:43:48.030 --> 00:43:50.390 align:middle line:84%
use that to signal
players their strategies,

00:43:50.390 --> 00:43:53.950 align:middle line:84%
and help induce the
non-cooperative equilibria

00:43:53.950 --> 00:43:57.590 align:middle line:84%
that we found to be
so beneficial earlier.

00:43:57.590 --> 00:44:02.350 align:middle line:84%
However, because we are
working with continuous prices,

00:44:02.350 --> 00:44:05.990 align:middle line:84%
or almost continuous, and
very large strategy spaces,

00:44:05.990 --> 00:44:09.830 align:middle line:84%
linear programming constraints
tend to grow very, very quickly.

00:44:09.830 --> 00:44:16.150 align:middle line:84%
And it's actually linear in
the number of outcomes, which

00:44:16.150 --> 00:44:17.823 align:middle line:84%
when we're talking
about outcomes,

00:44:17.823 --> 00:44:19.990 align:middle line:84%
that's the possible draws
from the joint probability

00:44:19.990 --> 00:44:20.532 align:middle line:90%
distribution.

00:44:20.532 --> 00:44:22.073 align:middle line:84%
Those were the
variables that we were

00:44:22.073 --> 00:44:23.910 align:middle line:84%
finding in our linear
programming problem.

00:44:23.910 --> 00:44:25.950 align:middle line:84%
And if you're talking
about the joint strategy

00:44:25.950 --> 00:44:30.790 align:middle line:84%
space in Dubey's game,
where each player has

00:44:30.790 --> 00:44:34.510 align:middle line:84%
many, many, many, many possible
limit orders they could submit,

00:44:34.510 --> 00:44:36.110 align:middle line:84%
the joint probability
distribution

00:44:36.110 --> 00:44:38.650 align:middle line:90%
grows extremely quickly.

00:44:38.650 --> 00:44:42.750 align:middle line:84%
So one example
for Dubey's game--

00:44:42.750 --> 00:44:47.120 align:middle line:84%
two goods, two players,
and very restricted

00:44:47.120 --> 00:44:48.345 align:middle line:90%
on the quantities and prices.

00:44:48.345 --> 00:44:49.720 align:middle line:84%
So for each good,
they're allowed

00:44:49.720 --> 00:44:54.680 align:middle line:84%
to submit a limit order for
quantity between 0 and 9

00:44:54.680 --> 00:44:56.400 align:middle line:90%
and for price between 0 and 9.

00:44:56.400 --> 00:44:56.900 align:middle line:90%
That's it.

00:44:56.900 --> 00:44:59.700 align:middle line:84%
There's no fractional
prices, just integer prices.

00:44:59.700 --> 00:45:03.960 align:middle line:84%
They have 10 options for each
good-- two players, two goods.

00:45:03.960 --> 00:45:06.240 align:middle line:84%
The strategy space
for each player

00:45:06.240 --> 00:45:12.280 align:middle line:84%
in this very limited setting is
100 million possible strategies.

00:45:12.280 --> 00:45:15.360 align:middle line:84%
So if we look at the joint
distribution between two

00:45:15.360 --> 00:45:19.960 align:middle line:84%
players, there's 10 to the
16th joint strategy outcomes.

00:45:19.960 --> 00:45:23.280 align:middle line:84%
So our linear programming
would be optimizing over 10

00:45:23.280 --> 00:45:25.600 align:middle line:90%
to the 16th variables.

00:45:25.600 --> 00:45:28.000 align:middle line:84%
If you add just a third
player, that number jumps to 10

00:45:28.000 --> 00:45:29.880 align:middle line:90%
to the 24th.

00:45:29.880 --> 00:45:33.840 align:middle line:84%
So even with polynomial
time algorithms across that,

00:45:33.840 --> 00:45:36.520 align:middle line:84%
our input is scaling
exponentially

00:45:36.520 --> 00:45:38.100 align:middle line:90%
with players and strategies.

00:45:38.100 --> 00:45:40.600 align:middle line:84%
So there's more
optimizations we need

00:45:40.600 --> 00:45:42.337 align:middle line:84%
to make in order to
make this tractable

00:45:42.337 --> 00:45:44.170 align:middle line:84%
because we want players
to be able to submit

00:45:44.170 --> 00:45:47.717 align:middle line:84%
nearly continuous prices,
varying quantities for goods,

00:45:47.717 --> 00:45:50.050 align:middle line:84%
and we certainly want there
to be more than two players.

00:45:50.050 --> 00:45:56.650 align:middle line:84%
And 10 to the 24th joint
strategy outcomes won't do.

00:45:56.650 --> 00:45:59.730 align:middle line:84%
We can move into a
different way of computing

00:45:59.730 --> 00:46:03.290 align:middle line:84%
correlated and coarse
correlated equilibrium, where

00:46:03.290 --> 00:46:06.090 align:middle line:84%
our linear programming that
we've been talking about

00:46:06.090 --> 00:46:11.390 align:middle line:84%
explicitly optimizes against all
of those joint probabilities.

00:46:11.390 --> 00:46:13.792 align:middle line:84%
The variables that we're
solving is the joint probability

00:46:13.792 --> 00:46:15.250 align:middle line:84%
distribution for
each player, which

00:46:15.250 --> 00:46:17.570 align:middle line:84%
is a very, very large
space because it

00:46:17.570 --> 00:46:20.370 align:middle line:90%
multiplies for each player.

00:46:20.370 --> 00:46:24.010 align:middle line:84%
We can use no-regret
learning instead,

00:46:24.010 --> 00:46:26.530 align:middle line:84%
which is a machine
learning-based algorithm.

00:46:26.530 --> 00:46:29.770 align:middle line:84%
And importantly, we
drop the optimization

00:46:29.770 --> 00:46:32.930 align:middle line:84%
over the joint probability
distributions and instead only

00:46:32.930 --> 00:46:36.770 align:middle line:84%
have to optimize over an
individual's strategies.

00:46:36.770 --> 00:46:40.130 align:middle line:84%
So it doesn't grow with
the number of players,

00:46:40.130 --> 00:46:42.490 align:middle line:84%
and it doesn't grow
exponentially as quickly

00:46:42.490 --> 00:46:45.623 align:middle line:90%
as the linear programming does.

00:46:45.623 --> 00:46:47.790 align:middle line:84%
So we're going to talk about
the no-regret learners,

00:46:47.790 --> 00:46:49.623 align:middle line:84%
get an understanding
of them, and then we'll

00:46:49.623 --> 00:46:52.890 align:middle line:84%
bring it back to the Bayes'
mechanism in a few slides.

00:46:52.890 --> 00:46:55.530 align:middle line:84%
So no-regret learning-- it's
a class of machine learning

00:46:55.530 --> 00:46:59.010 align:middle line:84%
algorithms, and it's
an online algorithm.

00:46:59.010 --> 00:47:02.550 align:middle line:84%
And an agent is trying to make
actions that maximize reward.

00:47:02.550 --> 00:47:05.190 align:middle line:84%
Like most machine learning
things, it takes an action,

00:47:05.190 --> 00:47:07.890 align:middle line:90%
gets feedback, updates itself.

00:47:07.890 --> 00:47:11.688 align:middle line:84%
It's online, which just means
that there's no training

00:47:11.688 --> 00:47:14.230 align:middle line:84%
phase like there is for a lot
of machine learning algorithms.

00:47:14.230 --> 00:47:17.850 align:middle line:84%
Instead, it's receiving
real feedback and rewards

00:47:17.850 --> 00:47:20.530 align:middle line:84%
at every time step and
making real decisions.

00:47:20.530 --> 00:47:24.170 align:middle line:84%
So for each time step-- we'll
denote it T here between 0

00:47:24.170 --> 00:47:25.250 align:middle line:90%
and big T--

00:47:25.250 --> 00:47:27.170 align:middle line:84%
it has to make a
decision without knowing

00:47:27.170 --> 00:47:28.670 align:middle line:90%
the full sequence of inputs.

00:47:28.670 --> 00:47:31.710 align:middle line:84%
And it has to update itself,
and it has to make a decision

00:47:31.710 --> 00:47:36.050 align:middle line:84%
and receive rewards
based off of that.

00:47:36.050 --> 00:47:38.590 align:middle line:84%
So we're going to have
our action space again.

00:47:38.590 --> 00:47:40.860 align:middle line:90%
We'll call the actions a now.

00:47:40.860 --> 00:47:42.900 align:middle line:90%
Still in the same strategy set.

00:47:42.900 --> 00:47:45.380 align:middle line:84%
And at each time
step, it's going

00:47:45.380 --> 00:47:47.480 align:middle line:84%
to submit an action
probabilistically.

00:47:47.480 --> 00:47:49.460 align:middle line:84%
So it's going to
keep some probability

00:47:49.460 --> 00:47:53.240 align:middle line:84%
vector over how often it
wants to play each strategy.

00:47:53.240 --> 00:47:55.480 align:middle line:84%
It's going to randomly
choose one, submit that,

00:47:55.480 --> 00:47:58.100 align:middle line:84%
and it's going to get feedback
from the mechanism on how good

00:47:58.100 --> 00:47:59.620 align:middle line:90%
it was.

00:47:59.620 --> 00:48:01.660 align:middle line:84%
When we talk about
how good it was,

00:48:01.660 --> 00:48:04.780 align:middle line:84%
we're actually not
going to use utility.

00:48:04.780 --> 00:48:07.760 align:middle line:84%
Just in how these algorithms
work, they work really nicely.

00:48:07.760 --> 00:48:09.720 align:middle line:84%
If you use the
inverse of utility,

00:48:09.720 --> 00:48:12.900 align:middle line:84%
which is going to be cost
here, and we're actually

00:48:12.900 --> 00:48:15.700 align:middle line:84%
going to normalize and scale
it just to help our algorithm

00:48:15.700 --> 00:48:16.560 align:middle line:90%
a little bit.

00:48:16.560 --> 00:48:19.900 align:middle line:84%
But in every case,
our cost here is just

00:48:19.900 --> 00:48:22.940 align:middle line:84%
a transformation, a linear
transformation, on utility.

00:48:22.940 --> 00:48:26.440 align:middle line:84%
So it's how far under-- so this
is our utility of our action--

00:48:26.440 --> 00:48:29.740 align:middle line:84%
it's how far under are we from
the best possible utility.

00:48:29.740 --> 00:48:30.723 align:middle line:90%
So you can imagine.

00:48:30.723 --> 00:48:32.140 align:middle line:84%
If the best utility
is really high

00:48:32.140 --> 00:48:34.182 align:middle line:84%
and we play an action that
is really low utility,

00:48:34.182 --> 00:48:38.820 align:middle line:84%
this is going to be big over
the largest possible utility

00:48:38.820 --> 00:48:39.700 align:middle line:90%
difference.

00:48:39.700 --> 00:48:40.740 align:middle line:90%
So you can imagine.

00:48:40.740 --> 00:48:43.820 align:middle line:84%
If we submit some action that
gives us the minimum utility,

00:48:43.820 --> 00:48:45.900 align:middle line:90%
our cost is going to be one.

00:48:45.900 --> 00:48:47.540 align:middle line:84%
If we submit some
action that gives

00:48:47.540 --> 00:48:49.580 align:middle line:84%
us really close to
our maximum utility,

00:48:49.580 --> 00:48:52.100 align:middle line:84%
this difference is
going to be small,

00:48:52.100 --> 00:48:54.220 align:middle line:84%
and our cost is going
to be close to 0.

00:48:54.220 --> 00:48:56.440 align:middle line:84%
So it's just a linear
transformation on the utility.

00:48:56.440 --> 00:48:59.180 align:middle line:84%
You can imagine really high
utility actions have really

00:48:59.180 --> 00:49:02.060 align:middle line:90%
low costs and vice versa.

00:49:02.060 --> 00:49:05.300 align:middle line:84%
Now we're talking
about these algorithms

00:49:05.300 --> 00:49:10.580 align:middle line:84%
in general, where there can
be any arbitrary mechanism

00:49:10.580 --> 00:49:13.580 align:middle line:84%
or adversary deciding
what these costs are.

00:49:13.580 --> 00:49:15.820 align:middle line:84%
And our cost vector
here is just C of t.

00:49:15.820 --> 00:49:17.620 align:middle line:84%
And at each time
period, they're allowed

00:49:17.620 --> 00:49:21.140 align:middle line:84%
to decide whatever they
want for some arbitrary cost

00:49:21.140 --> 00:49:23.820 align:middle line:84%
that you incur and
will update you.

00:49:23.820 --> 00:49:27.460 align:middle line:84%
And in a lot of cases, you
can design these adversaries

00:49:27.460 --> 00:49:30.220 align:middle line:90%
to play mean against you.

00:49:30.220 --> 00:49:32.385 align:middle line:84%
They can spy on your
strategies and what

00:49:32.385 --> 00:49:33.760 align:middle line:84%
they think your
probabilities are

00:49:33.760 --> 00:49:37.390 align:middle line:84%
and try and play things that
will make you incur high costs.

00:49:37.390 --> 00:49:40.270 align:middle line:84%
We are not super interested
in those applications.

00:49:40.270 --> 00:49:43.710 align:middle line:84%
We are specifically interested
in the multiplayer setting

00:49:43.710 --> 00:49:48.710 align:middle line:84%
where we have many agents all
trying to optimize individually,

00:49:48.710 --> 00:49:50.650 align:middle line:84%
submitting to some
central mechanism,

00:49:50.650 --> 00:49:52.970 align:middle line:84%
and that mechanism determines
utility for everybody.

00:49:52.970 --> 00:49:55.910 align:middle line:84%
So while these
algorithms are determined

00:49:55.910 --> 00:49:58.030 align:middle line:84%
for general
adversaries, they also

00:49:58.030 --> 00:50:00.530 align:middle line:84%
do work very nicely
in multiplayer games.

00:50:00.530 --> 00:50:03.030 align:middle line:84%
So they're applicable
to our mechanism here.

00:50:03.030 --> 00:50:07.810 align:middle line:84%
AUDIENCE: The random action a,
that's a single player's action.

00:50:07.810 --> 00:50:09.550 align:middle line:90%
SAMUEL BRUCE: Correct.

00:50:09.550 --> 00:50:11.250 align:middle line:84%
AUDIENCE: It says
probabilistically.

00:50:11.250 --> 00:50:15.680 align:middle line:84%
Can it depend on the actions
that other players have chosen

00:50:15.680 --> 00:50:16.930 align:middle line:90%
so as to get that correlation?

00:50:16.930 --> 00:50:18.590 align:middle line:90%
Or is it always like--

00:50:18.590 --> 00:50:20.770 align:middle line:84%
that probability, what
does it depend on?

00:50:20.770 --> 00:50:24.350 align:middle line:84%
Is it completely independent
of everything else?

00:50:24.350 --> 00:50:27.630 align:middle line:84%
SAMUEL BRUCE: Agents don't get
to learn the actions of others.

00:50:27.630 --> 00:50:30.730 align:middle line:84%
In fact, the information
they receive is very limited.

00:50:30.730 --> 00:50:34.920 align:middle line:84%
So in each turn, they just get
back the cost of their actions.

00:50:34.920 --> 00:50:37.620 align:middle line:84%
They don't even get to learn
their own utility function,

00:50:37.620 --> 00:50:39.500 align:middle line:84%
or why the costs
are as they are,

00:50:39.500 --> 00:50:41.240 align:middle line:84%
or what other
players were doing.

00:50:41.240 --> 00:50:43.080 align:middle line:84%
It's as limited as
possible, but they

00:50:43.080 --> 00:50:45.000 align:middle line:90%
keep an internal probability.

00:50:45.000 --> 00:50:47.000 align:middle line:84%
And we'll see in a couple
slides how they update

00:50:47.000 --> 00:50:50.680 align:middle line:90%
that based off of their costs.

00:50:50.680 --> 00:50:53.720 align:middle line:84%
So when we're talking about
evaluating our regret learning

00:50:53.720 --> 00:50:57.240 align:middle line:84%
algorithms, we can't
really evaluate them

00:50:57.240 --> 00:50:59.560 align:middle line:84%
against the gold
standard, which would

00:50:59.560 --> 00:51:02.080 align:middle line:84%
be the perfect algorithm that
makes all the right choices

00:51:02.080 --> 00:51:03.640 align:middle line:90%
at all the right times.

00:51:03.640 --> 00:51:07.760 align:middle line:84%
And a thought experiment we can
do to show why that's infeasible

00:51:07.760 --> 00:51:11.040 align:middle line:84%
is imagine you have
some adversary who's

00:51:11.040 --> 00:51:14.480 align:middle line:84%
picking your costs for
you, and at each time step,

00:51:14.480 --> 00:51:16.988 align:middle line:90%
they pick of all your actions.

00:51:16.988 --> 00:51:18.780 align:middle line:84%
They're all going to
have the maximum cost.

00:51:18.780 --> 00:51:20.322 align:middle line:84%
They're all going
to have a cost one.

00:51:20.322 --> 00:51:24.080 align:middle line:84%
And I'm going to pick one
action that has cost zero.

00:51:24.080 --> 00:51:25.260 align:middle line:90%
So that's the best action.

00:51:25.260 --> 00:51:27.260 align:middle line:84%
If you choose that
one, you have no cost.

00:51:27.260 --> 00:51:29.680 align:middle line:84%
But any other action you
choose has the maximum cost.

00:51:29.680 --> 00:51:32.960 align:middle line:84%
And every single time step,
I'm going to randomly change

00:51:32.960 --> 00:51:35.440 align:middle line:90%
which action is cost zero.

00:51:35.440 --> 00:51:37.680 align:middle line:90%
So there's no pattern.

00:51:37.680 --> 00:51:39.960 align:middle line:84%
There's no nothing
except one action

00:51:39.960 --> 00:51:42.000 align:middle line:90%
is better than all the others.

00:51:42.000 --> 00:51:44.500 align:middle line:84%
You can imagine the perfect
algorithm is going to know that,

00:51:44.500 --> 00:51:46.680 align:middle line:84%
and they're going
to find randomly

00:51:46.680 --> 00:51:49.840 align:middle line:84%
all of these possible
actions that have

00:51:49.840 --> 00:51:52.200 align:middle line:90%
zero cost at every time step.

00:51:52.200 --> 00:51:53.840 align:middle line:84%
And over our t
time steps, they're

00:51:53.840 --> 00:51:57.240 align:middle line:84%
going to incur a
total cost of 0--

00:51:57.240 --> 00:51:58.920 align:middle line:90%
perfect algorithm.

00:51:58.920 --> 00:52:02.080 align:middle line:84%
Even the best possible
algorithm online

00:52:02.080 --> 00:52:05.000 align:middle line:84%
that you could submit to
that, intuitively, it's

00:52:05.000 --> 00:52:09.000 align:middle line:84%
going to be randomly choosing a
strategy is as good as anything

00:52:09.000 --> 00:52:11.280 align:middle line:84%
else because they're
randomly choosing

00:52:11.280 --> 00:52:12.860 align:middle line:90%
which action has cost zero.

00:52:12.860 --> 00:52:15.600 align:middle line:84%
So you might as well
randomize to try and find it.

00:52:15.600 --> 00:52:17.360 align:middle line:84%
And over t time
steps, you're going

00:52:17.360 --> 00:52:22.040 align:middle line:84%
to have an expected cost of the
probability you get that random

00:52:22.040 --> 00:52:23.720 align:middle line:84%
guess wrong, which
is number of actions

00:52:23.720 --> 00:52:27.120 align:middle line:84%
minus 1 over all of your actions
times the number of time steps.

00:52:27.120 --> 00:52:29.400 align:middle line:84%
So you're going
to grow your cost

00:52:29.400 --> 00:52:31.410 align:middle line:84%
linearly with the
amount of time,

00:52:31.410 --> 00:52:33.250 align:middle line:84%
and they're going
to have cost zero.

00:52:33.250 --> 00:52:35.150 align:middle line:84%
It's not really a
useful comparison here.

00:52:35.150 --> 00:52:38.330 align:middle line:84%
And you can think of other
examples of adversaries, which

00:52:38.330 --> 00:52:41.490 align:middle line:84%
in retrospect, you can pick
some perfect algorithm that

00:52:41.490 --> 00:52:43.150 align:middle line:90%
would have had awesome cost.

00:52:43.150 --> 00:52:45.410 align:middle line:84%
But realistically,
with an algorithm

00:52:45.410 --> 00:52:47.310 align:middle line:84%
that's making
decisions as it goes,

00:52:47.310 --> 00:52:50.010 align:middle line:84%
it's not going to perform
well at all against that.

00:52:50.010 --> 00:52:51.870 align:middle line:84%
When we're evaluating
these algorithms,

00:52:51.870 --> 00:52:54.810 align:middle line:84%
this is a real concern
because we want to figure out

00:52:54.810 --> 00:52:55.990 align:middle line:90%
what makes a good algorithm.

00:52:55.990 --> 00:52:57.690 align:middle line:84%
And that clearly
isn't it, because it's

00:52:57.690 --> 00:52:59.410 align:middle line:90%
an impractical standard.

00:52:59.410 --> 00:53:03.010 align:middle line:84%
So instead, we're going to talk
about the notion of regret.

00:53:03.010 --> 00:53:06.570 align:middle line:84%
And what regret is
doing is it's comparing

00:53:06.570 --> 00:53:09.610 align:middle line:84%
what you did for all
of the past time steps

00:53:09.610 --> 00:53:12.715 align:middle line:84%
to what some other better
agent might have done.

00:53:12.715 --> 00:53:15.090 align:middle line:84%
But we're going to be very
restrictive on what the better

00:53:15.090 --> 00:53:16.730 align:middle line:90%
agent is allowed to do.

00:53:16.730 --> 00:53:18.690 align:middle line:84%
They're not allowed to
just arbitrarily pick

00:53:18.690 --> 00:53:20.490 align:middle line:84%
whatever they want
because they know what

00:53:20.490 --> 00:53:22.770 align:middle line:90%
the cost-minimizing action was.

00:53:22.770 --> 00:53:26.330 align:middle line:84%
We're going to have
some policy class G,

00:53:26.330 --> 00:53:28.450 align:middle line:84%
and G is just a
class of decisions

00:53:28.450 --> 00:53:30.990 align:middle line:84%
where we're going
to keep it simple.

00:53:30.990 --> 00:53:33.730 align:middle line:84%
So this is going to map
some simple subset of what

00:53:33.730 --> 00:53:35.730 align:middle line:90%
another agent could have done.

00:53:35.730 --> 00:53:37.210 align:middle line:84%
All the possible
iterations, we'll

00:53:37.210 --> 00:53:42.090 align:middle line:84%
call that G. Each individual
strategy of the opposing agent

00:53:42.090 --> 00:53:44.067 align:middle line:90%
we'll call pi.

00:53:44.067 --> 00:53:45.650 align:middle line:84%
So now what we're
going to do is we're

00:53:45.650 --> 00:53:48.030 align:middle line:84%
going to look at
different types of Gs,

00:53:48.030 --> 00:53:52.530 align:middle line:84%
different restrictions on what
the comparison is allowed to do.

00:53:52.530 --> 00:53:55.450 align:middle line:84%
And we want algorithms
that do well

00:53:55.450 --> 00:53:58.170 align:middle line:84%
against this restricted agent
as the number of time steps

00:53:58.170 --> 00:53:59.690 align:middle line:90%
increases.

00:53:59.690 --> 00:54:02.610 align:middle line:84%
So we're going to start
with external regret, which

00:54:02.610 --> 00:54:05.050 align:middle line:84%
is a specific version
of no-regret learning

00:54:05.050 --> 00:54:09.090 align:middle line:84%
where our G, our class of
alternative strategies,

00:54:09.090 --> 00:54:10.810 align:middle line:90%
is just fixed strategies.

00:54:10.810 --> 00:54:15.070 align:middle line:84%
So our comparison agent
is restricted to that.

00:54:15.070 --> 00:54:18.970 align:middle line:84%
They can only play for all the
past t time steps one fixed

00:54:18.970 --> 00:54:21.050 align:middle line:90%
action across every time step.

00:54:21.050 --> 00:54:25.230 align:middle line:84%
And G here is going
to maximize over that.

00:54:25.230 --> 00:54:31.220 align:middle line:84%
So we'll call RG is the cost of
the best possible fixed action.

00:54:31.220 --> 00:54:33.240 align:middle line:84%
Our quantity RG
here, we define it.

00:54:33.240 --> 00:54:36.260 align:middle line:84%
It's the minimum of all of your
actions in your strategy set

00:54:36.260 --> 00:54:38.320 align:middle line:84%
of the cost of playing
that fixed action.

00:54:38.320 --> 00:54:42.140 align:middle line:84%
That will be our class
here for external regret.

00:54:42.140 --> 00:54:44.500 align:middle line:84%
R algorithm, which
we'll call h, is

00:54:44.500 --> 00:54:48.740 align:middle line:84%
trying to minimize the
difference between our regret

00:54:48.740 --> 00:54:52.420 align:middle line:84%
and the regret of G, where
the regret again, is just

00:54:52.420 --> 00:54:56.500 align:middle line:84%
the sum of your cost over the
time steps from the beginning

00:54:56.500 --> 00:54:57.860 align:middle line:90%
until now.

00:54:57.860 --> 00:55:00.580 align:middle line:84%
So we want an algorithm
that does well

00:55:00.580 --> 00:55:02.880 align:middle line:90%
in this difference terms.

00:55:02.880 --> 00:55:04.100 align:middle line:90%
I.e.

00:55:04.100 --> 00:55:07.100 align:middle line:84%
Your regret from the
actions you played

00:55:07.100 --> 00:55:09.460 align:middle line:84%
are not much worse
than the regret

00:55:09.460 --> 00:55:13.220 align:middle line:84%
you would have gotten if you had
just played some optimal fixed

00:55:13.220 --> 00:55:15.900 align:middle line:90%
action in hindsight.

00:55:15.900 --> 00:55:20.380 align:middle line:84%
And this is a pretty restricted
class of comparison agents.

00:55:20.380 --> 00:55:23.380 align:middle line:84%
They're only allowed to play one
strategy for every time step.

00:55:23.380 --> 00:55:28.510 align:middle line:84%
However, algorithms which use
this comparison class actually

00:55:28.510 --> 00:55:30.190 align:middle line:84%
tend to have really
good properties.

00:55:30.190 --> 00:55:33.690 align:middle line:84%
So one specific algorithm
that's used a lot in practice,

00:55:33.690 --> 00:55:36.110 align:middle line:84%
it's called randomized
weighted majority.

00:55:36.110 --> 00:55:39.510 align:middle line:84%
It minimizes this quantity
of external regret.

00:55:39.510 --> 00:55:42.790 align:middle line:84%
And its difference
here between RH and RG

00:55:42.790 --> 00:55:47.710 align:middle line:84%
is bounded over t time steps
by the square root of t log n.

00:55:47.710 --> 00:55:50.510 align:middle line:84%
N here is the number
of strategies--

00:55:50.510 --> 00:55:52.710 align:middle line:84%
sorry, a little bit
of notation change.

00:55:52.710 --> 00:55:56.470 align:middle line:84%
But that means per
iteration, you're

00:55:56.470 --> 00:56:00.750 align:middle line:84%
actually decreasing your total
bounded regret as t goes on.

00:56:00.750 --> 00:56:04.590 align:middle line:84%
So as time continues,
the amount of regret

00:56:04.590 --> 00:56:09.470 align:middle line:84%
that you have over the best
fixed action on the highest end

00:56:09.470 --> 00:56:13.910 align:middle line:84%
approaches our fixed regret,
and it scales fairly well.

00:56:13.910 --> 00:56:16.997 align:middle line:84%
Now this is what the
algorithm actually looks like.

00:56:16.997 --> 00:56:18.330 align:middle line:90%
I'm going to walk us through it.

00:56:18.330 --> 00:56:20.710 align:middle line:84%
But the important thing
I want to stress here

00:56:20.710 --> 00:56:22.910 align:middle line:84%
is that it's actually
fairly simple.

00:56:22.910 --> 00:56:25.670 align:middle line:84%
And that's one of the best
parts of these algorithms

00:56:25.670 --> 00:56:28.630 align:middle line:84%
is they're fairly simple,
which means they run quickly

00:56:28.630 --> 00:56:29.990 align:middle line:90%
and they run efficiently.

00:56:29.990 --> 00:56:35.750 align:middle line:84%
Here we have w is the weight for
player assigning to strategy i

00:56:35.750 --> 00:56:39.450 align:middle line:84%
for time step T. L is going
to be our cost vector.

00:56:39.450 --> 00:56:43.950 align:middle line:84%
So L of t minus 1 is our full
cost vector over all actions.

00:56:43.950 --> 00:56:45.690 align:middle line:90%
I denotes the ith good.

00:56:45.690 --> 00:56:48.950 align:middle line:84%
It's very important to
know for these algorithms,

00:56:48.950 --> 00:56:52.390 align:middle line:84%
even though the player
submits one strategy

00:56:52.390 --> 00:56:57.030 align:middle line:84%
to the adversary, who gets
to decide what the payout is,

00:56:57.030 --> 00:56:59.150 align:middle line:84%
the adversary doesn't
just give them

00:56:59.150 --> 00:57:01.790 align:middle line:84%
back the cost of
the action they did.

00:57:01.790 --> 00:57:03.830 align:middle line:84%
The adversary actually
returns the cost

00:57:03.830 --> 00:57:06.270 align:middle line:90%
of all possible actions.

00:57:06.270 --> 00:57:07.910 align:middle line:84%
So you can imagine
the cost vector

00:57:07.910 --> 00:57:10.750 align:middle line:84%
you receive is not
individually the cost

00:57:10.750 --> 00:57:14.070 align:middle line:84%
you received for your action,
but a vector of the length

00:57:14.070 --> 00:57:17.150 align:middle line:84%
of your strategies, telling
you what the cost was

00:57:17.150 --> 00:57:20.510 align:middle line:84%
for all of your possible
strategies at that time period.

00:57:20.510 --> 00:57:23.720 align:middle line:84%
That's going to be important
for how the algorithm learns.

00:57:23.720 --> 00:57:26.000 align:middle line:84%
Here we're going
to fix our losses,

00:57:26.000 --> 00:57:28.800 align:middle line:90%
our costs to be either 0 or 1.

00:57:28.800 --> 00:57:33.460 align:middle line:84%
The algorithm works exactly
the same for arbitrary costs.

00:57:33.460 --> 00:57:35.820 align:middle line:84%
It's just the notation gets
a little more cumbersome.

00:57:35.820 --> 00:57:37.920 align:middle line:84%
So we're just going
to do 0, 1 loss.

00:57:37.920 --> 00:57:40.920 align:middle line:84%
And we'll update each
weight in each time step

00:57:40.920 --> 00:57:44.120 align:middle line:84%
by some learning rate
eta that will denote

00:57:44.120 --> 00:57:45.920 align:middle line:90%
when we start our training.

00:57:45.920 --> 00:57:49.800 align:middle line:84%
And then at each time
step, given our weights,

00:57:49.800 --> 00:57:54.440 align:middle line:84%
you will play each strategy i
with probability that weight

00:57:54.440 --> 00:57:56.760 align:middle line:90%
over the sum of all weights.

00:57:56.760 --> 00:57:58.600 align:middle line:84%
And you initialize
all of your weights

00:57:58.600 --> 00:58:01.120 align:middle line:84%
equal and all of your
probabilities equal

00:58:01.120 --> 00:58:03.827 align:middle line:84%
so that you randomly select
strategies at the beginning.

00:58:03.827 --> 00:58:06.160 align:middle line:84%
So the player is just going
to randomly pick a strategy,

00:58:06.160 --> 00:58:08.240 align:middle line:90%
submit it to the adversary.

00:58:08.240 --> 00:58:13.080 align:middle line:84%
The adversary is going to give
them a cost vector back of what

00:58:13.080 --> 00:58:15.960 align:middle line:84%
the cost would have been
for any of your strategies,

00:58:15.960 --> 00:58:18.340 align:middle line:84%
and it's going to
take that vector.

00:58:18.340 --> 00:58:24.040 align:middle line:84%
And for each good in the
vector, if our cost is 1--

00:58:24.040 --> 00:58:25.840 align:middle line:90%
so we have high cost--

00:58:25.840 --> 00:58:27.860 align:middle line:84%
we're going to decrease
our weight slightly.

00:58:27.860 --> 00:58:30.540 align:middle line:84%
We're going to decrease
it by 1 minus eta.

00:58:30.540 --> 00:58:32.640 align:middle line:84%
So you can imagine
if eta was 0.1,

00:58:32.640 --> 00:58:35.100 align:middle line:84%
we're going to multiply
our weight by 0.9.

00:58:35.100 --> 00:58:37.480 align:middle line:84%
So it's a little less
likely we'll play it again.

00:58:37.480 --> 00:58:40.100 align:middle line:84%
If the cost is zero,
it was a good action.

00:58:40.100 --> 00:58:42.160 align:middle line:84%
We're going to keep
our weight the same.

00:58:42.160 --> 00:58:46.800 align:middle line:84%
So you do this for all
strategies at every time step.

00:58:46.800 --> 00:58:49.240 align:middle line:84%
And slowly, the strategies
that consistently

00:58:49.240 --> 00:58:52.180 align:middle line:84%
have high cost, their weights
are going to decrease.

00:58:52.180 --> 00:58:54.600 align:middle line:84%
And the probability you're
going to play bad actions

00:58:54.600 --> 00:58:56.240 align:middle line:90%
is going to decrease.

00:58:56.240 --> 00:58:59.240 align:middle line:84%
And in practice, this
actually happens very quickly.

00:58:59.240 --> 00:59:03.460 align:middle line:84%
And even after just depending
on your learning rate,

00:59:03.460 --> 00:59:06.200 align:middle line:84%
but hundreds to
thousands of iterations,

00:59:06.200 --> 00:59:09.520 align:middle line:84%
you can get it to converge
on good strategies for kind

00:59:09.520 --> 00:59:12.320 align:middle line:90%
adversaries.

00:59:12.320 --> 00:59:15.993 align:middle line:84%
So back to economics
a little bit.

00:59:15.993 --> 00:59:18.160 align:middle line:84%
We're now going to think
about again the multiplayer

00:59:18.160 --> 00:59:21.450 align:middle line:84%
setting where the adversary
isn't some arbitrary

00:59:21.450 --> 00:59:24.050 align:middle line:90%
person trying to hurt you.

00:59:24.050 --> 00:59:27.170 align:middle line:84%
But we have a multiplayer game,
and you submit your strategy

00:59:27.170 --> 00:59:28.410 align:middle line:90%
to the mechanism.

00:59:28.410 --> 00:59:32.290 align:middle line:84%
Other agents using the same
no-regret learning algorithm

00:59:32.290 --> 00:59:35.030 align:middle line:84%
submit their strategies
to the mechanism.

00:59:35.030 --> 00:59:36.730 align:middle line:84%
The mechanism
computes the outcome

00:59:36.730 --> 00:59:40.257 align:middle line:84%
and informs each
agent on their costs,

00:59:40.257 --> 00:59:42.590 align:middle line:84%
and that would be the cost
for the action they submitted

00:59:42.590 --> 00:59:44.290 align:middle line:84%
but also the costs of all
the other actions they

00:59:44.290 --> 00:59:45.530 align:middle line:90%
could have submitted too.

00:59:45.530 --> 00:59:48.090 align:middle line:84%
So again, we're still getting
a lot of information here.

00:59:48.090 --> 00:59:52.410 align:middle line:84%
But in this multiplayer
setting, if we

00:59:52.410 --> 00:59:55.890 align:middle line:84%
take all of the
strategies of each player

00:59:55.890 --> 01:00:00.050 align:middle line:84%
and we take the average of
them over all t time steps,

01:00:00.050 --> 01:00:03.010 align:middle line:84%
we converge to a coarse
correlated equilibrium.

01:00:03.010 --> 01:00:05.850 align:middle line:84%
Just reading through
our proposition here,

01:00:05.850 --> 01:00:09.090 align:middle line:84%
after t iterations of
no-regret dynamics,

01:00:09.090 --> 01:00:11.330 align:middle line:84%
every player of a
cost-minimization game

01:00:11.330 --> 01:00:15.690 align:middle line:84%
has regret of at most epsilon
for each of its strategies,

01:00:15.690 --> 01:00:19.620 align:middle line:84%
where epsilon is our
time-averaged regret.

01:00:19.620 --> 01:00:22.220 align:middle line:84%
We had bounds on it
earlier, but it's just

01:00:22.220 --> 01:00:25.820 align:middle line:84%
the regret you faced over all
t time steps divided by t.

01:00:25.820 --> 01:00:30.300 align:middle line:84%
We take sigma here to
be the joint probability

01:00:30.300 --> 01:00:34.020 align:middle line:90%
space at time t, sigma t.

01:00:34.020 --> 01:00:36.420 align:middle line:84%
It's the outcome
distribution, and we have just

01:00:36.420 --> 01:00:38.460 align:middle line:90%
sigma as the time average.

01:00:38.460 --> 01:00:41.100 align:middle line:84%
So we take the joint
probability distribution

01:00:41.100 --> 01:00:43.460 align:middle line:84%
of all the agents
at each time step,

01:00:43.460 --> 01:00:46.740 align:middle line:84%
and we average
them over all time.

01:00:46.740 --> 01:00:50.100 align:middle line:84%
This sigma is a coarse
correlated equilibrium

01:00:50.100 --> 01:00:55.380 align:middle line:84%
approximated by the size of
the regret of each player.

01:00:55.380 --> 01:01:00.340 align:middle line:84%
So importantly, remember as
players continue to play,

01:01:00.340 --> 01:01:04.660 align:middle line:84%
their average regret, their
regret per iteration goes down.

01:01:04.660 --> 01:01:08.260 align:middle line:84%
So if we run our agents
for many, many iterations,

01:01:08.260 --> 01:01:10.820 align:middle line:84%
our approximation will
get better and better

01:01:10.820 --> 01:01:12.980 align:middle line:84%
to where the joint
strategies of the players

01:01:12.980 --> 01:01:15.180 align:middle line:84%
are very, very
closely approximating

01:01:15.180 --> 01:01:17.500 align:middle line:90%
a coarse correlated equilibrium.

01:01:17.500 --> 01:01:22.260 align:middle line:84%
And this is really
exciting because now

01:01:22.260 --> 01:01:24.220 align:middle line:84%
we no longer are
trying to optimize over

01:01:24.220 --> 01:01:26.380 align:middle line:84%
the joint probability
distribution.

01:01:26.380 --> 01:01:30.020 align:middle line:84%
We can take this many, many,
many, many dimensional space,

01:01:30.020 --> 01:01:30.680 align:middle line:90%
break it up.

01:01:30.680 --> 01:01:34.980 align:middle line:84%
So now each player is just
trying to individually maximize

01:01:34.980 --> 01:01:38.900 align:middle line:90%
over their n strategies.

01:01:38.900 --> 01:01:40.660 align:middle line:84%
And then the outcome,
when we play them

01:01:40.660 --> 01:01:43.740 align:middle line:84%
all against each other, still
converges on our equilibrium--

01:01:43.740 --> 01:01:47.580 align:middle line:84%
much less complex in terms
of algorithmic compute

01:01:47.580 --> 01:01:51.700 align:middle line:84%
and we still get reasonably
close to our equilibrium.

01:01:51.700 --> 01:01:54.900 align:middle line:84%
Now, to move from a coarse
correlated equilibrium

01:01:54.900 --> 01:01:57.020 align:middle line:84%
to a correlated
equilibria, we're

01:01:57.020 --> 01:01:59.220 align:middle line:84%
going to do something
very similar in nature

01:01:59.220 --> 01:02:02.620 align:middle line:90%
to how we did it earlier.

01:02:02.620 --> 01:02:09.060 align:middle line:84%
Remember that coarse correlated
was the class where players,

01:02:09.060 --> 01:02:11.040 align:middle line:84%
prior to learning
their strategies,

01:02:11.040 --> 01:02:12.680 align:middle line:90%
had to commit to the mechanism.

01:02:12.680 --> 01:02:15.350 align:middle line:84%
So they just needed to be in
expectation better than a fixed

01:02:15.350 --> 01:02:15.970 align:middle line:90%
action.

01:02:15.970 --> 01:02:17.850 align:middle line:84%
And indeed, the external
regret, remember,

01:02:17.850 --> 01:02:20.350 align:middle line:84%
we were comparing
to fixed actions.

01:02:20.350 --> 01:02:22.670 align:middle line:84%
However, in the
correlated equilibria,

01:02:22.670 --> 01:02:25.390 align:middle line:84%
you were given your
strategy and still

01:02:25.390 --> 01:02:27.850 align:middle line:84%
had to be better
than any deviation.

01:02:27.850 --> 01:02:31.010 align:middle line:84%
So you can almost think of,
oh, I'm given this strategy,

01:02:31.010 --> 01:02:33.990 align:middle line:90%
but I'd rather swap to that one.

01:02:33.990 --> 01:02:37.610 align:middle line:84%
And that's going to be a
dual with the idea here,

01:02:37.610 --> 01:02:39.630 align:middle line:84%
which is a slightly
larger comparison

01:02:39.630 --> 01:02:42.510 align:middle line:90%
class called swap regret.

01:02:42.510 --> 01:02:46.070 align:middle line:84%
No-swap regret learning
is the algorithms.

01:02:46.070 --> 01:02:50.070 align:middle line:84%
And it's going to
say exactly that.

01:02:50.070 --> 01:02:52.190 align:middle line:84%
It's the class of
all possible mappings

01:02:52.190 --> 01:02:59.870 align:middle line:84%
where at time t, what if every
time I had played i played j?

01:02:59.870 --> 01:03:02.710 align:middle line:84%
Instead of a fixed
deviation like

01:03:02.710 --> 01:03:07.710 align:middle line:84%
before from whatever you
did to any fixed action, now

01:03:07.710 --> 01:03:10.710 align:middle line:84%
we're going to say, OK, in
retrospect, anytime you played i

01:03:10.710 --> 01:03:12.350 align:middle line:90%
what if you had played j?

01:03:12.350 --> 01:03:15.110 align:middle line:84%
So it's any end-to-end
mapping for all

01:03:15.110 --> 01:03:16.370 align:middle line:90%
of your possible strategies.

01:03:16.370 --> 01:03:18.270 align:middle line:84%
You can map them
arbitrarily to having

01:03:18.270 --> 01:03:21.070 align:middle line:84%
played a different strategy
every time you played that one.

01:03:21.070 --> 01:03:24.630 align:middle line:84%
And this is going to be a
much bigger and more lenient

01:03:24.630 --> 01:03:27.110 align:middle line:84%
comparison class is going to
be much higher performing.

01:03:27.110 --> 01:03:30.190 align:middle line:84%
You can imagine actually, the
external regret comparison class

01:03:30.190 --> 01:03:34.070 align:middle line:84%
is a subset of this
one, just where you map

01:03:34.070 --> 01:03:37.750 align:middle line:90%
every action to a fixed action.

01:03:37.750 --> 01:03:39.470 align:middle line:84%
But now we can do an
arbitrary mapping,

01:03:39.470 --> 01:03:42.750 align:middle line:84%
so we can map every action to
whichever action we'd like.

01:03:42.750 --> 01:03:46.550 align:middle line:84%
And algorithms for
swap regret exist.

01:03:46.550 --> 01:03:48.510 align:middle line:84%
And again, they're
very tractable,

01:03:48.510 --> 01:03:51.430 align:middle line:84%
and they execute
in polynomial time.

01:03:51.430 --> 01:03:57.750 align:middle line:84%
And like the external
regret, if we time-average

01:03:57.750 --> 01:04:00.310 align:middle line:84%
the probabilities
for each player,

01:04:00.310 --> 01:04:03.430 align:middle line:84%
we come on a
correlated equilibrium.

01:04:03.430 --> 01:04:07.230 align:middle line:84%
So if we run an algorithm
that performs really

01:04:07.230 --> 01:04:11.600 align:middle line:84%
well against this slightly
harder class of comparisons,

01:04:11.600 --> 01:04:14.600 align:middle line:84%
then we converge to a
slightly better equilibrium.

01:04:14.600 --> 01:04:18.040 align:middle line:84%
We can look at algorithms for
both external regret and swap

01:04:18.040 --> 01:04:21.020 align:middle line:84%
regret, use them,
time-average strategies,

01:04:21.020 --> 01:04:23.900 align:middle line:84%
and get correlated and
coarse correlated equilibria,

01:04:23.900 --> 01:04:25.705 align:middle line:90%
which is really exciting.

01:04:25.705 --> 01:04:27.080 align:middle line:84%
You can estimate
these equilibria

01:04:27.080 --> 01:04:28.997 align:middle line:84%
without computing the
entire joint probability

01:04:28.997 --> 01:04:31.400 align:middle line:84%
distribution, which was
prohibitively expensive, as we

01:04:31.400 --> 01:04:33.160 align:middle line:90%
saw earlier.

01:04:33.160 --> 01:04:35.780 align:middle line:84%
And they work well in
multiplayer settings.

01:04:35.780 --> 01:04:38.400 align:middle line:84%
There's been a lot of
research and studies done

01:04:38.400 --> 01:04:39.980 align:middle line:90%
to prove that they converge.

01:04:39.980 --> 01:04:45.000 align:middle line:84%
They converge fairly quickly,
and they estimate equilibria

01:04:45.000 --> 01:04:45.580 align:middle line:90%
well.

01:04:45.580 --> 01:04:47.840 align:middle line:84%
Now, an interesting
thing about them

01:04:47.840 --> 01:04:51.720 align:middle line:84%
is we lose now the ability
to have a maximization

01:04:51.720 --> 01:04:53.940 align:middle line:84%
function, like we did
in linear programming.

01:04:53.940 --> 01:04:56.720 align:middle line:84%
Remember, we could maximize over
the sum of the agents' utilities

01:04:56.720 --> 01:05:00.200 align:middle line:84%
while doing the linear
programming at no extra cost.

01:05:00.200 --> 01:05:03.600 align:middle line:84%
Now that's kind of been
taken away from us,

01:05:03.600 --> 01:05:05.880 align:middle line:84%
but we'll have to do
some investigation

01:05:05.880 --> 01:05:08.930 align:middle line:84%
to see how much that
really costs us.

01:05:08.930 --> 01:05:11.650 align:middle line:84%
And again, just like there
were polynomial time algorithms

01:05:11.650 --> 01:05:15.070 align:middle line:84%
for simple Nash equilibria,
under simple games,

01:05:15.070 --> 01:05:18.730 align:middle line:84%
these algorithms
perform very, very well.

01:05:18.730 --> 01:05:20.570 align:middle line:84%
For two-player
zero-sum games, they're

01:05:20.570 --> 01:05:23.890 align:middle line:84%
guaranteed to converge on Nash
equilibria for each player.

01:05:23.890 --> 01:05:28.210 align:middle line:84%
And in fact, more generally, in
two-player general-sum games,

01:05:28.210 --> 01:05:29.990 align:middle line:84%
when players begin
with equal weights,

01:05:29.990 --> 01:05:33.250 align:middle line:84%
they're guaranteed to
converge on Nash equilibria.

01:05:33.250 --> 01:05:35.490 align:middle line:84%
There are lots of
situations, just

01:05:35.490 --> 01:05:38.370 align:middle line:84%
like with our normal
algorithms, where

01:05:38.370 --> 01:05:43.010 align:middle line:84%
if we simplify the game, then we
can find even better converging

01:05:43.010 --> 01:05:46.730 align:middle line:84%
properties for our
no-regret learners.

01:05:46.730 --> 01:05:48.330 align:middle line:90%
A little example here--

01:05:48.330 --> 01:05:51.270 align:middle line:84%
we're using our same players
that we were using earlier.

01:05:51.270 --> 01:05:53.290 align:middle line:90%
We have player 1 and player 2.

01:05:53.290 --> 01:05:55.950 align:middle line:90%
And we have some graphs here.

01:05:55.950 --> 01:05:58.530 align:middle line:84%
I ran this for
20,000 iterations,

01:05:58.530 --> 01:06:03.610 align:middle line:84%
and this map in the top left
is their probability of playing

01:06:03.610 --> 01:06:04.590 align:middle line:90%
each strategy.

01:06:04.590 --> 01:06:07.810 align:middle line:84%
So the blue line-- you can see
hidden in there somewhere--

01:06:07.810 --> 01:06:11.230 align:middle line:84%
is the probability that
player zero plays top.

01:06:11.230 --> 01:06:15.650 align:middle line:84%
Sorry, the letters are off, but
that's the blue line is for top.

01:06:15.650 --> 01:06:18.435 align:middle line:84%
The orange line is the
probability they play bottom.

01:06:18.435 --> 01:06:19.810 align:middle line:84%
The green line is
the probability

01:06:19.810 --> 01:06:23.410 align:middle line:84%
that player two plays left
and then red for right.

01:06:23.410 --> 01:06:26.530 align:middle line:84%
And as you can see,
the strategies,

01:06:26.530 --> 01:06:28.470 align:middle line:84%
the individual probabilities
for each player,

01:06:28.470 --> 01:06:29.890 align:middle line:90%
they never converge.

01:06:29.890 --> 01:06:31.830 align:middle line:84%
And that's very
common in these games.

01:06:31.830 --> 01:06:35.330 align:middle line:84%
They are not guaranteed to
converge under most situations.

01:06:35.330 --> 01:06:37.350 align:middle line:84%
They, in fact, go through
a lot of these phases.

01:06:37.350 --> 01:06:39.392 align:middle line:84%
You can see it where it
bounces up and down where

01:06:39.392 --> 01:06:41.990 align:middle line:84%
it's like they play one
strategy very dominantly,

01:06:41.990 --> 01:06:45.730 align:middle line:84%
and then the other player learns
to take advantage of that,

01:06:45.730 --> 01:06:47.570 align:middle line:90%
exploit it.

01:06:47.570 --> 01:06:50.350 align:middle line:84%
The one player receives high
costs from the exploitation,

01:06:50.350 --> 01:06:52.090 align:middle line:84%
so they completely
switch their strategy

01:06:52.090 --> 01:06:54.185 align:middle line:84%
and exploit the other
player and back and forth.

01:06:54.185 --> 01:06:56.810 align:middle line:84%
And you see these dynamics where
it's just going back and forth

01:06:56.810 --> 01:06:58.290 align:middle line:90%
and back and forth.

01:06:58.290 --> 01:07:02.770 align:middle line:84%
But if we take the time-average
of those strategies,

01:07:02.770 --> 01:07:07.700 align:middle line:84%
this mess seems to
converge to something.

01:07:07.700 --> 01:07:11.700 align:middle line:84%
And in fact, with these
probabilities for each player,

01:07:11.700 --> 01:07:13.640 align:middle line:84%
if we take the joint
probabilities--

01:07:13.640 --> 01:07:17.820 align:middle line:84%
so it's just probability of top
left here, just top times left--

01:07:17.820 --> 01:07:21.560 align:middle line:84%
we get strategies that again
converge and look like this--

01:07:21.560 --> 01:07:25.680 align:middle line:84%
somewhere between 0.0 and
0.1, a couple between 0.10.2,

01:07:25.680 --> 01:07:30.300 align:middle line:84%
and then the highest one is
bottom left with pretty high

01:07:30.300 --> 01:07:31.460 align:middle line:90%
probability.

01:07:31.460 --> 01:07:36.440 align:middle line:84%
This turns out to be the unique
Nash equilibrium of the game.

01:07:36.440 --> 01:07:40.220 align:middle line:90%


01:07:40.220 --> 01:07:43.060 align:middle line:84%
I ran it for 20,000 iterations,
but we're already starting

01:07:43.060 --> 01:07:44.580 align:middle line:90%
to converge fairly quickly.

01:07:44.580 --> 01:07:48.020 align:middle line:84%
So after just some
thousands of iterations,

01:07:48.020 --> 01:07:50.500 align:middle line:84%
we get fairly close to the
unique Nash equilibrium

01:07:50.500 --> 01:07:52.740 align:middle line:84%
of the game with just two
of these learners, which

01:07:52.740 --> 01:07:55.020 align:middle line:90%
is very impressive.

01:07:55.020 --> 01:07:58.005 align:middle line:84%
Remember, we've been
talking about algorithms

01:07:58.005 --> 01:07:59.380 align:middle line:84%
for the no-regret
learning, where

01:07:59.380 --> 01:08:02.880 align:middle line:84%
they learn at each time step the
cost of all of their actions.

01:08:02.880 --> 01:08:05.110 align:middle line:84%
And we'll just talk
about that for a second.

01:08:05.110 --> 01:08:08.350 align:middle line:84%
This is called the full
information setting.

01:08:08.350 --> 01:08:11.063 align:middle line:84%
And it is among these
classes of algorithms

01:08:11.063 --> 01:08:12.730 align:middle line:84%
the most information
they could receive.

01:08:12.730 --> 01:08:15.510 align:middle line:84%
They get at each time
step the expected cost

01:08:15.510 --> 01:08:17.290 align:middle line:84%
of any single one
of their strategies,

01:08:17.290 --> 01:08:20.109 align:middle line:84%
and they can use that to update
their weights very reasonably

01:08:20.109 --> 01:08:20.870 align:middle line:90%
well.

01:08:20.870 --> 01:08:23.590 align:middle line:90%
They get a lot of information.

01:08:23.590 --> 01:08:26.109 align:middle line:84%
A slightly less
information version

01:08:26.109 --> 01:08:29.149 align:middle line:84%
is the partial information
setting, where again, you

01:08:29.149 --> 01:08:32.710 align:middle line:84%
get to learn the entire cost
vector over all your strategies.

01:08:32.710 --> 01:08:36.390 align:middle line:84%
But now, instead of
getting perfect costs

01:08:36.390 --> 01:08:39.490 align:middle line:84%
for your strategies,
there's some noise in there.

01:08:39.490 --> 01:08:41.830 align:middle line:84%
In this case, you
only get to learn

01:08:41.830 --> 01:08:43.430 align:middle line:84%
the cost of your
strategies based off

01:08:43.430 --> 01:08:47.810 align:middle line:84%
of what other players did, not
what the expected cost might be,

01:08:47.810 --> 01:08:50.510 align:middle line:84%
as in the full
information setting.

01:08:50.510 --> 01:08:54.710 align:middle line:84%
And then most restrictive,
called the multi-armed bandit

01:08:54.710 --> 01:08:55.950 align:middle line:90%
setting.

01:08:55.950 --> 01:08:59.590 align:middle line:84%
The agent no longer receives
the cost vector in training.

01:08:59.590 --> 01:09:03.270 align:middle line:84%
They only instead get the
cost of their exact action

01:09:03.270 --> 01:09:05.170 align:middle line:90%
and the other player's actions.

01:09:05.170 --> 01:09:10.689 align:middle line:84%
So they just get to learn what
they did and how it played out.

01:09:10.689 --> 01:09:12.620 align:middle line:84%
You no longer get to
see any of what-ifs

01:09:12.620 --> 01:09:14.870 align:middle line:84%
that might have happened if
you had submitted a better

01:09:14.870 --> 01:09:16.790 align:middle line:90%
strategy or worse strategy.

01:09:16.790 --> 01:09:19.189 align:middle line:84%
You only get to learn
and update based off

01:09:19.189 --> 01:09:21.630 align:middle line:90%
of what actually happened.

01:09:21.630 --> 01:09:25.790 align:middle line:84%
And why would we use
other information models

01:09:25.790 --> 01:09:28.350 align:middle line:84%
when the full information
one performs so well?

01:09:28.350 --> 01:09:32.390 align:middle line:84%
It really depends on
your mechanism design.

01:09:32.390 --> 01:09:35.950 align:middle line:84%
So algorithms for those
limited in information

01:09:35.950 --> 01:09:38.790 align:middle line:84%
and multi-armed
bandit settings exist

01:09:38.790 --> 01:09:41.330 align:middle line:84%
but they converge slower,
which intuitively makes sense.

01:09:41.330 --> 01:09:43.770 align:middle line:84%
The agent is getting
worse or less information.

01:09:43.770 --> 01:09:46.830 align:middle line:84%
They're not going to be
able to converge as quickly.

01:09:46.830 --> 01:09:49.649 align:middle line:84%
However, when you're
designing the mechanism,

01:09:49.649 --> 01:09:54.109 align:middle line:84%
it may be extremely costly to
give an agent their entire cost

01:09:54.109 --> 01:09:54.830 align:middle line:90%
vector.

01:09:54.830 --> 01:09:57.510 align:middle line:84%
It may be prohibitively
difficult to figure out

01:09:57.510 --> 01:10:00.330 align:middle line:84%
what the cost of potential
actions would be.

01:10:00.330 --> 01:10:03.480 align:middle line:84%
Or you can think of some
real-world situations

01:10:03.480 --> 01:10:05.413 align:middle line:84%
where it's impossible
to know what

01:10:05.413 --> 01:10:07.080 align:middle line:84%
would have happened
if you had submitted

01:10:07.080 --> 01:10:08.300 align:middle line:90%
some alternative action.

01:10:08.300 --> 01:10:12.440 align:middle line:84%
You only really get to learn
what ended up executing.

01:10:12.440 --> 01:10:16.000 align:middle line:84%
So when you're designing
these algorithms that

01:10:16.000 --> 01:10:17.767 align:middle line:84%
determine the cost
for the players,

01:10:17.767 --> 01:10:19.600 align:middle line:84%
you have to really
consider what information

01:10:19.600 --> 01:10:23.040 align:middle line:84%
is available and easy to compute
and choose your information

01:10:23.040 --> 01:10:24.140 align:middle line:90%
setting accordingly.

01:10:24.140 --> 01:10:28.080 align:middle line:84%
Again, algorithms exist for
all three types of information.

01:10:28.080 --> 01:10:31.760 align:middle line:84%
However, the less information
you give the learning agent,

01:10:31.760 --> 01:10:34.120 align:middle line:90%
the slower it may converge.

01:10:34.120 --> 01:10:37.280 align:middle line:84%
Now looking back towards
Dubey's mechanism,

01:10:37.280 --> 01:10:40.080 align:middle line:84%
we're really worried about
these large action spaces.

01:10:40.080 --> 01:10:43.000 align:middle line:84%
That was what drew us to these
no-regret learners is, OK, we

01:10:43.000 --> 01:10:47.040 align:middle line:84%
have a lot of strategies,
continuous prices.

01:10:47.040 --> 01:10:48.220 align:middle line:90%
How can we handle that?

01:10:48.220 --> 01:10:50.240 align:middle line:84%
How we reconcile that
when the strategy

01:10:50.240 --> 01:10:52.120 align:middle line:90%
space is extremely huge?

01:10:52.120 --> 01:10:56.140 align:middle line:84%
And no-regret learners are
actually capable of doing that.

01:10:56.140 --> 01:10:58.960 align:middle line:84%
So algorithms exist
for no-regret learners

01:10:58.960 --> 01:11:02.520 align:middle line:84%
on completely continuous
action spaces like prices

01:11:02.520 --> 01:11:06.440 align:middle line:84%
and very large action spaces
like our possible strategy sets,

01:11:06.440 --> 01:11:09.560 align:middle line:84%
even under limited information
settings like our bandit

01:11:09.560 --> 01:11:13.160 align:middle line:84%
setting where you only learn the
cost of the action you actually

01:11:13.160 --> 01:11:14.000 align:middle line:90%
submitted.

01:11:14.000 --> 01:11:15.920 align:middle line:90%
So these algorithms exist.

01:11:15.920 --> 01:11:19.880 align:middle line:84%
They are fairly new and
improving very quickly,

01:11:19.880 --> 01:11:25.960 align:middle line:84%
but they can do in
relatively polynomial time

01:11:25.960 --> 01:11:27.800 align:middle line:84%
these optimizations
that converge

01:11:27.800 --> 01:11:30.200 align:middle line:84%
on correlate and
cross-correlated equilibria

01:11:30.200 --> 01:11:33.080 align:middle line:84%
without having to restrict
the strategy space at all,

01:11:33.080 --> 01:11:36.083 align:middle line:90%
which is very impressive.

01:11:36.083 --> 01:11:38.000 align:middle line:84%
Just to summarize regret
learning a little bit

01:11:38.000 --> 01:11:40.080 align:middle line:84%
because I know there's
a lot of information,

01:11:40.080 --> 01:11:42.000 align:middle line:84%
and they're a
little complicated.

01:11:42.000 --> 01:11:43.920 align:middle line:84%
They're online
algorithms, which means

01:11:43.920 --> 01:11:47.160 align:middle line:84%
that they have to take
in information and costs

01:11:47.160 --> 01:11:48.820 align:middle line:90%
and make decisions as it goes.

01:11:48.820 --> 01:11:50.580 align:middle line:90%
They don't train and then test.

01:11:50.580 --> 01:11:53.880 align:middle line:84%
It's all they update
as they perform.

01:11:53.880 --> 01:11:57.530 align:middle line:84%
Our external regret
learners compare error costs

01:11:57.530 --> 01:12:00.610 align:middle line:84%
to the best fixed strategy
in retrospect and optimize

01:12:00.610 --> 01:12:01.690 align:middle line:90%
against that.

01:12:01.690 --> 01:12:04.290 align:middle line:84%
Our swap regret
learners minimize cost

01:12:04.290 --> 01:12:08.170 align:middle line:84%
over the best fixed
mapping between strategies,

01:12:08.170 --> 01:12:14.770 align:middle line:84%
and the time-averaged strategies
for external regret and swap

01:12:14.770 --> 01:12:17.050 align:middle line:84%
regret converge
on correlated and

01:12:17.050 --> 01:12:19.810 align:middle line:90%
coarse correlated equilibria.

01:12:19.810 --> 01:12:21.970 align:middle line:84%
Now, they can make
these estimations

01:12:21.970 --> 01:12:25.170 align:middle line:84%
without computing the entire
joint probability distribution,

01:12:25.170 --> 01:12:28.170 align:middle line:84%
which is a huge
algorithm improvement

01:12:28.170 --> 01:12:31.970 align:middle line:84%
over our linear programming,
which had to optimize

01:12:31.970 --> 01:12:34.810 align:middle line:90%
the entire distribution.

01:12:34.810 --> 01:12:36.965 align:middle line:84%
Kind of bringing
it all together--

01:12:36.965 --> 01:12:39.090 align:middle line:84%
we've kind of gone in a
lot of different directions

01:12:39.090 --> 01:12:43.610 align:middle line:84%
here, but the general
roadmap for the work

01:12:43.610 --> 01:12:46.770 align:middle line:84%
that we're doing
together involves pieces

01:12:46.770 --> 01:12:47.710 align:middle line:90%
from all of this.

01:12:47.710 --> 01:12:52.330 align:middle line:84%
So we want to design regret
learning agents and a training

01:12:52.330 --> 01:12:55.940 align:middle line:84%
mechanism that can efficiently
estimate these equilibria

01:12:55.940 --> 01:12:57.060 align:middle line:90%
for Dubey's game.

01:12:57.060 --> 01:12:59.040 align:middle line:90%
For what strategies?

01:12:59.040 --> 01:13:00.580 align:middle line:84%
What limit order
should be submitted

01:13:00.580 --> 01:13:02.900 align:middle line:84%
to the mechanism at
each time period?

01:13:02.900 --> 01:13:06.020 align:middle line:84%
We want to use this
equilibrium as a distribution

01:13:06.020 --> 01:13:08.180 align:middle line:84%
for a coordinator,
and we want to be

01:13:08.180 --> 01:13:11.100 align:middle line:84%
able to deploy this
coordinator with the mechanism

01:13:11.100 --> 01:13:14.780 align:middle line:84%
for implementing Dubey's
market onto the blockchain.

01:13:14.780 --> 01:13:19.300 align:middle line:84%
And the goal is to get
algorithms in place

01:13:19.300 --> 01:13:23.580 align:middle line:84%
for these no-regret learners
that can effectively

01:13:23.580 --> 01:13:25.660 align:middle line:84%
compute what the
equilibria should look

01:13:25.660 --> 01:13:30.060 align:middle line:84%
like and importantly, what
it looks like under varying

01:13:30.060 --> 01:13:30.720 align:middle line:90%
conditions.

01:13:30.720 --> 01:13:33.860 align:middle line:84%
You could imagine if you
train your no-regret learners

01:13:33.860 --> 01:13:37.780 align:middle line:84%
under a very specific endowment
and a very specific endowment

01:13:37.780 --> 01:13:40.380 align:middle line:84%
of other players, it
may find solutions

01:13:40.380 --> 01:13:44.020 align:middle line:84%
that are very good for that
one iteration of the market.

01:13:44.020 --> 01:13:48.460 align:middle line:84%
But how to adapt those
learners when they face shocks

01:13:48.460 --> 01:13:50.020 align:middle line:90%
to their endowment?

01:13:50.020 --> 01:13:52.160 align:middle line:84%
They face shocks to
their utility functions.

01:13:52.160 --> 01:13:54.860 align:middle line:84%
Other players face shocks
that change the market.

01:13:54.860 --> 01:13:57.100 align:middle line:84%
How you adapt
these algorithms is

01:13:57.100 --> 01:13:59.940 align:middle line:84%
a very complicated
and interesting point

01:13:59.940 --> 01:14:03.740 align:middle line:84%
of study, which is what
we're focusing on right now.

01:14:03.740 --> 01:14:05.600 align:middle line:84%
But we want to build
these regret learners.

01:14:05.600 --> 01:14:07.308 align:middle line:84%
We want to be able to
estimate equilibria

01:14:07.308 --> 01:14:08.820 align:middle line:90%
for Dubey's mechanism.

01:14:08.820 --> 01:14:12.260 align:middle line:84%
And if we can efficiently
deploy them onto the blockchain,

01:14:12.260 --> 01:14:14.765 align:middle line:84%
then we'll have a functioning
example of how a lot

01:14:14.765 --> 01:14:17.140 align:middle line:84%
of the things we've talked
about in this class can be put

01:14:17.140 --> 01:14:21.900 align:middle line:84%
into practice, used effectively
in a real-world example

01:14:21.900 --> 01:14:23.580 align:middle line:90%
on the blockchain--

01:14:23.580 --> 01:14:28.660 align:middle line:84%
so very exciting for us
working towards it every day.

01:14:28.660 --> 01:14:32.380 align:middle line:84%
I guess I'll end off with just
some outstanding questions

01:14:32.380 --> 01:14:34.660 align:middle line:84%
that we're really
actively investigating,

01:14:34.660 --> 01:14:38.180 align:middle line:84%
just to think about what the
next steps past all of this is.

01:14:38.180 --> 01:14:40.780 align:middle line:84%
So what were some of the
things we're very curious about

01:14:40.780 --> 01:14:44.740 align:middle line:84%
is what is the efficiency
of the outcomes computed

01:14:44.740 --> 01:14:46.420 align:middle line:90%
by these no-regret learners?

01:14:46.420 --> 01:14:49.740 align:middle line:84%
Remember, we lose the ability to
maximize an objective function

01:14:49.740 --> 01:14:52.350 align:middle line:84%
when we lose the linear
programming solution.

01:14:52.350 --> 01:14:56.090 align:middle line:84%
So it's still ongoing exactly
for Dubey's mechanism,

01:14:56.090 --> 01:15:00.630 align:middle line:84%
how optimal and efficient
these learners are.

01:15:00.630 --> 01:15:02.450 align:middle line:84%
Even if they do
find an equilibrium,

01:15:02.450 --> 01:15:06.990 align:middle line:84%
are there ones that are better
for certain agents than others?

01:15:06.990 --> 01:15:09.910 align:middle line:84%
And then are there
ways that we can

01:15:09.910 --> 01:15:11.470 align:middle line:84%
modify our training
algorithm that

01:15:11.470 --> 01:15:13.830 align:middle line:90%
improve the outcomes we see?

01:15:13.830 --> 01:15:17.190 align:middle line:84%
And then again, like we talked
about shocks-- how can we

01:15:17.190 --> 01:15:18.330 align:middle line:90%
change the distribution?

01:15:18.330 --> 01:15:22.310 align:middle line:84%
How can we have the agents
be dynamic to new agents

01:15:22.310 --> 01:15:25.210 align:middle line:84%
entering the market,
shocks to their endowment,

01:15:25.210 --> 01:15:26.790 align:middle line:90%
shocks to their utilities?

01:15:26.790 --> 01:15:28.870 align:middle line:84%
When things change,
can these agents

01:15:28.870 --> 01:15:30.510 align:middle line:84%
predict that in
the market, instead

01:15:30.510 --> 01:15:34.910 align:middle line:84%
of just being very static for
a very certain configuration?

01:15:34.910 --> 01:15:38.270 align:middle line:84%
And then, lastly, what's
the optimal information

01:15:38.270 --> 01:15:39.490 align:middle line:90%
setting for these learners?

01:15:39.490 --> 01:15:41.430 align:middle line:84%
We talked about the
full information,

01:15:41.430 --> 01:15:45.590 align:middle line:84%
the limited information,
and the multi-armed bandits,

01:15:45.590 --> 01:15:48.210 align:middle line:84%
but there's trade-offs
when implementing them.

01:15:48.210 --> 01:15:51.270 align:middle line:84%
Because when you talk about
the ones with more information,

01:15:51.270 --> 01:15:54.630 align:middle line:84%
they converge faster, but
it may take much more time

01:15:54.630 --> 01:15:57.390 align:middle line:84%
per iteration because you have
to calculate the cost over all

01:15:57.390 --> 01:16:00.150 align:middle line:84%
of their possible
strategies, which can be hard

01:16:00.150 --> 01:16:02.990 align:middle line:90%
or prohibitively expensive.

01:16:02.990 --> 01:16:05.330 align:middle line:84%
On the other end, you have
the multi-armed bandits,

01:16:05.330 --> 01:16:08.350 align:middle line:84%
which may be very, very quick
per iteration because you just

01:16:08.350 --> 01:16:12.070 align:middle line:84%
need to see what the cost of the
realized state of the world was,

01:16:12.070 --> 01:16:14.070 align:middle line:84%
but it may take far
too long for them

01:16:14.070 --> 01:16:16.310 align:middle line:90%
to find appropriate strategies.

01:16:16.310 --> 01:16:18.670 align:middle line:84%
So where the
trade-off is in there,

01:16:18.670 --> 01:16:21.390 align:middle line:84%
how we can make that work--
that's also a question

01:16:21.390 --> 01:16:22.950 align:middle line:90%
we hope to answer.

01:16:22.950 --> 01:16:25.590 align:middle line:90%
I think that's just about it.

01:16:25.590 --> 01:16:28.073 align:middle line:84%
I hope that was
very informative,

01:16:28.073 --> 01:16:29.490 align:middle line:84%
and that we were
able to track it.

01:16:29.490 --> 01:16:32.430 align:middle line:84%
I know we talked about a ton of
different things but all fitting

01:16:32.430 --> 01:16:34.990 align:middle line:84%
together for this specific
limit-order market.

01:16:34.990 --> 01:16:37.230 align:middle line:84%
And hopefully, we can
get it implemented.

01:16:37.230 --> 01:16:38.780 align:middle line:90%
Thanks.

01:16:38.780 --> 01:16:50.000 align:middle line:90%