WEBVTT

00:00:00.060 --> 00:00:01.780
The following
content is provided

00:00:01.780 --> 00:00:04.019
under a Creative
Commons license.

00:00:04.019 --> 00:00:06.870
Your support will help MIT
OpenCourseWare continue

00:00:06.870 --> 00:00:10.730
to offer high quality
educational resources for free.

00:00:10.730 --> 00:00:13.340
To make a donation or
view additional materials

00:00:13.340 --> 00:00:17.215
from hundreds of MIT courses,
visit MIT OpenCourseWare

00:00:17.215 --> 00:00:17.840
at ocw.mit.edu.

00:00:26.840 --> 00:00:31.390
PROFESSOR: All right, well,
good afternoon and welcome back.

00:00:31.390 --> 00:00:35.400
We have an exciting fun-filled
program for you this afternoon.

00:00:35.400 --> 00:00:36.176
I'm David Gifford.

00:00:36.176 --> 00:00:37.550
I'm delighted to
be back with you

00:00:37.550 --> 00:00:40.970
again, here in computational
systems biology.

00:00:40.970 --> 00:00:43.140
Today we're going to talk
about chromatin structure

00:00:43.140 --> 00:00:45.540
and how we can analyze it.

00:00:45.540 --> 00:00:51.290
And to give you the narrative
arc for our discussion today,

00:00:51.290 --> 00:00:54.130
we're first going to
begin with looking

00:00:54.130 --> 00:00:56.750
at computational methods that
we can break the, quote unquote

00:00:56.750 --> 00:01:01.120
code, that describes
the epigenome.

00:01:01.120 --> 00:01:03.900
Now, epigenetic state is
extraordinarily important

00:01:03.900 --> 00:01:05.800
and one way you
can visualize this

00:01:05.800 --> 00:01:08.400
is that the genome is
like a hotel filled

00:01:08.400 --> 00:01:09.780
with lots of different rooms.

00:01:09.780 --> 00:01:13.360
And a lot of the doors are
locked and some of the doors

00:01:13.360 --> 00:01:13.990
are unlocked.

00:01:13.990 --> 00:01:15.890
And only in the doors
that we can go into,

00:01:15.890 --> 00:01:18.320
where the genome is
open and accessible

00:01:18.320 --> 00:01:22.270
can there actually be work
done, regulation performed

00:01:22.270 --> 00:01:26.254
and transcripts
and proteins made.

00:01:26.254 --> 00:01:28.420
So we're going to talk about
how to actually analyze

00:01:28.420 --> 00:01:30.777
epigenetic state.

00:01:30.777 --> 00:01:32.360
And then we're going
to talk about how

00:01:32.360 --> 00:01:35.150
to use epigenetic
information to understand

00:01:35.150 --> 00:01:39.260
the entire regulatory
occupancy of the genome.

00:01:39.260 --> 00:01:42.230
We've already talked about
ChIP-seq and the idea

00:01:42.230 --> 00:01:44.620
that we can understand where
individual regulators sit

00:01:44.620 --> 00:01:49.950
on the genome, and how they
regulate proximal genes.

00:01:49.950 --> 00:01:53.590
We're now going to see if we
can learn more about the genome.

00:01:53.590 --> 00:01:56.610
How it's state-- whether
it's open or closed.

00:01:56.610 --> 00:01:58.380
Is it self-regulated?

00:01:58.380 --> 00:02:00.440
And answer a puzzle.

00:02:00.440 --> 00:02:04.430
The puzzle is, if there
are hundreds of thousands

00:02:04.430 --> 00:02:06.390
of possible binary
locations that

00:02:06.390 --> 00:02:09.770
are equally good
for a regulator,

00:02:09.770 --> 00:02:12.680
why are only tens of
thousands occupied?

00:02:12.680 --> 00:02:15.420
And how are those sites picked?

00:02:15.420 --> 00:02:18.630
Because that level of regulation
is extraordinarily important

00:02:18.630 --> 00:02:21.860
to establish a basal
level of what genes

00:02:21.860 --> 00:02:24.140
are accessible and operating.

00:02:24.140 --> 00:02:29.540
And finally, we're going to
talk about how we can map,

00:02:29.540 --> 00:02:32.270
which regulatory
regions in the genome

00:02:32.270 --> 00:02:36.160
are affecting which genes.

00:02:36.160 --> 00:02:40.360
It turns out that about
1/3 of the regulatory sites

00:02:40.360 --> 00:02:44.440
in the genome skip over a
gene that's closest to them

00:02:44.440 --> 00:02:48.240
to regulate a gene
that's farther away.

00:02:48.240 --> 00:02:49.510
This is a million genomes.

00:02:49.510 --> 00:02:52.530
And so given that
rough approximation,

00:02:52.530 --> 00:02:54.380
how is it that we
can make connections

00:02:54.380 --> 00:03:00.720
between regulatory sites and
the genes that they control?

00:03:00.720 --> 00:03:03.190
Now, in computational
systems biology,

00:03:03.190 --> 00:03:05.340
we always talk a
lot about biology,

00:03:05.340 --> 00:03:08.880
but we also need to reflect
upon the computational methods

00:03:08.880 --> 00:03:11.090
that we're bringing to
bear on these questions.

00:03:11.090 --> 00:03:13.300
And so, today, we're
going to be talking

00:03:13.300 --> 00:03:14.550
about three different methods.

00:03:14.550 --> 00:03:17.520
We'll talk about dynamic
Bayesian networks as a way

00:03:17.520 --> 00:03:21.140
to approach, understanding
the histone code.

00:03:21.140 --> 00:03:24.580
We'll talk about how to
classify factor binding,

00:03:24.580 --> 00:03:27.180
using log likelihood ratios.

00:03:27.180 --> 00:03:29.280
And finally, we'll
turn to our friend,

00:03:29.280 --> 00:03:33.250
the hypergeometric
distribution to analyze

00:03:33.250 --> 00:03:35.200
which locations
in the genome are

00:03:35.200 --> 00:03:36.430
interacting with one another.

00:03:39.010 --> 00:03:43.680
So let's begin with
establishing a vocabulary.

00:03:43.680 --> 00:03:46.230
I'm sure some of you
have seen this before.

00:03:46.230 --> 00:03:48.070
This is the way
that chromatin can

00:03:48.070 --> 00:03:50.630
be thought of being organized
at different levels.

00:03:50.630 --> 00:03:53.440
There's the primary
DNA sequence,

00:03:53.440 --> 00:03:58.930
which can include
methylated CPGs.

00:03:58.930 --> 00:04:01.820
That's cysteine,
phosphate, guanine.

00:04:01.820 --> 00:04:09.470
And the nice thing about
that is that it's symmetrical

00:04:09.470 --> 00:04:15.830
so that when you have a CPG,
a methyltransferase during DNA

00:04:15.830 --> 00:04:18.050
replication can copy
that methy mark over.

00:04:18.050 --> 00:04:20.910
So it's a mark that's heritable.

00:04:20.910 --> 00:04:23.880
The next level down
are histone tails.

00:04:23.880 --> 00:04:29.310
On the amino terminus
of histones H3 and H4,

00:04:29.310 --> 00:04:32.060
different chemical
modifications can be made,

00:04:32.060 --> 00:04:33.960
and they serve as
sign posts, as we'll

00:04:33.960 --> 00:04:35.739
see, to give us
clues about what's

00:04:35.739 --> 00:04:37.780
going on in the genome in
that proximal location.

00:04:40.710 --> 00:04:43.340
The next level down
is, whether or not

00:04:43.340 --> 00:04:46.830
the chromatin is
compacted or not.

00:04:46.830 --> 00:04:48.630
Whether it's open or closed.

00:04:48.630 --> 00:04:50.360
And that relates
to whether or not

00:04:50.360 --> 00:04:54.260
DNA binding proteins are
actually on the genome.

00:04:54.260 --> 00:04:57.380
And finally, certain
domains of the genome

00:04:57.380 --> 00:05:00.160
can be associated with
the nuclear lamina.

00:05:00.160 --> 00:05:03.880
And so they're different levels
of organization of chromatin.

00:05:03.880 --> 00:05:08.620
And we'll be exploring
all of these today.

00:05:08.620 --> 00:05:13.690
So the cartoon
version of the way

00:05:13.690 --> 00:05:20.260
that the genome is
organized is that at the top

00:05:20.260 --> 00:05:22.080
we have a transcribed gene.

00:05:22.080 --> 00:05:24.480
And you can see that
there's an enhancer that

00:05:24.480 --> 00:05:29.310
is interacting with the RNA
polymerase II start site.

00:05:29.310 --> 00:05:31.120
And you can see
varied histone marks

00:05:31.120 --> 00:05:34.920
that are associated with
this activated gene.

00:05:34.920 --> 00:05:36.640
There are also marks
that are associated

00:05:36.640 --> 00:05:37.860
with that active enhancer.

00:05:40.790 --> 00:05:44.040
Down below, you see
an inactive gene.

00:05:44.040 --> 00:05:46.650
And you can see that there's
a boundary element that's

00:05:46.650 --> 00:05:50.240
bound by CTCF, which,
one of its function

00:05:50.240 --> 00:05:53.820
is to serve as a genomic
insulator, which insulates

00:05:53.820 --> 00:05:58.140
the effect of the enhancer
above from the gene below.

00:05:58.140 --> 00:06:01.790
So through careful biochemical
analysis over the years,

00:06:01.790 --> 00:06:10.620
these different marks have been
analyzed and characterized.

00:06:10.620 --> 00:06:15.340
And a general paradigm
for understanding

00:06:15.340 --> 00:06:19.410
how the marks transition
as genes are activated

00:06:19.410 --> 00:06:21.640
is shown here.

00:06:21.640 --> 00:06:24.870
So genes that are
fairly active and cycle

00:06:24.870 --> 00:06:27.880
between active and
inactive states typically

00:06:27.880 --> 00:06:31.750
have a high CPG content
in their promoters.

00:06:31.750 --> 00:06:33.750
And transition is
shown on the left.

00:06:33.750 --> 00:06:37.140
Where in the repressed
state on the bottom,

00:06:37.140 --> 00:06:41.990
they're marked by
H3K27 trimethyl marks.

00:06:41.990 --> 00:06:47.210
When they're poised, they have
both H3K4 trimethyl and H3K27

00:06:47.210 --> 00:06:48.490
trimethyl.

00:06:48.490 --> 00:06:55.320
And when they're active, they
only have H3K4 trimethyl.

00:06:55.320 --> 00:06:59.200
And on the right hand side are
genes that are less active.

00:06:59.200 --> 00:07:02.480
So in their completely shut down
state, they may have no marks,

00:07:02.480 --> 00:07:04.650
but the DNA is
methylated, silencing

00:07:04.650 --> 00:07:06.670
that region of the genome.

00:07:06.670 --> 00:07:11.280
And other marks then,
culminating in H3K4 trimethyl

00:07:11.280 --> 00:07:15.320
once again when they
become active at the top.

00:07:15.320 --> 00:07:20.040
So I'm summarizing
for you here, decades

00:07:20.040 --> 00:07:23.590
of research in histone marks.

00:07:23.590 --> 00:07:28.520
And it has been
summarized in figures

00:07:28.520 --> 00:07:33.627
like this, where you can
look at different classes

00:07:33.627 --> 00:07:35.960
of genetic elements-- whether
they be promoters in front

00:07:35.960 --> 00:07:40.330
of genes, gene bodies
themselves, enhancers,

00:07:40.330 --> 00:07:42.830
or the large scale
repression of the genome--

00:07:42.830 --> 00:07:44.840
and you can look
at the associated

00:07:44.840 --> 00:07:47.985
marks with those
characteristic elements.

00:07:52.180 --> 00:07:56.340
OK, so, how can we
learn this de novo?

00:07:56.340 --> 00:07:59.520
That is, you could
memorize, and of course it's

00:07:59.520 --> 00:08:01.694
important to
understand, for example,

00:08:01.694 --> 00:08:03.360
if you want to look
for active enhancers

00:08:03.360 --> 00:08:05.470
in the genome, that
looking for things

00:08:05.470 --> 00:08:12.657
like H3K4 monomethyl and H3K7
27 acetyl marks together,

00:08:12.657 --> 00:08:14.740
would give you a good clue
where the enhancers are

00:08:14.740 --> 00:08:17.520
in the genome that are active.

00:08:17.520 --> 00:08:20.150
But if we want to
learn all this de novo,

00:08:20.150 --> 00:08:23.622
without having to memorize it
or rely upon the literature,

00:08:23.622 --> 00:08:26.080
the great thing is that there's
a lot of data out there now

00:08:26.080 --> 00:08:29.990
that characterizes, or profiles
all these marks, genome-wide,

00:08:29.990 --> 00:08:32.080
in variety of cellular states.

00:08:32.080 --> 00:08:34.799
And there's the epigenome
roadmap initiative

00:08:34.799 --> 00:08:38.820
to look at this in hundreds
of different cell types.

00:08:38.820 --> 00:08:43.770
So, what is the histone code?

00:08:43.770 --> 00:08:48.600
That is, how can we
unravel the different marks

00:08:48.600 --> 00:08:51.650
present in the genome and
understand what they mean?

00:08:51.650 --> 00:08:54.810
Because the genome doesn't come
ready-made with those little

00:08:54.810 --> 00:08:57.560
cute labels that we had on
it-- enhancer, gene body,

00:08:57.560 --> 00:08:59.030
and so forth.

00:08:59.030 --> 00:09:00.900
So somehow, if we
want to understand

00:09:00.900 --> 00:09:03.360
the grammar of the
genome and its function,

00:09:03.360 --> 00:09:07.650
we're going to need to be
able to annotate it, hopefully

00:09:07.650 --> 00:09:11.210
with computational help.

00:09:11.210 --> 00:09:13.890
So here's a picture of
what typical data looks

00:09:13.890 --> 00:09:15.970
like along the genome.

00:09:15.970 --> 00:09:19.500
So, obviously you can't
read any of the legends

00:09:19.500 --> 00:09:20.480
on the left-hand side.

00:09:20.480 --> 00:09:22.063
If you want to look
at the slides that

00:09:22.063 --> 00:09:24.850
are posted on Stellar, you
can see the actual marks.

00:09:24.850 --> 00:09:26.780
But the reason I posted
this is because you

00:09:26.780 --> 00:09:28.821
can see the little pink
thing at the top-- that's

00:09:28.821 --> 00:09:32.360
where the RNA transcript has
been mapped to the genome.

00:09:32.360 --> 00:09:35.450
The actual annotated
genes are above.

00:09:35.450 --> 00:09:37.740
And then down below you
can see a whole collection

00:09:37.740 --> 00:09:41.510
of histone marks and other
kinds of chromatin information

00:09:41.510 --> 00:09:43.340
that have been
mapped to the genome

00:09:43.340 --> 00:09:46.710
and spatially
create patterns that

00:09:46.710 --> 00:09:52.880
are suggestive of the function
of the genomic elements,

00:09:52.880 --> 00:09:54.600
if they're properly interpreted.

00:09:54.600 --> 00:10:01.460
And below, you see in blue,
the binding of different TFs,

00:10:01.460 --> 00:10:04.670
as determined by ChIP-seq.

00:10:04.670 --> 00:10:08.600
So, what we would
like to do then,

00:10:08.600 --> 00:10:13.130
is to take this
kind of information

00:10:13.130 --> 00:10:15.810
and automatically learn,
or automatically annotate

00:10:15.810 --> 00:10:20.172
the genome as to its
functional elements.

00:10:20.172 --> 00:10:21.880
Let me stop here and
ask, how many people

00:10:21.880 --> 00:10:27.630
have seen histone mark
information before?

00:10:27.630 --> 00:10:28.820
OK.

00:10:28.820 --> 00:10:32.860
And how many people have
used it in their research?

00:10:32.860 --> 00:10:35.710
Not too many-- a couple people?

00:10:35.710 --> 00:10:37.410
OK.

00:10:37.410 --> 00:10:40.240
So it's getting
quite easy to collect

00:10:40.240 --> 00:10:45.690
and there are a couple of ways
of analyzing this kind of data,

00:10:45.690 --> 00:10:47.760
genome-wide.

00:10:47.760 --> 00:10:51.640
One way is that we could
run a hidden Markov

00:10:51.640 --> 00:10:55.670
model over these data
and predict states

00:10:55.670 --> 00:10:56.730
at regular intervals.

00:10:56.730 --> 00:10:59.460
For example, every 200
bases down the genome,

00:10:59.460 --> 00:11:02.920
and see how the HMM transition
from state to state and let

00:11:02.920 --> 00:11:08.330
the state suggest what the
underlying genome elements

00:11:08.330 --> 00:11:10.920
that we're doing.

00:11:10.920 --> 00:11:16.220
Another way is to use a
dynamic Bayesian network.

00:11:16.220 --> 00:11:19.790
So a dynamic Bayesian network
is simply a Bayesian network.

00:11:19.790 --> 00:11:22.760
We've talked about those before.

00:11:22.760 --> 00:11:25.810
And it models data
sampled along the genome.

00:11:25.810 --> 00:11:29.510
And so it's a directed
acyclic graph.

00:11:29.510 --> 00:11:31.580
There are tools out
there that allow

00:11:31.580 --> 00:11:34.850
us to learn these
models directly.

00:11:34.850 --> 00:11:40.140
And it allows us, as we'll
see, to analyze the genome

00:11:40.140 --> 00:11:45.450
at high resolution, and
to handle missing data.

00:11:45.450 --> 00:11:47.010
So we'll be talking
about Segway,

00:11:47.010 --> 00:11:50.470
which is a particular
dynamic Bayesian network that

00:11:50.470 --> 00:11:52.220
takes the kind of data
we saw on the slide

00:11:52.220 --> 00:11:58.500
before and essentially parses
it into labels that allow us

00:11:58.500 --> 00:12:01.640
to assign function to
different genomic elements.

00:12:01.640 --> 00:12:04.670
And it does this in
an unsupervised way.

00:12:04.670 --> 00:12:07.980
What I mean by that is
that it is automatically

00:12:07.980 --> 00:12:12.100
learning the states,
and then afterwards we

00:12:12.100 --> 00:12:14.950
can look at the states and
assign meaning to them.

00:12:17.660 --> 00:12:23.360
So here is the dynamic Bayesian
network that Segway uses.

00:12:23.360 --> 00:12:26.020
And let me explain
this somewhat scary

00:12:26.020 --> 00:12:28.970
looking diagram of lots of
little boxes and pointers

00:12:28.970 --> 00:12:31.160
to you.

00:12:31.160 --> 00:12:36.380
The genome is described
through the variables

00:12:36.380 --> 00:12:38.880
on the bottom-- the
observation variables,

00:12:38.880 --> 00:12:41.860
going from left to
right, where each base is

00:12:41.860 --> 00:12:44.440
a separate observation
variable which consists

00:12:44.440 --> 00:12:47.730
of the level of a
particular histone mark

00:12:47.730 --> 00:12:51.720
at a particular based position
as described by mapped

00:12:51.720 --> 00:12:54.420
reads to that location.

00:12:54.420 --> 00:12:56.645
The little square
box-- the little boxes

00:12:56.645 --> 00:12:59.050
that says "x" on it with the
other small print you can't

00:12:59.050 --> 00:13:01.880
read-- is simply an
indicator, whether or not

00:13:01.880 --> 00:13:03.890
the data is present.

00:13:03.890 --> 00:13:06.940
If the data is absent, we
don't try and model it.

00:13:06.940 --> 00:13:09.860
If that box contains a zero,
we don't model the data.

00:13:09.860 --> 00:13:13.960
If the box is one, then we
attempt to model the data.

00:13:13.960 --> 00:13:16.710
And the most important part of
the dynamic Bayesian network

00:13:16.710 --> 00:13:20.960
is the q box above, where
those are the states.

00:13:20.960 --> 00:13:25.330
And each state describes an
ensemble of different histone

00:13:25.330 --> 00:13:27.380
marks that are output.

00:13:27.380 --> 00:13:30.260
And so the key thing
is that for each state

00:13:30.260 --> 00:13:33.460
we learn what marks
it's outputting.

00:13:33.460 --> 00:13:35.240
And the model learns
this automatically

00:13:35.240 --> 00:13:37.920
through a learning phase.

00:13:37.920 --> 00:13:42.970
The boxes above
simply are a counter.

00:13:42.970 --> 00:13:47.420
And the counter allows us
to define maximum lengths

00:13:47.420 --> 00:13:51.720
for particular states, so
states don't run on forever.

00:13:51.720 --> 00:13:53.650
So unlike a hidden
Markov model that

00:13:53.650 --> 00:13:55.460
doesn't have that
kind of control,

00:13:55.460 --> 00:14:00.880
we can adjust how long we
want the states to last.

00:14:00.880 --> 00:14:05.880
So this model, if you
turned it 90 degrees

00:14:05.880 --> 00:14:10.220
and rotated it clockwise,
would be more familiar to you

00:14:10.220 --> 00:14:12.585
because all the arrows
would be flowing

00:14:12.585 --> 00:14:14.200
from the top of the screen down.

00:14:14.200 --> 00:14:17.990
There are no cycles in this
directed acyclic graph.

00:14:17.990 --> 00:14:21.271
And therefore, it can be
probabilistically viewed

00:14:21.271 --> 00:14:22.645
and learned in
the same framework

00:14:22.645 --> 00:14:24.970
that we learn a
Bayesian network.

00:14:24.970 --> 00:14:27.790
In fact, it is a
Bayesian network.

00:14:27.790 --> 00:14:29.660
The reason it's
called dynamic is

00:14:29.660 --> 00:14:33.310
because we are learning
temporal information,

00:14:33.310 --> 00:14:36.650
or in this case,
spatial information

00:14:36.650 --> 00:14:38.850
with these different
observations

00:14:38.850 --> 00:14:42.540
along the bottom of the model.

00:14:42.540 --> 00:14:44.760
Now before I go on,
perhaps somebody

00:14:44.760 --> 00:14:46.760
could ask me a question
about the details

00:14:46.760 --> 00:14:49.180
of these dynamic
Bayesian networks,

00:14:49.180 --> 00:14:53.330
because the ability to
automatically assign labels

00:14:53.330 --> 00:14:57.790
to genome function, given
the histone marks is really

00:14:57.790 --> 00:15:00.450
a key thing that's gone on
the last couple of years.

00:15:00.450 --> 00:15:01.401
Yes?

00:15:01.401 --> 00:15:03.325
AUDIENCE: Could you
re-explain that--

00:15:03.325 --> 00:15:06.700
what the labeled-- the second
[INAUDIBLE] was all about?

00:15:06.700 --> 00:15:07.610
PROFESSOR: Sure.

00:15:07.610 --> 00:15:16.300
So the Q label is right
here, these labels.

00:15:16.300 --> 00:15:19.050
And each of these
Q labels defines

00:15:19.050 --> 00:15:20.290
one of a number of states.

00:15:20.290 --> 00:15:23.196
For example, 24
different states.

00:15:23.196 --> 00:15:28.420
In a given state, describes
the expected output

00:15:28.420 --> 00:15:32.100
in terms of what histone marks
are present in that state.

00:15:32.100 --> 00:15:34.701
So it's going to
describe the means of all

00:15:34.701 --> 00:15:35.950
those different histone marks.

00:15:35.950 --> 00:15:38.570
24 different means,
let's say, of the marks

00:15:38.570 --> 00:15:41.090
it's going to output.

00:15:41.090 --> 00:15:46.160
And the job of fitting the model
is picking the right states,

00:15:46.160 --> 00:15:48.770
or a set of 24
states, each of which

00:15:48.770 --> 00:15:53.540
is most descriptive of its
particular subset of chromatin

00:15:53.540 --> 00:15:54.990
marks.

00:15:54.990 --> 00:15:59.100
And then defining how we
transition between states.

00:15:59.100 --> 00:16:04.000
So we not only need to
define what a state means

00:16:04.000 --> 00:16:07.150
in terms of the marks
that it outputs, but also

00:16:07.150 --> 00:16:11.290
when we transition from
one state to another.

00:16:11.290 --> 00:16:13.244
Does that make sense to you?

00:16:13.244 --> 00:16:16.632
AUDIENCE: So I know it
states the information that

00:16:16.632 --> 00:16:18.568
tells at each of the Q boxes.

00:16:18.568 --> 00:16:22.260
Is that a series
of probabilities?

00:16:22.260 --> 00:16:24.592
Or is it something else?

00:16:24.592 --> 00:16:26.970
PROFESSOR: It's actually
a discrete number, right.

00:16:26.970 --> 00:16:30.000
So it actually is a
single-- there's only

00:16:30.000 --> 00:16:31.520
a single state in each Q box.

00:16:31.520 --> 00:16:33.570
So it might be a
number between 1 and 24

00:16:33.570 --> 00:16:35.190
that we're going to learn.

00:16:35.190 --> 00:16:37.460
And based upon
that number, we're

00:16:37.460 --> 00:16:41.970
going to have a
description of the marks

00:16:41.970 --> 00:16:45.190
that we would expect to
see at the observation

00:16:45.190 --> 00:16:47.810
at that particular
genomic location.

00:16:47.810 --> 00:16:53.960
And so our job here is to
learn those 24 different states

00:16:53.960 --> 00:16:58.910
and what they output
in the training phase,

00:16:58.910 --> 00:17:00.870
and then once we've
trained the model,

00:17:00.870 --> 00:17:03.430
we can go back and look
at other held out data,

00:17:03.430 --> 00:17:04.899
and then we can
decode the genome.

00:17:04.899 --> 00:17:06.690
Because we know what
the states are, and we

00:17:06.690 --> 00:17:09.190
know what they are
supposed to be producing,

00:17:09.190 --> 00:17:13.010
we can use a Verterbi decoder
and go back and-- as we

00:17:13.010 --> 00:17:15.930
did with the HMM and we
learned the HMM-- go back

00:17:15.930 --> 00:17:19.550
and read off on the
histone mark sequence

00:17:19.550 --> 00:17:21.640
and figure out what
their relative states are

00:17:21.640 --> 00:17:25.569
for each base position
of the genome.

00:17:25.569 --> 00:17:27.079
Is that helpful?

00:17:27.079 --> 00:17:29.314
Yes?

00:17:29.314 --> 00:17:32.920
Any other questions about
dynamic Bayesian networks?

00:17:32.920 --> 00:17:33.790
Yes?

00:17:33.790 --> 00:17:36.086
AUDIENCE: How do you choose
the number of states?

00:17:36.086 --> 00:17:37.710
PROFESSOR: That's a
very good question.

00:17:37.710 --> 00:17:40.220
How do you choose
the number of states?

00:17:40.220 --> 00:17:43.320
Well, if you choose
too many states,

00:17:43.320 --> 00:17:45.300
they obviously don't
really become descriptive

00:17:45.300 --> 00:17:46.890
and you can become
over fit and then

00:17:46.890 --> 00:17:48.985
can start fitting
noise to your model.

00:17:48.985 --> 00:17:52.260
And if you choose too few
states, what will happen

00:17:52.260 --> 00:17:54.320
is, that states can
get collapsed together

00:17:54.320 --> 00:17:56.430
and they won't be
adequately descriptive.

00:17:56.430 --> 00:17:59.010
The answer is, it's more
or less trial and error.

00:17:59.010 --> 00:18:01.160
There really isn't
a principled way

00:18:01.160 --> 00:18:03.620
to choose the right
number of states

00:18:03.620 --> 00:18:05.580
in this particular context.

00:18:05.580 --> 00:18:06.930
Now, you could do--

00:18:06.930 --> 00:18:08.421
AUDIENCE: What's
the trial, then?

00:18:08.421 --> 00:18:10.906
You run it and you
get a set of things,

00:18:10.906 --> 00:18:12.727
and what do you do
with those labels?

00:18:12.727 --> 00:18:14.310
PROFESSOR: What do
you do with labels?

00:18:14.310 --> 00:18:17.260
AUDIENCE: Yeah, how
do you evaluate it?

00:18:17.260 --> 00:18:19.290
PROFESSOR: You
typically, in both

00:18:19.290 --> 00:18:23.690
of these cases-- both in the
case of chrome HMM and this--

00:18:23.690 --> 00:18:26.180
you rely upon the
previous literature.

00:18:26.180 --> 00:18:29.020
And we saw on that
slide earlier,

00:18:29.020 --> 00:18:31.780
what marks are associated
with what kinds of features.

00:18:31.780 --> 00:18:33.820
So you use the prior
literature and you

00:18:33.820 --> 00:18:36.720
use what the states are telling
you they're describing to try

00:18:36.720 --> 00:18:39.190
and associate those
states with what's

00:18:39.190 --> 00:18:41.540
known about genome function.

00:18:44.641 --> 00:18:45.657
All right, yes?

00:18:45.657 --> 00:18:47.656
AUDIENCE: Where does that
information concerning

00:18:47.656 --> 00:18:50.280
the distance between
states go again?

00:18:50.280 --> 00:18:51.711
Like, the counter?

00:18:51.711 --> 00:18:53.619
Like, how does that
control how long

00:18:53.619 --> 00:18:55.309
the states go on
and whether or not--

00:18:55.309 --> 00:18:57.850
PROFESSOR: What happens is that
the counter at the top, the C

00:18:57.850 --> 00:19:03.200
variables, influence the J
variables you can see there.

00:19:03.200 --> 00:19:04.860
When the J variable
terms to a 1,

00:19:04.860 --> 00:19:07.290
it forces the state transition.

00:19:07.290 --> 00:19:11.530
So the counters count
down and can then

00:19:11.530 --> 00:19:14.240
force a state
transition which will

00:19:14.240 --> 00:19:17.800
cause the Q variable to change.

00:19:17.800 --> 00:19:22.040
It's sort of a-- that particular
formulation of this model

00:19:22.040 --> 00:19:24.750
is a bit of a, sort
of Rube Goldberg kind

00:19:24.750 --> 00:19:26.570
of hackish kind of thing.

00:19:26.570 --> 00:19:29.135
I think to make it get
out of particular states.

00:19:33.210 --> 00:19:38.640
But it works, as we'll
see in just a moment.

00:19:38.640 --> 00:19:39.650
OK.

00:19:39.650 --> 00:19:45.480
So here's an example
of it operating.

00:19:45.480 --> 00:19:50.400
And you can see the different
states on the y-axis here.

00:19:50.400 --> 00:19:53.250
You can see the different
state transitions

00:19:53.250 --> 00:19:55.060
as we go down the genome.

00:19:55.060 --> 00:19:57.670
And you can see the
annotations that it's

00:19:57.670 --> 00:20:00.730
outputting, corresponding
to the histone marks.

00:20:00.730 --> 00:20:04.020
And so what this
is doing is it's

00:20:04.020 --> 00:20:08.130
decoding for us what it thinks
is going on in the genome,

00:20:08.130 --> 00:20:10.620
solely with reference
to the histone marks,

00:20:10.620 --> 00:20:15.180
without reference to primary
sequence or anything else.

00:20:15.180 --> 00:20:17.830
And this kind of
decoding is most useful

00:20:17.830 --> 00:20:22.340
when we want to discover things
like regulatory elements.

00:20:22.340 --> 00:20:27.600
When we want to look for H3K4
mono or dimethyl, and H3K27

00:20:27.600 --> 00:20:31.200
acetyl for example, and identify
those regions of the genome

00:20:31.200 --> 00:20:33.150
that we think are
active enhancers.

00:20:33.150 --> 00:20:33.650
OK.

00:20:37.129 --> 00:20:40.120
OK.

00:20:40.120 --> 00:20:48.310
So, any questions at all about
histone marks and decoding?

00:20:48.310 --> 00:20:50.610
Do you get the
general idea that you

00:20:50.610 --> 00:20:55.620
can assay these histone
marks through ChIP-seq using

00:20:55.620 --> 00:20:59.590
antibodies that are specific
to a particular mark.

00:20:59.590 --> 00:21:04.250
Pull down the histones that
are associated with DNA

00:21:04.250 --> 00:21:06.130
with that mark and map
them to the genome.

00:21:06.130 --> 00:21:10.910
So we get one track for
each ChIP-seq experiment.

00:21:10.910 --> 00:21:15.180
We can profile all the marks
that we think are relevant,

00:21:15.180 --> 00:21:18.470
and then we can look at
what those parks imply

00:21:18.470 --> 00:21:22.710
about both the static
structure of our genome,

00:21:22.710 --> 00:21:31.310
and also how it's being
used as cells differentiate

00:21:31.310 --> 00:21:34.191
or in different
environmental conditions.

00:21:34.191 --> 00:21:34.690
OK.

00:21:37.950 --> 00:21:39.550
OK.

00:21:39.550 --> 00:21:45.730
So, let's go on, then,
to the next step, which

00:21:45.730 --> 00:21:54.310
is that if we understand the
sort of epigenetics state,

00:21:54.310 --> 00:22:02.050
how is that established and
how is the opening of chromatin

00:22:02.050 --> 00:22:06.750
regulated and how is it that
factors find particular places

00:22:06.750 --> 00:22:09.460
in the genome to bind?

00:22:09.460 --> 00:22:13.770
So, the puzzle I talked
to you about earlier

00:22:13.770 --> 00:22:15.900
was that there are
hundreds of thousands

00:22:15.900 --> 00:22:18.170
of particular motifs
in the genome,

00:22:18.170 --> 00:22:20.730
but a very small
number are actually

00:22:20.730 --> 00:22:24.170
bound by regulatory factors.

00:22:24.170 --> 00:22:27.550
And you might think
that the difference

00:22:27.550 --> 00:22:31.430
is that the ones that are bound
have different DNA sequences.

00:22:31.430 --> 00:22:34.410
But in fact, on the
right-hand side, what we see

00:22:34.410 --> 00:22:38.194
is that identical DNA sequences
are bound differentially

00:22:38.194 --> 00:22:39.360
in two different conditions.

00:22:39.360 --> 00:22:41.280
Shown there are
sites that are only

00:22:41.280 --> 00:22:44.300
bound, for example,
in endodermal tissues

00:22:44.300 --> 00:22:46.860
or in ES cells.

00:22:46.860 --> 00:22:49.840
So it isn't the sequence
that's controlling

00:22:49.840 --> 00:22:54.010
the specificity of the
binding, it's something else.

00:22:54.010 --> 00:22:56.360
And we'd like to figure out
what that something else is.

00:22:56.360 --> 00:23:00.700
We'd like to understand
the rules that

00:23:00.700 --> 00:23:03.170
govern where those factors
are binding in the genome.

00:23:06.810 --> 00:23:12.400
So a set of factors are
known that bind to the genome

00:23:12.400 --> 00:23:13.140
and open it.

00:23:13.140 --> 00:23:15.200
They're called pioneer factors.

00:23:15.200 --> 00:23:18.220
There are some well known
pioneer factors like FoxA

00:23:18.220 --> 00:23:22.930
and some of the iPS
reprogramming factors.

00:23:22.930 --> 00:23:26.080
And the idea is that
they're able to bind

00:23:26.080 --> 00:23:28.990
to closed chromatin
and to open it up

00:23:28.990 --> 00:23:33.220
to provide accessibility
to other factors.

00:23:33.220 --> 00:23:36.220
So what we would
like to do, is to see

00:23:36.220 --> 00:23:39.570
if there's a way that we
could, both understand

00:23:39.570 --> 00:23:41.690
how to discover those
factors automatically,

00:23:41.690 --> 00:23:45.790
using a computational
method, and secondarily,

00:23:45.790 --> 00:23:49.190
understand where factors are
binding in a single experiment

00:23:49.190 --> 00:23:50.140
across the genome.

00:23:53.320 --> 00:23:57.279
So the results I'm going to
show you can be summarized here.

00:23:57.279 --> 00:23:58.820
I'm going to show
you a method called

00:23:58.820 --> 00:24:04.070
PIQ that can predict where
TFs bind from DNase-seq data

00:24:04.070 --> 00:24:05.525
that I'll describe in a moment.

00:24:05.525 --> 00:24:08.020
We'll identify pioneer factors.

00:24:08.020 --> 00:24:10.860
We'll show that certain of these
pioneer factors are directional

00:24:10.860 --> 00:24:14.560
and only operate in
one way on the genome.

00:24:14.560 --> 00:24:17.460
And finally, that the
opening of the genome

00:24:17.460 --> 00:24:22.350
allow subtler factors to come
in and to bind to the genome.

00:24:22.350 --> 00:24:27.280
So let's begin with
what DNase-seq data is,

00:24:27.280 --> 00:24:29.420
and how we can use
it to predict where

00:24:29.420 --> 00:24:30.670
TFs are binding to the genome.

00:24:33.700 --> 00:24:37.410
So DNase-seq is a
methodology for exploring

00:24:37.410 --> 00:24:40.320
what parts of the
genome are open.

00:24:40.320 --> 00:24:42.330
So here's the idea.

00:24:42.330 --> 00:24:48.190
You take your cell
and you expose it,

00:24:48.190 --> 00:24:52.280
once you've isolated the
chromatin to DNase-1 which

00:24:52.280 --> 00:24:55.670
will cut or nick
DNA at locations

00:24:55.670 --> 00:24:59.130
where the DNA is open.

00:24:59.130 --> 00:25:01.885
You then can collect the
DNA, size separate it

00:25:01.885 --> 00:25:02.842
and sequence it.

00:25:02.842 --> 00:25:04.300
And thus, you're
going to have more

00:25:04.300 --> 00:25:09.240
reads where the
DNA has been open,

00:25:09.240 --> 00:25:11.365
and less reads were it's
protected by proteins.

00:25:13.910 --> 00:25:16.810
So the cartoon below
gives you an idea

00:25:16.810 --> 00:25:20.670
that, where there are
histones-- each histone

00:25:20.670 --> 00:25:23.910
has about 147 bases of
DNA wrapped around it.

00:25:23.910 --> 00:25:28.230
Or where there are other
proteins hiding the DNA,

00:25:28.230 --> 00:25:32.010
you're going to cast
shadows on this.

00:25:32.010 --> 00:25:37.520
So we're going to be looking
at the shadows and also

00:25:37.520 --> 00:25:40.670
the accessible parts,
by looking directly

00:25:40.670 --> 00:25:41.845
at the DNase-seq reads.

00:25:45.040 --> 00:25:48.180
So if we sequence
deeply enough we

00:25:48.180 --> 00:25:53.140
can understand that
each binding protein has

00:25:53.140 --> 00:25:58.500
its own particular
profile of protection.

00:25:58.500 --> 00:26:01.330
So if you look at these
different proteins,

00:26:01.330 --> 00:26:05.010
they cast particular
shadows on the genome.

00:26:05.010 --> 00:26:09.370
I'm showing here a window
that's 400 base pairs wide.

00:26:09.370 --> 00:26:15.630
This is the average of thousands
of different binding instances.

00:26:15.630 --> 00:26:18.480
So this is not one binding
instance on the top row.

00:26:18.480 --> 00:26:21.550
You can see how CTCF
and other factors

00:26:21.550 --> 00:26:27.470
have particular shadows
they cast or profiles.

00:26:27.470 --> 00:26:29.214
Yes?

00:26:29.214 --> 00:26:33.038
AUDIENCE: How do you know
which factor was at which site?

00:26:33.038 --> 00:26:33.994
[INAUDIBLE].

00:26:33.994 --> 00:26:36.384
PROFESSOR: How do we know
which factor is at which site?

00:26:36.384 --> 00:26:38.140
By the motifs that
are under the site.

00:26:41.400 --> 00:26:42.810
And what's
interesting about CTCF

00:26:42.810 --> 00:26:47.160
is that you can actually see
how it phase the nucleosomes.

00:26:47.160 --> 00:26:51.090
You can see the, sort of,
periodic pattern in CTCF.

00:26:51.090 --> 00:26:55.200
And those dips are where
the nucleosomes are.

00:26:55.200 --> 00:26:58.570
There's a lot you can
tell from these patterns

00:26:58.570 --> 00:27:04.340
about the underlying molecular
mechanism of what's going on.

00:27:04.340 --> 00:27:08.530
Now, you can see at the very
bottom, the aggregate CTCF

00:27:08.530 --> 00:27:10.020
profile.

00:27:10.020 --> 00:27:13.060
And if all the CTCF
bindings looked like that,

00:27:13.060 --> 00:27:14.840
it'd be really easy.

00:27:14.840 --> 00:27:18.180
But above it, as I've shown
you what an individual CTCF

00:27:18.180 --> 00:27:21.040
site looks like, you can
see how sparse it is.

00:27:21.040 --> 00:27:23.520
We just don't get
enough re-density to be

00:27:23.520 --> 00:27:28.730
able to recover a beautiful
protection profile like that.

00:27:28.730 --> 00:27:30.810
So we're always working
against a lot of noise

00:27:30.810 --> 00:27:33.330
in this kind of
biological environment.

00:27:33.330 --> 00:27:35.220
And so our
computational technique

00:27:35.220 --> 00:27:37.030
will need to come up
with an adequate model

00:27:37.030 --> 00:27:39.150
to overcome that noise.

00:27:39.150 --> 00:27:41.970
But if we can, right,
the great promise

00:27:41.970 --> 00:27:44.990
is that with a single
experiment we'll

00:27:44.990 --> 00:27:47.910
be able to identify where all
these different factors are

00:27:47.910 --> 00:27:53.160
binding to the genome
from one set of data.

00:27:53.160 --> 00:28:00.100
So, just reiterating now,
if you think about the input

00:28:00.100 --> 00:28:04.460
to this algorithm-- we're
going to have three things

00:28:04.460 --> 00:28:06.200
that we input to the algorithm.

00:28:06.200 --> 00:28:09.990
We input the original
genome sequence.

00:28:09.990 --> 00:28:12.780
We input the motifs
of the factors

00:28:12.780 --> 00:28:16.550
that we care about, that
we think are interesting.

00:28:16.550 --> 00:28:20.510
And we input the
DNase-seq data that

00:28:20.510 --> 00:28:23.090
has been aligned to the genome.

00:28:23.090 --> 00:28:25.070
So those are the three inputs.

00:28:25.070 --> 00:28:27.310
And the output of
the algorithm is

00:28:27.310 --> 00:28:31.810
going to be the predictions
of which motifs are occupied

00:28:31.810 --> 00:28:35.780
by the factors,
probabilistically.

00:28:35.780 --> 00:28:39.490
And in order to do
that, for each protein

00:28:39.490 --> 00:28:43.070
we need to learn its
protection profile.

00:28:43.070 --> 00:28:45.370
And we need to
score that profile

00:28:45.370 --> 00:28:48.040
against each
instance of the motif

00:28:48.040 --> 00:28:50.440
to see whether or not we
think the protein is actually

00:28:50.440 --> 00:28:54.498
sitting at that
location in the genome.

00:28:54.498 --> 00:28:56.704
Any questions at all about that?

00:29:03.165 --> 00:29:04.160
No?

00:29:04.160 --> 00:29:07.010
OK.

00:29:07.010 --> 00:29:08.280
Don't hesitate to stop me.

00:29:08.280 --> 00:29:12.790
So the design goals for this
particular computational

00:29:12.790 --> 00:29:16.140
algorithm, as I said earlier,
is resistance to low coverage

00:29:16.140 --> 00:29:17.110
and lots of noise.

00:29:17.110 --> 00:29:20.010
To be able to handle
multiple experiment once,

00:29:20.010 --> 00:29:23.070
it has to work on the
entire mammalian genome.

00:29:23.070 --> 00:29:25.390
It has to have high
spatial accuracy

00:29:25.390 --> 00:29:31.890
and it has to have good
behavior in bad cases.

00:29:31.890 --> 00:29:36.970
So in order to model the
underlying re-distribution

00:29:36.970 --> 00:29:40.660
of the genome, what
we're going to do

00:29:40.660 --> 00:29:46.150
is something that
is in principle

00:29:46.150 --> 00:29:47.210
quite straightforward.

00:29:47.210 --> 00:29:49.501
Which is that we're going to
model all accounts that we

00:29:49.501 --> 00:29:52.890
see in the genome by a
Poisson distribution.

00:29:52.890 --> 00:29:55.860
So in each base of
the genome, the counts

00:29:55.860 --> 00:29:58.940
that we see there in
the DNase-seq data

00:29:58.940 --> 00:30:01.080
are modeled by a Poisson.

00:30:01.080 --> 00:30:06.160
And this is assuming that
there's no protein bound there.

00:30:06.160 --> 00:30:09.280
So what we're trying to do
is to model the background

00:30:09.280 --> 00:30:14.310
distribution of counts
without any kind of binding.

00:30:14.310 --> 00:30:17.930
And the log rate
of that Poisson is

00:30:17.930 --> 00:30:21.290
going to be taken from
a multivariate normal.

00:30:21.290 --> 00:30:24.070
And the particular structure
of that multivariate normal

00:30:24.070 --> 00:30:25.760
provides a lot of smoothing.

00:30:25.760 --> 00:30:28.500
So we can learn from
that multivariate normal

00:30:28.500 --> 00:30:31.040
how to fill in
missing information.

00:30:31.040 --> 00:30:33.760
It's very important
to build strength

00:30:33.760 --> 00:30:35.700
from neighboring bases.

00:30:35.700 --> 00:30:38.130
So, even though we may not
have lots of information

00:30:38.130 --> 00:30:39.750
for this base, if
we have information

00:30:39.750 --> 00:30:43.750
for all the bases around us,
we can use that information

00:30:43.750 --> 00:30:47.730
to build strength to estimate
what we should see at this base

00:30:47.730 --> 00:30:51.280
if it's not occupied.

00:30:51.280 --> 00:30:56.950
So the details of how we
learn the mean and the sigma

00:30:56.950 --> 00:30:59.436
matrix you see up
there for estimating

00:30:59.436 --> 00:31:01.310
the multivariate normal
are outside the scope

00:31:01.310 --> 00:31:03.440
of what I'm going
to talk about today.

00:31:03.440 --> 00:31:07.560
But suffice to say, they
can be effectively learned.

00:31:07.560 --> 00:31:13.210
And the second thing we need
to learn are these profiles.

00:31:13.210 --> 00:31:18.170
And so each protein is
going to have a profile.

00:31:18.170 --> 00:31:20.782
Here shown 400 bases wide.

00:31:20.782 --> 00:31:25.080
And it describes how that
protein, so to speak,

00:31:25.080 --> 00:31:26.455
casts a shadow on the genome.

00:31:28.960 --> 00:31:32.230
And we judge the significance
of these profiles--

00:31:32.230 --> 00:31:33.940
and remember that
one of my points

00:31:33.940 --> 00:31:36.250
was I wanted this to be robust.

00:31:36.250 --> 00:31:44.480
So I will not make calls for
proteins where I cannot get

00:31:44.480 --> 00:31:49.090
a robust profile that is
significant above background.

00:31:49.090 --> 00:31:52.740
And I also exclude the
middle region of the profile

00:31:52.740 --> 00:31:56.580
because it's been shown that
the actual cutting enzymes are

00:31:56.580 --> 00:31:59.260
sequence specific
to some extent.

00:31:59.260 --> 00:32:02.160
The DNase-1 cutting enzyme.

00:32:02.160 --> 00:32:04.440
And so we don't
simply want to be

00:32:04.440 --> 00:32:06.985
but picking up sequence
bias in our profile.

00:32:09.740 --> 00:32:14.590
So we learn these
profiles that describe

00:32:14.590 --> 00:32:19.220
for each particular
motif-- and typically we

00:32:19.220 --> 00:32:24.090
can take in hundreds of motifs,
over 500 motifs at once--

00:32:24.090 --> 00:32:27.440
for each motif, what its
protection looks like.

00:32:30.160 --> 00:32:34.350
So what we then have-- we're
going to learn this, actually,

00:32:34.350 --> 00:32:37.480
in an iterative process, but
what we're going to have is--

00:32:37.480 --> 00:32:41.720
now we have a model of what the
unoccupied genome looks like.

00:32:41.720 --> 00:32:47.860
And we have a model of the
reads that a particular protein

00:32:47.860 --> 00:32:50.260
at a motif location
is going to produce.

00:32:52.820 --> 00:32:59.960
And we can put those two
things together and the way

00:32:59.960 --> 00:33:05.290
that we do that is that we
have a binding variable.

00:33:05.290 --> 00:33:07.120
Showing there is delta.

00:33:07.120 --> 00:33:13.500
And we can either add or
not add the binding profile

00:33:13.500 --> 00:33:18.030
of a particular protein in
a location in the genome.

00:33:18.030 --> 00:33:20.960
And that will change the
expected number of counts

00:33:20.960 --> 00:33:23.670
that we see.

00:33:23.670 --> 00:33:32.060
So the key part of this is that
we use a likelihood ratio shown

00:33:32.060 --> 00:33:33.362
as the second probability.

00:33:33.362 --> 00:33:34.820
It's not really a
probability, it's

00:33:34.820 --> 00:33:39.600
a ratio, which is the
probability of a count, given

00:33:39.600 --> 00:33:43.530
that a protein j is
binding at that location,

00:33:43.530 --> 00:33:48.910
versus the probability of the
counts, were it not binding.

00:33:48.910 --> 00:33:52.870
And that quantity
is key because it's

00:33:52.870 --> 00:33:56.370
going to be-- once
we log transform it,

00:33:56.370 --> 00:33:59.594
will be a key component
of our test statistic

00:33:59.594 --> 00:34:01.260
to figure out whether
or not a protein's

00:34:01.260 --> 00:34:02.634
binding at a
particular location.

00:34:05.740 --> 00:34:11.630
And so the way that we go about
that is it we log that ratio

00:34:11.630 --> 00:34:14.670
and we add it to some other
prior information that gives us

00:34:14.670 --> 00:34:20.699
an overall measure
for whether or not

00:34:20.699 --> 00:34:23.260
the protein is binding
at a particular location.

00:34:23.260 --> 00:34:27.770
And then we can rank
these for all the motifs

00:34:27.770 --> 00:34:31.120
for that particular
protein in the genome.

00:34:31.120 --> 00:34:34.107
And then we can make
calls using a null set.

00:34:34.107 --> 00:34:35.940
So we could look in the
genome for locations

00:34:35.940 --> 00:34:39.179
that we know are not occupied,
compute a distribution

00:34:39.179 --> 00:34:43.800
of that statistic,
and then we can say,

00:34:43.800 --> 00:34:46.610
for what values of this
statistic that we observe,

00:34:46.610 --> 00:34:52.030
at the actual motif sites,
is it so unlikely that this

00:34:52.030 --> 00:34:53.110
would occur at random.

00:34:53.110 --> 00:34:57.540
At some desired p
value by looking

00:34:57.540 --> 00:35:01.670
at the area in the
tail of the null set.

00:35:04.290 --> 00:35:09.440
So, just summarizing, we
learn a background model

00:35:09.440 --> 00:35:14.460
of the genome, which
is a Poisson that

00:35:14.460 --> 00:35:18.460
takes log rates from
a multivariate normal.

00:35:18.460 --> 00:35:24.100
We learn patterns, or
profiles of protection,

00:35:24.100 --> 00:35:30.480
or the production of
reads for each motif.

00:35:30.480 --> 00:35:34.250
And at each motif location,
we ask the question

00:35:34.250 --> 00:35:37.260
whether or not, it's
likely that the protein

00:35:37.260 --> 00:35:42.400
was there and actually caused
the reads that we're seeing,

00:35:42.400 --> 00:35:44.310
using a log likelihood ratio.

00:35:49.210 --> 00:35:50.790
So what we're
integrating together,

00:35:50.790 --> 00:35:52.400
when we take all
these things, is

00:35:52.400 --> 00:35:55.280
that we're taking our
original DNA seq-reads,

00:35:55.280 --> 00:36:02.330
we're taking our TF-specific
specific binding profiles.

00:36:02.330 --> 00:36:07.130
We can build strength across
experiments for the background

00:36:07.130 --> 00:36:13.330
model and we can also learn,
to what extent, the strength

00:36:13.330 --> 00:36:18.930
of binding is influenced by
the match of the position--

00:36:18.930 --> 00:36:21.610
a specific weight matrix--
to a particular location

00:36:21.610 --> 00:36:23.970
in the genome.

00:36:23.970 --> 00:36:27.560
And then we can
produce binding calls.

00:36:27.560 --> 00:36:33.480
And when we do so,
it works quite well.

00:36:33.480 --> 00:36:38.710
So here you see three
different mouse ESO factors.

00:36:38.710 --> 00:36:43.820
And the area under
this receiver operating

00:36:43.820 --> 00:36:46.000
curve-- we've talked
about this before.

00:36:46.000 --> 00:36:48.190
Remember a receiver
operating characteristic

00:36:48.190 --> 00:36:50.990
curve-- has false
positives increasing

00:36:50.990 --> 00:36:55.120
on the x-axis and true positives
increasing on the y-axis.

00:36:55.120 --> 00:36:58.040
And if we had a perfect method,
the area under that curve

00:36:58.040 --> 00:37:01.480
would be 1.0.

00:37:01.480 --> 00:37:06.730
And so for this method,
the area under the ROC

00:37:06.730 --> 00:37:10.530
curve for these three
factors, using ChIP-seq data,

00:37:10.530 --> 00:37:16.869
is the absolute gold
standard, is over 0.9.

00:37:16.869 --> 00:37:18.410
And you might say,
well that's great,

00:37:18.410 --> 00:37:20.560
but how well does
it work in general?

00:37:20.560 --> 00:37:23.510
I mean, for example,
the On-code project

00:37:23.510 --> 00:37:26.241
has used hundreds and hundreds
of ChIP-seq experiments

00:37:26.241 --> 00:37:27.740
to profile where
factors are binding

00:37:27.740 --> 00:37:29.780
in different cellular states.

00:37:29.780 --> 00:37:32.620
If you take the DNase-seq data
from those matched cell types

00:37:32.620 --> 00:37:35.495
and you ask, can you reproduce
the ChIP-seq seq data?

00:37:38.670 --> 00:37:42.929
The answer is, a lot
of the time we can,

00:37:42.929 --> 00:37:44.220
using this kind of methodology.

00:37:44.220 --> 00:37:48.310
And that is, the
AUC mean is 0.93

00:37:48.310 --> 00:37:51.540
compared to 313 different
ChIP-seq experiments.

00:37:54.360 --> 00:37:59.090
So this methodology of
looking at open chromatin

00:37:59.090 --> 00:38:02.630
allows us to identify where
lots of different factors

00:38:02.630 --> 00:38:04.710
bind to the genome.

00:38:04.710 --> 00:38:11.680
And about 75 different
factors are strongly

00:38:11.680 --> 00:38:16.150
detectable using
this methodology.

00:38:16.150 --> 00:38:20.040
So it's detectable if
it has a strong motif,

00:38:20.040 --> 00:38:22.350
if it binds in
DNase-accessible regions

00:38:22.350 --> 00:38:25.100
and has strong
DNA-binding affinity.

00:38:25.100 --> 00:38:27.550
So I tell you this
just so you know

00:38:27.550 --> 00:38:30.450
that there are
new methods coming

00:38:30.450 --> 00:38:33.600
that allow us to take
a single experiment

00:38:33.600 --> 00:38:39.830
and analyze it and determine
where a large number of factors

00:38:39.830 --> 00:38:43.730
bind from that single
experimental data set.

00:38:46.290 --> 00:38:49.130
Now, a second question
we wanted to answer

00:38:49.130 --> 00:38:54.595
was, how is it that chrome,
opening and closing is

00:38:54.595 --> 00:38:56.170
controlled?

00:38:56.170 --> 00:39:01.640
And since we had a direct read
out of what chromatin is open,

00:39:01.640 --> 00:39:04.220
because reads are
being produced there,

00:39:04.220 --> 00:39:05.980
we could look in a
experimental system

00:39:05.980 --> 00:39:08.470
where we measured
chromatin accessibility

00:39:08.470 --> 00:39:11.150
through developmental time.

00:39:11.150 --> 00:39:15.910
And the idea was that as we
measured this accessibility,

00:39:15.910 --> 00:39:19.850
we could look at the
places that changed

00:39:19.850 --> 00:39:25.900
and determine what underlying
motifs were present that

00:39:25.900 --> 00:39:29.405
perhaps were causing the
genome to undergo this opening

00:39:29.405 --> 00:39:29.905
process.

00:39:32.760 --> 00:39:37.750
So we developed an
underlying theory

00:39:37.750 --> 00:39:42.340
that pioneer factors would bind
to closed chromatin as shown

00:39:42.340 --> 00:39:46.150
in the middle panel
and open it up,

00:39:46.150 --> 00:39:49.150
and that we could observe those
by looking at the differential

00:39:49.150 --> 00:39:51.320
accessibility of the genome
at two different time

00:39:51.320 --> 00:39:55.080
points that were related.

00:39:55.080 --> 00:39:59.960
And we couldn't observe pioneers
they didn't open up chromatin.

00:39:59.960 --> 00:40:03.720
And for non-pioneers--
obviously the left-hand panel--

00:40:03.720 --> 00:40:08.340
they would not, in
our design here,

00:40:08.340 --> 00:40:09.755
lead to increased accessibility.

00:40:13.050 --> 00:40:22.670
So we then looked at designing
computational indices that

00:40:22.670 --> 00:40:25.263
measured the--
oh, question, yes?

00:40:25.263 --> 00:40:27.195
AUDIENCE: When you
say pioneer factors,

00:40:27.195 --> 00:40:31.542
are you looking at what proteins
are pioneer factors, or are you

00:40:31.542 --> 00:40:34.239
looking at what sequences they
bind to that are [INAUDIBLE].

00:40:34.239 --> 00:40:35.780
PROFESSOR: So the
question is, are we

00:40:35.780 --> 00:40:38.080
looking at what
proteins are factors,

00:40:38.080 --> 00:40:40.042
or are we looking at
what sequence, right?

00:40:40.042 --> 00:40:42.000
What we're doing is,
we're making an assumption

00:40:42.000 --> 00:40:45.920
that the underlying sequence
denotes one or more proteins

00:40:45.920 --> 00:40:48.500
and thus, we are
hypothesizing, there's

00:40:48.500 --> 00:40:51.510
the proteins that are actually
binding to the sequence, that's

00:40:51.510 --> 00:40:52.860
causing that.

00:40:52.860 --> 00:40:55.730
And then later on, we'll go back
and test that experimentally,

00:40:55.730 --> 00:40:57.340
as you'll see in a second.

00:40:57.340 --> 00:41:00.340
OK?

00:41:00.340 --> 00:41:03.330
So here there are three
different metrics,

00:41:03.330 --> 00:41:06.460
which is the dynamic opening
of chromatin from one time

00:41:06.460 --> 00:41:10.690
point to the next, the
static openness of chromatin

00:41:10.690 --> 00:41:14.230
around a particular factor,
and a social index showing

00:41:14.230 --> 00:41:16.700
how many other factors
are around where

00:41:16.700 --> 00:41:20.150
a particular factor binds.

00:41:20.150 --> 00:41:24.660
And you can see that these
things are distributed in a way

00:41:24.660 --> 00:41:29.190
that certain of the factors have
a very high index in multiple

00:41:29.190 --> 00:41:30.225
of these scores.

00:41:33.390 --> 00:41:39.840
And thus, we were
able to classify

00:41:39.840 --> 00:41:44.910
a certain set of factors
as what we classified

00:41:44.910 --> 00:41:48.160
as computational pioneers,
that would open up the genome.

00:41:50.730 --> 00:41:54.000
Now, in any kind of
computational work,

00:41:54.000 --> 00:41:56.400
we're actually looking
at correlative analysis,

00:41:56.400 --> 00:41:57.890
which is never causal.

00:41:57.890 --> 00:41:58.390
Right.

00:41:58.390 --> 00:42:02.470
So we have to go back and we
have to test whether or not

00:42:02.470 --> 00:42:06.590
our computational
predictions are correct.

00:42:06.590 --> 00:42:12.940
So in order to do
that, we built a test

00:42:12.940 --> 00:42:15.880
construct where we
could put the pioneers

00:42:15.880 --> 00:42:20.240
in on the left-hand side
and ask, whether or not

00:42:20.240 --> 00:42:22.840
the pioneer would
open up chromatin

00:42:22.840 --> 00:42:26.600
and enable the expression
of a GFP marker.

00:42:26.600 --> 00:42:28.910
And the red bars
show the factors

00:42:28.910 --> 00:42:30.940
that we thought were pioneers.

00:42:30.940 --> 00:42:36.480
And as you can see, in
this case, all but one

00:42:36.480 --> 00:42:42.520
of the predictive pioneers
produces GFP activity.

00:42:42.520 --> 00:42:44.950
And this construct was
designed in an interesting way.

00:42:44.950 --> 00:42:48.960
We had to design it so that
the pioneers themselves

00:42:48.960 --> 00:42:51.930
were not simply activators.

00:42:51.930 --> 00:42:54.780
And so it was upstream of
another activator, which

00:42:54.780 --> 00:42:57.750
is a retinoic acid
receptor site.

00:42:57.750 --> 00:43:00.180
And so in the absence of
retinoic acid receptor,

00:43:00.180 --> 00:43:03.410
we had to ensure that when
we turned on the pioneer,

00:43:03.410 --> 00:43:06.110
GFP was not turned on.

00:43:06.110 --> 00:43:08.015
It was only with the
addition of the pioneer

00:43:08.015 --> 00:43:12.130
to open the chromatin
and the activator

00:43:12.130 --> 00:43:14.210
that we actually
got GFP expression.

00:43:17.750 --> 00:43:18.690
OK.

00:43:18.690 --> 00:43:24.630
So, through this
methodology we discovered

00:43:24.630 --> 00:43:31.520
about 120 different motifs
corresponding to proteins

00:43:31.520 --> 00:43:36.320
that we found computationally
open-- chromatin out.

00:43:36.320 --> 00:43:37.227
Yes?

00:43:37.227 --> 00:43:38.726
AUDIENCE: [INAUDIBLE]
concentrations

00:43:38.726 --> 00:43:41.720
of different pioneer
factors are different,

00:43:41.720 --> 00:43:44.215
wouldn't that show up
differentially [INAUDIBLE]?

00:43:49.205 --> 00:43:52.070
PROFESSOR: The question
is, if the concentration

00:43:52.070 --> 00:43:53.960
of different pioneer
factors was different,

00:43:53.960 --> 00:43:56.240
wouldn't that show
up differentially?

00:43:56.240 --> 00:43:59.170
And that's precisely, we
think how chromatin structures

00:43:59.170 --> 00:44:01.310
are regulated.

00:44:01.310 --> 00:44:06.320
That we think that the
concentration, or presence

00:44:06.320 --> 00:44:10.160
of different pioneer factors,
is regulating the openness

00:44:10.160 --> 00:44:12.330
or closeness of different
parts of the genome,

00:44:12.330 --> 00:44:16.770
based upon where their
motifs are occurring.

00:44:16.770 --> 00:44:19.916
Is that, in part,
answering your question?

00:44:19.916 --> 00:44:23.360
AUDIENCE: Yes, but,
if a concentration

00:44:23.360 --> 00:44:25.820
of a particular
pioneer factor is low,

00:44:25.820 --> 00:44:30.710
do they necessarily have lesser
binding sites on the genome?

00:44:30.710 --> 00:44:32.920
PROFESSOR: So you're
asking, how is

00:44:32.920 --> 00:44:34.640
the concentration
of a pioneer factor

00:44:34.640 --> 00:44:37.720
related to its ability
to open chromatin

00:44:37.720 --> 00:44:39.780
and whether or not a
higher dosage would

00:44:39.780 --> 00:44:40.820
open more chromatin?

00:44:40.820 --> 00:44:41.430
AUDIENCE: Yes.

00:44:41.430 --> 00:44:44.620
PROFESSOR: I don't have a
good answer to that question.

00:44:44.620 --> 00:44:46.120
Those experiments
haven't been done.

00:44:48.940 --> 00:44:55.680
However, one thing you may have
noticed about these profiles--

00:44:55.680 --> 00:44:58.650
remember these are the same
profiles that we talked

00:44:58.650 --> 00:45:03.200
about earlier of DNase-1
read reproduction

00:45:03.200 --> 00:45:05.220
around a particular factor.

00:45:05.220 --> 00:45:08.305
And what you might notice is
that some of these profiles

00:45:08.305 --> 00:45:08.930
are asymmetric.

00:45:12.030 --> 00:45:14.720
And that they appear to be
producing more region one

00:45:14.720 --> 00:45:16.306
direction than the
other direction.

00:45:19.620 --> 00:45:21.684
And so this is all
computational analysis, right.

00:45:21.684 --> 00:45:23.350
But when you see
something like that you

00:45:23.350 --> 00:45:24.970
say, well gee, why
is that going on?

00:45:24.970 --> 00:45:30.130
Why is it that for NRF-1 the
left-hand side has a lot more

00:45:30.130 --> 00:45:33.960
reads than the right hand side.

00:45:33.960 --> 00:45:38.280
Now, of course, the only reason
that we can produce an oriented

00:45:38.280 --> 00:45:42.530
profile like that is that the
NRF-1 motif is not palindromic,

00:45:42.530 --> 00:45:43.030
right.

00:45:43.030 --> 00:45:45.400
We can actually orient
it in the genome

00:45:45.400 --> 00:45:49.190
and so we know that the
more reads, in this case,

00:45:49.190 --> 00:45:50.850
are coming from
the five prime end

00:45:50.850 --> 00:45:55.070
then from the three prime end.

00:45:55.070 --> 00:45:56.890
So what do you think
would cause that?

00:45:56.890 --> 00:45:58.776
Does anybody have a--
when we first saw this,

00:45:58.776 --> 00:45:59.900
we didn't know what it was.

00:45:59.900 --> 00:46:01.858
But anybody have an idea
of what that could be?

00:46:08.440 --> 00:46:09.750
Oh, yes.

00:46:09.750 --> 00:46:12.060
AUDIENCE: It's the remodelers
that these transcription

00:46:12.060 --> 00:46:15.690
factors are calling in tend to
open the chromatin more on one

00:46:15.690 --> 00:46:17.440
side of the motif
than the other.

00:46:17.440 --> 00:46:19.160
PROFESSOR: Right,
so if the remodelers

00:46:19.160 --> 00:46:23.937
are working in some sort
of directional way, right.

00:46:23.937 --> 00:46:25.020
So that's what we thought.

00:46:25.020 --> 00:46:28.410
We didn't know whether
they were or not.

00:46:28.410 --> 00:46:35.470
And so we went back to our
assay and we tested the motifs,

00:46:35.470 --> 00:46:38.990
both in the forward and
the reverse direction.

00:46:38.990 --> 00:46:39.670
Right.

00:46:39.670 --> 00:46:41.070
To see whether or
not it mattered

00:46:41.070 --> 00:46:45.000
which way the motif
went into the construct,

00:46:45.000 --> 00:46:49.840
based upon selecting factors,
based upon a symmetry

00:46:49.840 --> 00:46:56.060
score that we computed for
their read profile, right?

00:46:56.060 --> 00:47:06.190
And what we found was that, in
fact, it was the case that when

00:47:06.190 --> 00:47:10.840
the motif was properly
oriented it would turn on GFP

00:47:10.840 --> 00:47:14.790
and was in the other
direction it would not.

00:47:14.790 --> 00:47:18.670
So it appeared, for the
factors that we tested,

00:47:18.670 --> 00:47:24.620
that they did have directional
chromatin opening properties.

00:47:24.620 --> 00:47:26.120
And so that's an
interesting concept

00:47:26.120 --> 00:47:28.060
that you actually
can have chromatin

00:47:28.060 --> 00:47:31.400
being opened in one direction
but not the other direction,

00:47:31.400 --> 00:47:33.550
because it admits
the idea of some sort

00:47:33.550 --> 00:47:38.770
of genomic parentheses,
where you could imagine

00:47:38.770 --> 00:47:41.730
part of the genome
being accessible where

00:47:41.730 --> 00:47:43.350
the other part is not.

00:47:47.000 --> 00:47:54.100
And overall this led us to
classifying protein factors

00:47:54.100 --> 00:47:56.230
that are operating in
genome accessibility

00:47:56.230 --> 00:47:59.300
into three classes.

00:47:59.300 --> 00:48:01.660
Here shown as two, where
we have pioneers which

00:48:01.660 --> 00:48:05.530
are the things that
open up the genome,

00:48:05.530 --> 00:48:08.190
and settlers that follow
behind and actually

00:48:08.190 --> 00:48:12.380
bind in the regions where
the chromatin is open.

00:48:12.380 --> 00:48:15.780
That is, it's much more likely
that those factors are going

00:48:15.780 --> 00:48:18.630
to bind where the doors
of the rooms are open,

00:48:18.630 --> 00:48:20.740
and the pioneers are
the proteins that

00:48:20.740 --> 00:48:24.865
come along and open the doors,
in particular, chromatin

00:48:24.865 --> 00:48:25.365
domains.

00:48:27.910 --> 00:48:30.980
And there were a couple of other
tests that we wanted to do.

00:48:30.980 --> 00:48:36.600
We wanted to test whether
or not we could knock out

00:48:36.600 --> 00:48:43.260
this pioneering activity by
taking a pioneer and just

00:48:43.260 --> 00:48:45.880
only including its
DNA-binding domain

00:48:45.880 --> 00:48:47.840
and knocking out the
rest of its domain

00:48:47.840 --> 00:48:52.250
which might be operative
in doing this chromatin

00:48:52.250 --> 00:48:53.225
remodeling.

00:48:53.225 --> 00:48:54.930
And then asked,
whether or not, when

00:48:54.930 --> 00:48:58.830
we expressed this sort
of poisoned pioneer,

00:48:58.830 --> 00:49:03.341
whether or not it would affect
the binding of nearby factors.

00:49:03.341 --> 00:49:06.050
And, in fact, when
you do express

00:49:06.050 --> 00:49:08.880
the sort of poison
pioneer, it does

00:49:08.880 --> 00:49:13.000
reduce the binding
of nearby factors.

00:49:13.000 --> 00:49:15.010
Here, we have a dominant
negative for NFYA

00:49:15.010 --> 00:49:16.830
and dominant negative for NRF1.

00:49:16.830 --> 00:49:23.780
It reduces the binding
of nearby factors.

00:49:23.780 --> 00:49:29.830
And finally, we wanted
to know, if we included

00:49:29.830 --> 00:49:33.820
a dominant negative for
the directional pioneer,

00:49:33.820 --> 00:49:37.100
if it actually
would preferentially

00:49:37.100 --> 00:49:39.730
affect the binding
of [INAUDIBLE] on one

00:49:39.730 --> 00:49:44.290
side of its binding
occurrences or the other side.

00:49:44.290 --> 00:49:46.210
And so we looked
at mix sites that

00:49:46.210 --> 00:49:48.800
were oriented with
respect to NFYA.

00:49:48.800 --> 00:49:53.690
And when we add
the NFYA, you can

00:49:53.690 --> 00:49:59.780
see that it actually-- the
dominant negative NFYA-- when

00:49:59.780 --> 00:50:04.550
the mix site is down of where
we think NFYA is opening up

00:50:04.550 --> 00:50:08.420
the chromatin, the binding
is substantially reduced.

00:50:08.420 --> 00:50:11.170
Whereas, when the
Myc site is not

00:50:11.170 --> 00:50:14.120
on the side where we think
that NFYA is opening,

00:50:14.120 --> 00:50:16.310
it doesn't really
have an effect.

00:50:16.310 --> 00:50:19.350
So this is further
confirmation of the idea

00:50:19.350 --> 00:50:22.330
that in vivo, these
factors are actually

00:50:22.330 --> 00:50:25.716
operating in a directional way.

00:50:25.716 --> 00:50:29.060
Now I tell you all
this because, you know,

00:50:29.060 --> 00:50:31.040
we do a lot of
computational analysis

00:50:31.040 --> 00:50:33.360
and it's important to
follow up and understand

00:50:33.360 --> 00:50:35.670
what the correlations tell us.

00:50:35.670 --> 00:50:37.240
So when you do
computational analysis

00:50:37.240 --> 00:50:40.507
and you see a very
interesting pattern,

00:50:40.507 --> 00:50:42.715
the thing to keep in mind
is, what kind of experiment

00:50:42.715 --> 00:50:46.340
can I design to
test whether or not

00:50:46.340 --> 00:50:48.270
my hypothesis is correct or not?

00:50:51.240 --> 00:50:56.910
We also did an analysis across
human and mouse data sets

00:50:56.910 --> 00:51:00.210
and found that
for a given motif,

00:51:00.210 --> 00:51:02.940
and thus, protein
family, it appeared

00:51:02.940 --> 00:51:04.980
that the chromatin
opening index was largely

00:51:04.980 --> 00:51:08.160
preserved, evolutionarily.

00:51:08.160 --> 00:51:10.910
So that there are
similar pioneers

00:51:10.910 --> 00:51:12.470
between human and mouse.

00:51:15.580 --> 00:51:20.000
Are there any questions
at all about the idea?

00:51:20.000 --> 00:51:22.590
So I told you, I mean, when you
go to cocktail party tonight,

00:51:22.590 --> 00:51:25.390
you say hey, you know, did
you know that DNase-seq

00:51:25.390 --> 00:51:28.420
is this really cool technique
that not only tells you

00:51:28.420 --> 00:51:31.200
whether or not chromatin is
open or not, but, you know,

00:51:31.200 --> 00:51:32.390
where factors bind?

00:51:32.390 --> 00:51:34.900
And some of those factors
open up the chromatin itself

00:51:34.900 --> 00:51:38.780
and, plus, get this,
some of the factors only

00:51:38.780 --> 00:51:42.500
do it in one direction, right.

00:51:42.500 --> 00:51:44.500
That'd be a good
conversation starter, right?

00:51:44.500 --> 00:51:47.320
That'd be the end of
the conversation, no.

00:51:47.320 --> 00:51:49.550
You get the idea, right.

00:51:49.550 --> 00:51:52.655
So are there any questions
about DNase-1 seq analysis?

00:51:55.220 --> 00:51:55.940
Yes?

00:51:55.940 --> 00:51:58.928
AUDIENCE: A little unrelated,
but I was just wondering--

00:51:58.928 --> 00:52:04.420
in the literature where people
have identified factors that

00:52:04.420 --> 00:52:07.850
neither directly reprogram
between different cell types,

00:52:07.850 --> 00:52:10.655
or go through some sort of
[INAUDIBLE] intermediate--

00:52:10.655 --> 00:52:11.280
PROFESSOR: Yes.

00:52:11.280 --> 00:52:13.488
AUDIENCE: There are a number
of transcription factors

00:52:13.488 --> 00:52:16.242
that have been
identified. [INAUDIBLE]

00:52:16.242 --> 00:52:17.661
but there are others.

00:52:17.661 --> 00:52:22.496
Do you often see, or always
see some of the pioneers

00:52:22.496 --> 00:52:24.400
that you've identified
in those cases.

00:52:24.400 --> 00:52:25.050
And then--

00:52:25.050 --> 00:52:25.410
PROFESSOR: Yes.

00:52:25.410 --> 00:52:27.076
AUDIENCE: And then,
a follow-up question

00:52:27.076 --> 00:52:29.695
would be, do you think that if
you took some of the pioneers

00:52:29.695 --> 00:52:32.120
that you generated that
were not known before

00:52:32.120 --> 00:52:35.922
and expressed them
in cell types,

00:52:35.922 --> 00:52:38.352
that they would open
up the chromatin

00:52:38.352 --> 00:52:40.550
sufficiently to potentially
reprogram the mistakes?

00:52:40.550 --> 00:52:41.258
PROFESSOR: Right.

00:52:41.258 --> 00:52:43.130
So the question
was, is it the case

00:52:43.130 --> 00:52:45.870
that known
reprogramming factors,

00:52:45.870 --> 00:52:47.596
at times are powerful pioneers?

00:52:47.596 --> 00:52:50.230
The answer is yes.

00:52:50.230 --> 00:52:53.130
The second question was,
now that you have a broader

00:52:53.130 --> 00:52:55.370
repertoire of pioneer
factors, and you

00:52:55.370 --> 00:52:59.230
can identify what they're
doing, is a possible to,

00:52:59.230 --> 00:53:02.580
in a principled way, engineer
the opening of chromatin

00:53:02.580 --> 00:53:05.375
by perhaps expressing those
factors to see whether or not

00:53:05.375 --> 00:53:07.625
you could match a particular
desired epigenetic state,

00:53:07.625 --> 00:53:09.950
let's say?

00:53:09.950 --> 00:53:12.500
Our preliminary results are yes
on the second count as well.

00:53:12.500 --> 00:53:16.430
That there appear to
be pioneer factors that

00:53:16.430 --> 00:53:18.490
operate, sort of at a
basal level that keep,

00:53:18.490 --> 00:53:23.525
sort of, the sort of usual
rooms open in the genome.

00:53:23.525 --> 00:53:25.150
And then there are
factors that operate

00:53:25.150 --> 00:53:27.600
in a lineage-specific
specific way.

00:53:27.600 --> 00:53:29.720
And when we express
lineage-specific pioneer

00:53:29.720 --> 00:53:34.160
factors, they don't completely
mimic but largely mimic

00:53:34.160 --> 00:53:35.650
the chromatin state
that's present

00:53:35.650 --> 00:53:41.560
in the corresponding
lineage committed cells.

00:53:41.560 --> 00:53:44.320
And so we think that for
principal reprogramming

00:53:44.320 --> 00:53:48.800
of cells, the basal level
of establishing matched

00:53:48.800 --> 00:53:51.300
open states is going to be
an interesting and important

00:53:51.300 --> 00:53:52.530
avenue to explore.

00:53:52.530 --> 00:53:54.620
Does that answer your question?

00:53:54.620 --> 00:53:55.120
Yeah.

00:53:57.830 --> 00:54:00.480
OK.

00:54:00.480 --> 00:54:09.880
So, now we're going to
turn to another-- well let

00:54:09.880 --> 00:54:13.190
me just first summarise
what I just told you about,

00:54:13.190 --> 00:54:14.920
which is that we
can predict where

00:54:14.920 --> 00:54:18.055
TFs bind from DNase-seq data.

00:54:18.055 --> 00:54:19.895
We can identify these
pioneer factors.

00:54:19.895 --> 00:54:21.580
Some of them are directional.

00:54:21.580 --> 00:54:24.860
And other factors follow
these pioneers and bind

00:54:24.860 --> 00:54:25.920
sort of in their wake.

00:54:25.920 --> 00:54:30.620
In where they are actually
open up the chromatin.

00:54:30.620 --> 00:54:35.630
And returning to our
narrative arc for today,

00:54:35.630 --> 00:54:37.920
we've talked about the
idea of histone marks.

00:54:37.920 --> 00:54:40.510
We've talked about the
idea of chromatin openness

00:54:40.510 --> 00:54:42.030
and closeness.

00:54:42.030 --> 00:54:45.270
And now I'd like to talk about
the important question of how

00:54:45.270 --> 00:54:49.940
we can understand which
regulatory regions are

00:54:49.940 --> 00:54:51.435
regulating which genes.

00:54:54.030 --> 00:54:56.540
Now the traditional
way to approach this,

00:54:56.540 --> 00:55:02.770
is that if you have a regulatory
region, the thing that you do

00:55:02.770 --> 00:55:04.890
is you look for
the closest gene.

00:55:04.890 --> 00:55:11.100
And you go, aha, that's the one
that that regulatory region is

00:55:11.100 --> 00:55:13.060
controlling.

00:55:13.060 --> 00:55:15.040
This applies not only
for regulatory regions

00:55:15.040 --> 00:55:15.950
but for snips, right.

00:55:15.950 --> 00:55:19.760
If you find a snip
or a polymorphism

00:55:19.760 --> 00:55:22.630
you are likely to
assume that it's

00:55:22.630 --> 00:55:25.910
regulating the closest gene.

00:55:25.910 --> 00:55:29.770
It could have an effect
on the closest gene.

00:55:29.770 --> 00:55:36.670
But there are other ways of
approaching that question

00:55:36.670 --> 00:55:39.260
with molecular protocols.

00:55:39.260 --> 00:55:45.220
And drawing you once again
a cartoon of genome looping,

00:55:45.220 --> 00:55:50.180
you can see how an enhancer is
coming in contact with the Pol

00:55:50.180 --> 00:55:52.420
II holoenzyme apparatus.

00:55:52.420 --> 00:55:56.080
And this enhancer will
include regulators

00:55:56.080 --> 00:56:00.920
that will cause Pol II
to begin transcription.

00:56:00.920 --> 00:56:05.590
And if somehow we could
capture these complexes

00:56:05.590 --> 00:56:11.990
so that we could examine them
and figure out what bits of DNA

00:56:11.990 --> 00:56:15.130
are associated with
one another, we

00:56:15.130 --> 00:56:19.340
could map, directly, what
enhancers are controlling what

00:56:19.340 --> 00:56:24.310
genes, when they're
active in this form.

00:56:24.310 --> 00:56:30.490
So the essential idea of a
variety of different protocols,

00:56:30.490 --> 00:56:36.960
whether it be protocols
like high c or ChIA-PET

00:56:36.960 --> 00:56:39.850
that we're going to
talk about are the same.

00:56:39.850 --> 00:56:43.570
The difference is that
in the case of ChIA-PET,

00:56:43.570 --> 00:56:45.790
we're only going to look
at interactions that

00:56:45.790 --> 00:56:48.780
are defined by a
particular protein.

00:56:48.780 --> 00:56:51.460
So what we're going to do in
the slides I'm going to show you

00:56:51.460 --> 00:56:53.625
today, is we're
going to only look

00:56:53.625 --> 00:56:56.560
at interactions that are
mediated through RNA polymerase

00:56:56.560 --> 00:56:57.940
II.

00:56:57.940 --> 00:57:00.190
And those are particularly
interesting interactions

00:57:00.190 --> 00:57:03.000
as you can see,
because they involve

00:57:03.000 --> 00:57:06.360
actively transcribed genes.

00:57:06.360 --> 00:57:09.860
So if we could capture all
the RNA polymerase II mediated

00:57:09.860 --> 00:57:16.650
interactions, we'd
be in great shape.

00:57:16.650 --> 00:57:23.070
So, we have a lot of very
talented biologists here.

00:57:23.070 --> 00:57:29.100
So would anybody like to make
a suggestion for a protocol

00:57:29.100 --> 00:57:32.075
for actually revealing
these interactions?

00:57:34.910 --> 00:57:39.260
Does anybody have any ideas
how you'd go about that?

00:57:39.260 --> 00:57:41.165
Or what enzyme
might be involved?

00:57:44.980 --> 00:57:46.510
Any ideas?

00:57:46.510 --> 00:57:49.953
Don't be bashful now.

00:57:49.953 --> 00:57:50.452
Yes.

00:57:50.452 --> 00:57:55.380
AUDIENCE: How about fixing
everything in place where it is

00:57:55.380 --> 00:57:58.967
and then getting
[INAUDIBLE] through DNA.

00:57:58.967 --> 00:57:59.550
PROFESSOR: OK.

00:57:59.550 --> 00:58:02.160
Fixing everything
where it is in place.

00:58:02.160 --> 00:58:03.090
That's good.

00:58:03.090 --> 00:58:06.902
So we might cross link this
whole thing, for example.

00:58:06.902 --> 00:58:07.800
OK.

00:58:07.800 --> 00:58:11.710
And then any other
ideas what we would do?

00:58:11.710 --> 00:58:13.817
That's done, this
protical-- yes.

00:58:13.817 --> 00:58:19.670
AUDIENCE: Well, [INAUDIBLE] that
you've going to be [INAUDIBLE].

00:58:19.670 --> 00:58:23.736
And then digesting the
DNA that's coming out,

00:58:23.736 --> 00:58:26.345
and then that lingers
to the DNA that

00:58:26.345 --> 00:58:30.950
are closest together
in the sequence.

00:58:30.950 --> 00:58:32.060
PROFESSOR: OK.

00:58:32.060 --> 00:58:35.160
So I think what you're
suggesting goes something

00:58:35.160 --> 00:58:36.365
like this.

00:58:36.365 --> 00:58:38.680
All right.

00:58:38.680 --> 00:58:44.770
Which is, that imagine that
we cross link those complexes

00:58:44.770 --> 00:58:47.940
and we precipitate them.

00:58:47.940 --> 00:58:55.310
And then what we do is we,
in a very dilute solution,

00:58:55.310 --> 00:58:58.990
we ligate the DNA together.

00:58:58.990 --> 00:59:02.430
And so we get two kinds
of ligation products.

00:59:02.430 --> 00:59:05.060
On the left-hand side we
get self-ligation products

00:59:05.060 --> 00:59:08.500
where a DNA molecule
ligates to itself.

00:59:08.500 --> 00:59:12.200
And on the right-hand side we
get inner ligation products,

00:59:12.200 --> 00:59:17.960
where the piece of DNA
that the enhancer was on,

00:59:17.960 --> 00:59:22.790
ligates to the pieces of DNA
that the RNA polymerase was

00:59:22.790 --> 00:59:25.540
transcribing the gene on.

00:59:25.540 --> 00:59:28.500
And those inter-ligation
bits of DNA,

00:59:28.500 --> 00:59:32.330
the ones that are red and blue,
are really interesting, right.

00:59:32.330 --> 00:59:35.287
Because they contain both
the enhancer sequence

00:59:35.287 --> 00:59:36.370
and the promoter sequence.

00:59:39.030 --> 00:59:44.690
And all we need to do now is
to sequence those molecules

00:59:44.690 --> 00:59:52.080
from the ends and figure out
where they are in the genome.

00:59:52.080 --> 00:59:53.026
Yes?

00:59:53.026 --> 00:59:56.428
AUDIENCE: How much variation
would there be in the sequence?

00:59:56.428 --> 01:00:00.302
I guess I'm just wondering-- the
RNA polymerase is not static,

01:00:00.302 --> 01:00:00.802
is it?

01:00:00.802 --> 01:00:03.718
In terms of its interaction
with the intenser and the gene.

01:00:03.718 --> 01:00:07.262
I just don't know what
would be capturing in this--

01:00:07.262 --> 01:00:07.970
PROFESSOR: Right.

01:00:07.970 --> 01:00:08.820
AUDIENCE: [INAUDIBLE]
doesn't just

01:00:08.820 --> 01:00:10.590
touch at the beginning
and then [INAUDIBLE].

01:00:10.590 --> 01:00:10.980
PROFESSOR: Right.

01:00:10.980 --> 01:00:12.646
And I think that's a
very good question.

01:00:12.646 --> 01:00:19.140
And in fact, a PhD thesis was
just written on this topic.

01:00:19.140 --> 01:00:22.030
Which is, when you have
proteins that are moving down

01:00:22.030 --> 01:00:24.590
the genome, in
some sense, you're

01:00:24.590 --> 01:00:27.020
looking at a blurred picture.

01:00:27.020 --> 01:00:31.540
So how do you
de-blur the picture

01:00:31.540 --> 01:00:34.320
so that it's brought
sharply into focus?

01:00:34.320 --> 01:00:38.340
And so a compute is something
called a point spread function

01:00:38.340 --> 01:00:43.210
which describes how things are
spread out down the genome.

01:00:43.210 --> 01:00:46.480
And then you invert that to get
a more focused picture of where

01:00:46.480 --> 01:00:49.810
the protein is actually,
primarily located.

01:00:49.810 --> 01:00:50.650
But you're right.

01:00:50.650 --> 01:00:52.714
Things like RNA
polymerase II are not

01:00:52.714 --> 01:00:54.255
thought of as
point-binding proteins.

01:00:54.255 --> 01:00:56.630
They're actually proteins
in motion most time

01:00:56.630 --> 01:00:57.950
when they're doing their work.

01:00:57.950 --> 01:01:01.950
AUDIENCE: [INAUDIBLE]
that it's polymerizing,

01:01:01.950 --> 01:01:04.489
does that it mean that it's
still continually bound

01:01:04.489 --> 01:01:05.280
to the [INAUDIBLE]?

01:01:05.280 --> 01:01:06.950
PROFESSOR: No.

01:01:06.950 --> 01:01:08.660
Although, I don't
think we really

01:01:08.660 --> 01:01:14.350
understand all of the
details of that mechanism.

01:01:14.350 --> 01:01:17.440
But, suffice to say
that what I can do

01:01:17.440 --> 01:01:20.090
is I can start showing
you data and from the data

01:01:20.090 --> 01:01:25.040
we can try and
understand mechanism.

01:01:25.040 --> 01:01:27.147
These are all great
questions, right.

01:01:27.147 --> 01:01:27.646
Yes.

01:01:27.646 --> 01:01:30.194
AUDIENCE: When we did the
citations and ligation,

01:01:30.194 --> 01:01:32.900
you're going to get a lot
of random ligation, right?

01:01:32.900 --> 01:01:34.400
PROFESSOR: A lot
of random ligation?

01:01:34.400 --> 01:01:37.225
AUDIENCE: Yeah, between DNA
sequences that aren't aren't, I

01:01:37.225 --> 01:01:38.890
guess, as close?

01:01:38.890 --> 01:01:41.240
Or you shouldn't really be
ligating certain things?

01:01:41.240 --> 01:01:44.850
PROFESSOR: Well, this picture is
a little bit deceiving, right?

01:01:44.850 --> 01:01:48.020
Because there's actually another
complex just like the one

01:01:48.020 --> 01:01:51.230
at the top, right
to its left, right?

01:01:51.230 --> 01:01:56.620
And you could imagine those
things ligating together.

01:01:56.620 --> 01:01:59.750
And so now you're going to
get ligation products that

01:01:59.750 --> 01:02:01.190
are noise.

01:02:01.190 --> 01:02:02.260
They don't mean anything.

01:02:02.260 --> 01:02:04.267
AUDIENCE: Do you just
throw those out, I guess?

01:02:04.267 --> 01:02:06.850
PROFESSOR: Well, the problem is,
you don't know which ones are

01:02:06.850 --> 01:02:08.300
noise and which ones aren't.

01:02:08.300 --> 01:02:09.920
Right?

01:02:09.920 --> 01:02:13.260
Now, there are some clever
tricks you can play.

01:02:13.260 --> 01:02:17.730
One clever trick is
to change the protocol

01:02:17.730 --> 01:02:23.500
to do these kinds
of reactions, not

01:02:23.500 --> 01:02:29.430
in solution, but
in some sort of gel

01:02:29.430 --> 01:02:32.690
or other thing that
keeps the products apart.

01:02:32.690 --> 01:02:35.770
The other thing you
can do is estimate

01:02:35.770 --> 01:02:38.720
how bad the situation is.

01:02:38.720 --> 01:02:40.330
And how might you do that?

01:02:40.330 --> 01:02:44.650
What you do is, you
take one set of-- you

01:02:44.650 --> 01:02:47.860
take your original preparation
and you split it into two.

01:02:47.860 --> 01:02:48.810
OK.

01:02:48.810 --> 01:02:54.180
And you color this one red and
this one blue using linkers,

01:02:54.180 --> 01:02:55.310
right.

01:02:55.310 --> 01:02:59.130
And then you put them together
and you do this reaction.

01:02:59.130 --> 01:03:01.090
And then you ask,
how many molecules

01:03:01.090 --> 01:03:03.550
have the red and the
blue linkers on them.

01:03:03.550 --> 01:03:06.050
And then you know those are
bad ones because they actually

01:03:06.050 --> 01:03:08.890
came from different
complexes, right.

01:03:08.890 --> 01:03:13.770
And so by estimating the amount
of critical chimeric products

01:03:13.770 --> 01:03:18.260
you get, from that split and
then recombined approach,

01:03:18.260 --> 01:03:21.750
you can optimize the protocol to
reduce the chimeric production

01:03:21.750 --> 01:03:22.430
rate.

01:03:22.430 --> 01:03:26.010
Current chimeric production
rates are about 20%.

01:03:26.010 --> 01:03:27.080
Something of that order.

01:03:27.080 --> 01:03:27.580
OK.

01:03:27.580 --> 01:03:30.070
It used to be 50%,
that's really bad.

01:03:30.070 --> 01:03:31.050
OK.

01:03:31.050 --> 01:03:38.520
So you can try
and optimize that.

01:03:38.520 --> 01:03:42.090
Now, if the protocol
has these issues--

01:03:42.090 --> 01:03:44.660
you have a moving protein
that was brought up here,

01:03:44.660 --> 01:03:46.600
right, that you're
trying to capture.

01:03:46.600 --> 01:03:51.210
You've got a lot of noise
coming from the background

01:03:51.210 --> 01:03:55.340
of these reactions, right.

01:03:55.340 --> 01:03:57.510
Why are we doing this?

01:03:57.510 --> 01:04:00.870
Well, it's the only
game in town right now.

01:04:00.870 --> 01:04:04.030
If you want to have
a mechanistic way

01:04:04.030 --> 01:04:07.900
of understanding what enhancers
are communicating with what

01:04:07.900 --> 01:04:11.960
genes, this and its
family-- I broadly

01:04:11.960 --> 01:04:14.545
call this a family
of protocols--

01:04:14.545 --> 01:04:16.640
is really the only way to go.

01:04:16.640 --> 01:04:18.440
OK.

01:04:18.440 --> 01:04:23.260
The interesting thing
is that when you do,

01:04:23.260 --> 01:04:25.280
you get data like this.

01:04:25.280 --> 01:04:27.540
And so, what you're
looking at here

01:04:27.540 --> 01:04:29.880
is exactly the same
location in the genome.

01:04:29.880 --> 01:04:33.590
It's about 600,000 bases
across from left to right.

01:04:33.590 --> 01:04:34.375
OK.

01:04:34.375 --> 01:04:38.270
And at the very bottom,
you see the SOX2 gene.

01:04:38.270 --> 01:04:41.840
And you have three
different cellular states.

01:04:41.840 --> 01:04:44.070
The top state--
our motor neurons

01:04:44.070 --> 01:04:49.520
have been programmed through
the ectopic expression

01:04:49.520 --> 01:04:52.240
of three transcription factors.

01:04:52.240 --> 01:04:56.520
The second set of
interactions are motor neurons

01:04:56.520 --> 01:04:59.800
that have been
produced by exposure

01:04:59.800 --> 01:05:02.720
to small molecules
over a 7-day period.

01:05:02.720 --> 01:05:05.030
And the bottom set
of interactions

01:05:05.030 --> 01:05:09.824
are from mouse ES cells
that are plueripotent.

01:05:09.824 --> 01:05:11.240
And what's interesting
is that you

01:05:11.240 --> 01:05:16.560
can see how-- I'm
going to point here.

01:05:16.560 --> 01:05:20.550
You can see here-- this is the
SOX2 gene down at the bottom.

01:05:20.550 --> 01:05:24.580
And you can see here--
this regulatory region

01:05:24.580 --> 01:05:30.460
is interacting heavily with
the SOX2 gene at the ESL state.

01:05:30.460 --> 01:05:35.360
And above here, I have
put SOX2 ChIP-seq data.

01:05:35.360 --> 01:05:40.310
So you can actually see that
SOX2 is regulating itself.

01:05:40.310 --> 01:05:46.320
And up here, we have the
same SOX2 gene locus.

01:05:46.320 --> 01:05:50.710
And OLIG2 is a key regulator
of this motor neuron fate.

01:05:50.710 --> 01:05:53.600
And you can see that it
appears that OLIG2 is now

01:05:53.600 --> 01:05:56.490
regulating SOX2.

01:05:56.490 --> 01:06:00.760
And we don't have as complete
dependence upon the SOX2 locus

01:06:00.760 --> 01:06:02.760
as we had before.

01:06:02.760 --> 01:06:05.730
And up here in the induced
motor neuron state,

01:06:05.730 --> 01:06:08.180
LHX4 is one of the
reprogramming factors

01:06:08.180 --> 01:06:12.560
and you can see how it is
interacting with SOX2 here

01:06:12.560 --> 01:06:15.480
and over here.

01:06:15.480 --> 01:06:19.890
So what this methodology
allows us to do,

01:06:19.890 --> 01:06:25.380
is to tie these regulatory
regions to the genes

01:06:25.380 --> 01:06:32.750
that they are regulating,
albeit it with some issues.

01:06:32.750 --> 01:06:37.876
So, we'll talk about the
issues in just a second.

01:06:37.876 --> 01:06:45.290
Are there any questions at all
about the idea of capturing,

01:06:45.290 --> 01:06:50.000
in essence, the folding of the
genome with this methodology

01:06:50.000 --> 01:06:53.620
to link regulatory
regions to genes?

01:06:56.830 --> 01:06:57.350
Yes?

01:06:57.350 --> 01:06:58.646
AUDIENCE: I have a question.

01:06:58.646 --> 01:07:02.678
So in each of
those charts you've

01:07:02.678 --> 01:07:05.392
got parts describing regions
that are interacting.

01:07:05.392 --> 01:07:06.017
PROFESSOR: Yes.

01:07:06.017 --> 01:07:07.450
AUDIENCE: Is that correct?

01:07:07.450 --> 01:07:08.510
PROFESSOR: Yes.

01:07:08.510 --> 01:07:12.290
The little loops underneath
are the actual read pairs

01:07:12.290 --> 01:07:14.560
that came out of the sequencer.

01:07:14.560 --> 01:07:16.780
And the green dotted
lines are the interactions

01:07:16.780 --> 01:07:18.384
I'm suggesting are significant.

01:07:21.040 --> 01:07:23.200
So I'm showing you
the raw data and I'm

01:07:23.200 --> 01:07:28.360
showing you the hypothesized
or purported interactions

01:07:28.360 --> 01:07:32.168
with the green dotted lines.

01:07:32.168 --> 01:07:32.668
Right?

01:07:35.500 --> 01:07:36.640
Right?

01:07:36.640 --> 01:07:42.810
AUDIENCE: So how is you raw
sequencing then transformed

01:07:42.810 --> 01:07:46.720
into this set of interactions?

01:07:46.720 --> 01:07:49.890
PROFESSOR: How is the raw
sequencing data-- remember

01:07:49.890 --> 01:07:52.970
that what came out
of the protocol

01:07:52.970 --> 01:07:59.500
were molecules on the
right-hand side that

01:07:59.500 --> 01:08:05.307
had little bits of DNA from two
different places in the genome.

01:08:05.307 --> 01:08:07.453
AUDIENCE: I'm
sorry, I meant, how

01:08:07.453 --> 01:08:11.508
did you determine-- because
I'm assuming each of these arcs

01:08:11.508 --> 01:08:14.615
has to have a single base start
side and a single base end

01:08:14.615 --> 01:08:15.114
site.

01:08:15.114 --> 01:08:15.580
PROFESSOR: Correct.

01:08:15.580 --> 01:08:16.955
AUDIENCE: However,
your reads are

01:08:16.955 --> 01:08:19.199
going to span-- your
joined paired reads are

01:08:19.199 --> 01:08:21.620
going to span a number of bases.

01:08:21.620 --> 01:08:23.451
So you have a number
of bases coming

01:08:23.451 --> 01:08:25.457
from the red part
and a number of bases

01:08:25.457 --> 01:08:26.540
coming from the blue part.

01:08:26.540 --> 01:08:28.732
PROFESSOR: We've got
20, 20 something, yeah.

01:08:28.732 --> 01:08:32.099
AUDIENCE: How do you determine
which of these red bases

01:08:32.099 --> 01:08:34.504
and which of these blue
bases are your start

01:08:34.504 --> 01:08:36.430
and end points for
the [INAUDIBLE].

01:08:36.430 --> 01:08:38.830
PROFESSOR: Well, you are
looking at a 600,000 base pair

01:08:38.830 --> 01:08:41.830
window of the
genome and we're not

01:08:41.830 --> 01:08:43.813
quite at the resolution
of 28 bases yet.

01:08:43.813 --> 01:08:44.354
AUDIENCE: OK.

01:08:44.354 --> 01:08:46.460
PROFESSOR: So, you know--

01:08:46.460 --> 01:08:49.960
AUDIENCE: So this is not
necessarily single base pair

01:08:49.960 --> 01:08:52.044
resolution, but this
is a region resolution?

01:08:52.044 --> 01:08:55.109
Is that correct?

01:08:55.109 --> 01:08:57.300
PROFESSOR: Once
again, the question

01:08:57.300 --> 01:09:00.580
of how to improve the spatial
resolution of these results

01:09:00.580 --> 01:09:02.720
is a subject of active research.

01:09:02.720 --> 01:09:05.990
And once again, you
can deconvolve things

01:09:05.990 --> 01:09:08.590
like the shearing to
actually get things

01:09:08.590 --> 01:09:12.779
down to within, say, 10 to
100 base pairs resolution.

01:09:12.779 --> 01:09:13.667
AUDIENCE: OK.

01:09:13.667 --> 01:09:14.250
PROFESSOR: OK?

01:09:14.250 --> 01:09:15.770
AUDIENCE: Got it.

01:09:15.770 --> 01:09:19.659
PROFESSOR: But you can't
identify the exact motif

01:09:19.659 --> 01:09:21.620
that the things land on, right.

01:09:21.620 --> 01:09:23.920
They can get in the
ballpark, so to speak, right.

01:09:23.920 --> 01:09:27.970
You can figure out where
you need to look for motifs.

01:09:27.970 --> 01:09:32.850
And so one thing
that we and others do

01:09:32.850 --> 01:09:34.760
is look at these
regions and we ask

01:09:34.760 --> 01:09:38.740
what motifs are present
into these regions.

01:09:38.740 --> 01:09:42.060
Or if you have match DNase-seq
data, you can go back

01:09:42.060 --> 01:09:44.979
and you can say, aha,
I have DNase-seq data.

01:09:44.979 --> 01:09:48.180
I have this data and
I know that there's

01:09:48.180 --> 01:09:50.180
something going on at
that region of the genome.

01:09:50.180 --> 01:09:51.971
What proteins do I
think are sitting there,

01:09:51.971 --> 01:09:55.330
based upon the protection
profiles I see.

01:09:55.330 --> 01:09:55.830
Right.

01:09:55.830 --> 01:09:57.454
So you can take an
integrative approach

01:09:57.454 --> 01:09:59.740
where you use different
data types to begin

01:09:59.740 --> 01:10:02.035
to pick apart the
regulatory network.

01:10:02.035 --> 01:10:05.600
Where you see the connections
directly molecularly,

01:10:05.600 --> 01:10:08.070
and you see the
regulatory proteins

01:10:08.070 --> 01:10:11.101
that are binding
at those locations.

01:10:11.101 --> 01:10:11.600
OK?

01:10:11.600 --> 01:10:13.225
Was that helpful?

01:10:13.225 --> 01:10:13.725
Good.

01:10:13.725 --> 01:10:14.557
Good questions.

01:10:14.557 --> 01:10:15.390
Any other questions?

01:10:15.390 --> 01:10:15.925
Yes?

01:10:15.925 --> 01:10:17.550
AUDIENCE: Would you consider
Hi-C and 5C and all of those

01:10:17.550 --> 01:10:19.059
to be the same
family of technique?

01:10:19.059 --> 01:10:19.850
PROFESSOR: I would.

01:10:19.850 --> 01:10:25.520
They're all, sort of the same
family and they're improving.

01:10:25.520 --> 01:10:28.860
I'm about to tell you why
this doesn't work very well.

01:10:28.860 --> 01:10:32.770
But, that said, it's the
best thing we have going.

01:10:32.770 --> 01:10:34.220
Right.

01:10:34.220 --> 01:10:37.060
5C is not any to any.

01:10:37.060 --> 01:10:38.920
It's to one to any.

01:10:38.920 --> 01:10:43.740
This protocol, when you do
one experiment with this,

01:10:43.740 --> 01:10:47.170
it tells you all the interacting
regions in the genome.

01:10:47.170 --> 01:10:49.070
Right.

01:10:49.070 --> 01:10:51.170
I believe 5C-- help
me if I'm wrong.

01:10:51.170 --> 01:10:52.900
You pick one anchor
location and then

01:10:52.900 --> 01:10:54.290
you can tell all the
regions and genomes that

01:10:54.290 --> 01:10:56.150
are interacting with
that anchor location.

01:10:56.150 --> 01:10:57.150
AUDIENCE: Isn't that 3C?

01:10:57.150 --> 01:10:58.137
PROFESSOR: What?

01:10:58.137 --> 01:10:59.220
AUDIENCE: 3C's one to one.

01:10:59.220 --> 01:11:00.916
4C's one to any.

01:11:00.916 --> 01:11:01.790
AUDIENCE: And 5C is--

01:11:01.790 --> 01:11:02.987
AUDIENCE: 5C's any to any.

01:11:02.987 --> 01:11:04.435
PROFESSOR: And 5C's any to any?

01:11:04.435 --> 01:11:04.935
OK.

01:11:04.935 --> 01:11:06.580
I stand correct.

01:11:06.580 --> 01:11:07.280
Thank you.

01:11:10.751 --> 01:11:11.250
Yeah.

01:11:14.501 --> 01:11:15.000
OK.

01:11:17.660 --> 01:11:20.950
You didn't critique
my bond type.

01:11:20.950 --> 01:11:23.091
See I was trying to
get you and you didn't.

01:11:23.091 --> 01:11:23.590
OK.

01:11:26.920 --> 01:11:28.240
And other questions about this?

01:11:31.520 --> 01:11:32.810
OK.

01:11:32.810 --> 01:11:34.940
What could go wrong?

01:11:34.940 --> 01:11:35.964
What could go wrong?

01:11:35.964 --> 01:11:37.630
Well, I can tell you
what will go wrong.

01:11:37.630 --> 01:11:44.290
What will go wrong is that it
has a low true positive rate.

01:11:44.290 --> 01:11:44.790
OK.

01:11:47.860 --> 01:11:49.200
And how can you tell that?

01:11:49.200 --> 01:11:53.080
You do the experiment
twice and you

01:11:53.080 --> 01:11:55.680
get thousands of interactions
from each experiment in exactly

01:11:55.680 --> 01:11:59.710
matched conditions and
there's a very small overlap

01:11:59.710 --> 01:12:02.230
between the conditions.

01:12:02.230 --> 01:12:04.830
Oops.

01:12:04.830 --> 01:12:08.180
So, that's a pretty
big oops, right?

01:12:08.180 --> 01:12:11.350
Because you would like it
to be the case that when

01:12:11.350 --> 01:12:15.460
you do an experiment multiple
times, you get the same answer.

01:12:15.460 --> 01:12:19.560
So let us just
suppose that you get

01:12:19.560 --> 01:12:23.562
10,000 interactions
in experiment one.

01:12:23.562 --> 01:12:27.380
10,000 interactions in
experiment two, but only

01:12:27.380 --> 01:12:31.380
2,000 of them are the same.

01:12:31.380 --> 01:12:33.750
What could possibly
be going wrong?

01:12:37.190 --> 01:12:37.760
Any ideas?

01:12:40.264 --> 01:12:42.430
If you're looking at the
data, what would you think?

01:12:48.730 --> 01:12:50.080
Well?

01:12:50.080 --> 01:12:50.580
Yeah?

01:12:50.580 --> 01:12:52.371
AUDIENCE: [INAUDIBLE]
could be really high,

01:12:52.371 --> 01:12:54.994
so you're just seeing
a couple of things

01:12:54.994 --> 01:12:56.782
that are above the background.

01:12:56.782 --> 01:12:58.130
And they don't necessarily--

01:12:58.130 --> 01:12:58.838
PROFESSOR: Right.

01:12:58.838 --> 01:13:00.500
So is it maybe
that, you know, it's

01:13:00.500 --> 01:13:02.280
just tough to get
these interactions out.

01:13:02.280 --> 01:13:06.162
And so you got a lot
of background trash.

01:13:06.162 --> 01:13:07.620
And the things that
are significant

01:13:07.620 --> 01:13:12.185
are tough to pick out.

01:13:12.185 --> 01:13:12.685
Yeah?

01:13:12.685 --> 01:13:16.509
AUDIENCE: Maybe it's a real
biological noise issue?

01:13:16.509 --> 01:13:20.733
So rather than the technique,
actually any given time that

01:13:20.733 --> 01:13:24.461
the interactions as so diverse
that when you take the snap

01:13:24.461 --> 01:13:25.460
shot you can't--

01:13:25.460 --> 01:13:26.518
PROFESSOR: I like
that explanation

01:13:26.518 --> 01:13:28.180
because it's very pleasing
and makes me feel good.

01:13:28.180 --> 01:13:29.470
And I would be hopeful
that that would be true

01:13:29.470 --> 01:13:31.845
that there's enough biological
noise that that's actually

01:13:31.845 --> 01:13:32.885
what I'm observing.

01:13:32.885 --> 01:13:34.810
It doesn't make me feel
too warm and fuzzy,

01:13:34.810 --> 01:13:36.830
but you know, I'd
go with that, right.

01:13:40.194 --> 01:13:41.860
The other thing you
might think is, gee,

01:13:41.860 --> 01:13:43.720
if we just sequenced
that library more,

01:13:43.720 --> 01:13:46.670
we'd get more interactions
out them, right?

01:13:46.670 --> 01:13:49.120
So you go off and you compute
the library complexity

01:13:49.120 --> 01:13:53.210
of your library and you go,
oops, that's not going to work.

01:13:53.210 --> 01:13:55.980
There just isn't enough
diversity in the library.

01:13:55.980 --> 01:13:59.080
Meaning that the underlying
biological protocol did not

01:13:59.080 --> 01:14:02.400
produce enough of those
interesting inner ligation

01:14:02.400 --> 01:14:06.560
events to allow you to reveal
more information about what's

01:14:06.560 --> 01:14:07.840
going on.

01:14:07.840 --> 01:14:09.890
OK.

01:14:09.890 --> 01:14:15.130
Now if I ask you to judge the
significance of an interaction

01:14:15.130 --> 01:14:16.630
pair here.

01:14:16.630 --> 01:14:19.360
Let's think about
this using what

01:14:19.360 --> 01:14:22.060
we know already
from the subject.

01:14:22.060 --> 01:14:23.290
OK.

01:14:23.290 --> 01:14:25.375
So I'm going to draw a picture.

01:14:28.480 --> 01:14:31.470
So I have my genome.

01:14:31.470 --> 01:14:37.400
And let's just say that I have
a location, CA and a location CB

01:14:37.400 --> 01:14:43.690
and I have a pile of ends that
wind up in those two locations.

01:14:43.690 --> 01:14:45.200
OK.

01:14:45.200 --> 01:14:51.080
And what I would like
to know is-- and I have,

01:14:51.080 --> 01:14:56.230
let me just see what
variable I used for this.

01:14:56.230 --> 01:15:01.900
And I have a certain number of
interactions between a and b.

01:15:01.900 --> 01:15:05.840
That is I have a certain
number of reads that

01:15:05.840 --> 01:15:09.080
cross between these two
locations in the genome.

01:15:09.080 --> 01:15:11.511
And I'd like to know whether
or not this number of reads

01:15:11.511 --> 01:15:12.135
is significant.

01:15:14.860 --> 01:15:17.240
OK.

01:15:17.240 --> 01:15:18.652
How could I estimate that?

01:15:21.810 --> 01:15:24.500
Any ideas?

01:15:24.500 --> 01:15:29.140
Oh, I'm also going
to tell you that n

01:15:29.140 --> 01:15:38.160
is the total number
of read ends observed.

01:15:42.700 --> 01:15:43.200
OK.

01:15:46.040 --> 01:15:49.530
Well, here is the idea.

01:15:49.530 --> 01:15:54.240
I've got n total
read ends, right?

01:15:54.240 --> 01:15:57.770
I've got ca read ends here.

01:15:57.770 --> 01:16:00.720
I've got cv read
ends here, and I

01:16:00.720 --> 01:16:05.490
have iab that are overlapping.

01:16:05.490 --> 01:16:07.890
So now, this is just our old
friend, the hypergeometric,

01:16:07.890 --> 01:16:08.390
right.

01:16:08.390 --> 01:16:11.290
We can ask what is the
probability of that happening

01:16:11.290 --> 01:16:12.960
at random?

01:16:12.960 --> 01:16:18.460
This many interactions or
fewer would happen at random.

01:16:18.460 --> 01:16:20.510
And if it's very
unlikely, we would

01:16:20.510 --> 01:16:23.330
reject the null hypothesis
and accept that there's really

01:16:23.330 --> 01:16:25.730
an interaction going on here.

01:16:25.730 --> 01:16:27.390
OK?

01:16:27.390 --> 01:16:31.130
So, just to be more
precise about that.

01:16:31.130 --> 01:16:32.380
This is what it looks like.

01:16:32.380 --> 01:16:34.170
You've seen this before.

01:16:34.170 --> 01:16:37.300
That the probability of
those interactions happening

01:16:37.300 --> 01:16:40.610
on a null model, given a total
number of interactions end in

01:16:40.610 --> 01:16:45.070
ca and cb is given by
the hypergeometric.

01:16:45.070 --> 01:16:46.640
OK.

01:16:46.640 --> 01:16:50.690
So that's one way of going
about assessing whether or not

01:16:50.690 --> 01:16:52.672
the interactions we
see are significant.

01:16:56.920 --> 01:17:00.156
Now, let me ask you a
slightly different question.

01:17:00.156 --> 01:17:00.655
Right.

01:17:04.290 --> 01:17:13.900
Imagine that I have-- and
I'm being very generous here.

01:17:13.900 --> 01:17:20.330
Imagine that I have
two experiment-- that's

01:17:20.330 --> 01:17:21.365
the wrong size bubbles.

01:17:24.610 --> 01:17:25.870
I don't want to mislead you.

01:17:28.772 --> 01:17:30.480
One of your friends
comes to you and say,

01:17:30.480 --> 01:17:32.430
"I've done this
experiment twice."

01:17:32.430 --> 01:17:34.690
Twice, OK.

01:17:34.690 --> 01:17:41.900
"And each time I get
1,000 interactions.

01:17:41.900 --> 01:17:46.600
So each one gives
you 1,000, let's say.

01:17:46.600 --> 01:17:54.290
And I have 900 that are common
between the two replicates.

01:17:54.290 --> 01:17:57.150
And your friend says,
"how many interactions

01:17:57.150 --> 01:18:02.190
do you think there
are in total?"

01:18:02.190 --> 01:18:03.328
How could we estimate that?

01:18:07.720 --> 01:18:11.570
Well, what's interesting
about this problem

01:18:11.570 --> 01:18:19.470
is that what we're
asking is what's n?

01:18:19.470 --> 01:18:19.970
Right.

01:18:19.970 --> 01:18:22.670
What's the total number of
interactions of which we're

01:18:22.670 --> 01:18:26.810
observing this set and this set
of which 900 is overlapping.

01:18:26.810 --> 01:18:29.860
There's the hyperlink
geometric again.

01:18:29.860 --> 01:18:36.420
So all we need to do is to find
the maximum value, the best

01:18:36.420 --> 01:18:43.800
value for n that predicts the
observed overlap given that we

01:18:43.800 --> 01:18:49.460
have two experiments
of size, with m

01:18:49.460 --> 01:18:54.510
and n different observations,
and we have an overlap of k.

01:18:54.510 --> 01:18:55.335
OK.

01:18:55.335 --> 01:18:57.820
Does that makes
sense to everybody?

01:18:57.820 --> 01:19:02.760
Of how to estimate the
total number of interactions

01:19:02.760 --> 01:19:06.016
out there making a
set of assumption that

01:19:06.016 --> 01:19:07.199
they're all equally likely.

01:19:10.990 --> 01:19:12.916
Any questions about that at all?

01:19:19.540 --> 01:19:20.440
OK.

01:19:20.440 --> 01:19:26.670
And, just so you know, you can
approximate this, this way.

01:19:26.670 --> 01:19:30.710
Which is that the maximum
likelihood estimate

01:19:30.710 --> 01:19:32.690
of the total number
of interactions

01:19:32.690 --> 01:19:35.280
is approximately
n times n over k,

01:19:35.280 --> 01:19:38.780
as seen by the
approximation on the bottom.

01:19:38.780 --> 01:19:40.780
OK?

01:19:40.780 --> 01:19:43.850
Just so that you can
approximate how many things

01:19:43.850 --> 01:19:45.810
are out there that you
haven't seen when you've

01:19:45.810 --> 01:19:48.460
done a couple of replicates.

01:19:48.460 --> 01:19:50.200
OK, you guys have
been totally great.

01:19:50.200 --> 01:19:52.033
We've talked about a
lot of different things

01:19:52.033 --> 01:19:54.830
today in chromatin
architecture and structure.

01:19:54.830 --> 01:19:57.365
Sort of the DC to light
version of chromatin structure

01:19:57.365 --> 01:20:00.670
and architecture lecture.

01:20:00.670 --> 01:20:03.720
Next time we're going
to talk about building

01:20:03.720 --> 01:20:05.677
genetic models of EQTLs.

01:20:05.677 --> 01:20:07.135
And the time after
that we're going

01:20:07.135 --> 01:20:09.220
to talk about human genetics.

01:20:09.220 --> 01:20:10.380
Thank you so much.

01:20:10.380 --> 01:20:11.730
Have a great, long weekend.

01:20:11.730 --> 01:20:13.850
We'll see you next Thursday.