As DNA reveals its secrets, scientists are assembling a new picture of humanity
statnews.com
statnews.com
As many in the field can tell you, "labeling" genomes is hard, too. You can stick labels like "died at 39 of AVCR" or "had 5 primary tumors by age 28" or "insane multi-substance abuser with extra toes" on a person, but that doesn't really encapsulate all their traits. I claim that genome analysis is a great candidate for semi-supervised learning. Looks like Ben Paten already had that thought...
Haussler, Paten, McVean, and the usual suspects are working on a tractable graph representation to replace (say) hg39, i.e. instead of a new "reference genome", there should be thousands. This makes more sense when you look at how common structural variants (say, inversions) are, and then when you do something like ask Haussler at ASHG "how do we represent inversions in this thing?" you realize how fcking hard it really is to get it right.
McVean & co. made it work for the major histocompatibility complex, which immunologists can explain better than I can, but that's perhaps the most diverse bit of genome there is. It's riddled with ancient repeat element insertions and generally fascinating. It's also a source of Little Problems like organ transplant rejections, graft vs. host disease in stem cell transplants, maybe bits of schizophrenia and neurodegenerative disease... in short, if it can be done for the MHC, it's probably doable for the whole enchilada. It's not easy, though. In fact it's so difficult to get right that a sub-field of accurate HLA typing and immunotyping evolved in parallel to large-scale genome analysis, because it's really, really important not to fuck this up. So that's a note of caution from historical precedent.
Nonetheless, it turns out that your pals at Google spearheaded the effort to have a Hadoop-like "data lake" for genomes (sign up for the GA4GH mailing lists if you like to watch professors bikeshed an API, and occasionally produce incredible insights by accident). Maybe this is going to converge in an interesting way. It won't happen overnight, but it will happen, and the mathematicians will be vindicated.
* Seven Bridges is a little genomics company with a funny name. Unless you're familiar with Eulerian paths and the Seven Bridges of Konigsberg, in which case it's sort of obvious why they chose their name. Nice people, other than the patent, which infuriated most everyone else.
1) in the absence of a reference genome, does Burrows-Wheeler not apply?
2) Could you recommend any good articles to start from?
You can look at Pall's work to see how the current approaches may evolve into something compressible:
https://github.com/GFA-spec/GFA-spec
As far as articles? McVean's, without a doubt!
http://www.nature.com/ng/journal/v47/n6/full/ng.3257.html
An implementation paper for graph assembly HLA typing is at:
http://biorxiv.org/content/early/2015/12/24/035253
Interesting times ahead, with people recognizing that GxE and GxExE matters far more than G alone.
http://bioinformatics.oxfordjournals.org/content/29/13/i361....
The other thing that would be neat is that then you'd have a direct tie-in to ancestral recombination graphs and could, in principle, get IBS/IBD for the same cost as high-confidence genotyping for any two individuals. Come to think of it, there's probably a way to recast this as shortest paths and get all admissible traversals between a population of genotyped individuals (given an ARG) for the same price as any two. Hmmm. This is a little disturbing.
Why wouldn't it have been obvious 16 years ago that a thoughtfully designed data model was necessary, possibly using graphs, to account for variability and other attributes of the genome?
Surely it was foreseeable that tooling would be crucial and that a solid software foundation would be invaluable to enable efficient and flexible processing for years to come?
Who were the lead developers supporting the original public genome project and what were they thinking?
The rate of technology change for sequencing capabilities in the past 16 years makes Moore's law look like the rate of change in battery technology.
16 years ago I didn't think it would be possible to sequence individuals in the clinic before 2050 or so. Now we are building the technology to analyze the varitation in a million genomes.
I guess your question is kind of like, "why didn't computer scientists build systems like Kubernetes or Mesos in the 80s?" The problems and challenges were just different 16 years ago, and there was more than enough to work on between then and now. We don't need genome graphs at this instant, but we will in the coming years. And it's likely that newer representations will come about, too, as more math and theory is invented.
Sequencing was expensive and complicated, so even by sequencing several people the best they could do was make one composite image of the whole genome. Representing things as a graph would have provided no benefit at the time, and so people didn't do it.
Now we have thousands of public sequenced genomes. We need a new model for how we manage genomes. We are also learning how much is missed in the standard linear model of the genome, and need a way to incorporate new information into the reference. This all takes time.
Graphs are a good ways of encoding using our prior knowledge of genomes, but they are also difficult for researchers who have grown up on linear systems to understand.
Looking back and expecting otherwise is probably a case of hindsight-bias. Also worth emphasising is that the technology (and methodology) is moving very fast.
Very interesting company. They had a typical white board interview process.
What they were doing didn't seem that technically hard, and they were more concerned about prior credential ( like most bioinformatics company ) instead of what you were able to do.
They also seem to think JavaScript is not worth their time :( The problem they gave me was algorithmic and I just used the tool that was available to me. They apparently write most of it in C++ due to "Speed".
The only reason I even got a face-to-face, even though I didn't even has a Masters Degree was due to having done the assignment better than their Phd candidates ( their words ).
This piece seems like a submarine article for their proprietary platform.
How many cells do you need in a population so that there is at least one variant at each position (ie a SNP)? This can be for any species or cell line for which the info is available.
For a human, you would need a LOT more than a single human to have a variant at each position. Something like HIV can actually have a small enough genome, a high enough error rate, and a large enough population in an infected patient, that it can realistically have a unique variant at each position in its genome within a single human host.
"For example, the intestinal epithelium contains approximately 10^6 independent stem cells, each of which generates transient daughter cells every week or two. Thus, the intestinal epithelium of a 60-y-old is expected to harbor >10^9 independent mutations. This implies that, not far beyond the age of 60 y, nearly every genomic site is likely to have acquired a mutation in at least one cell in this single organ." http://www.pnas.org/content/107/3/961.full
Edit:
>"A reasonable error rate of DNA replication (not under stress) is about 1 write error in 1 billion reads"
I think it is that you are using the 10^-9 value to be per genome while that reference uses it as per basepair
These estimates ignore indels and SVs but empirical evidence suggests that the "everyone over 50 has a 50/50 chance of at least one adult stem cell harboring a mutation in at least one interesting gene". My personal favorite is
http://www.nature.com/nature/journal/v518/n7540/abs/nature13...
but Druley's follow up was equally awesome:
http://www.nature.com/articles/ncomms12484
and the recent survey of adult stem cell mutations is nice:
http://www.nature.com/nature/journal/vaop/ncurrent/full/natu...
One take-away from all this is that, while 95% of a sensitively surveyed population of 50-60 year olds had at least one stem cell with at least one known preleukemic mutation, it is equally clear that most people aren't walking around with anything resembling an acute leukemia. The natural conclusion is that in individuals with a competent immune system and diverse enough pools of healthy stem cells, it's not that big of an issue. Only when bad luck and/or stresses to which the mutants are adapted (e.g. TP53 mutations in therapy-related leukemia) afflict people, or the natural diversity of their stem cell populations collapses (as with really old people and individuals whose immune system actively attacks their stem cells, as in severe aplastic anemia) do you see the sort of massive, life-endangering takeover that we recognize clinically as disease.
Furthermore, nearly all of us are born with 5-10 predicted-to-be-lethal variants in our genomes. Clearly, we're also not dead, so our conception of "lethal" can't be quite right. There is an enormous amount of complexity in how real live multicellular organisms deal with variation and mutation, something we're really only just starting to grasp, and of course all of that then interacts with the person's environment to manifest (or not) their genetic tendencies. We build models of reality because the actual thing is too complicated to be tractable; it's important never to confuse the two :-)
But my point is that I think a key epidemiological variable should be the age of peak incidence. This has been largely missed due to various common practices like:
1) Binning into 5/10 year age groups
2) Looking at age adjusted data
3) Truncating the age-specific incidence data at 70-85
years old because the later data is deemed unreliableThanks, I have been thinking along those lines for a few years now after looking at the age-specific incidence of a bunch of different cancers from SEER. You see that many cancers peak consistently year after year at a given age, while the height of the curve may change drastically. The same was true when I looked at some data from other countries, although I never followed up very much on that aspect.
Then if you read the paper which spawned the multi-stage model of cancer that has been widely adopted[1], you see they make some assumptions for computational reasons that are unnecessary in these days of cheap computing power:
pt ~ 1-(1-p)^t, if p<<1,
where,
p = probability a required mutation occurs during a given time interval
t = number of elapsed time intervals (ie age)
Then by the product rule of probability they derive that, if cancer is due to accumulation of errors (usually considered to be mutations), the incidence at a given age would be: I(t) = k*p1*p2*...*pn*t^n = k*(p'*t)^n
where
I(t) = incidence at age t
n = number of required mutations
p' = geometric mean of the probabilities for mutations 1:n
k = a constant determined by the number of cells in each tissue,
the proportion of times that a detectable tumor forms from
the carcinogenic cell, and possibly the sequence in which the
mutations occur
If you use the non simplified version of their theory you would instead get: I(t) = k*(1 -q^t)^n
where
q = 1-p'
In contrast to the model that was simplified for computational reasons, this has a turnover. By setting the second derivative to zero you can get the age at which the peak incidence should occur as a function of number of required mutations (n) and geometric mean of the probabilities each mutation occurs (q = 1-p'): t_peak = log(1/n, base = q)
From this you will see either the multi-stage model is totally wrong, the error rate must be much higher than commonly thought, and/or the cell division rate of the error-accumulating cells must be much higher than commonly thought (the age is usually taken as a stand in for number of divisions). The last two possibilities suggest that we are constantly generating these cancerous cells and they are being cleared somehow.In normal stem cells, it appears that attrition and immune clearance gets rid of damaged cells when they cycle and senescent cells all the time (subject to some variation, not entirely age related, at least in our volunteers). We may expect higher rates in filter organs, but liver cancer isn't too common, and the paper I referenced earlier shows that this can't be just an issue of fewer divisions (I despise the oversimplified Tomasetti & Vogelstein paper because the facts simply don't support it). Colorectal is probably more common because the crypts are "facing out" ala melanocytes, thus prone to accumulating lots of environmental damage.
Anyways, the latter of your possibilities (proliferative mutants divide faster and error more often than normal counterparts) makes the most sense -- the eventual "winner" in a tumor is the cell that produces the most progeny and resists apoptosis due to stress the best. It's probably not a coincidence that these are traits which adapt a mutated cell to survive chemotherapy as well. However, spawning nonself mutations willy-nilly is a great way to attract immune attention -- particularly if you haven't blown the immune system away by nuking it with chemotherapy. :-/
maybe, or you can use the full Armitage and Doll model I described above and replace t with something like a discrete exponential decay where N(t) = number of divisions since zygote as a function of time. Ie N(t) = N0(1 - k)^t + 1 where N0 = N_birth - N_adult
That is, take the difference between division rate at birth and division rate as adult and fit a constant k between zero and one. It is just a first approximation at best because data on division rate by age in various tissues doesn't seem available...
This depends on the error + division rates (along with number of cells) in that tissue at that age though. From my research there not really such data available on any of those terms. Also, the errors need not be somatic mutation. For example, chromosomal missegregation may be much more common and potent since it can mess up the expression of many genes at once:
"Nevertheless, the rate of chromosome missegregation in untreated RPE-1 and HCT116 cells is 0.025% per chromosome" https://www.ncbi.nlm.nih.gov/pubmed/18283116
I'm just saying there are a number of other assumptions being made here, and if we get rid of the standard ones the Armitage-Doll model is capable of fitting the data surprisingly well.
[1] https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?rs=rs6564851
Some assemblers (SPAdes?) have started to support this format, but most downstream software only uses the FASTA format (the non-graph version from a single genome/representation).
Heng Li (somewhat now the godfather of bioinformatics) wrote a blog post about various implementations here: https://lh3.github.io/2014/07/25/on-the-graphical-representa...
A novel data structure to store the whole graph: https://almob.biomedcentral.com/articles/10.1186/s13015-016-... Software is here: https://www.uni-ulm.de/in/theo/research/seqana.html
This one maps the whole DNA space using markers: http://www.nature.com/articles/ncomms7914
Here's a paper that looks at the total genetic space of several individuals, but with read mapping alone, no graphs: http://genomebiology.biomedcentral.com/articles/10.1186/s130...
Most of this is based on open source or openly available software. A big company that's closed source is NRGene from Israel, I've read good things about their DeNovoMAGIC/PanMAGIC but I'm unsure how that stuff works exactly (apart from massive short read coverage).
This is pretty close to my area...