The genomic era arrives, and this time is probably real
economist.com
economist.com
I'm not trying to belittle recent advances in genomics, on the contrary I believe that cheap sequencing and technologies like CRISPR are going to change the world. However, its the insinuation that the last 20+ years of hard work were a fraud that upsets me. We wouldn't be were we are now without the last 20 years of genomics research. /rant
Actually, cells have a lot of state. See [1] for a basic, slightly old, CS-biased explanation.
That's how a single genome yields an incredibly varied repertoire of cell types. We need to work a lot on understanding those epigenomes to then get a chance to pinpoint genetic causes of disease.
As someone with a CS background who has worked almost 5 years in bioinformatics, one thing to realise is if you think something obvious is not done, either you're missing the right search term on pubmed or there's some horrible lab issue that makes it really hard (and we need a new assay/tech)
A few decades ago ago we read our first line of source code from a computer from the year 2,000,000,000 - some pretty advanced technology that was just handed to us. Two decades ago we did our very first read-through of an entire human. The reading itself was not particularly productive. We already knew some of the most important lines of code, from the decades before, and we annotated a lot more sections that are critical for basic function. We've even traced a few bugs back to the source code - especially those bugs where 1) the source was unusually easy to read [1], or 2) the bug was really bad [2]. As of this year we've developed compilers to enable us to write and run small, multi-line programs of our own composition [3]. We're doing really well in trying to read, comprehend, manipulate, and act on some of the most sophisticated technology ever presented to us. More sophisticated by orders of magnitude than the silicon in Nvidia's new graphics card (16nm 2d dry feature size [4], vs 0.1nm, 3D wet feature size [5]) . Reverse engineering it takes a bit of time, and yet we've already put that knowledge to significant use.
Further, once we do get our working knowledge up and running there will be a significant outpouring of function from having reverse-engineered such technology. But of course there is a lag time in the millions of lab-hours being poured into understanding that technology before we can actually use it. Not dissimilar to the rapid advances in science-fiction when a race reverse-engineers an advanced alien technology. How long did it take us to build The Machine in Sagan's 'Contact'? We will not just cure cancer, but we will be capable of all sorts of new feats. Purchasing the textbook doesn't make one a master - it's the study, the hard work, and the practice.
[1] https://en.wikipedia.org/wiki/Human_papillomavirus
[2] https://en.wikipedia.org/wiki/BRCA_mutation
[3] https://news.ycombinator.com/item?id=11417689
[4] http://wccftech.com/rumor-nvidia-pascal-gtx-1080-gddr5x-gtx-...
[5] https://en.wikipedia.org/wiki/Resolution_%28electron_density...
freudian slip? ;-)
Do I need to dig out this story or do you mean Carl Sagan's 'Contact'?
For a scientist working incrementally, accumulating data, sieving it for information, fighting for grants, doing administration while making sure family gets three squares, articles like this may cause a wince. But they probably influence support for more funding. Pro scientists I know do not have time to read these articles and will never see them unless it's from someone else, but budding scientists, career-changers, influencers, all get misinformed.
In short, news on science causes long-term damage for short-term gains.
BLAST (Stephen Altschul, Warren Gish, Webb Miller, Eugene Myers, and David J. Lipman)
BWA (Heng Li, Richard Durbin)
Samtools (Heng Li)
GATK (from the GATK Team; DePristo, M., Banks, E., Poplin, R., Garimella, K., Maguire, J., Hartl., C., Philippakis, A., del Angel, G., Rivas, M.A, Hanna, M., McKenna, A., Fennell, T. Kernytsky, A., Sivachenko, A, Cibulskis, K., Gabriel, S., Altshuler, D. and Daly, M. A)
I'm missing quite a few from my more niche field; but feel free to add more.
The costs of assaying a genome has dropped from $100 million to $1 thousand.
Identifying correlations with clinical metadata, understanding the post translational interactions and understanding the full roles of proteins and other molecules remains an intimidating and difficult job.
http://www.genomicsengland.co.uk/the-100000-genomes-project/
Edit: I realize it's a small part, but it's the first step, and it's already relatively inexpensive, which is a good sign.
I'm not saying that it's not worth it in some cases, but the sequencing is only a small part of the overall costs.
Interpretation, on the other hand, is quite different. Currently, interpretation still requires a human to look at everything. And regardless of how good annotations and algorithms get, I have yet to hear of anyplace where the final annotations aren't manually curated (maybe automated, but still requires a human in the loop). This aspect of Precision Medicine doesn't scale quite as well as sequencing or computation...
It is a comprehensive database of dosing guidelines, drug labels, and annotations for thousands of combinations of genes, variants, drugs, and diseases.
(We're actually helping them with a redesign, and if you'd like to be a beta tester, contact me.)
The VCF files are reasonably formatted, the raw sequencing data is in the FASTQs, which are huge and hard to deal with. Go nuts!
Here is the consent template: http://www.1000genomes.org/sites/1000genomes.org/files/docs/...
The post I responded to sounded like the genomic data would be uploaded without consent to a giant public database.
Even with consent the greater concern is that employers, medical providers, national health services, and health insurance companies could, by sequencing part of your genome, match it to a public database identifying you and your entire genome. This could then be used to raise prices, deny health care, or deny employment.
De-identification is not enough if the material is indicative enough that efforts to re-identify it can succeed.
Kudos to those advancing science by making their genomes public, but it's a risk I would not take.
With new technology comes new responsibility - we shouldn't shy away from the tech because we don't want the burden of enforcing regulation.
Update: Please ignore my comment. The original post was changed and my upvotes have turned to downvotes. Now I can't delete this comment.
In Robert Gordon's The Rise and Fall of American Growth, a great deal of attention is focused profitably on the changes and progress in medicine and health outcomes from 1870 to present. Most notably, Gordon divides the period at 1950, noting that life expectencies improves 2x more before 1950 than afterward (and from my own explorations, far more prior to 1920 than after).
There's been exceedingly little progress in medicine since 1970, at which point major cancer, heart disease, and virtually all infectuous disease treatments existed. We've seen many more treatments and more so imaging and diagnostic capabilities since ... but virtually no changes in outcomes.
Healthcare is a land of very, very, very rapidly diminishing returns, and for which shoring up treatment and preventive care for the most under-served pays hugely greater dividends than highly invasive or heroic treatments at the top end. Much of the progress in health outcomes since 1970 appears to be in minority populations -- that is, the under-served.
Gordon's book was published this year, its information is quite current. I see little reason to suspect massive improvements in actual outcomes -- impacts rather than change -- in the nine months or so since it was put to bed.
But more recent improvements in healtcare aren't just trivial stuff either. The five-year survival rates for childhood leukemia, for instance, have been on a steady increase from 40% to 80% since 1970. Recent monoclonal antibody therapies against inflammation/autoimmune diseases have provided actually revolutionary changes in the health situation of many.
Going from chronic pain and having a hard time functioning at school/work, to living a relatively normal life, is a huge thing. But it's not captured by life expectancies.
For me, the big story in genomics over the past few years is one of massive failure. We spent a lot of time and effort sequencing tumors in the hopes that the genetics would tell us something interesting and lead to cures. It did not. We now have a relatively complete cabinet of the major variants that drive tumors. Most of these variants we already knew about before we did these genomics studies (through older sequencing methods from decades prior), and most of what we learned tells us nothing new.
Genetics is at this point mostly garbage information. Why? Because we don't know what it means. We still don't understand gene expression, we don't understand signal transduction, and we can't understand the effects of mutations without painstaking characterization. All of these things mean our sudden wealth of knowledge of genetic variation tells us fuck-all about biology.
We learned quickly that 'cancer' is to the twentieth century 'the fevers'. We learned there are many reasons and causes for the catch-all, 'cancer'. Some we now know how to cure effectively. Some we now now how to prevent, effectively. Some we now know what to do to cure, but don't have good tools yet. And many we still do not understand. That is a significant improvement in a very short time span.
But there's a LOT more to learn than 'how to cure cancer'. We very clearly understand gene expression and signal transaction to a first order, and have bits of the second order down too. There might be more - but that doesn't make that first order inneffective or wrong. There's lots you can do with a first order understanding - see what Newtonian physics did for the world - it was correct only to the first order. The idea that we know 'fuck-all' about biology because of dna sequencing is just ridiculous. The ability to read and write dna is fueling one of the fastest growing segments of human progress today.
The proof in the pudding is this: you can get your exome sequenced today and report hundreds of tumor variants, but there is a bare handful of variants that will result in a change in your treatment, and most of those lesions were well-studied before genomics took off.
>We very clearly understand gene expression and signal transaction to a first order, and have bits of the second order down too. There might be more - but that doesn't make that first order inneffective or wrong.
Yes, in fact, our first order understanding IS ineffective. This is what I study every day, so perhaps I'm too close to this, but we literally do not understand the effect of most (99.9%) of genetic variants on gene expression. Sure, there's lots you can do with a first order understanding, but materially, what happens when you sequence a tumor (or a germline, for that matter), is that researchers get a list of mutations, stare at it mystified for a while, and then shrug and move on, because there's really nothing you can do to understand what these things mean.
Yes, this stuff is the way of the future, and it's important to do it for our greater knowledge and understanding in the future. But we're far, far away from this point right now.
I think there's probably a fifteen-year lag in this stuff being useful. The insights we're having now are because of the sequencing of the human genome; the work we're doing now will pay off in another decade and a half when we've figured out how to deal with gene expression.
>do the full genome sequences yet actually read every base in your genome even?
Read, sequence: yes. Align: no.
They are not "blueprints" much less some sort of engineering document.
More realistically, they are a simple list of materials (in this case a listing of proteins).
They only describe the materials that make up a structure, not what that structure is nor how it functions.
Its the equivalent of someone saying: 14,234 tons of steel, 23,000 tons of concrete, 8000 tons of glass...etc. Now, what does it make? Why does it make it?
I think of genes as the initial configuration in Conway's game of life. Small initial (genetic) variations can cause large differences after some generations. Some don't matter at all and just die out. But no matter what is the end result, it's not a blue print or a description, it just _is_, and the result flows from it using some basic rules.
Epigenetics as a concept can also be abused into a sort of neo-Lamarckism.
But good luck growing a healthy organism without its epigenetic landscape properly initialized
This is beside the point, as parent already explained what he meant in his first 2 sentences: "Physics. More specifically, 'conditions'".
So, yes, the laws of physics might be universal, but gravity is X here and N on the moon, pressure is Y here and M under the sea, temperature, etc...
So according to one definition, it's even worse than your 14,234 tons of steel example: it's closer to "steel is a mixture of iron and carbon and trace amounts of other dopants, glass is silicon dioxide, concrete is ..." Without even listing the amounts. But by another definition, the genes are, if not a blueprint, at the very least a recipe and ingredient list. "Butter, flour, water, sugar, apples, cinnamon, cloves. Cut 125g butter into flour; if too coarse, continue cutting. Add water to ice. Add water from ice into flour..."
The information is clearly there—we have a long, unbroken history of offspring being very similar to their parents, despite having started with only a single, undifferentiated cell. But with the new tools available to us, we're starting to have a hope of understanding it, even if we're using more brute force than we'd like.
> Its the equivalent of someone saying: 14,234 tons of steel..
I feel like that's taking things a bit too far in the other direction. With our current understanding, it's like saying:
14,234 tons of steel that's capable of self-assembling (e.g., collagen), 23,000 tons of concrete that automatically folds into a specific structure (~everything other than disordered proteins).
That's not to say that we fully understand how proteins fold/function... but it's a bit more than a static block of steel.
Tumor sequencing is a HUGE boon to cancer treatment and research. Patients lives have been drastically improved by targeted treatment.
Sure, genomics gives us a comprehensive way to sequence, but my point is that it did not yield the sudden, vast sweep of insights that we expected it would.
Whether this was the best use of scientific resources, I donno.