Google Genomics: store, process, explore and share genomic data
cloud.google.com
cloud.google.com
The biggest challenge we face is data protection. Sequencing data is not considered protected health information so theoretically HIPAA regulations do not apply, but this may change in future so nobody is willing to take responsibility. If sequencing data is considered PHI, HIPAA requires a BAA between the institution and Google. The process of getting these in place was so convoluted and difficult that it was just easier to use another provider more closely associated with our institution.
Frankly, CPU-hours and TB-years are still within 5-10% of the total cost providing sequencing results to the patients. What is really important is that the cloud allows us to determine exactly our cost-per-analysis without paying for excess capacity or maintenance.
What matters also is reliability, redundancy, and peak capacity, where the cloud (both Google and the other provider) win hands down.
Disclaimer: I work on Compute Engine.
Taking this to the extreme, you might even fire up a Windows VM so you could run some analysis or reporting tool there (say, Tableau), within GCP and then only need to egress the Dashboard or Report that it generated. You'd still have the raw data, but you would only be paying to store it. Especially if you're only keeping it around for archival purposes you could put it in a Nearline bucket and really minimize your costs.
(note: I work on GCE)
Disclaimer: I work on Compute Engine.
Disclaimer: I work on Google Genomics.
We ended up moving towards AWS because they were willing to sign. It sucked from my perspective because we already had such a tie in with Google with our email/ collab platform.
In the couple years since then our agreements have only gotten stronger.
The best way for Google to address this is to sign something with Internet2 and then the institutions can use that contract instead of negotiating separately.
http://googlecloudplatform.blogspot.com/2015/11/introducing-Custom-Machine-Types-the-freedom-to-configure-the-best-VM-shape-for-your-workload.html
Which would let you adjust your RAM to vCPU ratio (as low as ~1 GiB per vCPU and up to 6.5).Disclaimer: I work on Compute Engine.
(The other terms were a bit easier to google)
Here are the other terms:
HIPAA: the acronym for the Health Insurance Portability and Accountability Act that was passed by Congress in 1996
PHI: Protected health information
Edit: link to relevant definition http://searchhealthit.techtarget.com/definition/HIPAA-busine...
Basically an agreement between organizations that there is an understanding that PHI (protected health information) is at stake and that both parties are aware of their obligations to protect it.
I'm confident a Google version could be spun up if there were the appropriate demand.
I bet your university has hundreds of computers idling at night - why not use them?
[1] - http://www.fiercebiotech.com/story/google-ventures-backs-dna...
Not that I blame them.
Finally, the question: this is public data, (sometimes held in confidence but technically only for a "short" period), google is a private company, why should we place our trust in a for-profit organization to keep our public data publicly available without charging access fees? Admitting freely that having dealt for decades with the old model, i do have an open mind about this approach.
We thought the same with Freebase and the Knowledge Graph. It will happen again.
Ultimately, the "hard part" about genomics is not big-data requiring Spanner and BigTable to get anything done. I actually wrote a blog post about this this week:
http://blog.goldenhelix.com/grudy/genomic-data-is-big-data-b...
Both BAM and VCF files can be hosted through a plain HTTP file-server and be meaningfully queried through their BAI/TBI indexes. Visualization tools like our GenomeBrowse or the Broad's IGV can already read S3 hosted genomic files directly without having an API layer and very efficiently (gzip compressed blocks of binary data). So, I see the translation of the exact same data into API-only accessible storage system, where I can't download the VCF and do quick and iterative analysis on it more of a downside that plus.
Disclaimer: I build variant interpretation software for NGS data at Golden Helix. Our customers are often small clinical labs who size of data and volume are not driving them to the cloud.
It looks to be solving the same problems as DNAnexus, Seven Bridges, BaseSpace etc as a way to wrap open source tools in more user-friendly ways.
But it's orchestrating the production of smaller set of data that still needs the next step of human interpretation, report writing, family-aware algorithms and most complex annotations (the problem space Golden Helix is in).
In other words, the automatable bits that is not the hard part that I mentioned in my blog post.
This seems more a way to pipe people into using Google Cloud Services for storing data than it seems a useful way to do bioinformatics. As with most of these cloud-genomics things, it's not useful to have canned analyses that I can't extend (for example, I am so far from giving a shit about transition/transversion ratios and would never think to store that information in my variant table) What's the big plus?
But I think google is onto something. If they manage to deconstruct BAM files and store them in "format-free" way that will allow search, aggregation, and computation/reasoning across hundreds of thousands of patients/samples it may accommodate the next big wave of personal genomics data.
Disclaimer: I work on Google Genomics, and on the GA4GH APIs.
Disclaimer: I work on Google Genomics.
Disclaimer: I work in a neurobio lab but I wish I worked for Google Genomics.
Another features that comes to mind for enabling this sort of work is GCE's support for attaching a persistent disk volume to multiple VMs (read-write for one, read-only for others) -- have each pipeline stage consume from a read-only PD containing the output of the previous stage and write its results to a PD shared with the VM(s) that handle the next stage. Have the VMs communicate with each other for control plane operations ("hey, this directory contains the complete results of stage <x>, one of you stage <y> workers can pick it up" or "I have completed processing the stage <x> outputs for job <5>; they can be removed or archived to GCS."). This strategy also allows independent scaling of the workers responsible for each stage based on the computational requirements of that stage, since a given PD can be shared read-only with many VMs (with limited performance impact in the many-reader case, especially if the readers are not hammering the same parts of the volume).
Another option for truly pipeline oriented processing (assuming the underlying workers can generate data suitable for it) is to leverage something like Cloud Dataflow -- then they'll handle the work balancing for you.
(note: I work on GCE)
Clound Dataflow on the other hand is difficult to use in practice - one of the main reasons is the large footprint of intermediate files (sometimes 10x the size of the input)
Currently we are running monolithic pipelines on the cloud where one instance does the job from start to finish (if it get's preempted - I restart). I plan the next version to be in some way "split-apart" but have not found a good strategy/tooling - any advice?
Here are two examples from the iobio project.
A heads-up display describing a BAM (alignment) file streaming between two remote servers:
http://bam.iobio.io/?bam=http://s3.amazonaws.com/iobio/NA128...
Streams another BAM from the 1000 Genomes project through two variant callers, with dynamically-updated statistics about the results. You can modify the variant quality threshold for each caller and the results are recomputed in real time:
http://iobio.io/demo/variantcomparer/
Edit:
Very nice visualization of interesting statistics from a VCF file, provided by URL or uploaded:
http://vcf.iobio.io/?vcf=http://s3.amazonaws.com/vcf.files/A...
You decide who can see the data you store, and the default is private. That's spelled out in lots of words in the legal terms somewhere; sorry it wasn't more clear on the website.
Disclaimer: I work for Google Genomics.