The impact of Docker containers on the performance of genomic pipelines
peerj.com
peerj.com
I'm not sure about your comment about Big Data. If it turns out every hospital will have a next-gen sequencer, it seems to me the pipelines are really important, and having someone else working out the logistics sounds good to me. I'm curious as to what you consider the tougher big data problem. From where I sit, I consider this to be a complex data problem, and the issues of data size are simply solved by better/faster computing.
There will be (are in fact) one-click, cloud-hosted solutions for analysis, but given how quickly tools have evolved in this space, there will always be groups wanting to run on their own hardware so as to experiment with the latest new developments.
In any case, many parts of the field simply don't have the software engineering discipline to pull off proper "big data" workflows. Advances in commodity hardware, stronger programming tools for ad-hoc work, and "cloudification" toolchains will probably delay a lot of what used to require proper engineering effort from maturing.
Not to mention there's plenty of fertile ground solving problems which by now can be answered with merely "annoyingly non-small" rather than "big data" techniques.
The big win I've found with Nextflow is that once you've written a workflow, you have a lot of flexibility in the execution environment: Have all the tools already installed on your workstation or large compute instance? Use the local executor to saturate the box with concurrently running jobs. Don't have or want all those tools installed? Use the local executor with Docker images. Have access to a traditional compute cluster (e.g. LSF, SGE, Torque, etc.)? Use the cluster executor with Docker images.
A couple other resources worth checking out:
Toil workflow engine https://github.com/BD2KGenomics/toil
Common Workflow Language (CWL) specification https://github.com/common-workflow-language/common-workflow-...
Hope you don't mind a plug here for ADAM, a genomics analysis platform with specialized file formats built using Apache Avro, Apache Spark and Parquet.