Amazon Genomics CLI
aws.amazon.com
aws.amazon.com
(QLDB)
It's funny, but after many years in the field, and many generations of workflow definitions, I'm still now happy with any of the options. Snakemake is the closest to being usable for me for both prototyping and longer term work does, but they are all still fragile and inflexible in weird ways.
It's a tough problem to solve, which I think is evidenced by the large number of solutions. I've even tried looking outside at non-bioinformatics tools and never been very happy. And for all their flaws, the bioinformatics tools are better at typical bioinformatics workflows than tools developed for other domains.
When I read that sentence, I was surprised not to see the classic suggestions from tech/data science:
Prefect, Airflow, Luigi, Apache NiFi, Jenkins, AWS Step Functions, or Pachyderm
Of course, that didn’t end up being the case. If you can’t program at an elementary enough level to define simple workflow graphs using an API in a non-DSL (like those you mention), then the language is irrelevant—you’re just not technically sophisticated enough to be writing your own workflows. So now we have these DSLs that non-programmers still can’t use and programmers all hate. A lose-lose.
There's another part of it as well. Bioinformatics workflows tend to be just enough different from more standard workflows that there is friction using off the shelf tooling.
For one, the DAG nodes are often mostly/fully represented by command line tools expecting a POSIX style file system and making assumptions/asserting opinions on where files live, to where it writes outputs, etc. Bioinformatics workflow orchestrators can understand this and provide optimizations in the DSL to express how to manage the file movement. In contrast, I find many people with a more standard data/biz workflow mentality think of DAG nodes as being queries, running blocks of code, etc.
Lifecycles vary as well. The institutions who have these workflows will either be pure research or some degree of research/production hybrid. There's an advantage to being able to use the same software on both side of that hybrid. When your research workflow is ready to be blessed to production, having to translate it into a different system can be expensive in terms of both time and bugs.
Another aspect is that these workflows, until not too many years ago, were often run on HPC compute clusters using software like SGE, SLURM, PBS, etc. These DSLs can provide optimizations for tweaking parameters in a way that more mainstream tools do not.
Bioinformatics and Other workflow management systems just have different focuses, more than they have different end user skills. In the 2000s I tried to shoe-horn so many CS Systems type workflow onto using clusters, and every single one had abstractions that simply did not fit what I needed. Typically they seem to require knowledge of how many tasks a job must be split up into, for example, instead of dynamically adapting the workflow to the inputs as desired. Or they hard code the method does launching processes to a certain type of compute environment. So when somebody says "workflow," I assume they mean something entirely different then what I mean until proven otherwise.
For example, if I go to the Apache NiFi docs to see if it's suitable, the Getting Started page starts with a GUI, which is a huge red flag, I don't want to do workflow entry with a GUI because a pre-defined graph will only cover a fraction of my use cases. But worse then that, when it describes where to start creating a processor, it says you need to select from a list of processors. Now this is an even BIGGER red flag because I know absolutely none of the commands I need will be available, and I'll need to add tons of metadata for whatever I want, most likely in some sort of bad GUI, or even worse, XML or some other bad data standard that software engineering adopted due to cargo-cult design. The bigger danger is that the system assumes that processors must be written in some certain language, or assumes that data itself will be introspectable, or something that will require massive amounts of work.
Every one of these systems embodies a huuuuuuge number of assumptions about what it means to run a workflow, nobody ever states what those assumptions are, or probably even realizes how different their needs are from others' needs. At least with the workflow systems from other people in bioinformatics, I know that the assumptions they make will work with the tools and data sizes in the field, and won't have some deal-breaker I only find out about after three hours of making workflows.
Speaking for myself, we didn't target wetlab biologists with WDL but we did target "non-programmers" as MontyCarloHall described. This was influenced by the groups from which the original developers originated prior to working on the DSL.
The degree to which we were successful can be debated, but hey that was the idea at least ;)
If you look at a small WDL tutorial, such as this:
https://www.rc.virginia.edu/userinfo/howtos/rivanna/wdl-bioi...
I think that this is more complex than, say, defining Afro data types, as it involves calling tasks like functions, string substitutions, filling in variables from a config file, etc. There’s a certain base level of complexity that can be hard to abstract away, and I don’t think this has yet got to the simplicity of, say, pivot tables in Excel.
100% and a conclusion I came to far too long into my tenure in the bioinformatics workflow domain. It is also why I soured a bit on some of the GA4GH efforts in the same space (WES & TES in particular).
I still very much agree w/ the goals & ideals (respectively), but in my experience the middle ground between "I don't care, just give me the defaults" and "please do not abstract away any of the complexity, I need that" is quite small. And that zone would be the sweet spot for these efforts.
The maddening part for us has been that Snakemake, which like for the parent poster has been the best workflow management system for us, is not as widely supported across commercial platforms , and is somewhat less big PaaS enabled than e.g. nextflow.
On the one hand Snakemake does support kubernetes, but on the other hand we are a bioinformatics research group and don’t have k8s engineers :/
(Scaling) Optimal workflow management for genome informatics groups is still very much an open problem.
I'd be curious if the underpinnings of the support they have for these tools allows more toolsets to be brought over or if the integration is highly specialised.
a) They can sell a lot of compute time, storage space, etc
b) It's an academic heavy field, and selling proprietary tools tend to be tough sledding
In my old university, the people working on the genome studies had to sign a draconian NDA and they had a dedicated air-gapped compute cluster for it. If your genomics study can produce a patent for a new diagnostic method or maybe even a drug precursor molecule, that'll be way more valuable than everything you spent on security.
And already 8 years ago, hacking attempts by what was thought of as foreign government actors was a common experience in life science research. When I had to give a talk about our methods, we scrambled all the gene IDs for fear of the slides getting leaked.
So based on how people working in that field behave, I presume they would never ever agree to upload the data to servers outside of their control. Accordingly, I can only imagine this service to be used accidentally by inexperienced students (who are then on the hook for violating their NDA).
That restricts you to your own infrastructure then. There are tradeoffs in both directions.
The are also dozens of online databases for all sorts of data from biology.
Reading between the lines, you seem to attach certain ideas to the term „genetics“?
Now, it is technically possible to have this data in AWS with assurances of encryption and some sane security measures like not leaving it open to public. However, the big catch is - the compute part of the cluster is on-premise. We did the math of moving the compute to AWS, and it was several orders of magnitude expensive(consider that the group did not need to pay for power, network links, and maintenance of the server room by virtue of being part of the university). The only “extra cost” of being on-premise was my salary, which is substantially small portion of the increase by going to AWS. Also, guess what going to AWS doesn’t magically eliminate the need for a guy whose responsibility it is to ensure the cluster is operated, and for users to come in for help, and to manage the ecosystem in general.
The argument of moving to the cloud is more complicated than it appears on the surface. I believe the HPC shops that have a lot of compute and very little need for elasticity will be the last to head to the cloud.
These workflow orchestration tools are ubiquitous in the bioinformatics space. A lot mores than things like Airflow and the like.
I ask, because I've shared my genome with Google, and it's online for anybody to see.