Common Workflow Language
commonwl.org
commonwl.org
In this case we seem to be talking about "Scientific Computing Workflow" or some sort of data pipeline. (As opposed to some sort of "Enterprise" workflow)
People seem to use workflow for all of these things:
#1 Business process modeling and automating platforms - Camunda, BPMN, jBPM etc.
#2 Assembling together software functions (also called tasks, sometimes) for getting an output data - Airflow, Cadence, CWL is related too.
#3 Moving one entity (bug, document, deal) through different stages of a state machine, until the entity reaches an end state - Sharepoint workflow, Zoho Orchestly, Comala workflows
#4 Orchestrating several apps together for automating something, usually as reactions for events - Zapier
When developers talk about workflows, it's mostly about #2. When business people talk about workflows it's mostly about #1. When managers/leads of a team talk about workflows it's mostly about #3.
We should accept that Workflow as a word is too ambiguous, and come up with what kind of workflow category the tool belongs to, before branding it as an "one stop solution for all things workflow".
So, more along the lines of https://en.wikipedia.org/wiki/Dataflow_programming. A good problem space to bring declarative/functional/logic/concatenative concepts and tools, IMO, as opposed to the ubiquitous “hit it with a rock” of the usual C* suspects. Be great to see real progress here, but my bullshit detector is immediately dialed to MAX.
The great problem with “workflow systems” is that they very easily disappear up their own ass: “Potemkin solutions” that solve every problem except the customer’s own. (See: https://thedailywtf.com/articles/The_Inner-Platform_Effect). “Enterprise workflows” in particular I think are just a way for indolent code monkeys to shirk learning how the business works for themselves. Looks all very impressive without ever solving shit.
..
FWIW I do think Alan Kay was onto something with his Nile strategy (https://news.ycombinator.com/item?id=10535364). Pity there wasn’t followthrough (bit of a mayfly, Kay). Ditto Iverson and APL, although APL’s sidelining was more down to a nascent Programming Profession redefining Programming™ to be what they do for users rather than what users (mathematicians in this case) do for themselves.
In which case, why didn’t you just use bash?
Yeah, I know bash stinks. But this is not my fist rodeo* so I’m struggling here to see how CWL stinks less. See again: Inner-Platform Effect, Greenspun’s Tenth Rule.
--
* i.e. I already know how easily “workflow engines” go up their own arse because I’ve done it myself. And CWL fails the same sniff test. Not a good start.
The only bit that sounds at all novel or interesting is the dispatcher; and even that is really just an expression of `<load_balancer> | ssh`. At which point, Unix Philosophy tells us we should implement <load_balancer> as a small simple single-purpose Unix command which can easily pipe to other Unix tools. Itch scratched; everyone can now go get on with their actual work.
So if that is the case, the precisely what problem is all CWL’s Castles-in-the-sky YAML crap actually solving, other than bored developers’ need to keep entertained? Especially when [from what I can tell] the project doesn’t even provide you a dispatcher component but instead tells everyone to take a spec and write their own.
How many wheels need to be reimplemented before someone involved declares it a pig in a poke? And how many more before the rest can accept this?
Bioinformatics software isn't perfectly written software, there are a number of weird behaviour that a simple unix pipe doesn't solve. There are engines that support CWL, and other existing engines have been adding CWL support.
I'm not saying that there aren't other frameworks out there for doing analysis, or that this is the best way but this is an option that IS working for researchers.
Edit: workflow -> scientific workflow | data pipeline
> the project doesn’t even provide you a dispatcher component but instead tells everyone to take a spec and write their own.
Close...
Software that supports CWL are SaaS vendors, FOSS projects, and various HPC schedulers that all have their own incompatible data management and dispatch/scheduling systems. If you want to write an analysis that runs on more than one of these platforms, you need some abstraction for it. CWL is one such an abstraction.
This matters because maybe you've developed a research pipeline that integrates a bunch of different tools written in different languages and want to run it on somebody else's data, and you need to run it on their infrastructure because copying 12 terabytes of HIPAA-restricted data from their LSF cluster to your Google cloud instance isn't an option.
"Just use bash" is what people who adopt CWL are trying to get away from. It is nearly impossible to write portable parallel / distributed analysis in bash, and the result is brittle scripts with more coordination code than code that actually does scientific work. Because CWL is declarative, the CWL engine handles all the coordination, scheduling and data staging for your particular infrastructure.
You may not have any of these needs, but suggesting that we're just bored developers creating castles in the sky is really unhelpful.
The goal is to provide a way to describe dataflow processes that is highly portable, auditable, and reproducible. This is incredibly important in research, clinical, and regulatory domains where you need to be able to show how you came up with a result.
It's not a general purpose language on purpose, and operational concerns like notifications are the domain of specific implementations (engines).
I agree the syntax is horrible (and I designed most of it) but it also makes it easy to write programs that read and write CWL, enabling an ecosystem. For example, here is a transpiled languages that emits CWL:
Just plain old bash seems to work just fine, though. It is simple and gives you a lot of portability to HPC clusters and general purpose cloud. Not hardwiring to a configuration file and not using an, in practice, un-portable workflow language saves you a lot of pain.
A truly portable cloud/hpc module for python would be a good solution.
Any information about which language is leading the field?
CWL seems to be more popular than WDL, as far as I can tell, but Nextflow is probably more popular than both.
At least the Nextwave homepage (https://www.nextflow.io/) starts with a clean concrete written example, even if it does leave you to deduce its actual meaning for yourself.
You can find some CWL workflows on dockstore [1].
[0] https://www.ga4gh.org/ [1] https://dockstore.org/search?descriptorType=CWL&searchMode=f...
See also: Greenspun’s Tenth Rule, “Yo Dawg…”.
Those who can’t, implement Standards.
Look no further than their own hand-crafted examples (confusingly called "Episodes") for compelling arguments for why to NOT use Common Workflow Language.
Your comment almost gets there, but not quite—and then it sinks itself lower by breaking this one:
"Don't be snarky."
https://news.ycombinator.com/newsguidelines.html
Can you please take the guidelines a bit more to heart when posting here? We'd appreciate it.