In this case we seem to be talking about "Scientific Computing Workflow" or some sort of data pipeline. (As opposed to some sort of "Enterprise" workflow)
In this case we seem to be talking about "Scientific Computing Workflow" or some sort of data pipeline. (As opposed to some sort of "Enterprise" workflow)
People seem to use workflow for all of these things:
#1 Business process modeling and automating platforms - Camunda, BPMN, jBPM etc.
#2 Assembling together software functions (also called tasks, sometimes) for getting an output data - Airflow, Cadence, CWL is related too.
#3 Moving one entity (bug, document, deal) through different stages of a state machine, until the entity reaches an end state - Sharepoint workflow, Zoho Orchestly, Comala workflows
#4 Orchestrating several apps together for automating something, usually as reactions for events - Zapier
When developers talk about workflows, it's mostly about #2. When business people talk about workflows it's mostly about #1. When managers/leads of a team talk about workflows it's mostly about #3.
We should accept that Workflow as a word is too ambiguous, and come up with what kind of workflow category the tool belongs to, before branding it as an "one stop solution for all things workflow".
So, more along the lines of https://en.wikipedia.org/wiki/Dataflow_programming. A good problem space to bring declarative/functional/logic/concatenative concepts and tools, IMO, as opposed to the ubiquitous “hit it with a rock” of the usual C* suspects. Be great to see real progress here, but my bullshit detector is immediately dialed to MAX.
The great problem with “workflow systems” is that they very easily disappear up their own ass: “Potemkin solutions” that solve every problem except the customer’s own. (See: https://thedailywtf.com/articles/The_Inner-Platform_Effect). “Enterprise workflows” in particular I think are just a way for indolent code monkeys to shirk learning how the business works for themselves. Looks all very impressive without ever solving shit.
..
FWIW I do think Alan Kay was onto something with his Nile strategy (https://news.ycombinator.com/item?id=10535364). Pity there wasn’t followthrough (bit of a mayfly, Kay). Ditto Iverson and APL, although APL’s sidelining was more down to a nascent Programming Profession redefining Programming™ to be what they do for users rather than what users (mathematicians in this case) do for themselves.
In which case, why didn’t you just use bash?
Yeah, I know bash stinks. But this is not my fist rodeo* so I’m struggling here to see how CWL stinks less. See again: Inner-Platform Effect, Greenspun’s Tenth Rule.
--
* i.e. I already know how easily “workflow engines” go up their own arse because I’ve done it myself. And CWL fails the same sniff test. Not a good start.
The only bit that sounds at all novel or interesting is the dispatcher; and even that is really just an expression of `<load_balancer> | ssh`. At which point, Unix Philosophy tells us we should implement <load_balancer> as a small simple single-purpose Unix command which can easily pipe to other Unix tools. Itch scratched; everyone can now go get on with their actual work.
So if that is the case, the precisely what problem is all CWL’s Castles-in-the-sky YAML crap actually solving, other than bored developers’ need to keep entertained? Especially when [from what I can tell] the project doesn’t even provide you a dispatcher component but instead tells everyone to take a spec and write their own.
How many wheels need to be reimplemented before someone involved declares it a pig in a poke? And how many more before the rest can accept this?
Bioinformatics software isn't perfectly written software, there are a number of weird behaviour that a simple unix pipe doesn't solve. There are engines that support CWL, and other existing engines have been adding CWL support.
I'm not saying that there aren't other frameworks out there for doing analysis, or that this is the best way but this is an option that IS working for researchers.
Edit: workflow -> scientific workflow | data pipeline
> the project doesn’t even provide you a dispatcher component but instead tells everyone to take a spec and write their own.
Close...
Software that supports CWL are SaaS vendors, FOSS projects, and various HPC schedulers that all have their own incompatible data management and dispatch/scheduling systems. If you want to write an analysis that runs on more than one of these platforms, you need some abstraction for it. CWL is one such an abstraction.
This matters because maybe you've developed a research pipeline that integrates a bunch of different tools written in different languages and want to run it on somebody else's data, and you need to run it on their infrastructure because copying 12 terabytes of HIPAA-restricted data from their LSF cluster to your Google cloud instance isn't an option.
"Just use bash" is what people who adopt CWL are trying to get away from. It is nearly impossible to write portable parallel / distributed analysis in bash, and the result is brittle scripts with more coordination code than code that actually does scientific work. Because CWL is declarative, the CWL engine handles all the coordination, scheduling and data staging for your particular infrastructure.
You may not have any of these needs, but suggesting that we're just bored developers creating castles in the sky is really unhelpful.