> The dependency graph specifies what
> steps depend on what steps. If you don't
> know it, you don't even know how to
> start evaluating the workflow, because
> you don't know which step to build
> first. I don't understand this statement
> at all. Could you please elaborate or
> give me an example?
I can see this is really really hard to grok if you're basing everything on the idea of a DAG, and so many tools are that it's very natural to think you couldn't do it any other way. Think of it as imperative vs declarative if you like. In Bpipe the user declares the pipeline order explicitly (as you've seen) - so that's the first part of the answer to your question. Bpipe knows which part to execute first because the user said to explicitly. But this isn't used for figuring out dependencies - dependencies arise as actual commands are executed. Back to our famous example: fix_names = {
exec "sed 's/Neverbrown/Evergreen/g' $input > $output"
}
extract_evergreen = ...
run { fix_names + extract_evergreen }
We run it like this: bpipe run pipeline.groovy input.csv
If you run it once, Bpipe builds input.fix_names.csv. If you run it twice, Bpipe is clever enough not to build input.fix_names.csv again! How is that if it doesn't know about the dependency graph?! Well, it does it "just in time". It executes the "fix_names" pipeline stage (or "method") and that calls the "exec" command. The "exec" command sees that all the inputs referenced ($input variables) are older than the outputs referenced ($output variables). So it knows it doesn't have to rebuild those outputs, and skips executing the command. So what about transitive dependencies? If C depends on B which depends on A, (so dependencies are A => B => C) what happens if you delete file B? Technically you don't need to build C because it's still newer than A, but Bpipe can't see it any more. Well, Bpipe knows this too because it keeps a detailed manifest on all the files created. So when the call to create B is executed it can see that although B was deleted, it did exist and in its last known state was newer than input files, so there's no need to rebuild it, as long as downstream dependencies are OK.So in this way Bpipe handles dependencies for you. What it does not do is figure out which order to execute things in. It does them in exactly the order you tell it. This is one of those things that conventional tools solve which isn't actually that important (in my uses) but which occasionally is very annoying - I actually want to control the order of things sometimes. I want to be able to tell it "do this first, then that, then the next thing" regardless of dependencies. Usually it's pretty obvious what the right order things should be in and there are other externalities that influence how I like to do it ("I know this part uses a lot of i/o so try to do it in parallel with another bit that's mainly using CPU", or "Let's run this part last because it will be after hours and the other jobs will have finished"). Having the tool think this stuff up by itself can save you a bit of time but it can lose you a lot because you don't have the ability to really control what's going on.