We're not getting anywhere. Just give me goddamn examples! :) Please! Examples!
> I can see this is really really hard to grok if you're basing everything on the idea of a DAG, and so many tools are that it's very natural to think you couldn't do it any other way.
There is no other way. BPipe is based on the idea of a DAG. You just don't see it.
> In Bpipe the user declares the pipeline order explicitly.
And this is a big mistake. The reason is simple - explicit order is very hard to manage once you have multiple inputs and outputs, and as a consequence, complicated (instead of linear) dependency relationships.
What you don't seem to realize, is that by "declaring the pipeline order explicitly" you create a dependency graph. It's a part of your workflow definition. Your workflow contains the full definition of the dependency graph. Even if it didn't, you would still use it. There is no other way.
This is what I meant when I said - you create your dependency graph in "run". And this is a bad idea.
> dependencies arise as actual commands are executed.
What does it mean exactly? That the first command will somehow tell Bpipe what to run next? If not, then I don't understand this statement at all.
> How is that if it doesn't know about the dependency graph?! Well, it does it "just in time".
It does not matter if you calculate the dependency graph before you run the first command, or as you run the commands. It makes absolutely no difference. The only difference is whether it is computable or not. If you say it's not computable until run-time, please elaborate on that.
> So in this way Bpipe handles dependencies for you.
So far I see that this is very standard and doesn't differ in any way from what Drake or any other tool does. The only thing that differs, and I am repeating myself, is how you define your dependency graph - through input and outputs, or in "run". So far it seems that "run" is quite unfortunate. But please give me examples.
> So in this way Bpipe handles dependencies for you. What it does not do is figure out which order to execute things in. It does them in exactly the order you tell it.
This is a meaningless statement. Drake also executes steps in the order you tell it. The only difference is how you tell it. In Drake, you tell it through specifying a list of steps each step depends on individually (once again, it doesn't matter that filenames are used for that - Drake also supports tags, or it could be some other identifiers). In Bpipe, you tell it in "run", collectively and sequentially. Drake's way supports the whole variety of graphs, while Bpipe's way - only a very limited subset. And for this limited subset, Drake can give you (I think) a syntax just as good if not better than Bpipe's. If you don't quite understand what I'm talking about, give me an example, and I will demonstrate.
> I actually want to control the order of things sometimes.
This is fine, the only question is how. You say Bpipe's way is convenient. I say give me an example and I'll show you that Drake's way is not any less convenient. I'm sorry to keep repeating myself, I thought I stressed the importance of examples quite a bit in my previous email and I want to stress it again. Examples, please!
> I want to be able to tell it "do this first, then that, then the next thing" regardless of dependencies.
This statement is self-contradictory. You don't seem to realize that by telling it "do this first, then that" you are defining dependencies. It's fine, and it's OK, and it can be convenient, but you can't say regardless of them.
Again - give me examples! Our conversation is becoming useless without examples.
You did not, but I'll just grab whatever you threw my way:
fix_names = {
exec "sed 's/Neverbrown/Evergreen/g' $input > $output"
}
extract_evergreen = ...
run { fix_names + extract_evergreen }
$ bpipe run pipeline.groovy input.csv
Drake can support this perfectly: _ <- $[in]
exec "sed 's/Neverbrown/Evergreen/g' $INPUT > $OUTPUT"
$[out] < _
........
$ drake -v out=pipeline.groovy,in=input.csv
Isn't that much nicer? What disadvantages you can see?Tell me what is it that you would like to do with this script, and I'll tell you a better way to do it in Drake. Is it multiple versions of run that you want to have? Easy. Are you concerned about inserting a step in the middle? Trivial. Tell me why Drake's code is worse, and I'll listen. So far it seems like it's better because it's shorter and more flexible at the same time.
> Having the tool think this stuff up by itself can save you a bit of time but it can lose you a lot because you don't have the ability to really control what's going on.
What exactly are you losing?
I am sorry if I sound irritated. I am. I've just been begging for examples, and you keep talking in abstract, and it would be fine, but you're making a lot of mistakes. So, instead of looking at concrete things that would make my point apparent to you (or the opposite, prove that I'm wrong), I keep pointing to flaws in your reasoning, which frankly, is irrelevant. One picture is worth a thousand words.
I really want your feedback. But please give me examples.