27 karma · joined January 24, 2013
If you're serious about it, please submit a feature request (https://github.com/Factual/drake/issues), and describe more specifically what you would like to be able to do in your case.
Thank you for a great thought.
Artem.
http://stackoverflow.com/questions/2973445/gnu-makefile-rule...
Clojure is a Lisp. Lisp stands aside all other programming languages, first of all, because it supports syntactic abstraction (a.k.a. "code is data"). Hardcode addicts (I'm not one of them) say there are only two programming languages - Lisp and non-Lisp.
Here's a good comparison of Scala and Clojure: http://stackoverflow.com/questions/1314732/scala-vs-groovy-v...
When we made the decision to switch to Clojure, several things affected it, in no particular order: - we had some people who were already very proficient in Lisp - we liked how expressive and compact it was - Lisp is considered to possess immense expressive power (see http://www.paulgraham.com/lisp.html) - we were enamoured by Cascalog (http://nathanmarz.com/blog/introducing-cascalog-a-clojure-ba...), and it's written in and for Clojure. This one payed off very well. - Lisp has a reputation of being great at manipulating data: lists, graphs, etc.
Here's a good answer from one of our engineers: http://www.quora.com/Clojure/Why-would-someone-learn-Clojure
As for libraries, both Clojure and Scala are JVM-based, and Clojure has a very good syntax for Java interop, so all Java libraries are available to us. But, of course, Clojure community also spits out libraries like crazy, for example, take a look at this marvel which we use in Drake for parsing: https://github.com/joshua-choi/fnparse.
Sorry, guys, D was our working codename, and it slipped off my tongue, I guess, more than several times. :)
Drake supports the ability to run stages in parallel (at least in theory) - it's been speced out (https://docs.google.com/document/d/1bF-OKNLIG10v_lMes_m4yyaJ...), just not implemented yet. But of course, once you have the entire dependency graph, it's easy to know what can be run in parallel and what cannot.
As for distributing computations, our approach is that it lies outside of Drake's scope. Drake doesn't know what's going on in steps. But you can always implement a step that would use distributed computation, for example, by submitting a Hadoop job, or in any other way. The only requirement Drake has is for the step to be synchronous, i.e. do not return before all the computation is complete. But even that can be changed for some cases.
Other examples of intractable problems in Make would be timestamped dependency resolution between local and HDFS files. If Make can't look at HDFS, it can't say if the step needs to be built or not. I don't think you can fix it with external commands.
But generally, search for intractable problems is a futile one. Remember, everything you can code in Java, you can code in a Turing machine. :)
You raise some interesting points (for example, a frequently changing code), which we ran into as well. Our current approach to it is not as fundamental, and basically includes ability to force re-build any target and everything down the tree and methods, and you can also add your binaries as a step's dependency.
I'm sure as we and other people use the tool, we'll have better ideas. For example, Drake could automatically sense that the step's definition has changed and offer to rebuild or dismiss.
Other points you raised are also definitely worth thinking about.
http://www.youtube.com/watch?feature=player_detailpage&v...
In short, we don't feel like it's an either or question. We want to have Drake as a command-line frontend to the core functionality, but we would love to see/have other frontends developed as well. Currently, there's no Clojure DSL for Drake, but I think it'd be totally awesome.
The reason we started from command-line is because our workflows are heterogenous, and we also didn't want to limit Drake to developers and associate it with coding. Clojure can be quite a big learning curve if you only need it to specify steps and link them together through file dependencies.
We had an important design goal in mind: Drake should be as simple as writing a shell script. If it's not, our experience shows that most workflow start as trivial shell-scripts with one or two steps, and by the time it grows into something unmanageable, it's kinda too late. :)
On a related note, Drake supports Clojure code inlining for manipulation of the parse tree. It's not an equivalent, just a somewhat related feature. It allows you to modify the steps, dependencies, and anything else in the parse tree directly from Clojure.
http://news.ycombinator.com/item?id=5111527
I suspect most of the points I made would be applicable to redo as well, if not more so. Trivial things don't require Drake. Heck, they often times don't require Make as well - just put it in a linear shell script if the steps are not too expensive. It's when things are getting complicated you need something like Drake.
The concept of Make is not unique. Everything that has dependencies and executes steps is similar to Make in concept. Drake is no exception, and it can be replaced with Make, but no more so than Rake, Ant or Maven can be replaced by Make. That is, if it's trivial - yes. Just a bit more complicated - no.
Some things are merely painful to implement with Make, some are just impossible:
- multiple outputs
- no-input and no-output steps
- HDFS support
- Hadoop's partial files support (part-?????)
- forced execution of any subbranch, up or down the tree or any individual targets (crucial for debugging and development)
- target exclusions
- protocol abstraction - inline Python is just one example
- tags
- branching
- methods
These are just what's implemented already. Other things are planned such as: - automated data versioning (backup and revert)
- parallelization
- real-time status console
- retries, email notifications
- etc.
Requirements for building executables and working with large, complicated and expensive data workflows are quite visible different, and the most important thing about Drake is that it provides the platform for convenient features (such as versioning or email notifications) to be implemented. And once they are, every data workflow can take advantage of them.I guess, if Make was really, really extendable, we could have considered it as a platform for all this. But it's not, and hacking all of that into Make's source code in C would be, I'm sure, a much greater pain than writing Drake.
Artem.