HNHacker News
TopNewBestAskShowJobs

aboytsov

27 karma · joined January 24, 2013

submissionscomments
aboytsov··on Introducing Drake, a kind of ‘make for data’
No, make -B mytarget rebuilds either mytarget only or mytarget and everything mytarget depends on. A more common scenario is when you need to rebuild mytarget and everything that depends on it. Without rebuilding other parts of the workflow that you don't need.
aboytsov··on Introducing Drake, a kind of ‘make for data’
Thanks! This much I know. But it doesn't answer my question. Let me repeat it: could you please give me a command to re-build a particular target and everything that depends on it?
aboytsov··on Introducing Drake, a kind of ‘make for data’
This is an awesome idea. Currently Drake only supports timestamped and forced evaluations, but it would be great to have an evaluation abstraction where you could provide your own implementation of whether a target's changed and/or whether a target is to be considered fresher/younger than another target. Timestamped would compare modification times, forced would return true, and it could be extended indefinitely.

If you're serious about it, please submit a feature request (https://github.com/Factual/drake/issues), and describe more specifically what you would like to be able to do in your case.

Thank you for a great thought.

Artem.

aboytsov··on Introducing Drake, a kind of ‘make for data’
I'd like to hear your thoughts on this:

http://stackoverflow.com/questions/2973445/gnu-makefile-rule...

aboytsov··on Introducing Drake, a kind of ‘make for data’
Sorry, I might be very ignorant of make - could you please give me a command to re-build a particular target and everything that depends on it?
aboytsov··on Introducing Drake, a kind of ‘make for data’
It's hard to compare Clojure and Scala. Scala is a multi-paradigm programming language with strong OOP support and functional support. It's arguably more verbose than Clojure but looks much more similar to Java.

Clojure is a Lisp. Lisp stands aside all other programming languages, first of all, because it supports syntactic abstraction (a.k.a. "code is data"). Hardcode addicts (I'm not one of them) say there are only two programming languages - Lisp and non-Lisp.

Here's a good comparison of Scala and Clojure: http://stackoverflow.com/questions/1314732/scala-vs-groovy-v...

When we made the decision to switch to Clojure, several things affected it, in no particular order: - we had some people who were already very proficient in Lisp - we liked how expressive and compact it was - Lisp is considered to possess immense expressive power (see http://www.paulgraham.com/lisp.html) - we were enamoured by Cascalog (http://nathanmarz.com/blog/introducing-cascalog-a-clojure-ba...), and it's written in and for Clojure. This one payed off very well. - Lisp has a reputation of being great at manipulating data: lists, graphs, etc.

Here's a good answer from one of our engineers: http://www.quora.com/Clojure/Why-would-someone-learn-Clojure

As for libraries, both Clojure and Scala are JVM-based, and Clojure has a very good syntax for Java interop, so all Java libraries are available to us. But, of course, Clojure community also spits out libraries like crazy, for example, take a look at this marvel which we use in Drake for parsing: https://github.com/joshua-choi/fnparse.

aboytsov··on Introducing Drake, a kind of ‘make for data’
Haha, this is so funny. :)

Sorry, guys, D was our working codename, and it slipped off my tongue, I guess, more than several times. :)

aboytsov··on Introducing Drake, a kind of ‘make for data’
Great points.

Drake supports the ability to run stages in parallel (at least in theory) - it's been speced out (https://docs.google.com/document/d/1bF-OKNLIG10v_lMes_m4yyaJ...), just not implemented yet. But of course, once you have the entire dependency graph, it's easy to know what can be run in parallel and what cannot.

As for distributing computations, our approach is that it lies outside of Drake's scope. Drake doesn't know what's going on in steps. But you can always implement a step that would use distributed computation, for example, by submitting a Hadoop job, or in any other way. The only requirement Drake has is for the step to be synchronous, i.e. do not return before all the computation is complete. But even that can be changed for some cases.

aboytsov··on Introducing Drake, a kind of ‘make for data’
The most crucial thing that Make lacks is multiple outputs and precise control over execution. When you're debugging/developing a large and expensive workflow, you absolutely must have the ability to say things like: - run only this step, I'm debugging it - I've changed implementation of this step, re-build it and everything that depends on it - build everything except this branch, it's expensive and I don't need to rebuild it that often (example: model training)

Other examples of intractable problems in Make would be timestamped dependency resolution between local and HDFS files. If Make can't look at HDFS, it can't say if the step needs to be built or not. I don't think you can fix it with external commands.

But generally, search for intractable problems is a futile one. Remember, everything you can code in Java, you can code in a Turing machine. :)

aboytsov··on Introducing Drake, a kind of ‘make for data’
Thank you very much. We're really looking forward to other people using this tool.

You raise some interesting points (for example, a frequently changing code), which we ran into as well. Our current approach to it is not as fundamental, and basically includes ability to force re-build any target and everything down the tree and methods, and you can also add your binaries as a step's dependency.

I'm sure as we and other people use the tool, we'll have better ideas. For example, Drake could automatically sense that the step's definition has changed and offer to rebuild or dismiss.

Other points you raised are also definitely worth thinking about.

aboytsov··on Introducing Drake, a kind of ‘make for data’
We love Clojure. Lisp is an extremely powerful language, and Clojure brings all this to the practical JVM world. And Lisp is quite good in operating on lists and graphs, which is a big part of Drake.
aboytsov··on Introducing Drake, a kind of ‘make for data’
Nice. Surprisingly, we weren't aware of Makeflow and kinda missed it completely. On the first look, it seems like Drake is quite a bit more feature-rich than Makeflow. Please see the designdoc and/or the tutorial video for details.
aboytsov··on Introducing Drake, a kind of ‘make for data’
Thanks! We feel that in practice, there's quite a lot of differences between Drake and most Make-like systems. See this response for details: http://news.ycombinator.com/item?id=5111527
aboytsov··on Introducing Drake, a kind of ‘make for data’
This is a great question. Our approach to this is described here:

http://www.youtube.com/watch?feature=player_detailpage&v...

In short, we don't feel like it's an either or question. We want to have Drake as a command-line frontend to the core functionality, but we would love to see/have other frontends developed as well. Currently, there's no Clojure DSL for Drake, but I think it'd be totally awesome.

The reason we started from command-line is because our workflows are heterogenous, and we also didn't want to limit Drake to developers and associate it with coding. Clojure can be quite a big learning curve if you only need it to specify steps and link them together through file dependencies.

We had an important design goal in mind: Drake should be as simple as writing a shell script. If it's not, our experience shows that most workflow start as trivial shell-scripts with one or two steps, and by the time it grows into something unmanageable, it's kinda too late. :)

On a related note, Drake supports Clojure code inlining for manipulation of the parse tree. It's not an equivalent, just a somewhat related feature. It allows you to modify the steps, dependencies, and anything else in the parse tree directly from Clojure.

aboytsov··on Introducing Drake, a kind of ‘make for data’
Please see my response to Make comparison:

http://news.ycombinator.com/item?id=5111527

I suspect most of the points I made would be applicable to redo as well, if not more so. Trivial things don't require Drake. Heck, they often times don't require Make as well - just put it in a linear shell script if the steps are not too expensive. It's when things are getting complicated you need something like Drake.

aboytsov··on Introducing Drake, a kind of ‘make for data’
Drake supports "protocol" abstraction, which is much more than just specifying an interpreter. Python is a trivial protocol, not much more complicated than shell. There are slightly more complicated protocols, for example, "eval", which runs the first line as a shell command before putting everything else in $CMDS environment variable. There could be protocols for running an HBase query, a Pig query, Cascalog query, or an SQL query. Some of these things could involve building a JAR file and giving it to Hadoop binary. Currently only a handful of protocols is implemented, but more are described in the spec.
aboytsov··on Introducing Drake, a kind of ‘make for data’
The example in the blogpost is understandably trivial, and it can be implemented in almost any Make-like system.

The concept of Make is not unique. Everything that has dependencies and executes steps is similar to Make in concept. Drake is no exception, and it can be replaced with Make, but no more so than Rake, Ant or Maven can be replaced by Make. That is, if it's trivial - yes. Just a bit more complicated - no.

Some things are merely painful to implement with Make, some are just impossible:

  - multiple outputs
  - no-input and no-output steps
  - HDFS support
  - Hadoop's partial files support (part-?????)
  - forced execution of any subbranch, up or down the tree or any individual targets (crucial for debugging and development)
  - target exclusions
  - protocol abstraction - inline Python is just one example
  - tags
  - branching
  - methods
These are just what's implemented already. Other things are planned such as:

  - automated data versioning (backup and revert)
  - parallelization
  - real-time status console
  - retries, email notifications
  - etc.
Requirements for building executables and working with large, complicated and expensive data workflows are quite visible different, and the most important thing about Drake is that it provides the platform for convenient features (such as versioning or email notifications) to be implemented. And once they are, every data workflow can take advantage of them.

I guess, if Make was really, really extendable, we could have considered it as a platform for all this. But it's not, and hacking all of that into Make's source code in C would be, I'm sure, a much greater pain than writing Drake.

Artem.

← PreviousPage 2 of 2