Our experience using Clojure to speed up Beanstalk
blog.beanstalkapp.com
blog.beanstalkapp.com
Basically same loop to go through all revisions in a repo was 15-20x times faster in Clojure than in Ruby. Very similar calls, pretty much same algorithm. But something that Ruby does inside the bindings was not very efficient.
For Git the problem was a bit different. Internally Grit (the Git API library for Ruby) tries to read huge pack files (potentially hundreds of MBs) in pure Ruby. This just can't be fast.
There's also additional overhead of not using any ORM for operating with our caches in Clojure version. That probably contributed to overall performance improvement, but 20x for svn / 40x for git measurement was taken even before we got to saving the results into DB and it stayed that way later on.
In summary: benchmark, profile, find hotspots, optimize. Works in every language. >:3
But I agree, we've used Clojure(and JVM) in exact point where it would bring the most speed up to our app.
Foo::Bar.method() vs (foo.bar/method)
The ruby version does three loads from the heap, at runtime, while clojure does zero.
invokedynamic in JDK 7 will help narrow the gap, but the fact remains that Clojure was designed with more performance in mind than ruby.
Median speedup for the same problem in Clojure was 12x over Ruby. and 14x for Haskell; ... and 25x in Java .. and 35x in C++...
In our case on the pretty much same loop the difference between Ruby and Clojure was 20-30 times. Please note that using Clojure allowed us direct access to Java libraries, so essentially this is Java performance that we've been able to tap into using Clojure.
Another plus for Clojure is that it allows us to easily have a multi-threaded implementation. JRuby allows native threads too, but Clojure make concurrency really easy, while JRuby just wraps Java threads with nothing more, so I would have to use regular locking and stuff.
The only example you chose shows programs that manually unroll loops are not accepted. (And that's stated explicitly on the Help page. http://shootout.alioth.debian.org/help.php#unroll)
It's a sad nuisance when people opine without checking what's actually shown on a public website.
It's sad to see fact based comments being down voted on HN.
Well, I could simply ignore the question. Or give the answer I got a few years ago, what I did...
> It's sad to see fact based comments being down voted on HN.
Fact based? More like facts querying. Which is the reason I upvoted every single of your comments...
And yet people have down voted.
(Except for meteor-contest.)
>>pervasive laziness... requiring a lot of space/time complexity for other algorithms<<
And GHC provides strictness analysis and explicit strictness to avoid reducing performance.
Do argue - but argue better.
Unfortunately, to argue better does require learning what the website says about the benchmarks game - and for most that's far too much effort.
An interesting comparison might have been Clojure vs their previous code on JRuby.
1. Micro-optimizations, mostly due to JVM and its excellent JIT (the garbage collectors are quite impressive, too, if what you need is predictable response time).
2. Architectural gains: thanks to the Clojure's excellent concurrency support I can make much better use of multiple cores. I get more parallellism, hence better performance on same hardware.
The first kind is cool, because you get it "for free". The second kind is the real game-changer, because non-parallel software only gets you so far in terms of performance, and writing concurrent software is Hard. Clojure makes it much, much easier.
But overall I wouldn't say that Clojure is a performance daemon on a single CPU. You can get performance similar to carefully written Java code. This is good, but you can always do better with C or hand-written assembly on critical sections. But that's not the main advantage: the big thing is that I can write correct Clojure code fast, it runs well enough, and I can easily make use of multiple cores. You can debate micro-benchmarks all you want, but what really counts for me is how quickly (and correctly) I can get from zero to production code that runs fast enough.
Also if you're into Web development, I can recommend playing with Noir[2] framework and Korma[3] for SQL abstraction. Heroku also support deploying Clojure apps out of the box, so you can easily use a free tier to get something out there.
(defpage "/welcome" []
"<html><head></head><body>Hi there!</body></html>") (html
(head (title "My page")
(script :type "text/javascript" :src "/static/page.js"))
(body (p "Stuff")))
Actually for a while I had a JS tool which I was using called "build.js" which would just build DOM components like this. (It was therefore JSON-serializable too, but I never really had an occasion to use that.)There is something nice about HTML which is a bit lost here: HTML (and LaTeX for that matter) allow unquoted text, with fewer escape characters. They are markup languages, which s-expressions crucially lack. (On the other hand, XML & family lack the ability of Lisps to make the first token anything other than a symbol.) There was briefly a plot called NML / Enamel and some others -- DTML and TML I believe -- which would instead write:
<html |
<head | <title | My page >
<script <type|text/javascript> <src|/static/page.js>>>
<body | <p | Stuff>>>
This is actually an emulation of a C-type syntax with a Lisp-type semantics: the idea is that you have in some sense two syntactically different channels into your expressions, one which comes before the pipe character | and one that comes after. The stuff that comes after is allowed to be marked-up text; the stuff that comes before is some sort of node list, and perhaps has certain conventions (one could imagine instead using a Clojure-style `:type "text/javascript"`, which would limit you to what XML attributes can do -- short text only, symbolic keywords).That is, one could hypothetically rewrite this in some C-ish syntax which would look like:
html {
head { title { My page }
script(type {text/javascript} src {/static/page.js}) {}}
body { p { Stuff }}}
and again, if you wanted to limit yourself to XML, the parens above could then say instead `type: "text/javascript" src: "/static/page.js"` but it doesn't have to be that way.Some crazy ideas for anyone building a new language to think about.
Could you post the Clojure code somewhere? It would be very interesting to see clean, fast Clojure code written to solve a real-world problem (as opposed to some toy example).
I think the reason for the performance difference is pretty clear. According to the article and the grit documentation, all the git api calls were either done in pure ruby, or by shelling out to `git`. Also, reading between the lines, it sounds like they required some information/relationships on commits that was non-trivial to retrieve using a basic git shell command. So by switching out the runtime and the algorithm, they get a huge performance increase.
With jgit, they can more easily traverse the graph directly and efficiently. I'm really curious to see what kind of performance they could get from using FFI+jruby and raw c calls to libgit/libgit2 (http://libgit2.github.com/).
Didn't know what those changes entailed, but this is interesting. Thanks!
Scheme
(let ((x 2) (y 3))
(let ((x 7) (z (+ x y)))
(* z x)))
Clojure (let [x 2 y 3]
(let [x 7 z (+ x y)]
(* z x)))
I think the Clojure version is easier to read without a paren-matching editor, though Scheme's rigorous minimalism does have its charm.You also missed out a few of the syntax quote stuff.
So, other tokens are #{} ^ #() #_ #'x ` ~ ~@
Full details: http://clojure.org/reader
Except in wanting to mess around with Clojure, you didn't use the best tool for the job.
The Clojure solution now uses parallel implementations of Git and SVN to solve the problem, rather than the core code of SVN and Git. And now you also have a one-off daemon written in Clojure. It doesn't have the same support structure, ops requirements, or anything, as your Ruby code. Virtually no one uses clojure, so hiring and training are different, etc. You've incurred a lot of overhead for something not that great.
The best tool for this job was to improve the Ruby version by way of C extensions, or write a new C command that does this work for you, linking in the Git and SVN code directly. This has little to no new concepts, is straightforward, and would have given you the best compatibility and performance.
Clojure caching took me like a month max.
Obviously, you can choose whatever you like, it's your company. I'm just explaining what the best solution was for the many students and young programmers who visit hacker news. Clojure is not it for all the reasons I mentioned. And if people think C is hard, practice it. Read Zed's book and do Project Euler problems with it until you feel comfortable with it. It's an essential tool everyone must know how to reach for.
It might not impress their programmer friends, but it'll impress their accountant.
If impressing accountants is the goal of what we do, then the whole thing should have been outsourced overseas.
I read their requirements in the blog post. They need fast access to SVN and Git internals to store off metadata. Based on that it's really not hard to determine what to use to solve this problem. There may not be one exact right answer (in fact, I suggested two), but there is a direction that makes sense and one that does not.
The real disservice to those coming up in this industry is reading blog post after blog post on Hacker News with people looking for "excuses" to use random languages they're interested in playing with. The answer to most of these problems is to use something really boring and straightforward, like C, C++ or Java. I guess I've worked too many places where legacy code becomes a huge burden because one guy wants to play around with X one time. The funny thing is, I see these same guys go onto the next place and do it there.
Can you elaborate on why you think it's a bad choice?
That's still a widely held view, and maybe even a majority, but probably not on HN. (see pg's Beating the Averages essay: http://www.paulgraham.com/avg.html)
Not following what you mean. Do you mean that I do expect people to learn stuff, or that I don't expect them to learn stuff in order to do their job?
Favoring "little to no new concepts" at least suggests a bias against certain kinds of learning. Sometimes learning a new concept is the best way to solve a problem.
I'm also not sure I'd compare it to Clojure and suggest it as a better tool for this job.