Why we are choosing Clojure as our main programming language
appvise.me
appvise.me
As for the GIL, the author didn't even consider multiple processes -- he explained them away as a solution for some workloads. A Web app is one giant workload where this model makes sense: multiprocessing is the approach you should be taking with a Web application. Tie a request to one core, and spawn enough WSGI applications for the number of cores you have (and then some). That's an elegant solution to this problem, since a request doesn't need to fart around with other requests in most cases. If your app can't handle multiple copies of itself running, how do you expect it to scale?
This reads a little bit like a waffling, like a person who considers a completely different environment and rewrite of the entire stack a solution to some thorns in Python. I'm all for picking an alternative Web stack, but:
> I had written a database loader to import Apple’s Enterprise Partner Feed (EPF) and a web crawler in Python and next up was the web interface.
> It all seemed like a smooth sailing but in the back of my head I was beginning to have doubts about my decisions.
> Why we are choosing Clojure as our main programming language
This being the third blog post for an as-yet-unreleased product, I'd say worry about delivering a product instead of justifying a rewrite of your work thus far on your blog. If I were considering funding your startup, this blog post would be a fairly bad sign to me.
At any rate, a computer language is just a tool to implement an idea, and focusing strongly on the language of choice is busywork itself.
I totally agree with you on the point regarding worrying about delivering a product instead of rewiting the work. That's why the services already written will remain in place and I'll there fore be running Python alongside the Clojure/JVM stack.
Yikes, really? That sounds even worse than a clean break. You don't shed your perceived problems with Python and instead get to manage two codebases and stacks. Your operations team will absolutely love you in the future.
Based upon a casual read of what your service does, I'm really stressing to think of a CPU-bound situation that can't be pooled into a multiprocessing pool. All of the heavy work I can think of your service doing -- particularly crawling -- is going to end up network or I/O-bound, isn't it? I'm coming up empty and working with a theory. Perhaps you can share what your CPU-bound process is?
Of course most things can bee pooled into a multiprocessing pool. It's just a matter of the pain you have to endure while doing so. For some tasks it's a straight forward process while for other tasks it can be pretty painful. The crawler is IO-bound but most of the collaborative filtering algorithms are CPU-bound and having nice concurrency constructions in the language is a bonus.
And if that's too heavy, there's always stuff like GEvent. Asynchronous I/O and multiprocessing solves 80% of all problems you may have.
I do wish that Python / Ruby would be GIL-free, but for my startup I choose these platforms because they've got no parallel when it comes to rapid-development of web applications, because no matter how cool this and that language is, nothing beats the productivity gained by robust, battle-tested web frameworks (not to mention other stuff, like NLTK or SciPy).
As an early-adopter myself I might try out a couple of projects in Clojure, but I wouldn't bet my business on it unless I saw a clear need that Clojure satisfies and that dwarfs all other disadvantages, like platform immaturity.
Saying that Asynchronous I/O and multiprocessing solves 80% of the problems is hand-waving away the complexity such approaches might add to a project. Clojure simply gives you more and better tools then Ruby and Python currently do for dealing with concurrency.
You might not bet your business on Clojure, but certainly other people have for a couple years now and they don't seem to take issue with it.
EDIT: removed the defensive bit. My experience, as this thread shows, is that people get touchy about the limitations of the GIL.
EDIT: RE: Don't take me wrong, I sometimes hate the limitations that the GIL brings.
I'm continually looking at alternative implementations, like Pypy or Rubinius or JRuby, or other languages, like Haskell or Clojure, but then I end up doing a lot of yak shaving and nothing gets done.
You say that like it's a bad thing. ;)
I really hate this argument for its disingenuity. Some things (including message passing) are naturally fastest with shared memory. Multiple processes are not an equivalent substitute.
Clojure, for instance, leverages shared memory to implement its excellent, lightweight STM and related high-level threading constructs, and shared memory is what allows Clojure to easily implement cheap, MVCC functional data structures.
Multiple processes aren't an answer, here. To implement the same high-level constructs efficiently requires re-introducing shared memory through more complex, less efficient mechanisms, and often leaves the problem of sharing access to higher-level constructs (objects, instead of data) unsolved.
Multiple processes are a poor work-around for a lack of support for concurrency, not a solution. They make sense if you're sandboxing, but make little sense for implementing a high-performance concurrent server.
To pre-empt the erlang discussion -- erlang's message passing model does scale, but Erlang's runtime does require support for real shared memory concurrency in order to maximize single-machine performance.
This isn't a disingenuous answer at all, since it goes on to point out that A Web app is one giant workload where this model makes sense: multiprocessing is the approach you should be taking with a Web application.
As someone with a strong Java background, I have to continually force myself to be aware of my bias against concurrency-via-multiple-processes.
If you have a stateless application (and if at all possible, statelessness is a good thing [1]), then multiple processes scale vertically reasonably well (on the same machine), but having a multiple-process architecture makes scaling horizontally (across machines) natural and easy.
It's true that shared access to higher-level constructs in a multiple process environment isn't solved, but (a) shared access to data is an anti-pattern [2], and (b) Memcache plus some kind of serialization works pretty well when you do need shared data.
[1] Yes, applications exist where statelessness doesn't make sense. In this case, though the author was originally developing on AppEngine, so I doubt there is much state in the app.
[2] I'm not saying you should never have shared data. I am saying you should avoid it as much as possible.
Everything from database connections to user data can be cached locally, shared across connections, and does not require multiple processes each with their own large heap.
I'd like to elaborate on more examples (such as comet and efficient handling of a large number of blocked connections while allowing unblocked connections to proceed concurrently while maintaining cached local state) ... But I'm on an iPad and this keyboard is driving me crazy.
But I think caching is something best done in system/platform code, not in your application.
Like you say, there are plenty of examples where sharing state makes sense. But I still believe that state in application code is something that be avoided, and that most of the time it's best to rely on platforms that do state management for you.
For example, session support in web platforms is a great example of something that supplies (simulated) state, and is supplied by the platform.
I have a hard time with the notion that web apps should be "stateless." It seems to be an argument borne out of limitations of the frameworks/platforms being used, and repudiated by the fact that such platforms do share state, but are forced to use less efficient external mechanisms (such as network requests to memcached, the local database, etc), rather than leveraging the advantages of data locality as available in a non-multiprocess system.
We've often taken advantage of sticky sessions to allow individual servers to maintain and share state across requests while ensuring that the state could be reconstructed by another server should it become necessary. It is simply more efficient to do so, and efficiency in implementation directly translates to dollars spent on operational costs, as well as effects on human observable response times.
But I think it's an illuminating discussion anyway.
But even if you go through the initial pains of figuring out how things fit together, many tools are still missing. An easy to use templating library, for example. (several exist already but they are still early stages /are being refactored). Also things like the gems and plugins ecosystem in RoR - where you can pretty much find gems for so many things out there. (e.g.: tagging, authentication, etc). There simply isn't a comparison.
So, while I really love Clojure and I believe the community around it is growing - in both size and its contributions, I would say that, let's face it, writing full web apps in it right now cannot be compared to the productivity you would get in say RoR, due to the immaturity of the tools. That being said, I still see Clojure being immensely useful in other scenarios (for example: A.I. algorithms, high performance data processing, etc) and as such I would use it mostly in those settings, while integrating it with a more mature 'front-end' framework.
Probably works with JSP too...
We jumped from Clojure/Compojure to Erlang/WebMachine with much better success so far. I did in Erlang in two days what Clojure took me two weeks with numerous false starts to accomplish, mostly because the libraries were so poorly documented or incomplete. I spent more time digging through code and trying to assemble a framework to build upon than I did writing useful business logic.
I found the Erlang infrastructure to be well-structured and strongly documented, mostly because of it's maturity. What I surmise is that Clojure is just too young, and its web framework lacks a strong commercial force driving it. Rails has 37signals, Webmachine has Basho, Lift has sites like Twitter and Foursquare. Clojure needs something similar to push it forward.
All arguments and emotions aside, I'd say if you're starting to build a new web site, and you're not considering Rails or Lift, you're doing yourself a major disservice. Like others have said, JRuby nicely avoids some of the issues discussed for scripting languages.
As far as I know, nothing is stopping you from using any of the existing templating systems in the Java ecosystem. They don't have to be written in Clojure to use them with Clojure.
You're absolutely on the right track. As startups struggle to hire, new startups should be looking to off loading more and more of their work to more productive programming languages.
But the thing is if you go with windows, you're eventually going to have to support linux too just to get a the wealth of open-source codebases like say Redis. I mean its certainly possible to run Redis on windows via cygwin but you're 32-bit limited, and its a pain to actually install and get everything working.
On Fedora it's "yum install redis", and you're done.
The F# mono implementation is somewhat slower than the .Net implementation though, so there is that caveat.
On the other hand, F# does use less memory.
http://shootout.alioth.debian.org/u32q/benchmark.php?test=al...
http://shootout.alioth.debian.org/u32/benchmark.php?test=all...
For whatever it's worth, if you like Clojure, use Clojure. If you like F# use that. If you like Perl, Java, PHP, use those. But, if you're going to consider other options hopefully the FUD doesn't get in the way ;).
Most FP communities (incl clojure) have devote non-negligible blocks of time to code review and benchmarking to make sure that at the very least poorly written code isnt' submitted (and the right hotspot knobs are on).
http://groups.google.com/group/clojure/browse_frm/thread/d27...
I pointed out how comical it is for someone to declare they dislike (who knows why?) the benchmarks game, and then present the benchmarks game to others as a reliable source of information.
(However, although you provided a different URL you did not point out that you thought the parent had linked to the wrong comparison.)
You didn't just "correct" the link to show F# and Clojure.
You changed the link from quad-core to single core and that reduced the difference shown between F# on Mono and Clojure.
Look, the results aren't usefully different: http://shootout.alioth.debian.org/u32q/benchmark.php?test=al...
You do seem to be using the what you dislike and what you are not presenting as reliable to suggest "F# on mono and Clojure (which is slower than plain Java) look pretty similar".
If you really dislike the benchmarks game, don't look at the benchmarks game and don't show it to other people :-)
iirc Andy Fingerhut started benchmarking Clojure programs more than a year before Clojure was even included in the benchmarks game.
https://github.com/jafingerhut/clojure-benchmarks
Here's another thread "Comparing clojure speed to java speed" which has nothing what-so-ever to do with the benchmarks game -
http://groups.google.com/group/clojure/browse_thread/thread/...
(Incidentally, those "hotspot knobs" made the Clojure programs slightly slower but forced collection of the temporary objects that were showing up as much greater memory use than the Java programs.)
I never attemped to actually use Mono, so please correct me if I'm wrong, but if I understood correctly, Mono has to duplicate all .NET frameworks, libraries and tools, which means Mono
a) is not complete (e.g. Silverlight, VisualStudio)
b) will hardly ever keep the pace of development of Microsofts implementation (simply due to resources).
Mono may be viable if you are happy with a subset of the .NET ecosystem, but I'd really feel more comfortable if I have access to all of Java with Clojure. Chances are you do need that library...
Mono is a complete implementation of the C# specification, additionally, the Mono project has ported many .Net libraries.
Moonlight is the Mono version of Silverlight. Visual Studio is an IDE, and doesn't have anything to do with Mono vs .Net (in the same way IntelliJ IDEA has nothing to do with Java portability). You can write code in Visual Studio and compile it with Mono with no problems. You can also use MonoDevelop on OSX and Linux.
The time between Microsoft releasing new versions of .Net / C# and the Mono implementation is very small, usually weeks but sometimes only days. Unless you need to work on the bleeding edge right now, I don't think that it really makes much of a difference.
Yes, you can write C# in a way that isn't portable, especially if you use libraries that are OS specific. You can do that in Java as well (or any other language). Just look at all the libraries that require epoll or kqueue - those won't run on Windows no matter what language they were written in.
It's unfair to say that Mono requires you use a subset of the .Net ecosystem. If you want to write cross-platform code, you will always be constrained to a subset of the libraries.
C# is a nice language, especially with Linq (which is in Mono). You should spend a weekend with it sometime to form an opinion :D. MonoDevelop works fine on the mac and is free.
I will agree though that F# is a beautiful language, and am hoping on getting everyone else at work on board so we can start using it more for our development.
Our entire server codebase, baring a few external libraries, is Scala (working with Lift) which has been an awesome experience but there is a nagging doubt in my mind that if/when we need to start looking to add in developers we will either need to invest in cross-training a java dev or end up paying out probably more than we could/should afford to get a seasoned java dev who trained themselves in Scala already. If over a short period of time your choice gets some major traction then it will work in your favour, but if not then you could be out in the cold or risking employing someone with no real provable history.
Remember in business there is no "cheap" or "expensive". There's only "worth the money" and "not".
Why? Because learning cool tech and Getting Stuff Done are two very different things.
While it is true that most developers suck at writing code, knowing Scala doesn't tell me whether you:
1. Have enough discipline to do the boring, tedious parts of your job 2. Know how to prioritize tasks 3. Can write easily maintainable code 4. Can work well with others.
Technical ability is only one part of an employee.
"Personnel is one area in which OCaml has been an unmitigated success for us. Most importantly, using OCaml helps us find, hire, and retain great programmers."
How long it takes to set up a dev environment and push to cloud service?
With appengine, 1) I sign up in 5 minutes, 2) clone an app in python or Java, 3) use my favorite free ide to edit the code 4) appcfg.py update
Also, if I don't have any of the tools like git I just type: sudo apt-get git-core
Heroku on AWS is just as easy.
Can you provide a link showing how easy it is to get started with F#?
Also step 2 sounds like magic and a recipe for disaster. Clone an app? What, with all its settings? Random library includes you weren't expecting?
Don't confuse well trod paths with new roads, closure and F# are both new roads, you'll need to do a bit of work yourself.
Also visual studio has a free edition these days.
Anyway, my point is don't compare apples to oranges when the author said he can't use apples.
It's my understanding it's not so easy to do that kind of learning and experimentation on Windows.
Go to http://www.asp.net/get-started.
Although I would recommend going down the MVC route as asp.net forms are sucky. http://www.asp.net/mvc
I know there's an anti-MS tendency round here, but fast experimentation is just as easy with MS these days as it is with everyone else.
The only gripe I've got with MS these days is that they seemed obsessed with videos, which are irritating as you can't go at your own pace (i.e. faster) and it's a nightmare when you just want to find that way of doing x that you remember seeing in the video but not at which point.
It's much worse with Java. Instance startup times are a lot higher.
> Clone an app? What, with all its settings? Random library includes you weren't expecting?
In Python, at least, you just edit app.yaml and change app name and version. If there is a settings file, you also edit it. And libraries that are not provided by GAE should be included in the app, so, you are bundling dependencies Java-style.
> Also visual studio has a free edition these days.
Unfortunately running Windows takes away many nice things for developers.
>Unfortunately running Windows takes away many nice things for developers.
This. Having essentially one choice in monolithic IDE which doesn't really provide anything novel you can't get elsewhere, without a nicely integrated POSIX shell and all the useful stuff that comes with it is a net lose IMO.
http://www.mono-project.com/Release_Notes_Mono_2.10#Language...
Now I just need to find the time to try it out...
to somebody not already a C#'er is the cost of Visual Studio: You really want that concurrency profiler, which I think is only in the Ultimate SKU (for which MS is giving away licenses in Bizspark, dreamSpark).
And a fair number of Csharpers I've met recently (admittedly a small sample) will tell you FP and the parallel/concurrent libs in C# are good enough: TPL, the .NET 5 Async lib (Basically they'll tell you about all the C# stuff in Petricek's book, without looking into what F# can do for them.
I'll concede that sometimes there are specific cases on which you need to use a new or non-mainstream language or environment. (I'm 99% certain that this business is not one of those cases.) I'll concede, too, that being on the cutting edge is pretty cool and gets you lots of hacker cred.
But if you want your business to survive the departure of the founding team, you need to consider whether you can solve the problem with a mainstream environment. If you don't, don't expect your invention (at least in its current form) to carry any legacy.
Case in point:
People may think Paul Graham is a fucking genius for selling Viaweb to Yahoo!, despite being written in Lisp, but I assure you none of his code still lives on there, not even a fork. His brilliance is as a businessman for getting Yahoo! to fork over $<lots> to him, not for the technical merits of Viaweb itself.
Viaweb was purchased because it was successful. According to PG, it was successful because they could implement major features in a weekend.
That the original author can major features quickly does not necessarily imply that an ongoing concern after the departure of the original author will also be able to implement major features quickly.
Choosing your environment with your successors in mind usually increases the value of your business (assuming, of course, that your valuation has a rational basis, which tech companies aren't always good at determining).
If using Java makes you 10% slower, but your code is more maintainable by Yahoo - Yahoo will buy your competitor who won market share with their superior features and quick response to customers.
NB: Choose Java because it is "enterprisy"
Consider the Crash Bandicoot folks. Sure after Naughty Dog got bought, Sony couldn't figure out how to deal with the Lisp codebase and future projects were in C++. But Lisp let them build the first major platform game on the PSX, beating Sony itself to the market, and none of the stuff afterwards would've happened without that.
Steve Jobs had it right on this: artists ship.
You forget that he had something to sell. He explains why his choice of lisp helped him with that detail.
You say: "...if you want your business to survive the departure of the founding team, you need to consider whether you can solve the problem with a mainstream environment." Could you go into more detail about what you mean by a mainstream environment and what advantages it would bring in this case?
But if the business is going to continue to operate, the new development team needs to be capable of tending to the product, as it will need bug fixes, security patches, and perhaps be updated to scale better.
If the product is built on an uncommon platform, finding new developers, especially senior ones, to tend to the product will be difficult and expensive. If a sufficient number of qualified developers cannot be found, the business may have little choice but to rebuild the product using a more mainstream technology. In the meantime, the business takes on a higher risk of continued operation if the current implementation (which is no longer maintainable) is surpassed by its competitors.
TL;DR: short-term optimizations create technical debt, which screws your business in the long run.
Would you regard the use of a non-mainstream language to be, in of itself, technical debt?
It depends. If the code adequately serves the business purpose its was designed for and does not need modifications, then no. Otherwise, yes.
But the answer to the question can change over time. What is not technical debt today could be tomorrow.
How is that a valid argument? "not enough sex going on here" ... if you choose a language because it's sexy and not because of its utility, you are making a poor business decision.
Digg spent a long time revamping their system in new languages and on new databases and look how that turned out (not that the languages were the core part of their failure, but still)
I actually thought he meant it in a Darwinian/Dawkinsian(?) sense. That would have been a pretty neat turn of phrase. But yes you're right, he clearly means sexy. Stupid word.
It is perfectly possible to configure a servlet environment once with 15 lines of XML and then never look at it again.
Look at Google Guice for example, which does awesome Java based configuration and composition of components.
Deployment is also a breeze. Place WAR file on a server, let operations people grab it and drop it in a servlet container. Done.
Python need to reach a point where it is that simple. I see people around me handle this with Python but it takes a lot of effort to get it right.
I asked the teacher about your comment on String manipulation. Yes, it is pretty in-efficient. There are libraries to make manipulation easier, and, you can always go down to binary types, which is much more performant.
We chose to use Erlang for a variety of reasons, and String manipulation isn't a problem for us (it's definitely a pre-mature optimization point (for us) at the moment).
1> [104,105].
"hi"
2> [$h,$i].
[104,105]
3> "hi".
"hi"
Which is efficient for some uses (i.e. iterating over UTF32 characters) and inefficient for others (high memory usage).You can always use:
* atoms - for interned strings or enums
* binaries - for memory efficiency (i.e. UTF8 byte sequences)
* IO-lists - for efficient appending and IO.
What I would like is a per-module compiler directive/pragma, which will turn every "string" into <<"string">>, while @"string" will remain syntax sugar for list.A typical deployment instruction file (project.clj) looks like this and with "lein deps" you're all set up in a few seconds.
(defproject leiningen "0.5.0-SNAPSHOT" :description "A build tool designed to not set your hair on fire." :url "http://github.com/technomancy/leiningen :dependencies [[org.clojure/clojure "1.1.0"] [org.clojure/clojure-contrib "1.1.0"]] :dev-dependencies [[swank-clojure "1.2.1"]])
But take a look at https://github.com/technomancy/leiningen
int main(int argc, char **argv) {
printf("whee!\n");
return 0;
} (defproject leiningen "1.5.0-RC1"
:description "A build tool designed not to set your hair on fire."
:url "https://github.com/technomancy/leiningen"
:license {:name "Eclipse Public License"}
:dependencies [[org.clojure/clojure "1.2.0"]
[org.clojure/clojure-contrib "1.2.0"]
[lancet "1.0.0"]
[jline "0.9.94"]
[robert/hooke "1.1.0"]
[org.apache.maven/maven-ant-tasks "2.0.10" :exclusions [ant]]]
:disable-implicit-clean true
:eval-in-leiningen true)A "lein uberjar" will roll your whole program up into one .jar file, ready for deployment.
Oddly enough, C has superior runtime checks (dynamic library loading using major version numbers) and install-time checks (*nix package dependencies) compared to dynamic languages. It is very, very strange to me that other language communities have neither embraced alternatives such as OSGi nor worked to transfer responsibility to native package managers such as dpkg or RPM. C et al. under Linux have set a standard that other language communities don't seem interested in matching, much less exceeding. As far as I know, the standard answer for deploying a security update to a Java library is to rebuild and redeploy the entire application that depends on it.
Out-of-the-box it's a development tool, but it has plugins for certain types of deployment: https://github.com/technomancy/leiningen/wiki/Plugins
You can create tar/jar/war files, deploy artifacts to remote mvn repositories, push to the Google App Engine or Elastic Beanstalk, etc. Deployment to generic unix servers is handled by Pallet, which integrates well with Leiningen: https://github.com/pallet/pallet-lein
> Does it let different applications share libraries when possible?
This goes strongly against the culture of the JVM for various reasons that are outside the scope of Clojure itself.
> Can you push a security fix for a library to servers in the field without completely rebuilding and redeploying every application that uses the library.
Sure, this is pretty easy to do with Swank, but the specifics are going to vary widely based on the type of deployment.
>> carefully isolating your app from other apps in an application container. I.e., don't even try to solve the problem of sharing libraries between applications.
I would argue that this is the most reliable solution, if you can do it.What version of a library is installed, if any: dpkg -l
What's the latest version available to be installed: apt-get update, apt-cache search
Update the library for all applications that link it: apt-get upgrade
These tools are important for administration, security, and troubleshooting. It mystifies me that Java developers and sysadmins managing Java servers don't demand them. Instead, they're willing to muck around in a web interface, go back to their dev box to look at Maven scripts, or go searching through the filesystem just to see what software is installed and running. Even if you're not averse to doing that by hand, how scriptable is it?
If you have a large Java environment with many different services running on dozens or hundreds of boxes, how do you get a report of which boxes have a particular version of a library installed? I know our sysadmins don't know how, and I know our Java developers don't care. They could develop the tools themselves, but they have no interest. The sysadmins do not do web pages; they are not going to spend all day going click-click-click to update a few dozen servers. They have told the Java developers not to expect the same level of support for their applications as our C++ programmers get because Java is an unmanageable platform, and the developers don't care. I really, really do not understand why our Java programmers are not writing scripts to automate any of these basic tasks.
I guess it is important to discuss the tools of the trade, but I wonder how long it will be before asking someone what technology they used to build an app would be like asking a musician what brand her instrument was. ("Hey, great song - is that a Stratocaster?) At the end of the day, does your app enable your business to make money? Heck, I've built a business with revenues in the millions using MS Access as one of the main tools! (A long story there!)
What I have learned, which other commenters have already pointed out, is that issues like maintainability after you are gone, finding resources quickly, the size of the user community, etc. are all equally important and should not be overlooked.
wat
Probably the author underestimated node.js though. Using the author's language, it has a lot of "sex".