ClojureC, a compiler for Clojure that targets C as a backend
github.com
github.com
What worries me about this and similar efforts (like https://github.com/takeoutweight/clojure-scheme) is that clojure's standard library design assumes that the underlying runtime will be do some kind of polymorphic method inlining.
For example: the sequence library is all defined in terms of ISeq, which basically requires a "first" and "rest" to be defined for the data structure in question. These are polymorphic: there are different implementations of these for different data structures (list, vector, map, etc). So a dispatch step is required to choose the right one. In clojure-jvm, this is implemented using a java interface; this means the jvm will inline calls to said methods when they're being used in a tight loop. And if you use the standard library, calls to 'first' and 'rest' are going to be inside nearly all of your inner loops.
Compare this to a normal lisp or scheme: 'first' and 'rest' (or 'car' and 'cdr', whatever) are monomorphic. They only work on the linked-list data structure. So compiling these directly down to C functions makes perfect sense and incurs no performance penalty.
So in summary: clojure assumes theres a really smart JIT which is helping things along. This means it's not as suitable for alternate compilation targets as you might want it to be.
I wonder if there's something clever you could do here. Vtables could be reordered based on expected usage, certainly. Clojure can already do some measure of type inference, so this could be used for AOT inlining when it's available. Even if it's not, perhaps several versions of a call could be speculatively generated based on what the compiler does know already. The normal polymorphic inline caching technique could perhaps be abused to apply here. But it's hard to see how any of this can work in absence of a profile or heavy hinting.
(not a compiler writer, just interested in the problem)
For a fair amount of cases you can do static analysis to get good guesses at likely types, even for cases where you can't be sure. E.g. speculatively even looking near call sites by method name to see if you can guess the type of objects that will get passed in looks to get you a reasonable chance at guessing at the top contenders to let you speculatively generate inlined versions without creating too much junk. But to get the most performance out of this you're likely to need to be prepared to do some very basic JIT.
That is devious and fantastic.
> But to get the most performance out of this you're likely to need to be prepared to do some very basic JIT.
Yeah. But the attractive targets here are places where you can't have a JIT: embedded systems and iOS.
You still can't have memory pages marked write+execute, which is what you need for a JIT.
The idea of speculatively looking at method names comes from testing that to create vtables ahead of time for Ruby classes, to avoid hash tables in the common-case.
As it turns out, most method on most Ruby classes are the ones inherited from Object or other standard classes, and the number of classes is usually fairly constrained, so again speculatively looking at method names in the compile-time available source and allocating sparse vtables for the most common names results in relatively little waste.
And it reduces typical method lookup to a vtable lookup for common methods, with expensive method dispatch becoming much more rare. There's the tradeoff between theoretical horrible blowup in vtable waste from apps dynamically adding tons of methods and tons of classes, with a unique vtable slot required for each method name across all classes, vs. falling back to doing hash-table lookups all the way up the inheritance chain for "unusual" method names ones you reach certain thresholds for waste.
You do incur the cost of propagating vtable changes down the inheritance tree when methods are dynamically redefined in other places than leaves, but it is fairly rare to see apps where this happens at a very high rate, and the number of subclasses usually fairly small, so it is likely to be quite cheap. Doing it that way is something I first saw in a technical report by (now) prof. Michael Franz from '93 or '94 on "Protocol Extension" for Oberon.
You can probably also get some decent gains by adding heuristics to give preference to names that appears to be used in loops when picking names for the vtables to reduce the need of any JIT'ing.
That said: It looks like a really interesting project. I do wish they would phrase the project description as a "Ruby VM" instead of a separate programming language, through. There should not be any need to fork the entire language just to provide better performance.
I just wanted to raise your attention to it, but you are right, it would be better to have a proper Ruby compiler available instead.
What is there uses vtable's exclusively - I effective punted on the slow path (and so on adding methods at runtime) completely, but keep track of how much of the vtable allocations is wasted space. If/when I get there, the goal is to use various mechanisms like this to determine when to fall back on a slow path, and couple both with polymorphic inline caching when suitable.
EDIT:
I don't see MRI as very interesting to work on, largely because interpreters aren't much fun, and ironically given the amount of time I spend using Ruby, I prefer compilers to be as static as possible. I also prefer my compilers to be bootstrapped in their target language. Hence my "ideal" Ruby compiler would be written in pure Ruby, do a ton of static analysis, with minimal fallback to JIT when users user features that are too dynamic to analyse fully ahead of time
E.g. there's a ton of annoying uses of eval() in Ruby code where a more complete meta-programming API would make it trivial for a compiler to do full ahead of time static analysis, so one big thing an AOT Ruby compiler really need to do is to provide a library of compiler specific meta-programming facilities with a fallback that uses eval() as needed, and either convince people to use it, or provide monkey-patches for a number of popular projects. Some of these uses don't even need eval() in the first places, but uses it just as a quick shortcut because it's simpler...
Just to make clear, I'm not sure when or even if my compiler project will ever get to a state where it's even remotely useable to compile Ruby. I started it out without even having decided to compiler Ruby, mostly to write about various parts of the process of writing a compiler that I find interesting. I find compiling Ruby as incredibly fascinating from a theoretical point of view because of the complexity involved, but unfortunately working on it takes a lot more time and effort than thinking about the problems.
Sure, MRI is boring, but it's the best implementation right now (though some may argue that JRuby is better), and it's in desperate need of VM innovations.
Any new compiler/VM starting from scratch will be years away from being available for use in production environments. By the time it's finished, we'll all be using Go. (Sigh.)
E.g. nothing will stop MRI from having to interpret thousands of lines of code each time because it can't draw a line between runtime and compile time, while for an ahead of time compiler for Ruby, finding a pragmatic line between what needs to be executed at runtime vs. compile time is essential (consider for example the tendency to do stuff like getting the list of files in a directory and require all of them in turn).
(I agree about JRuby. I also wonder why Rubinius, which showed so much promise at the beginning, has stagnated. Is it simply the lack of developers?)
Regarding Rubinius, writing compilers for dynamic languages is hard. Most textbooks you'll find cover techniques most suitable for statically typed languages (the best resource I know for starting to catch up on compiling dynamic languages is actually the Self papers). So you need more than an unusual level of interest in writing compilers to be likely to try to tackle a language like Ruby which is tricky even for dynamic languages (e.g. my favorite pet problem to meditate on: What constitutes 'compile time' vs. 'runtime' for ahead of time compiled Ruby?), and even more to actually persevere until you start getting proper results where you can get decent results in days with a simpler language.
It's made worse because of Ruby's horrendous grammar. And I mean that from a compiler writers perspective - as a developer I love to use Ruby to a large extent because the complexities of the grammar means it reads and writes better 95% of the time. But MRI's bison based parser was 6k-7k lines with ugly parser/lexer interplay last time I checked... There are full compilers substantially smaller than that for other languages...
To me, these complexities are part of what makes it fascinating. I firmly believe you can parse Ruby fully with a much, much simpler parser for example. A lot of the ugliness can be abstracted away, and C parser code is rarely good examples of succint code.
I did start playing with MRI years ago, specifically the parser, actually, and started chopping out redundant pieces, but got frustrated and bored with it. That's part of the problem - it's one thing to play around with a toy compiler like I've done, and another entirely to put in the effort to push a major change to MRI through to production quality given the number of years of accumulated history encapsulated in it. Doing the latter as a hobby is a daunting task.
However my applications run on systems with very limited resources (and a Harvard-ish architecture, e.g separate data/program memories) and I wonder how far tools like ClojureC could go with regard to these constraints.
This requires you to execute the application with profiling turned on.
Then you compile the application a second time with the additional input of the profile results, this way the optimizer gets additional information that helps it to do decisions similar to a JIT.
It is all a matter of adding support for this.
It might be interesting to translate Clojure into a subset of Scheme that is widely supported -- and use one of the mature Scheme -> C compilers (Gambit, Chicken, Bigloo) to generate the target executable.
If that's not your problem with the JVM, what is?
At the same time, splitting your application into a client/server architecture is not a hack but an engineering decision. There are times when this decision is natural e.g. Music Player Daemon (MPD)[1]. For most CLI applications, there's no clear benefit (but the general approach has no clear downside either - the code overhead of this approach can be brought very low).
Certainly, in a production application you would want to secure the messaging channel (Nailgun doesn't).
[1] A music playing server: http://www.musicpd.org/. Some of the clients happen to be command line: http://mpd.wikia.com/wiki/Clients#Command_Line_Clients
I suppose if you're coming from a C-based runtime like Ruby or Python, you have to adapt your workflow. I imagine that a lot of Clojure people come from Java and have already shaped their workflow around the JVM.
As such, I'd say efforts such as these are greatly welcomed.
Like it or not, this is what made Clojure successful in the enterprise, at least when compared against other Lisps.
I see lots of Java shops now having Clojure code and talking about it at JUGs and how it enables their business.
You just need to have a look at InfoQ, Skills That Matter, Devoxx or Jax for ongoing talks.
I believe it to be nonsense.
Does he think that it's technically more powerful? Again he should be able to prove that if that's the case.
Otherwise he's just giving a shitty opinion, and should say that.
I think the claim is nonsense because with inline assembler there is nothing that you cannot express in C that you can with LLVM. So the decision between the two is opinion.
It's not like it's some controversial opinion what he said -- it's both self evident and common place. It's you who offers the more controversial opinion (and in a rude way, to top).
>I think the claim is nonsense because with inline assembler there is nothing that you cannot express in C that you can with LLVM. So the decision between the two is opinion.
It's not about "expression", and nobody argued that you can express more in LLVM.
This is missing the point by miles!
It's about having more structure and less of an ad-hoc pipeline, which helps with better tooling, error prevention, etc.
(Not only what you wrote is wrong, but even if the original argument was about expression, your opinion would still be wrong. Two things offering equivalent expressive power, does not mean that they are just as good to use in practice at all. Might as well ask "why invent new languages, when assembly can express everything").
The only benefit to using C for something like this is portability, which is something else altogether.
As someone who has written more than one compiler, I don't see how it is self-evident at all. It's also not at all that common-place compared to generating C or asm output textually.
> It's about having more structure and less of an ad-hoc pipeline, which helps with better tooling, error prevention, etc.
Those provide some benefits, sure. At the cost of massive amounts of complexity in the case of LLVM.
> The only benefit to using C for something like this is portability, which is something else altogether.
Now it is you who are wrong. Other people have already pointed out, for example, that C provides an easy-to-read intermediate format, and is simple to generate, as other benefits. Not having to deal with a massive C++ codebase is another.
You may disagree that these other benefits are worth it, but for me at least they are (just taking a break from a compiler that generates textual asm because I find even that preferable to dealing with LLVM).
So portability and less dependencies, plus easier.
Outputting text to be interpreted as code is far more low level and error prone than targeting an AST via an API like LLVMs.
And you loose a lot of high quality tooling that you could take advantage of.
Can you prove that? If not, it's just a baseless opinion.
Not everything can (or should) be proved of the drop of a hat in a discussion list -- that doesn't make everything without a formal proof "baseless" opinion.
If you cannot see the self-evidentness (sic) of a STRUCTURED API to produce an AST makes it easier to avoid mistakes compared to spitting out text to compile as a C program, then I'm not sure any proof would help anyway.
It's like asking me to prove why using an XML processor to crete and save a DOM tree would produce more error free results than manually compiling tags as strings.
Or why parsing a JSON file and working on the nodes is less error prone than using regular expressions to extract values from the JSON as a big string.
Isn't that just a charming response to get?
I don't see any proof fot that. Where's your proof?
Not everything can (or should) be proved of the drop
of a hat in a discussion list
If you say something that doesn't make sense, you're saying that one shouldn't have to back that up? It's like asking me to prove why using an XML processor
to crete and save a DOM tree would produce more error
free results than manually compiling tags as strings.
It is in fact the equivalent of asking you to prove that an XML parser after serializing a DOM tree and then parsing that same document produces a non-equivalent DOM tree to the original.I think the reason pat_punnu is balking at what you have said is because of this: You can think of C as a serialized form of an AST. In order for parsing to produce a non-equivalent AST to the one used to serialize it, the C grammar must be non-deterministic. The C grammar is not non-deterministic, therefore what you said does not make sense and pat_punnu (somewhat rudely) asked you to back up what you were saying.
Asking people to back up what they have claimed is part of intelligent discourse. This happens often on HN and is one of the reasons I like this community because when the person who makes a non-intuitive or seemingly wrong remark turns out to be correct, I learn something.
UPDATE: I should point out that I'm responding specifically to this sub-thread of the tree and not arguing about whether or not this should target the LLVM IR. I think that would be nice.
Read what you quoted from me, and what you ask. Where do I say that people should not back up things they say that "don't make sense"?
Where do I even say they do not have to "back up" the things they say? I merely say that they do not have to PROVE everything. You can back stuff up with some arguments and counterargurments, you don't need to provide some "proof".
>I think the reason pat_punnu is balking at what you have said is because of this: You can think of C as a serialized form of an AST. In order for parsing to produce a non-equivalent AST to the one used to serialize it, the C grammar must be non-deterministic. The C grammar is not non-deterministic, therefore what you said does not make sense and pat_punnu (somewhat rudely) asked you to back up what you were saying.
And the reason I'm balking at this is that you examine the case AFTER C has been generated. I'm not talking about that stage (when reading back C to generate an AST). I (and pat_punnu) and talking at the previous stage of spitting out the C code to disk in the first place. I'm saying that a structured way to do that (LLVM API) is safer than merely creating strings yourself.
So your: "It is in fact the equivalent of asking you to prove that an XML parser after serializing a DOM tree and then parsing that same document produces a non-equivalent DOM tree to the original"
takes this from several steps ahead. I (and pat_punny) were concerned with the generation of the document in the first place.
That is resting on the assumption that using the LLVM API is less error prone to a typical compiler developer than creating strings.
You are also assuming that spitting the C code to disk needs to be done in an unstructured way.
Neither of these assumptions are self evident.
It is not a given that there are "more moving parts" in generating C output from a compiler than in using LLVM.
While that's a true enough statement by itself, your snipe conveniently skips half of the process in question.
No, but it's a given that the LLVM moving parts have been already written, and are tested by millions.
Your moving parts in your own solution, you'd have to write yourself.
> Your moving parts in your own solution, you'd have to write yourself.
That's not always a bad thing for error rates, if the alternative is figuring out to use a massive library correctly.
You are wrong here. Unproved statements are not necessarily baseless opinions. They may be well supported opinions.
I think you meant to ask him, "Can you support that?"
The reality of the matter is that interacting with non-java linked libraries is a real pain from every JVM language know of. Two years ago this was posted here [https://github.com/jasonjckn/llvm-clojure-bindings], but since then LLVM has gone through two major restructurings so it didn't work out of the box and I estimated that it would take less effort to build my own naive infrastructure than to patch this one & integrate it.
As a result I'm generating code in terms of lists of newline terminated assembly statement strings that I can just print or write to a file when I'm done. While I agree that C is a sub-optimal output format in that you have to compile the output, it is also the clear lingua franca for systems programming and assembly generation these days. Generating C gives you interesting options like linking to other C codebases or your own C code the same way that cljs gives you the option of interacting with "native" javascript libraries as well as clojurescript toolkits.
[0] https://github.com/halgari/mjolnir
[1] https://github.com/strangeloop/clojurewest2013/tree/master/s...
My Firefox history indicates that I have read the Mjolnir page before, but I don't recall why I didn't use it at the time. Taking another look :-P. Thanks for the link!
With Datomic, the inference engine is completely re-written in datalog. This allows for a massive code clean-up, and the code in that branch is much cleaner.
That said I'd love to see a serious Clojure-in-Clojure targeting LLVM.
Some other advantages of C over LLVM IR:
ABI compatibility with C automatically. In LLVM IR, being ABI-compatible with C is often a considerable headache and you have to do different things on each platform.
You can use C libraries by #including their header files (the way many C libraries were designed to be used), instead of hardwiring all that information (macro expansions, enum values, typedefs, inline functions, struct definitions, etc.) from their header files in your code generator.
You have the option of switching C compilers; you're not as locked into a single backend.
With a bit of care, you can make your output much more readable. LLVM IR demands either SSA form (most front-ends don't want to do this) or herds of allocas, loads, and stores everywhere. In C, you just say "int x;" to declare an int, and just "x" to refer to it.
You can use C features like bitfields, designated initializers, short-circuit operators (&&, ||), compound assignment (+=, -=, etc.), and so on. You can do all these things in LLVM IR, but you have to lower them yourself.
It's a fine strategy to start with, there are more important things than fucking around with LLVM IR at this point, they can switch to LLVM or a native generator if they get the thing effective and off-the ground.
Wow, that's a pretty skimpy list of dependencies. But..
"Make sure you're using Leiningen 2."
..argh, installing that on ubuntu that requires 110 packages. All that just for a build system?
Don't install lein using the OS packaging system (apt, rpm, yum, etc.).
Instead, just grab the `lein` script (linked to from http://leiningen.org/ ), put it into your ~/bin, set it executable, and you're all set.
Maybe multithreading is just not yet implemented?
> So that kills one of the main reasons to use Clojure in the first place.
I could see ClojureC being very useful for when you want:
* small footprint
* easy C-library interop
* fast start-up time
and where you don't necessarily need multithreading.(BTW, I'd be curious to hear what you think might be the differences in use cases between ClojureC and mjolnir/clojure-metal.)