Some other advantages of C over LLVM IR:
ABI compatibility with C automatically. In LLVM IR, being ABI-compatible with C is often a considerable headache and you have to do different things on each platform.
You can use C libraries by #including their header files (the way many C libraries were designed to be used), instead of hardwiring all that information (macro expansions, enum values, typedefs, inline functions, struct definitions, etc.) from their header files in your code generator.
You have the option of switching C compilers; you're not as locked into a single backend.
With a bit of care, you can make your output much more readable. LLVM IR demands either SSA form (most front-ends don't want to do this) or herds of allocas, loads, and stores everywhere. In C, you just say "int x;" to declare an int, and just "x" to refer to it.
You can use C features like bitfields, designated initializers, short-circuit operators (&&, ||), compound assignment (+=, -=, etc.), and so on. You can do all these things in LLVM IR, but you have to lower them yourself.
That said I'd love to see a serious Clojure-in-Clojure targeting LLVM.
The reality of the matter is that interacting with non-java linked libraries is a real pain from every JVM language know of. Two years ago this was posted here [https://github.com/jasonjckn/llvm-clojure-bindings], but since then LLVM has gone through two major restructurings so it didn't work out of the box and I estimated that it would take less effort to build my own naive infrastructure than to patch this one & integrate it.
As a result I'm generating code in terms of lists of newline terminated assembly statement strings that I can just print or write to a file when I'm done. While I agree that C is a sub-optimal output format in that you have to compile the output, it is also the clear lingua franca for systems programming and assembly generation these days. Generating C gives you interesting options like linking to other C codebases or your own C code the same way that cljs gives you the option of interacting with "native" javascript libraries as well as clojurescript toolkits.
[0] https://github.com/halgari/mjolnir
[1] https://github.com/strangeloop/clojurewest2013/tree/master/s...
My Firefox history indicates that I have read the Mjolnir page before, but I don't recall why I didn't use it at the time. Taking another look :-P. Thanks for the link!
With Datomic, the inference engine is completely re-written in datalog. This allows for a massive code clean-up, and the code in that branch is much cleaner.
It's a fine strategy to start with, there are more important things than fucking around with LLVM IR at this point, they can switch to LLVM or a native generator if they get the thing effective and off-the ground.
I believe it to be nonsense.
Does he think that it's technically more powerful? Again he should be able to prove that if that's the case.
Otherwise he's just giving a shitty opinion, and should say that.
I think the claim is nonsense because with inline assembler there is nothing that you cannot express in C that you can with LLVM. So the decision between the two is opinion.
It's not like it's some controversial opinion what he said -- it's both self evident and common place. It's you who offers the more controversial opinion (and in a rude way, to top).
>I think the claim is nonsense because with inline assembler there is nothing that you cannot express in C that you can with LLVM. So the decision between the two is opinion.
It's not about "expression", and nobody argued that you can express more in LLVM.
This is missing the point by miles!
It's about having more structure and less of an ad-hoc pipeline, which helps with better tooling, error prevention, etc.
(Not only what you wrote is wrong, but even if the original argument was about expression, your opinion would still be wrong. Two things offering equivalent expressive power, does not mean that they are just as good to use in practice at all. Might as well ask "why invent new languages, when assembly can express everything").
The only benefit to using C for something like this is portability, which is something else altogether.
As someone who has written more than one compiler, I don't see how it is self-evident at all. It's also not at all that common-place compared to generating C or asm output textually.
> It's about having more structure and less of an ad-hoc pipeline, which helps with better tooling, error prevention, etc.
Those provide some benefits, sure. At the cost of massive amounts of complexity in the case of LLVM.
> The only benefit to using C for something like this is portability, which is something else altogether.
Now it is you who are wrong. Other people have already pointed out, for example, that C provides an easy-to-read intermediate format, and is simple to generate, as other benefits. Not having to deal with a massive C++ codebase is another.
You may disagree that these other benefits are worth it, but for me at least they are (just taking a break from a compiler that generates textual asm because I find even that preferable to dealing with LLVM).
So portability and less dependencies, plus easier.
Outputting text to be interpreted as code is far more low level and error prone than targeting an AST via an API like LLVMs.
And you loose a lot of high quality tooling that you could take advantage of.
Can you prove that? If not, it's just a baseless opinion.
Not everything can (or should) be proved of the drop of a hat in a discussion list -- that doesn't make everything without a formal proof "baseless" opinion.
If you cannot see the self-evidentness (sic) of a STRUCTURED API to produce an AST makes it easier to avoid mistakes compared to spitting out text to compile as a C program, then I'm not sure any proof would help anyway.
It's like asking me to prove why using an XML processor to crete and save a DOM tree would produce more error free results than manually compiling tags as strings.
Or why parsing a JSON file and working on the nodes is less error prone than using regular expressions to extract values from the JSON as a big string.
Isn't that just a charming response to get?
I don't see any proof fot that. Where's your proof?
Not everything can (or should) be proved of the drop
of a hat in a discussion list
If you say something that doesn't make sense, you're saying that one shouldn't have to back that up? It's like asking me to prove why using an XML processor
to crete and save a DOM tree would produce more error
free results than manually compiling tags as strings.
It is in fact the equivalent of asking you to prove that an XML parser after serializing a DOM tree and then parsing that same document produces a non-equivalent DOM tree to the original.I think the reason pat_punnu is balking at what you have said is because of this: You can think of C as a serialized form of an AST. In order for parsing to produce a non-equivalent AST to the one used to serialize it, the C grammar must be non-deterministic. The C grammar is not non-deterministic, therefore what you said does not make sense and pat_punnu (somewhat rudely) asked you to back up what you were saying.
Asking people to back up what they have claimed is part of intelligent discourse. This happens often on HN and is one of the reasons I like this community because when the person who makes a non-intuitive or seemingly wrong remark turns out to be correct, I learn something.
UPDATE: I should point out that I'm responding specifically to this sub-thread of the tree and not arguing about whether or not this should target the LLVM IR. I think that would be nice.
Read what you quoted from me, and what you ask. Where do I say that people should not back up things they say that "don't make sense"?
Where do I even say they do not have to "back up" the things they say? I merely say that they do not have to PROVE everything. You can back stuff up with some arguments and counterargurments, you don't need to provide some "proof".
>I think the reason pat_punnu is balking at what you have said is because of this: You can think of C as a serialized form of an AST. In order for parsing to produce a non-equivalent AST to the one used to serialize it, the C grammar must be non-deterministic. The C grammar is not non-deterministic, therefore what you said does not make sense and pat_punnu (somewhat rudely) asked you to back up what you were saying.
And the reason I'm balking at this is that you examine the case AFTER C has been generated. I'm not talking about that stage (when reading back C to generate an AST). I (and pat_punnu) and talking at the previous stage of spitting out the C code to disk in the first place. I'm saying that a structured way to do that (LLVM API) is safer than merely creating strings yourself.
So your: "It is in fact the equivalent of asking you to prove that an XML parser after serializing a DOM tree and then parsing that same document produces a non-equivalent DOM tree to the original"
takes this from several steps ahead. I (and pat_punny) were concerned with the generation of the document in the first place.
That is resting on the assumption that using the LLVM API is less error prone to a typical compiler developer than creating strings.
You are also assuming that spitting the C code to disk needs to be done in an unstructured way.
Neither of these assumptions are self evident.
It is not a given that there are "more moving parts" in generating C output from a compiler than in using LLVM.
While that's a true enough statement by itself, your snipe conveniently skips half of the process in question.
No, but it's a given that the LLVM moving parts have been already written, and are tested by millions.
Your moving parts in your own solution, you'd have to write yourself.
> Your moving parts in your own solution, you'd have to write yourself.
That's not always a bad thing for error rates, if the alternative is figuring out to use a massive library correctly.
You are wrong here. Unproved statements are not necessarily baseless opinions. They may be well supported opinions.
I think you meant to ask him, "Can you support that?"