C--
cminusminus.org
cminusminus.org
The cminusminus domain is no longer valid (though it has more modern CSS), also it lacks links to all the informative papers!
C-- is very similar overall to LLVM IR, though there are crucial differences, but overall you could think of them as equivalent representations you can map between trivially (albeit thats glossing over some crucial details).
In fact, a few people have been mulling the idea of writing a LLVM IR frontend that would basically be a C-- variant. LLVM IR has a human readable format, but its not quite a programmer writable format!
C-- is also the final rep in the ghc compiler before code gen (ie the "native" backend, the llvm backend, and the unregisterized gcc C backend).
theres probably a few other things I could say, but that covers the basics. I'm also involved in GHC dev and have actually done a teeny bit of work on the c-- related bits of the compiler.
relatedly: i have a few toy C-- snippets you can compile and benchmark using GHC, in a talk I gave a few months ago https://bitbucket.org/carter/who-ya-gonna-call-talk-may-2013... https://vimeo.com/69025829
I should also add that C-- in GHC <= 7.6 doesn't have function arguments, but in GHC HEAD / 7.7 and soon 7.8, you can have nice function args in the C-- functions. See https://github.com/ghc/ghc/blob/master/rts/PrimOps.cmm for GHC HEAD examples, vs https://github.com/ghc/ghc/blob/ghc-7.6/rts/PrimOps.cmm for the old style.
Slides: http://www.cs.tufts.edu/~nr/c--/download/c--exnslides.ps.gz
Audio: http://wino.eecs.harvard.edu:8080/ramgen/nr-pldi00.rm
The Manual: http://www.cs.tufts.edu/~nr/c--/extern/man2.pdf
The manual contains the specifications and a few code examples. It looks like it's easy to learn, but it's a little different than other languages.
Edit: Ok, I've found the following SO thread: http://stackoverflow.com/questions/3891513/how-does-c-compar...
As for how they compare, the answers to this question, on why Haskell didn't use LLVM (in 2009, it didn't) get into it a bit: http://stackoverflow.com/questions/815998/llvm-vs-c-how-can-...
That said, so much work goes into LLVM supporting new platforms, optimizations, etc, that it's probably easier to hack around LLVM's limitations than use C--, etc.
But I'm a Luddite, though; I like self-hosting native compilers and I'm not terribly fond of the idea of having to pack a 20MB blob of C++ code with my code either. (There's a lot of apps that could profit from dynamic translation at run-time, but using LLVM for that feels unnecessarily heavyweight. I wish the algorithms and passes from LLVM were available in form of some high-level DSL that you'd be able to translate into whatever language you use and automatically adapt to whatever data structures you use in your code. It doesn't sound exactly impossible.)
In a world where my phone has 2gb of RAM, I don't understand why 20mb for an optimizing compiler is in any way onerous or unreasonable.
<label> ::= <letter> [ [ <ldh-str> ] <let-dig> ]
<ldh-str> ::= <let-dig-hyp> | <let-dig-hyp> <ldh-str>
<let-dig-hyp> ::= <let-dig> | "-"
<let-dig> ::= <letter> | <digit>
<letter> ::= any one of the 52 alphabetic characters A
through Z in upper case and a through z in lower case
<digit> ::= any one of the ten digits 0 through 9I'm involved in GHC dev (and thus incidentally c-- dev as it exists in GHC). And i may be spending a lot of time helping improve ghc's code gen over the coming year (which is essentially the most widely used c-- compiler on the planet per se)
your remark intrigues me!
http://www.georgehernandez.com/h/xComputers/Programming/Medi...
If this has any truth (perhaps a different c--?) I'd like to know which one is being referred to.
It's quite an obvious name, I'd expect there to be many more called C--!
Yes, the graph talks about http://sourceforge.net/projects/cmmscript/ not the C-- used in/extracted from GHC.
There's at least one other C-- used in the OCaml compiler.
An overview here: http://cr.yp.to/qhasm/20050129-portable.txt
EDIT: changed, I said decrement and meant increment.
for( i=LEN-1; i>=0; --i ) {...}
I've carried this into a lot of other languages, as it's equally readable to the alternatives, when ordering isn't important. In (mainly older) optimised code where ordering is important, you'll often still see this, with a second incrementing variable so that the exit condition retains the compare to zero (avoiding a variable comparison); although that part is separate from the use of the pre-decrement.
I'm not sure if either of these make much difference with modern CPUs and compilers. Certainly not worth worrying about for the most part.
It's untrue that this idiom started with C++. This style is preferred by K&R, which predates optimising compilers. Since then it's use among C programmers has probably been force of habit, but it is definitely idiomatic. It's use in K&R makes it about as idiomatic as it is possible to be.
for (i = 0; i < n; i++)
...
This is the kind of code I run across most commonly (almost exclusively) in pre-1990s C. It's also the style used in the old Unix sources. For example, take a look at the source code to 'nohup' or 'mount' from 5th Edition Unix, 1974: http://minnie.tuhs.org/cgi-bin/utree.pl?file=V5/usr/source/s... http://minnie.tuhs.org/cgi-bin/utree.pl?file=V5/usr/source/s...They do also use predecrement/preincrement, but only in assignments or comparisons, where it actually semantically matters that the increment/decrement is "pre":
while(*--np == '/')
*np = '\0';
In contexts where it doesn't matter, like the 3rd clause of a for loop, or just incrementing a variable as a standalone operation, they always default to x++. for (fahr = 0; fahr <= 300; fahr = fahr + 20)
...
Each for loop after the introduction of ++ uses preincrement (where applicable) for a while, and then the style shifts to postincrement.EDIT: specifically, increments are deliberately prefix ("For the moment we will stick with the prefix form") until the full introduction of increment and decrement operators in section 2.8, after which they are postfix.
That trend didn't last very long, though.
Now as for current ideas and projects similar in concept to c--, that is interesting to think about.
Maybe the startup lesson is sometimes, if you try to intermediate yourself as a middleman, even if you do a good job of it, and appear to be a good idea, it just doesn't work. It would be interesting to dissect the c-- experience and figure out why.
Imo, that hypothesis did have some legs, but people are now instead usually using LLVM in various ways to accomplish it. LLVM's intermediate representation isn't really properly cross-platform, but it can be sort of hacked to be used for that purpose.
My impression is that the GHC people, at least, still think that C-- is a nicer IR than LLVM-IR for their purposes. But they are slowly moving more things to LLVM anyway, because the LLVM project as a project has, in the meantime, built up a lot more momentum and infrastructure. In the early 2000s this wasn't obvious, but in 2013 it's clear that LLVM has an ecosystem, institutional support, resources to maintain ports, etc., while C-- didn't manage to get the same traction.
From Xavier Leroy, one of the lead Ocaml developers [1]:
I think I'm the one who coined the name "C--" to refer to a low-level,
weakly-typed intermediate code with operations corresponding roughly
to machine instructions, and minimal support for exact garbage
collection and exceptions. See my POPL 1992 paper describing the
experimental Gallium compiler. Such an intermediate code is still in
use in the ocamlopt compiler.
I had many interesting discussions with Simon PJ and Norman Ramsey
when they started to design their intermediate language. Simon liked
the name "C--" and kindly asked permission to re-use the name.
However, C-- is more general than the intermediate code used by
ocamlopt, since it is designed to accommodate the needs of many source
languages, and present a clean, abstract interface to the GC and
run-time system. The ocamlopt intermediate code is somewhat
specialized for the needs of Caml and for the particular GC we use.
[1] http://article.gmane.org/gmane.comp.lang.caml.inria/9436/First they are talking about the problem and then they present the solution.
I also like the words that are marked bold.
This is how interaction design should be done (imho).
I know that Python compiles to C and that Clojure compiles to JVM (or even to JavaScript).
My cartoon:
scripting lang --> programming lang --> native code
Honestly, I have never experimented with Assembly language much except for COOL (http://en.wikipedia.org/wiki/Cool_(programming_language)) and TOY (http://introcs.cs.princeton.edu/java/52toy/).In CPython, Python compiles to bytecode, which is then interpreted by the Python interpreter (which itself is written in C).
You could pick C, but C is actually still fairly complex and has lots of undefined behavior. You could pick JVM, but then you'll pick up the entirety of the JVM architecture which is quite large and may include many things you're not interested in.
C-- is another choice. It's decidedly lower-level than C (and thus far lower than JVM), higher level than assembly, and was crafted, as far as I know, under the deep influence of how to compile the pure functional language Haskell (or more specifically, it's System FC style and STG underlying languages).
In the mean time, LLVM took a similar place in this hierarchy and has probably taken off much more than C-- has. GHC, a Haskell compiler, in fact is moving its compilation pathway that way.
For one, I think C-- still looks a lot like C, i.e. still an imperative language. CIL is a stack language (also with a built-in object system).
Idris uses llvm-general to have a simple llvm backend. Also llvm general is probably the nicest and most thorough llvm binding you'll find.