C as an Intermediate Language (2012)
yosefk.com
yosefk.com
Let's start with an excellent quote from Wittgenstein. "The limits of my language mean the limits of my world."
Using C as your intermediate language means that your expressiveness is limited to valid C programs. This is workable but only if your language can be mapped to C in _useful_ ways.
For example, let's say your language has behavior similar to scheme's tail-call. How would you get this behavior from a C compiler? You will never be able to make this reliably across optimization levels, etc.
Guaranteed tail-calls are the tip of the iceberg, there are a lot more features which cannot be reasonably mapped onto C.
Real compiler IRs increase your expressivity beyond what the C language designers decided was important.
I would cite GC as the most prominent example of a feature that is incompatible with compiling to C. It is impossible to write a competitive high performance tracing GC on top of C. Mandatory register spills at safe points and conservative scanning have huge downsides.
Posts on C get blindly upvoted on HN and proggit. I guarantee it. Me thinks it's like the history channel for programmers; learning about their forerunners and how they had to bang rocks together to make fire.
It's also how I've gotten all of my karma. Unfortunately, the abyss has also looked into me, as I now write C code for a living...
1. Write an interpreter to host the compiler (probably in C).
2. Write a cross-compiler.
3. Write an X-to-C compiler.
I insist on C here not because of performance, but because it's ubiquitous.
Well, there's still the unreasonable way: compile in continuation passing style, and trampoline every call. I shall not be held responsible for any performance problem this advice may cause however…
To take a simple example, I have here a 2428-line C program generated by compiling Linus Åkesson's Game of Life in BF (http://www.linusakesson.net/programming/brainfuck/) into C using Daniel B. Cristofani's dbf2c.b (http://www.hevanet.com/cristofd/brainfuck/dbf2c.b), which is a BF compiler written in BF. Compiling these 2428 lines of C to machine code using tcc 0.9.25 takes 20ms on my 1.6GHz Atom netbook. Most of this is about 16ms of tcc overhead (startup and shutdown time); the rest is compiling several hundred thousand lines of C per second with tcc. You should get several million lines of C per second with tcc on a modern machine.
This isn't optimized code, about equivalent to gcc -O0, typically about 3×–5× slower than optimized code. But that's enormously better than interpretation overhead.
(Using decent optimization levels with GCC makes it take several seconds to compile, because GCC's optimizer doesn't deal well with enormous functions.)
dbf2c.b, the C-generating BF compiler, is 892 bytes of BF code when stripped. Now, I'm not saying you should write your compilers in deliberately obfuscated programming languages in as few bytes as possible; I'm saying that the fact that this is even possible at all should give you really good feelings about how easy it is to compile things to C.
Also, C may have easier debugging support, but it's significantly worse than having exact control over the contents of your DWARF DIEs. Making it easy to have a bad debugging experience isn't a plus as I see it.
You're probably right that LLVM also has better optimization than C because of things like the possibility of doing better aliasing.
I agree that BF is unrepresentative of anything in the real world. Thank goodness.
I can certainly see how debugging support could be better if you can go beyond the support C gives you.
Despite this, there are many people out there who view languages that compile to C as being inferior. I still don't understand it.
A minimal porting effort for the best of all worlds.
You might not get the best optimization possible if you have to leave it up to the C compiler with inline assember... But that mix worked well for C and C++.
And rust can still optimize things like deciding which integer optimizations need the overflow check, and only add it where static analysis failed.
How often will idiomatic rust code need overflow checks?
Also, if the UB operations are defined by target assembler (instead of in C++) then you may not even need extra overflow checks, depending on the assembler semantics. This would only be needed in the places where static analysis failed, of course.
One could also use the native compiler options, if available, to avoid the issue alltogether (like -fwrapv).
No, you're talking about all signed addition.
> How often will idiomatic rust code need overflow checks?
You'd be surprised. Go look at the implementations of containers in the standard library.
> Also, if the UB operations are defined by target assembler (instead of in C++) then you may not even need extra overflow checks, depending on the assembler semantics. This would only be needed in the places where static analysis failed, of course.
I don't know what this means exactly. If you're saying you should implement a static analysis to try to eliminate unneeded overflow checks, then this is a huge burden.
> One could also use the native compiler options, if available, to avoid the issue alltogether (like -fwrapv).
If you're dependent on GCC/clang extensions to C, then you're not compiling to C. You're compiling to a front end to GIMPLE/LLVM. At that point there's no benefit over just compiling to GIMPLE or LLVM in the first place.
No, I'm talking about the ones that won't overflow. Some can be statically shown to not need a runtime check.
> If you're dependent on GCC/clang extensions to C, then you're not compiling to C.
I'm not saying that. I'm saying require either the target compiler (which the user is in control of) to provide a flag to give signed overflow the proper semantics (which would be useful if the user happens to be using clang or gcc), or fall back to requiring some asm for the target system which provides the needed semantics.
This would all be up to someone to configure for the target machine. Not a big deal once per architecture.
> At that point there's no benefit over just compiling to GIMPLE or LLVM in the first place.
Well if GIMPLE or LLVM was available for all of the targets that C compilers are, then I'd agree.
But they aren't. So I don't.
https://gcc.gnu.org/onlinedocs/gcc/Integer-Overflow-Builtins...
could be implemented in plain C if you're careful. It's not like the other extensions like __builtin_constant_p or asm that'd leave you stuck with GNU C.
And besides, GNU C is good enough for lots of people!
* nim
* Vala (GObject backend)
* Purescript (technically it has a C++ backend, but it is worth mentioning here for the very clean C++ it produces!)
I think all of the new languages are amazing these days. I am learning Rust, Clojure, Haskell, and many more.
Buy I can not use any of them at work. I'm in a position to influence teams of smart programmers, but the only language that I would be able to use would be one that fits in to the existing infrastructure, which is natively compiled shared libraries on esoteric unix platforms.
So no LLVM. No JVM. No Haskell. I could possibly get away with a lisp or scheme that compiled portable C if I wanted. But that doesn't excite me as much.
I would kill for a Rust to human-readable C++ transpiler. I think it could be used immediately by many.
I may have to write one.
This is why I think rust is a winner. I bet most rust code could be compiled to c++ very cleanly.
You get rust's compiler with all of the safety checks, and you get the platform support of C++.
It would be a dream.
The x86 assembly or LLVM intermediate representation that comes out of most compiler is not very human readable, but I don't see anyone having any problem with it whatsoever. Readability of generated code was never the point of C using C as a back-end. The point is portability and tooling.
By the way /u/jjnoakes can suggest fancy languages and compilers, if only they can generate native code on obscure platforms (I guess, something other than x86 and ARM). Generating C code (even utterly unreadable C code) would solve his problem.
All things that require work if I cross compile, but all things I get for free if I generate C or C++.
I would prefer readable C or C++ as much as possible because it gives me debugability if the code generator goes wrong and it gives me an out if I want to move away from the language some day.
<sarcasm> Now this might require a little bit more work. </sarcasm>
Or you port it every time you need a new platform.
But does it hurt to strive for better? I don't think it does.
For me, C as an intermediate vs LLVM language is a "worse is better" or "good enough" issue.
LLVM is far more powerful but more complex. C is simpler and already works everywhere.
I just need the compiler to run on my systems. Since it is written in rust, I either need it ported to my systems (too much work) or I need my rust code transpiled to C or C++ so I can build it myself (seems doable).
There are no shortcuts to a proper implementation of Rust. Implement an LLVM backend for your architecture. It doesn't make sense for Rust to maintain a backend that is doomed to be forever broken in various ways.
Rust won't have to maintain the backend. Who said they would? They might choose to if it becomes useful - it would give them things they dint have now (C++ interoperability, and new platform support for almost free), but no one here asked them to maintain anything.
I am also not sure what you mean by shortcuts or broken forever in various ways.
If out of tree, the backend will constantly break due to massive compiler churn. You need it in tree.
If Rust has plans on building a large amount of their internal optimizations on MIR, then I would think MIR would be fairly stable too, which makes this even less of an issue. It's going to be a toy project for now. If it works out, it'll be maintained against a specific version of the Rust compiler (stability is important in industry). And if that works out even more, perhaps Rust will accept patches to put it in-tree.
Of course, in-tree would be better. But it isn't a requirement.
But none of that is important at this point.
I'm not sure how it all works but I may take a closer look.
How would it handle my object file, shared library file, and executable file formats? Or my calling conventions?
So, how does that prevent you from using natively compiled languages like haskell?
Ignoring that, if I don't compile to C or C++, then I'm not sure how easy it will be to use. Are you suggesting Haskell can generate object files, shared objects, and binaries for my systems without compiling to C or C++ first? How would it get the file formats and calling conventions right? Can I specify them somehow to Haskell (doubtful)?
Of course, that's what ghc does by default.
>How would it get the file formats and calling conventions right?
The same way every other compiler does. What systems are we talking about?
ghc by default can generate object files for systems it doesn't know about? That sounds positively magical. How does it work?
> The same way every other compiler does.
Every other compiler that targets native code for these systems has first-hand knowledge of the file formats. If Haskell doesn't require that knowledge, I would love to hear about how it works.
Here's an example. Say Haskell only supported Windows on x86_64 and I'd like to build my Haskell program for Linux on PowerPC. Are you telling me Haskell on Windows could generate Linux on PowerPC executables and shared libraries without first-hand knowledge of the instruction set, calling conventions, and the ELF file format?
And I'm not being particularly evasive either. I gave you an example which fits my situation perfectly. I'd prefer not to reveal the names of the actual systems I need to build for, since it would fairly uniquely identify me.
Are you willing to work with the hypothetical I put forward? Or are you dodging that in favor of attacks on me personally?
I mean why would my first post be "Boy I wish I could use Haskell on my platform" if... I could use Haskell on my platform?
Now I could put a ton of effort into it... for each platform... but I don't have that kind of time to maintain Haskell for N platforms.
It would be much easier to compile via portable C or C++ that was human readable. But I would expect, since Haskell has lazy evaluation and a GC, that there is almost no way to map to readable C or C++. Thunks flying around, continuations, ...
Correct.
>It isn't supported and it doesn't work
What is "it"?
>I mean why would my first post be "Boy I wish I could use Haskell on my platform" if... I could use Haskell on my platform?
Because you incorrectly believe you can't.
>It would be much easier to compile via portable C or C++ that was human readable
Then do that, like I told you repeatedly. The fact that you have one of the developers of a haskell compiler trying to help you and you are entirely hostile and evasive makes it hard to take you seriously.
Why would you think that when I've said exactly the opposite? Who is being hostile and evasive here?
> What is "it"?
Also already mentioned above (file formats, instruction sets, calling conventions).
> Because you incorrectly believe you can't.
I wish that was true.
> Then do that, like I told you repeatedly.
If you show me how to generate portable C or C++ that's human readable from Haskell, I'd be happy to.
You keep claiming Haskell solves my problems, except every time we look closely at your recommendations, they are full of holes.
And I'm hostile and evasive.
Because you refuse to provide even the most basic information. Like what exactly you tried to use even.
>Also already mentioned above (file formats, instruction sets, calling conventions).
Mentioning is not helpful, telling is. So tell. What format? What instruction set? What compiler?
>You keep claiming Haskell solves my problems
I claimed nothing. I asked you a simple question to try to help you. And look at your responses. You still haven't answered a single basic question.
>And I'm hostile and evasive.
Yes. You refuse to answer even the most basic question that is needed to help you. You do not want help, you want to make a vague, baseless complaint and then avoid having that complaint proven baseless by refusing to provide any information.
That's all that matters to me. I'm not sure what crusade you are on... but if you had read the comments you replied to you would have avoided wasting my time and yours offering up Haskell as a solution when, based on information contained in the very posts you replied to, you should have been able to figure out that such a position was a misrepresentation.
Whether it was intentional or not I will refrain from speculating.
>Haskell isn't magical and can't compile to my platforms if it never heard of them, and it doesn't generate portable and human readable C
Haskell isn't a compiler. Different compilers support different platforms. You insist "haskell" has never heard of your platform, but I guarentee you can't find a platform that I can't compile haskell code on successfully. You are simply being dishonest. And yet again, yes, multiple haskell compilers produce human readable C code as output.
> I [...] tried to help you to use haskell.
No, you made claims like Haskell natively supported compiling to the targets I am using without even knowing what those targets are. That's not helpful to anyone.
> You have done nothing but deflect and derail
Also not true. I provided plenty of information and offered up an analogous situation. You decided to ignore that, and ignore the other things I said, and you continue to attack me instead.
> multiple haskell compilers produce human readable C code as output
If this was true all you had to do (8 posts up) was mention this. Why all the cloak and dagger? Why all of the ad hominem attacks and misrepresentation? Why not just offer up a simple solution that you know about?
I'd appreciate a list of which Haskell compilers produce portable human readable C code. I'd love to use them, if they exist.
Next time someone says "I'd love if Haskell could compile to human readable C or C++" and you know of an implementation of such, it'd be helpful to mention that, instead of tilting off on whatever personal crusade you imagined up against me.
I did not. I asked what they were. You still have not answered.
>That's not helpful to anyone
Neither is lying.
>I provided plenty of information
You have provided one piece of information, after many many posts: ghc. What OS and platform do you think it doesn't support?
You did. Do you need a link to it?
https://news.ycombinator.com/item?id=11707328
> You still have not answered
That's right, I refused to, over and over. The answers are irrelevant.
> You have provided one piece of information
That's right, and zero pieces of information were required for you to be helpful. You claimed you were trying to be, but your actions speak much louder than your words.
> Neither is lying.
You claim repeatedly that Haskell compiles to readable C. It doesn't, which is why you can't list any single implementation which does.
Thanks for trolling.
Good bye.
That does not say anything of the sort. It says there are haskell compilers that compile to native binaries and interoperate in a native code environment. You simply refuse to specify your platform so you can try to maintain the illusion that someone would still think you aren't full of it. You aren't even fooling yourself.
>That's right, I refused to, over and over. The answers are irrelevant.
Not if you actually wanted help they aren't. But since you actually just want to shitpost, all your refusal does is make it clear you are dishonest.
Do I have to read it for you too?
Me: "Are you suggesting Haskell can generate object files, shared objects, and binaries for my systems without compiling to C or C++ first?"
You: "Of course, that's what ghc does by default."
And now you claim you were not saying "anything of the sort" when I point out that you claimed "Haskell natively supported compiling to the targets I am using without even knowing what those targets are"?
> Not if you actually wanted help they aren't.
100% they are. Unless you can explain how me naming the systems I support has any relationship to whether or not some Haskell implementation has the ability to generate human-readable C code?
> But since you actually just want to shitpost
The only shitpost here is you backtracking your claims repeatedly.
Unless you listed those Haskell implementations you claimed existed somewhere and I missed it?
I'd appreciate a link if so.
If not: Good bye again.
You: "Of course, that's what ghc does by default."
And it does that. And that has nothing to do with the random nonsense about magic and knowing nothing about formats you made up afterwards. Your dishonesty is remarkable.
>Unless you can explain how me naming the systems I support has any relationship to whether or not some Haskell implementation has the ability to generate human-readable C code?
It has a pretty serious relationship to whether or not some haskell implementation has the ability to generation objects and libraries and binaries for it. You know, the thing you are still harping on dishonestly?
That's right. And as i just said, zero relationship to the question you keep failing to answer. Where are these multiple Haskell-to-readable-C implementations you tried to make up?
> And it does that.
So now you reversed your position again and you are claiming that Haskell can generate binaries for my system without knowing what my system is?
Keep it classy. I'll let you talk yourself in circles. You don't need me for that.
Start here:
https://news.ycombinator.com/item?id=11708601
(Let me know if I have to read the thread back to you again).
So which is it going to be next? You did or you didn't claim that Haskell supports my platforms?
Inquiring minds want to know!
Also, Rust is interoperable to the exact same extent C is. Making shared libraries exposing a C ABI is a snap.
The c code it generates is basically the ocaml bytecode interpreter and the bytecode of my program.
Not what I was thinking of.
Can you link to some information about that?
But on the platforms it compiles to, Ocaml can generate object files that look like they came from C. I use this feature right now to integrate a compiler in an otherwise C++ project. Here's a link to the FFI documentation: http://caml.inria.fr/pub/docs/manual-ocaml/intfc.html
But I run on Debian stable on x86-64. Not exactly an exotic platform.
Thanks.
So far as my experiments went, a 3000 native score Dhrystone test compiles to C at 1300, and that C compiles to 850 worth of JS. Which suggests that a direct-to-JS compiler might be worth the efforts...
Though on embedded systems you'd also want more constraints for the backend, like do not use dynamic memory, perhaps being able to specify the output code is MISRA compliant.
Another language that is widely used by hobbyists and in production is D which, as far as I know, produces C code too.
EDIT: I should say there are a few D compilers, DMD, LDC, and some other one by GNU(?).
I think that DMD is the main compiler, but I could be wrong. LDC generates LLVM code.
Edit: A cursory look seems to suggest they do translate to IR: https://github.com/ldc-developers/ldc/tree/master/ir