IR is better than assembly (2013)
idea.popcount.org
idea.popcount.org
Very noble goal, but I can imagine that it can take a lot more time to do that than just writing a bunch of assembler instructions.
Perhaps there could be some intermediate approach, where LLVM can learn from a IR/assembly pair and improve itself (?)
They generally seem pretty willing if it's simple, and if it isn't then it'd probably be a pretty involved (but interesting!) side project for you anyway.
Oof, Reflections on Trusting Trust just got more interesting...
The biggest reasons to drop to assembly is because there are high level constructs that the compiler is very unlikely to recognize and optimize effectively. Things like AES-NI, hardware RNGs, and similar.
http://stackoverflow.com/questions/6981810/translation-of-ma...
Obviously not cross os, but might be good for bare metal stuff. I've gotten libraries in the past compiled with weird ABIs. This sounds really neat.
I'm not sure I'd go hand writing IR code, though. It's pretty easy to just write C code with vector extensions, etc to produce the IR I'd be after. When I do need to write assembler code these days, it's typically to get access to some privileged instructions in kernel space. Most other instructions are available in C code via __builtin_foo_bar functions.
You would not go writing entire apps with it anyway, just a few inner loops or so.
I'd still use assembly-looking C with extensions and intrinsics for that, though.
Doesn't AS/400 use an IR approach as well? Which let IBM seamlessly migrate the underlying CPU a few(?) times now?
The purpose of every IR is to remove the ambiguities and language complexities of programs. By simplifying programs into series of statements such as "%3 = op $type %1, %2", generic optimisers can be built easily. Certain language specific optimizations can be written for the frontend of the compiler as they have knowledge of the language being compiled. Generic LLVM-IR may not be optimised to deal with issues such as devirtualization in C++ (though there is work being done in that area).
LLVM's IR undergoes fairly occurrent changes to better handle "new" problems.
Correct. Libfirm[0] is the only compiler I am aware of that attempts to use a "firm" IR.
Sorry if the tone of "only compiler I am aware of" came of as snotty it was meant as an expression of my naivety on the subject.
Yes it uses an IR approach. MI code is essentially a byte code to which programs are compiled and then in turn the OS compiles them to the underlying machine code. AS/400 (or IBM i as it is officially now called) did use this to help in the switch from a proprietary/undocumented CISC CPU architecture (apparently similar but not identical to the IBM mainframe instruction set) to PowerPC/POWER. However, binaries are stored with both MI code and machine code together, and there is an option "Delete Program Observability" which removes the MI code section, which makes that migration strategy impossible – but, in practice, many people didn't choose that option, and if you did, so long as you or the vendor still have the source code it is just a recompile to fix it.
AS/400 has two program models – OPM (Original Program Model) and ILE (Integrated Language Environment). Basically, the original object code format, runtime library, etc, were designed for use with RPG/COBOL/PL/I, and they didn't work very well with languages such as C and C++, so ILE was created to remedy that deficiency. The relevance to IR, is that OPM and ILE actually use two different MI formats - Old MI (OMI) for OPM and New MI (NMI) for ILE. (NMI is also called "W-code".) IBM has publically documented most of the details of OMI and provides a public API to convert OMI code to machine code. By contrast, my understanding is they've chosen to keep the details of NMI confidential, and there are no public APIs to convert it to machine code.
In these systems both something like IR and final assembly is stored in the binary. The OS recompiles the IR for the current CPU if necessary and replaces the assembly in the binary. That way there is no compilation overhead unless architectures are changed.
It also wouldn't tie me to any particular library - I think the only actively maintained one is the C++ one.
The bitcode format is rather complicated, but at least there is backwards compatibility.
LLVM is the library, so you are bound to it anyway.
Further by having code for emiting IR you are basically just copying functionality that already exists.
It's not guaranteed to be stable, but it's not "highly unstable" either. Not too many breaking changes have been introduced in the past few years and it's unlikely that you'd hit those parts when hand writing it.
Not that I'd recommend hand writing LLVM IR.
It's a shame that we don't have an actual portable IR, which would not be tied to toolchain version or contain target specifics. LLVM IR can't be used as such for a portable IR, which is why we've seen efforts like SPIR-V (Vulkan GPU IR) and WebAssembly (IR for browsers), both of which are very similar to LLVM IR (and a lot of work was duplicated).
All of these things may change in the future (apart from the reducible control flow restriction), but, at the moment, the goal of WebAssembly is just to have a viable compilation target for C++ in the browser. Google was championing LLVM IR for this purpose in the form of pnacl, but that was not enough of a compromise to work on the web platform. :(
Most projects that aren't written in C++ thenselves use the C api. It has fewer features than the full C++ api, but is more stable.
IR is short for "Internal Representation". Most complex compilers have at least one level of IR, usually multiple ones that are progressively lower level.
The point is that an IR carries more information than machine code, and so potentially allows more specific optimizations.
This should at least be mentioned in the article.
Rust and Swift, for example, both use LLVM, but have their own intermediate levels of IR ('Mir' for Rust).
LLVM IR is already stripped of a lot of information that might be important for certain higher level optimizations. For example, numbers are all unsigned, and there are different operations for signed and unsigned arithmatic.
Many if not most compilers have one or more intermediate representations, but most of them are not as rigorously specified as LLVM's.
If you're starting from a clean slate, what's the benefit of writing IR? Why not use C? After all, IR won't really give you complete control over generated code, and it's still an abstract VM (albeit that obviously allows writing IR that will only sensibly compile on a specific arch - e.g. system register acceses and so on).
Working in C means you get restricted to only doing things which C can do, and you're out of luck if you want to do things that C can't do: unaligned accesses, tail calls, saturating arithmetic, overflow detection, exceptions, stuff like that. IR allows you to do all this, at the expense of being considerably more complex and painful to use.
Of course, if your compiler don't want to do any of this, then C is a perfectly valid choice, as it doesn't tie you to the LLVM toolchain --- see Nim, for example. But as soon as you venture outside C's comfort zone, working in IR starts paying off.
So again, I see all the value in IR for compilers. Can someone give me at least a reasonable use case for beginning to write in IR, or even using IR for optimizing cerain code paths?
http://lists.llvm.org/pipermail/llvm-dev/2011-October/043724...
Any LLVM experts have thoughts on this or my original goal within the context of LLVM's current situation?
In many cases it's just drop-in replacement for gcc.
Pretty much the only way is to just read the C api header and figure things out from there.
It does contain comments, which are mostly helpful, but sometimes incomplete or out of date.
Use whatever you want in production, I don't care. But don't discourage people from learning assembly. It's a worthwhile task.
Some CPU have specific idioms that are not only hard to translate but requires to be used fluently. Like natural language.
Btw, I never uses any software relying on a name of a myth that was a pure failure such as Babel or death star. It makes me feel like people intend to fail.
However I am still looking for a use case to write IR directly, or in place of bits of inline assembly.