Compiler Design in C
holub.com
holub.com
The book starts with parsing (I prefer PEG or Pratt parsers for their simplicity (and tool independence) to be honest and skipped some of that chapter) but then goes into semantic analysis, type checking, code generation, optimization passes, even mentions basic type inference.
There is some code online at http://www.cs.princeton.edu/~appel/modern/c/ .
It's a good if somewhat outdated book if you're interested mostly in parsing and lexing, but for all the claims it makes in the preface about being practical instead of theoretical and all the source code presented throughout, I found the lack of actual Asm code generation (or any mentions of this compiler being able to compile itself) disappointing.
Parser/lexer generators also seem to have fallen out of favour for the creation of actual compilers, both big and small, at least for C-like languages; techniques based on recursive-descent (RD) are quite popular now.
On the "big, production-quality" side, gcc used a generated parser but moved to a handwritten RD-based one, and Clang always used RD. EDG's front-end, used in Intel's and other commercial compilers, is also handwritten RD. On the small "toy compiler"/hobbyist/experimentation side, there's TCC/OTCC, CC500 ( https://news.ycombinator.com/item?id=8576068 ), C4 ( https://news.ycombinator.com/item?id=8558822 ), SubC ( http://www.t3x.org/subc/ ), and many others, all based on RD parsers.
In fact I can't think of any C compilers at the moment that are using a generated parser/lexer...
It's important to view this book in the context of its time. At that time, there were no books that showed you both the theory and complete running code for even this much. The compiler books (Dragon, 1st ed.; etc.) showed toy snippets but not a full lexer, parser, and code generator. There were only the articles in Dr. Dobb's about Small-C to show the way. So, at the time, Holub's book was a godsend. It married theory and implementation in a way that simply had not been done before. Hanson's and Fraser's 1995 book on a retargetable C compiler was a similar milestone, although it came out several years later.
After those two landmark books, explanations of full compiler implementations were no longer a rarity.
And I agree for all its thoroughly-written tutorial approach, I don't find myself going back to it very often.
This explains the SubC compiler I mentioned in the original post.
It was one of the first books about compiler design that I got hold of, back when Internet access was only available at the university.
It's interesting that most of the content is still just as relevant today. Coding styles change over time, and definitely the popularity of languages has changed, but the theory is still just as useful.
Off topic, but we had the 'International' edition of this book. Many publishers at the time did this, I'm not sure why. The explanation was that it was to make it more affordable for us, although the prices were still high. They were soft cover editions with plain covers (this book had a red cover with only the title and author, in white). The linked page is the first time I've seen the real cover. This is probably an early example of regional pricing, much like DVD region codes. They could charge more in USA and less in other territories I guess.
Really helped me get a deep understanding of C. Sadly, I haven't touched C for such a long that I've forgotten most of it. I highly recommended it for anyone who has a basic knowledge of C and wants to get deeper.
[1] http://www.amazon.com/C-Companion-Allen-I-Holub/dp/013109786...
There are several versions of the LLVM Kaleidoscope language tutorial where you build a compiler that supports numeric operations, it is often based on Pratt parsers (at least the official C++ version[1]), but there's a C version on github [2].
In my opinion, you could get pretty far by reading the "Compiler design in C" or "Modern compiler implementation in C" up to the point of optimization and then applying the code generation from kaleidoscope tutorials. LLVM takes care of many optimizations.
[1] http://llvm.org/docs/tutorial/LangImpl1.html [2] https://github.com/benbjohnson/llvm-c-kaleidoscope
If you want to go all in on SSA optimization, there is a book called Static Single Assignment Book and its written by a hole list of compiler writers. Its not finished but there is still a lot of information.
You can find it here: http://ssabook.gforge.inria.fr/latest/book.pdf
Or you can go with the classic, Advanced Compiler Design & Implementation. See here, http://www.amazon.com/Advanced-Compiler-Design-Implementatio...
All of them will teach you a lot about LLVM.
[1] http://wingolog.org/archives/2011/07/12/static-single-assign...
[2] http://wingolog.org/archives/2014/01/12/a-continuation-passi...
"There are advantages and disadvantages to an intermediate-language approach to compiler writing[...] Intermediate languages give you flexibility as well. A single lexical-analyzer/parser front end can be used to generate code for several different machines by providing different back ends that translate a common intermediate language to a machine-specific assembly language. Conversely, you can write several front ends that parse several different high-level languages, but which all output the same intermediate language. This way, compilers for several languages can share a single optimizer and back end."
Related article: http://www.globalnerdy.com/2007/09/14/reimagining-programmin...
[1] http://galvin.info/2007/03/13/history-of-the-operating-syste...
http://www.amazon.com/Forth-Programmers-Handbook-3rd-Edition...
I guess it at least has its uses if you want your compiler to be really fast.
On the other hand, my projects in this vein are C++ projects but they implement functional languages or Scheme. I have implemented one parser that creates a syntax tree from a simple equational language, and several interpretive backends, such as SKI-combinator-based graph reduction and a TIM-based interpreter. I do care about the speed of the compilers I use for production code, but not for this work. The Scheme compiler I wrote was written to run on my own Scheme interpreter which was written in C++.
The reason I think most hobby compilers I know of don't go the way I've gone is that it takes tons and tons of research to find out how functional languages are implemented, and the Appel books are a bit daunting, too.
What exactly is "algebraic rectification"?
While it is generally true that having a formal semantics aids greatly in analysis, it is worth noting that a very large amount of program analysis work is targetted towards C. (And mind you, flexible, high level languages bring with it their own troubles. Analysis in the presence of higher order functions is not a panacea at all)