Most of the legwork in compilers like gcc involves optimizing the output code and targeting multiple architectures -- writing a dumb translator for a single architecture is a far more tractable problem.
Most of the legwork in compilers like gcc involves optimizing the output code and targeting multiple architectures -- writing a dumb translator for a single architecture is a far more tractable problem.
Honestly, it's worded very poorly. "Compliant with the latest 2011 standard (C++11)" suggests all of this. I have a hard time believing this isn't some kind of joke - there's literally no-one alive that could write all of this in the timeframe of a course.
To quote my adviser, "writing a parser from scratch has no value except as a character building exercise."
Anyone who wanted to write their own parser should probably just use a packrat parser (I find it much simpler).
Whoever is organizing this course, frankly, is either way out of their depth or holds their 1985 compilers course in far too high esteem.
But I think I see where you are going with this. Using a parser generator for compiling, in the words of Dave Conroy (author of MicroEmacs and many other things), "A Parser Generator makes the hard part harder and the easy part easier."
Edit: Wait--1985? Ah, that is the problem. I used the First dragon book, not the second. I only got the second after I did the compiler work. Was it much worse than the first?
First, note that basically the first half of the book is about parser implementation. Then, it basically just teaches you that there's only one parser and it's called LALR. Of course, they didn't include anything like packrat, but even worse, they pretend like a LALR generator is still better than LR, even though LR was extended to be tractable for large grammars years ago.
And even with all that, it's far too high level to be of use in actually engineering a compiler (which, by the way, is a good intro book).
LALR represents a nice tradeoff: both LR and packrat involve much much larger state tables than LALR parsers.
The ordering is also confusing because the grammar parsers go from less powerful to more powerful. However, LALR is after LR.
Edit: Oh, and for C++ you need a GLR, which isn't even covered in dragon.
I actually found it way too low level. And not meaty enough.
Do you have an opinion about the Holub book? Also I have a collection of Davidson papers about code generation that I haven't looked at since back then.
Remember, a compiler is a translator from text to (often)text. You're going to be dealing with a lot of strings and probably allocations for your AST. You should think about using a language which doesn't make string manipulation and memory management like pulling teeth.
In general, I'd say the ML family is best for writing a general-purpose compiler. SML even gives your compiler a formal semantics for free.
Also, if you do want to implement a traditional Yacc like parser generator then it is a pretty good resource for that as well (having done it). Finally, while writing a parser might be a "character building exercise" sometimes there is also no getting around it.
How can you say there are better references for lexing and parsing and leave out that the entire first half of the book is lexing and parsing?
There are a ton of books that are better at every single thing you would want. Engineering a Compiler. Modern Compiler Implementation in ML. The compiler Handbook.
A guess my point is this: I have learned a great deal from the Dragon book. I think it is a solid book that has taught me a lot. There may be better books out there but I haven't read one yet (Advanced Compiler Design is great but it really is only about optimization and analysis you need an undergrad book to supplement it).
Finally, I have encountered worse books on the subject of compilers. So yes, this is a book that I would recommend and continue to recommend.
ps. You mis-characterize the length of the lexing and parsing coverage. It starts on page 109 and ends on page 302, the content goes to 964. Chapter wise: 3-4 lexing->parsing, (chapter 1-2 are really an introduction and illustrative example so they don't count). Chapters 5-8 cover the rest of what you need to get a working compiler + some other stuff. Chapters 9-12 (page wise 583-964) cover optimization and analysis in depth. So really, nearly 40% is optimization while about 20% is syntax analysis. This book has a lot of good material most of it isn't to do with syntax analysis and the syntax analysis is for the most part high quality.
Handwritten parsers have better error reporting and might even be faster.
If you glance in my profile you will see I'm well aware of the practices of production compilers and also how few people ever even look inside one.
This whole "you should never write your own parser" thing is so often parroted out as "wisdom"... But I honestly believe that most of the people that say this don't realise how easy it is to "roll your own".
Skip the parser, learn the useful stuff.
Of course you asked for examples... let me give that a try:
1) data from your favorite application that's been end-of-lifed and you're thinking of replacing with a competitor's tool 2) data from a later version of your favorite application that you want to use with an earlier version, because you don't want to upgrade 3) configuration information from some part of your IT infrastructure that you need to refer to as you restructure and upgrade 4) a big config file for some software, that contains an error somewhere and "grep" won't find it. Maybe it's a semantic error, for example. 5) a config file for some ancient crufty software you're replacing, but the config file is huge and contains a lot of institutional knowledge, so you want to automatically translate it to the new system's setup.
Being able to generate even simple parsers gives you a lot of power. It's not as uncommon as you might imagine.
In the security business, one is often asked to assess some not-very-well specified protocol, or some protocol for which there is no documentation. So to deal with it you 1) fuzz the hell out of it to make the end point fall over or 2) hexdump the protocol and write pieces of it in ruby or python to get messages through so that you can fuzz the hell out of it in a structured way.
And if there was some need to write a parser, you can bet it ain't gonna be LALR, it will be hand-crafted, likely recursive descent.
To reply to each of your points:
1) If you are lucky, this XML. I don't need to know how to write a parser if the data is XML. If it is some sort of Java serialization, dejad is your friend--no parser required. If it is binary, you are going to use the protocol reversing route mentioned above.
2) See #1
3) Maybe just insert parentheses around the whole bit of data, and insert more strategically, and you are all but done.
4) See #3 or #1.
5) See #4.
If I was working on a team, and I saw someone writing a parser for a data-related problem, I would seriously question what they are doing.
I do think you overestimate the work required to write a parser for simpler formats - for someone familiar with one of the popular parser generators this can be a handful of hours, and the quality of results should be much higher than an ad hoc method. This can be a good design decision.
There are simpler ways, as I point out above.
Parsers, from a practical perspective, are child's play for people who seek to eventually master a compiler.
[1] http://drhanson.s3.amazonaws.com/storage/documents/compact.p...
Agreed. I can never believe the people who still praise it.
If you have a functional bend, Simon Peyton Jones' book (https://research.microsoft.com/en-us/um/people/simonpj/Paper...) is worth reading, too. His book, however, is not a complete treatment. It assumes you know e.g. how to write a parser, and concentrates on the challenges unique to lazy functional languages.
I think I just read the best potentially unintentional compiler pun I have ever seen.
The first, you'll at least need something capable of parsing context-free languages. My recommendation is to start here[1].
It has still taken a number of the best C++ programmers and compiler designers years to accomplish, and they haven't finished yet.
The other part of the issue is building the internal representation so that halfway decent code can be generated, not even thinking about optimization.
And, like the syntax and grammar of the language, the semantics of C++ are quite complex.
Indeed, building Saturn V is nothing compared to flying men to the Moon and back. Does not mean you can build a Saturn V from scratch. And people who are promising to teach you how either are geniuses or just in denial.
Quote:
"The concern that has long been expressed by the FSF (which owns the copyrights on GCC) is that a general plugin mechanism would make it possible for companies to traffic in binary-only GCC modules. Rather than contribute a new analysis or optimization tool - or a new language - to the community, companies might have an incentive to distribute their work separately under a restrictive license. That runs very much counter to what the FSF is trying to accomplish, so opposition from that direction is not particularly surprising."
Whatever you think of RMS's stance on plugins, or gcc's plugin system, they have little relationship with the ease or otherwise of adding C++11 features to gcc.
The latter has much more to do with the difficulty of understanding the gcc C++11 front end, the difficulty of understanding the fine details of the C++11 standard sufficiently to implement it, and the amount of manpower available from people who can do both those things (or who have the time to learn).
gcc certainly does have a lot of historical baggage in its code base (though this is slowly improving with time), but given its rather complete support for C++11 (on par with clang certainly, and far ahead of MS's compiler and most other proprietary C++ compilers), they're not doing so bad...
I have a (basic) understand of clang, which is helped by the fact that there is a very clear, simple and DOCUMENTED boundary into LLVM, which I can ignore the other side of. The interface between gcc front ends and backends is none of clear, simple or documented.
I'm not sure how seriously to take that.
I don't see what this has to do with RMS or plugins.
> clang, which is helped by the fact that there is a very clear, simple and DOCUMENTED boundary into LLVM
(1) People adding C++11 features to clang are not going to be dealing with LLVM, they're going to be modifying and extending clang's existing C++ parser. So however nice the clang-LLVM front-end-middle-end interface is, that's not going to have much impact on this job. Rather, what's important is the quality of clang's internal algorithms and data-structures (and those in gcc's c++ front-end). If clang does better there (dunno), that's great for them, but it has nothing to do with RMS's plugin position.
(2) RMS is not against clean code, nice interfaces, good data structures, and good documentation. His concerns (whether you agree with them or not) are the degree to which interfaces are expressed in a way that circumvents the GPL. Good interfaces don't circumvent the GPL;
So it's perfectly fine to clean up and document gcc's data structures and interfaces (and indeed, this is already happening, and has been for a long time). RMS isn't going to stop you.
Chris Lattner, "The Design of LLVM" http://www.drdobbs.com/architecture-and-design/the-design-of...
(Seems possibly relevant to the boundary issue.)