C++ Grandmaster Certification
cppgm.org
cppgm.org
Most of the legwork in compilers like gcc involves optimizing the output code and targeting multiple architectures -- writing a dumb translator for a single architecture is a far more tractable problem.
Honestly, it's worded very poorly. "Compliant with the latest 2011 standard (C++11)" suggests all of this. I have a hard time believing this isn't some kind of joke - there's literally no-one alive that could write all of this in the timeframe of a course.
To quote my adviser, "writing a parser from scratch has no value except as a character building exercise."
Anyone who wanted to write their own parser should probably just use a packrat parser (I find it much simpler).
Whoever is organizing this course, frankly, is either way out of their depth or holds their 1985 compilers course in far too high esteem.
This whole "you should never write your own parser" thing is so often parroted out as "wisdom"... But I honestly believe that most of the people that say this don't realise how easy it is to "roll your own".
Skip the parser, learn the useful stuff.
Parsers, from a practical perspective, are child's play for people who seek to eventually master a compiler.
Of course you asked for examples... let me give that a try:
1) data from your favorite application that's been end-of-lifed and you're thinking of replacing with a competitor's tool 2) data from a later version of your favorite application that you want to use with an earlier version, because you don't want to upgrade 3) configuration information from some part of your IT infrastructure that you need to refer to as you restructure and upgrade 4) a big config file for some software, that contains an error somewhere and "grep" won't find it. Maybe it's a semantic error, for example. 5) a config file for some ancient crufty software you're replacing, but the config file is huge and contains a lot of institutional knowledge, so you want to automatically translate it to the new system's setup.
Being able to generate even simple parsers gives you a lot of power. It's not as uncommon as you might imagine.
In the security business, one is often asked to assess some not-very-well specified protocol, or some protocol for which there is no documentation. So to deal with it you 1) fuzz the hell out of it to make the end point fall over or 2) hexdump the protocol and write pieces of it in ruby or python to get messages through so that you can fuzz the hell out of it in a structured way.
And if there was some need to write a parser, you can bet it ain't gonna be LALR, it will be hand-crafted, likely recursive descent.
To reply to each of your points:
1) If you are lucky, this XML. I don't need to know how to write a parser if the data is XML. If it is some sort of Java serialization, dejad is your friend--no parser required. If it is binary, you are going to use the protocol reversing route mentioned above.
2) See #1
3) Maybe just insert parentheses around the whole bit of data, and insert more strategically, and you are all but done.
4) See #3 or #1.
5) See #4.
If I was working on a team, and I saw someone writing a parser for a data-related problem, I would seriously question what they are doing.
I do think you overestimate the work required to write a parser for simpler formats - for someone familiar with one of the popular parser generators this can be a handful of hours, and the quality of results should be much higher than an ad hoc method. This can be a good design decision.
There are simpler ways, as I point out above.
I think I just read the best potentially unintentional compiler pun I have ever seen.
But I think I see where you are going with this. Using a parser generator for compiling, in the words of Dave Conroy (author of MicroEmacs and many other things), "A Parser Generator makes the hard part harder and the easy part easier."
Edit: Wait--1985? Ah, that is the problem. I used the First dragon book, not the second. I only got the second after I did the compiler work. Was it much worse than the first?
First, note that basically the first half of the book is about parser implementation. Then, it basically just teaches you that there's only one parser and it's called LALR. Of course, they didn't include anything like packrat, but even worse, they pretend like a LALR generator is still better than LR, even though LR was extended to be tractable for large grammars years ago.
And even with all that, it's far too high level to be of use in actually engineering a compiler (which, by the way, is a good intro book).
LALR represents a nice tradeoff: both LR and packrat involve much much larger state tables than LALR parsers.
The ordering is also confusing because the grammar parsers go from less powerful to more powerful. However, LALR is after LR.
Edit: Oh, and for C++ you need a GLR, which isn't even covered in dragon.
Do you have an opinion about the Holub book? Also I have a collection of Davidson papers about code generation that I haven't looked at since back then.
Remember, a compiler is a translator from text to (often)text. You're going to be dealing with a lot of strings and probably allocations for your AST. You should think about using a language which doesn't make string manipulation and memory management like pulling teeth.
In general, I'd say the ML family is best for writing a general-purpose compiler. SML even gives your compiler a formal semantics for free.
I actually found it way too low level. And not meaty enough.
Also, if you do want to implement a traditional Yacc like parser generator then it is a pretty good resource for that as well (having done it). Finally, while writing a parser might be a "character building exercise" sometimes there is also no getting around it.
How can you say there are better references for lexing and parsing and leave out that the entire first half of the book is lexing and parsing?
There are a ton of books that are better at every single thing you would want. Engineering a Compiler. Modern Compiler Implementation in ML. The compiler Handbook.
A guess my point is this: I have learned a great deal from the Dragon book. I think it is a solid book that has taught me a lot. There may be better books out there but I haven't read one yet (Advanced Compiler Design is great but it really is only about optimization and analysis you need an undergrad book to supplement it).
Finally, I have encountered worse books on the subject of compilers. So yes, this is a book that I would recommend and continue to recommend.
ps. You mis-characterize the length of the lexing and parsing coverage. It starts on page 109 and ends on page 302, the content goes to 964. Chapter wise: 3-4 lexing->parsing, (chapter 1-2 are really an introduction and illustrative example so they don't count). Chapters 5-8 cover the rest of what you need to get a working compiler + some other stuff. Chapters 9-12 (page wise 583-964) cover optimization and analysis in depth. So really, nearly 40% is optimization while about 20% is syntax analysis. This book has a lot of good material most of it isn't to do with syntax analysis and the syntax analysis is for the most part high quality.
[1] http://drhanson.s3.amazonaws.com/storage/documents/compact.p...
Handwritten parsers have better error reporting and might even be faster.
If you glance in my profile you will see I'm well aware of the practices of production compilers and also how few people ever even look inside one.
Agreed. I can never believe the people who still praise it.
If you have a functional bend, Simon Peyton Jones' book (https://research.microsoft.com/en-us/um/people/simonpj/Paper...) is worth reading, too. His book, however, is not a complete treatment. It assumes you know e.g. how to write a parser, and concentrates on the challenges unique to lazy functional languages.
The first, you'll at least need something capable of parsing context-free languages. My recommendation is to start here[1].
It has still taken a number of the best C++ programmers and compiler designers years to accomplish, and they haven't finished yet.
The other part of the issue is building the internal representation so that halfway decent code can be generated, not even thinking about optimization.
And, like the syntax and grammar of the language, the semantics of C++ are quite complex.
Indeed, building Saturn V is nothing compared to flying men to the Moon and back. Does not mean you can build a Saturn V from scratch. And people who are promising to teach you how either are geniuses or just in denial.
Quote:
"The concern that has long been expressed by the FSF (which owns the copyrights on GCC) is that a general plugin mechanism would make it possible for companies to traffic in binary-only GCC modules. Rather than contribute a new analysis or optimization tool - or a new language - to the community, companies might have an incentive to distribute their work separately under a restrictive license. That runs very much counter to what the FSF is trying to accomplish, so opposition from that direction is not particularly surprising."
Whatever you think of RMS's stance on plugins, or gcc's plugin system, they have little relationship with the ease or otherwise of adding C++11 features to gcc.
The latter has much more to do with the difficulty of understanding the gcc C++11 front end, the difficulty of understanding the fine details of the C++11 standard sufficiently to implement it, and the amount of manpower available from people who can do both those things (or who have the time to learn).
gcc certainly does have a lot of historical baggage in its code base (though this is slowly improving with time), but given its rather complete support for C++11 (on par with clang certainly, and far ahead of MS's compiler and most other proprietary C++ compilers), they're not doing so bad...
I have a (basic) understand of clang, which is helped by the fact that there is a very clear, simple and DOCUMENTED boundary into LLVM, which I can ignore the other side of. The interface between gcc front ends and backends is none of clear, simple or documented.
> clang, which is helped by the fact that there is a very clear, simple and DOCUMENTED boundary into LLVM
(1) People adding C++11 features to clang are not going to be dealing with LLVM, they're going to be modifying and extending clang's existing C++ parser. So however nice the clang-LLVM front-end-middle-end interface is, that's not going to have much impact on this job. Rather, what's important is the quality of clang's internal algorithms and data-structures (and those in gcc's c++ front-end). If clang does better there (dunno), that's great for them, but it has nothing to do with RMS's plugin position.
(2) RMS is not against clean code, nice interfaces, good data structures, and good documentation. His concerns (whether you agree with them or not) are the degree to which interfaces are expressed in a way that circumvents the GPL. Good interfaces don't circumvent the GPL;
So it's perfectly fine to clean up and document gcc's data structures and interfaces (and indeed, this is already happening, and has been for a long time). RMS isn't going to stop you.
Chris Lattner, "The Design of LLVM" http://www.drdobbs.com/architecture-and-design/the-design-of...
(Seems possibly relevant to the boundary issue.)
I'm not sure how seriously to take that.
I don't see what this has to do with RMS or plugins.
Oh, and I'd like to say that knowing everything about C++11 makes you a language lawyer, not a Grandmaster. Grandmaster is more than knowing the language. It's about using it right.
The website is missing some information however: who's behind it? It says:
The CPPGM Foundation was formed by a software company that
recognized the value to programmer productivity that a good
knowledge of language mechanics had to new developers to
their team. The C++ Grandmaster Certification began
development as an internal training program, and the
foundation was founded to offer it publicly.
but fails to mention what that company is. Additionally, the domain is registered anonymously...Edit: After some digging, it seems that the only information available on the people behind this is the fact that they emailed the press release to the comp.lang.c++ newsgroup from a residential IP-address in Switzerland. (NNTP-Posting-Host header)
In this case, putting a flashy name will lessen the scheme, as it will attract people that will do it with purely career-oriented motives. Or, as another post noted, it might just be a scam.
:)
1.) There is not a single name of anyone involved in this endeavor.
2.) The sing-up confirmation is a simple alert box? It seems like an XHR request does go out but no email is sent to the email address provided. Also they don't even check for email uniqueness. That seems somewhat...strange.
If it is a joke, I will be pretty sad :-(.
The IP pointed to by the A record does have a service running on the SMTP port.
Hopefully nothing involving project planning.
(If it's a joke we might organize a study group ourselves^^)
However, there is something a little off about this proposal. First, the size of this effort is really quite substantial, even neglecting optimization. Secondly, the phrase The C++ Grandmaster Certification began development as an internal training program, and the foundation was founded to offer it publicly suggests some compiler-writing company heavily involved in the C++ space. How many of those are there really? I mean, it has been 20 years since anyone made any money producing C++ compilers. All for-profit companies do it as a side effect. VC++, for example, in the 90s had 50 people working on just the compiler itself, not counting the Visual part.
So it presents a secondary challenge, which is 1) is this a real company 2) what really is the end goal?
Edit: Also, the bootstrapping question is not well addressed. What do we have to start with? Regular C? can I do the first phase in Lisp or Arc or Factor?
Am I allowed to look at other source, like that of g++ or clang or llvm or objective C?
Finally, the apparent copyright terms seem at least unacceptable, if not downright goofy.
I can't find it now, but they published a paper that said "There is no such thing as C" in which they discuss the widely varying implementations and expectations of C, different enough that special flags had to be invented on their tool to get weird programs to pass.
...with an end date somewhere in 2023, it might be feasible.
As part of my GSOC project (auto-generating Common Lisp bindings for C++ libraries), I tried to implement parts of the Itanium C++ ABI (http://refspecs.linux-foundation.org/cxxabi-1.83.html#vtable). Like the language itself, the spec heaps complexity onto the compiler in a pointless effort to save the occasional load or arithmetic operation here and there. Yet ironically, it's faster and simpler to just have a simpler v-table layout and put a PIC in front of it: http://www.jot.fm/issues/issue_2009_01/article4.pdf (see page 233).
shivers
I don't disagree with anything you said here, but I admit I am feeling a pull.
Out of curiosity, what was the compiler for?
They were building an 8085 powered computer with their own OS. They hacked the PL/M compiler--actually rewrote it in Fortran to emit the appropriate object code and the rest of their tool chain. They started an effort to rewrite the compiler, but it failed. I came in, and being familiar with XPL (From the book "A Complier Generator"), having finished a port of it for the Xerox Sigma 5 at a previous gig. One other person on the team had finished a PhD in computational complexity and the other had project experience writing COBOL compilers.
The project stared with a desire to change the language a bit, so we designed a new language. It wasn't too different from PL/M as it had to be mechanically translatable. The language was named, embarrassingly, Syclops. We started out with batch jobs on a 360/35, writing it in PL/1. They shortly got a PDP-10, which I miss, and we commenced to rewrite the compiler in Bliss-36. So it was essentially a cross-compiler, spitting out 8085 machine code with a very fancy assembler listing, including timings for each basic block.
The project was wildly successful, and opened the door to switching processers. For a while, they considered the 6809, but ultimately stuck with the 8085. The CEO was very pleased, at least, to have the opportunity to consider the switch.
My favorite activity was taking bug reports from the programmers who always looked at the generated code. The would come in with the listing, saying that the code was wrong. I would go over the code, and the reaction from them always was "wait--that is odd, the code is right. Weird."
My boss at the time like to say "A compiler should produce code that an assembly-language programmer would be fired for writing."
EDIT - Looking a little more, I can't tell exactly but I feel like the people that put this site up should put some more explanation up about who is doing it and ensure it is on the up and up. Surely, they are aware of this HN post so I'm waiting. :)
That will show them!
[] (this bit of reasoning deserves more explanation than I've given...)
A certificate or title isn't worth anything unless the organization who emits it has the authority to emit them. And if you can truthfully claim to have written a fully-compliant C++11 compiler all by yourself... you're already so badass that throwing in an extra title or certificate won't make a difference.
I had a real-time operating systems course a few years ago in which we developed a simple real-time kernel running directly on an arm9 microcontroller board. The toolchain of cross-compiling with gcc and programming the device (and creating the required linker file for the architecture and bootup assembly) was fortunately done for us with source available, and we did a high-level overview of how openOCD works and a whirlwind tour near the end of the course of how we might compile a program independent of our RTOS kernel's program and load the separate program from SD card with the RTOS and run it. We talked about elf files and related topics but didn't go too deep--it'd be nice to revisit some of the topics and especially in the context of Linux. (I would prefer an arm device over x86 though.)
I went to a C++ standards committee meeting once. They were having this fascinatingly complex discussion about temporary object lifetimes; it was so amazing how everyone there understood C++ so thoroughly. I'm hoping that taking this course will give me a better appreciation for their art.
- For those criticizing the amount of work ... I tend to agree with you all. However, in the faq this is addressed, take that for what it's worth. However, it's free so it seems like, even if you fail at building a fully working compiler, you could learn a lot, so I say good for them!
I believe it would be possible for an expert C++ programmer (and this is clearly who they're targeting) to write an essentially compliant ("fully" sounds difficult, and I expect some fudge there) compiler in a year; two, perhaps, if it's a nights and weekends effort. This estimate comes straight out of my bum. It's an intuition. I'm not including the standard library: that would be a multi-year job even for a small team of truly excellent C++ programmers. But if we focus on the compiler, and allow a little fudge in compliancy, I think there are some individuals who could do it in a year. So I don't believe the scope of the task implies a joke.
The lack of information on the site raises worries. Do the organizers lack the confidence to reveal themselves? Or are they scamming? They shouldn't leave this a question.
Then I remember that this is the Internet, and if it's 'free' you're usually not looking at the product; but you can see it in the mirror.
i.e. is there a good book "C++ for C programmers who hate the thought of it"
http://yosefk.com/c++fqa/index.html isn't a book, but the author shares your disgust for C++ and it's fairly detailed. As he says somewhere, he knows C++ better than it deserves to be known. I also don't like C++ as much as C (and I'd like to replace ever having to use either of them with Rust, hence Rust is my language-of-the-year to learn for 2013), and I found the FQA immensely useful for understanding the craziness and defectiveness of C++.
I'd also recommend http://www.amazon.com/The-Standard-Library-Tutorial-Referenc... I own the first edition and found it useful.
What? One of the prerequisites is 2+ years experience with C++ (or similar language), and on completion one shall demonstrate, "a complete, exhaustive knowledge of the C++ language and C++ standard library".
OK, world class senior software engineers and a base point of 2 years experience generally do not go together.
Perhaps the author meant, having had at least 2 years C++ experience at one time in one's career.
In the process I learned about the joys of code reuse, surely what a true grandmaster would use.
When can I expect my certificate in the mail?