Results of the Grand C++ Error Explosion Competition (2014)
tgceec.tumblr.com
tgceec.tumblr.com
It works far better than attempting to repair the AST into some plausible state.
It's analogous to the propagation of NaN values in floating point code.
"public intd Test() { ... typing new line }"
then you will not provide hints when writing that new line due to "intd" being invalid type?
Foo getFoo(int);
getFoo().???; // want code completion!
`getFoo()` is missing an argument, but you still want to complete Foo's members. If you type `getFoo().getBar()` you want go-to-definition to work on `getBar`.In clang, we use heuristics to preserve the type in high-confidence cases. (here overload resolution for `getFoo` failed, but there was only one candidate function). This means you can get cascading errors in those cases (but often that's a good thing, especially in an IDE - tradeoffs).
I also use it during type checking. There is a special "error" type. If an expression produces a type error, the yielded type of that expression becomes "error". Surrounding expressions that consume that type will see the error type and suppress any other type errors they might otherwise produce. That way, you only see the original type error.
https://github.com/dlang/dmd/blob/master/compiler/src/dmd/mt...
(Hi, Bob! I was actually just using your recursive shadowcasting algorithm as a reference yesterday.)
I feel like I just got an Erdos number or something.
We added a slightly-cursed version of this to clang. The goal was: include more broken code in the AST instead of dropping it on the floor, without adding noisy error cascades.
The problem is, adding a special case to all "surrounding expressions that consume that type" is literally thousands of places. It's often unclear exactly what to do, because "consume" means so many things in C++ (think overload resolution and argument-dependent lookup) and because certain type errors are used in metaprogramming (thanks, SFINAE). So this would cost a lot of complexity, and it's too late to redesign clang around it.
But C++ already has a mechanism to suppress typechecking! Inside a template, most analysis of code that depends on a template parameter is deferred until instantiation. The implementation of this is hugely complicated and expensive to maintain, but that cost is sunk. So we piggy-backed on this mechanism: clang's error type is `<dependent type>`. The type of this expression depends on how the programmer fixes their error :-)
And that's the story of how C gained dependent types (https://godbolt.org/z/szGdeGhrr), because why should C++ have all the fun?
(This leaves out a bunch of nuance, of course the truth is always more complicated)
In C++, I solved this problem by simply matching { } in the template body, and accumulating a list of tokens within the { }. Then, when instantiated, the template parameter values were known, and the template syntax could then be semantically analyzed. It was simple and effective.
But I was informed that C++ required the syntax parsing and semantics for non-dependent types without instantiation. I asked why, and the answer was "to check for errors without needing to instantiate it." I responded with "of what use is checking it if it is never used or tested?" And that was the end of that.
> The implementation of this is hugely complicated and expensive to maintain
I quietly revolted and refused to implement that disaster. AFAIK there was never a problem with deferring parsing/semantic until instantiation.
These days it supports both. (IIRC the default is legacy/nonstandard, you select the standard behavior with /fpermission-, and VS adds /fpermission- to newly generated projects)
https://devblogs.microsoft.com/cppblog/two-phase-name-lookup...
(I expect it's possible to construct cases where this difference is observable)
I think checking templates in isolation has value. We use statically typed languages in part to make more error classes locally-verifiable. But bolting that into a mostly-textual system is a mess.
(Checking templates in isolation is particularly valuable in IDEs, which tend to share logic with compiler frontends. IDEs only need that much power to do a passable job because the language is so complex, so I don't know which way this argument points)
[1] https://clang.llvm.org/doxygen/classclang_1_1RecoveryExpr.ht...
I laugh to hide my pain.
2>&1
as part of the redirect?(Probably this is an ideal LLM question though!)
I likely got that memorized from the time that I wrote a shell (/bin/sh) tutorial back in the day.
|& is your friend:
g++ foo.cc |& less
g++ foo.cpp 2>&1 | tee ./log
or, the shorthand of 2>&1 |: g++ foo.cpp |& tee ./logA non-problem if you use any IDE which has been able to group error under foldable menus (such things have existed on Linux for longer than some of the people you are teaching may have been on this planet)
Suppose you are given a task of adding some new functionality to an existing code base. You have been told that the guy who wrote it was “really smart” and that his code is of “enterprise quality”. You check out the code and open a random file in an editor. It appears on the screen. After just one microsecond of looking at the code you have lost your will to live and want nothing more than to beat your head against the table until you lose consciousness.
This entry could be that code.
Back that day, it was possible because parsers tried to repair bad input to try to keep going, in hopes of diagnosing as many real errors as possible, so as to reduce the number of iterations. Iterations used to be expensive: punching corrections onto cards, etc.
If the parser repairs bad syntax, it can cause more errors. If a repair involves insertion, there is a risk of getting into an infinite loop of diagnostics, even.
It's kind of anachronistic to have people playing this with C++.
Why is this 2014? Did that error explosion competition die out?
Maybe tumblr isn't where you find C++ people.
That's... several orders of magnitude larger than I'd have guessed.
clang 14 terminates preprocessing as soon as any header exceeds maximum nesting, so only ~6kB of errors are produced by it.
Interesting, for the 'Biggest error, category Bare Hands' entry, clang++ gives me a much longer error message than g++:
$ g++ var.c++ |& wc -c
8 810948 5670500
$ clang++ var.c++ |& wc -c
20 5232857 36625300
Clang++ has about 6.5 times as many characters in the error message as g++ here.Though looking at the actual error messages, that seems to be down to a deeper default maximum instantiation depth on clang++ (1024 there vs 900 in g++).
Random samples / "this repeated X times" etc would work fine. Having nothing makes the whole thing seem pointless
#include ".//.//.//.//jeh.cpp"
#include "jeh.cpp"Clang++ gives you a few lines of spam and one line of decent error message.
G++ still barfs reams and reams of errors.