Compiler Errors for Humans
elm-lang.org
elm-lang.org
Many years ago I wrote the front end of a compiler-like system (it was for formal specifications, not for runnable code) and dealt with some of these problems. Whenever a type problem was detected, the error was reported and the type of the failed object was changed to an internal error type. For any error in which an error type was involved, no message was generated. This avoided error cascading, something GCC inflicted on its users for decades.
For parse errors, display the line in error and mark correctly the item involved in the error. Don't just display the index into the source stream at the point the error was detected; that's often a token or two beyond the problem. Work back to the point at which things stopped making sense to the parser. You have to carry source position info with each token, but it's worth it.
For errors which represent an inconsistency between several parts of the source code, show all the places that conflict, not just one side of the conflict. Rust is good about this. They have to be, because the borrow checker reports inconsistencies between different code sections, not just declarations and uses.
C is horrible in this respect; all nesting looks the same. Because of that it took clang lots of effort to get good error recovery, for example to correctly report a missing semicolon at the end of a header file instead of reporting an error in the file including it (http://blog.llvm.org/2010/04/amazing-feats-of-clang-error-re... (NB: that page is 5 years old, and does not represent the current state of gcc))
gcc makes that job even harder by allowing nested function definitions. That means that accidentally forgetting a single '}', as in
void f( int i) {
if( i > 2) {
exit(EXIT_FAILURE);
}
void g() {}
[...]
void h() {}
will make the compiler think that g() and h() are nested function definitions. So, the error you get is a “missing closing brace" on the last line of your file. Without nested function support, the error would be reported on the 'void g() {}' line. Still incorrect, but potentially thousands of lines closer to the source.The Algol compiler (helped by Algol's syntax) I used years ago was way better. It frequently gave errors of the form
"x undeclared. Assumed real"
(if, say, you call sin(x) without declaring x) or "semicolon missing after 'end' (inserted)”
Compilation would still fail after such errors, but you often would get meaningful error messages for the entire program. That was quite a boon if the compilation is run in batch, but it still is a good idea nowadays.About a year ago there was a bug in the elm compiler that occasionally caused the types in a type mismatch error message to be swapped. For example the compiler told you:
Something weird is happening with this value:
x
Expected Type: number
Actual Type: String
but it actualy meant the opposite: Expected Type: String
Actual Type: number
Evan said it wasn't so easy to fix. You can see the issue here: https://github.com/elm-lang/elm-compiler/issues/216At the time I was trying to learn elm (and functional programming) and I didn't know it was a bug so it gave me a lot of troubles: I could never fully trust the compiler.
Then I left elm and focused on haskell but now I think I see how he solved the bug:
As I infer the type of values flowing through your program, I see a conflict
between these two types:
String
number
Anyway these new messages are very helpful, especially when compared to haskell error like: Could not deduce (b ~ c)
from the context (Num b, Num c)
bound by the type signature for
f :: (Num b, Num c) => [a] -> (b, c)
at restriction.hs:(4,1)-(5,32)
`b' is a rigid type variable bound by
the type signature for f :: (Num b, Num c) => [a] -> (b, c)
Great job.I think that's a good reason to prefer a language with local type inference, but no type inference across function boundaries. It saves you from too much boiler plate to type and has (IMHO) better developer ergonomics.
-no need to have a separate definition of generic types as in fn f<T>(a:T): fn f(a:$T) tells you already which type is generic.
> It is kind of shocking how much better things get when you focus on the user.
So true. We've seen this in Rust as well: every bit of time we spend working on better diagnostics makes Rust so much more wonderful to use. Even a small reduction in jargon helps new people out a ton. Diagnostics aren't particularly exciting, but they're really important and your users will love you for them.I'm in that camp. Both compilers function properly for what I need to do, but clang's output is a lot friendlier. The choice to switch was easy to make.
Chandler Carruth - Clang: Defending C++ from Murphy's Million Monkeys GoingNative 2012 - Day 2 - Clang https://youtu.be/NURiiQatBXA
http://cs.brown.edu/~sk/Publications/Papers/Published/mfk-mi... http://cs.brown.edu/~sk/Publications/Papers/Published/mfk-me...
In one particular case I recall racking my brain to figure out what I was doing wrong at work, and then I went home and rebuilt with the new error messages, and immediately saw the problem.
Highly recommend! :D
[0] We use Elm in production at http://noredink.com - and by the way, we're hiring!
Nevertheless the error message point is good. It's ridiculous how atrocious error messages can be. (And they probably seem even worse if you have to find the file+line yourself.) This is actually pretty straightforward to implement, and supported directly with flex+bison, so it's a shame that it's not more commonplace.
Perhaps in the early days of gcc it was seen as too much overhead? Anyway, should be de rigeur for anything being embarked upon today.
Can you clarify what your objection is?
A core dump may be the most effective way to debug a program that dumped core when crashing in a way that you can't reproduce, but that's because it's the only way to debug a program that dumped core when crashing in a way that you can't reproduce.
Sorry for the lack of clarity, my statement was incorrect as written.
With C coredumps, you get a snapshot of the entire system at the moment it crashed. You can literally walk through everything that wasn't corrupted beyond reading. If the system is, for example, a multiplayer game, you can literally navigate the entire gameworld like a frozen snapshot. This can be very helpful for finding well-hidden bugs.
We're professional developers. It is worth the tiny bit of learning curve in order to get more productivity out of our tools!
Also, thanks for taking a look at Elm! I think of it as a member of "the ML-family" of languages and I draw from a lot of lessons from working with these tools and seeing what issues have come up for other folks, both within the typed functional world and not.
I would be delighted if the same level of developer-friendliness were present in other pieces of our toolchain, most importantly in our version control systems. I routinely feel that while we have version control data structures suitable for the 2010s, the user interfaces on version control systems are about two decades behind where they should be.
I think I deserve a medal for understanding enough git to get by. Specifically, a purple heart.
Automated testing helps implement these things too.
Well done, so far! I hope this kind of effort is eventually undertaken by many compilers (and their authors), as everyone benefits from simplifying the debugging/refactoring process.
That said, "eyeball-grepping" a specific set of messages, along with which bits are important, is a learned skill. I have found myself many times trying to fix the wrong thing because I skimmed and wound up with a wrong understanding of what the error was. And then found a more careful read told me precisely what was needed.
First the code doesn't work like described and than the error messages don't help you to navigate around those problems.
Moral of the story: sometimes, explaining why you are doing something may require hyperlinks, just saying it's wrong is not convincing enough.
The error messages in XL and Tao3D are terrible. I'm ashamed :-)
It complicates parsing
Perhaps it is because compiler messages suck but my brain has been trained to quickly pick out the line numbers and file names.
Compiler is relatively simple piece of machinery, it isn't really smart. So maybe it shouldn't pretend it is? It doesn't have personality, it doesn't think anything, it doesn't actually even do anything: it is just a set of transformations from one set of bytes to another. In the worst case it has some heuristically-statistical optimizations, but it still is some process that can be accurately described in a relatively short amount of time. I don't want it to have an opinion and actually I know it doesn't have one. I know all has happened is that I made a mistake and I just want to know exactly what went wrong. I don't want to read compiler's essay about how it spent its holidays, because I know it cannot write essays and didn't have holidays anyway.
At the moment I don't use Elm, but if I did, I would be getting such messages tens, hundreds times a day. Every time I have to search for something one could call "the real problem" in the bigger text — I do work. I spend energy. Sometimes I can do so while being mentally exhausted or in a hurry. Or both. The more short, exact and concise message will be, the easier to understand the problem it will be and the greater my gratitude towards the author of this message will be. And it's not only about length of the message, human-like natural language constructs are inherently more complicated and diverse than short messages and labels — that's the reason why we invent labels and such after all.
All the good answers must be the direct answers to the question one might ask. When I'm about to read the error message of the compiler, I'm not thinking "Dear friend Compiler, what have you been doing in the meanwhile and what do you think about life?". No, I'm asking "What the hell did happen?". What did happen, and not what "did compiler do" or "how things work" or whatever.
One more problem is that Elm compiler might work just nice, but it won't always be right. I don't believe it and you probably shouldn't rely on it. The less creative it will be in its search for "what might have happened", the more likely it is I won't be distracted by its opinion. Even as skeptical and suspicious as I am, I am prone to get used to how some tool behaves. If it usually is right and tries to deceive me by behaving smarter and more human like than it is — one day I will believe it and spend way more time to figure out the problem than I would have if I treated is as a compiler and not my mate.
So, finally, what is the real purpose of the error message? I love artistic people, but unfortunately it isn't to show how thoughtful and creative the author of this piece of technology is. The real purpose is to help me find my mistake. There are 2 general types of mistake I can make: a simple one, like a typo, syntax error, forgotten type, variable definition, you know what I mean — and a complicated one, like unexpectedly finding a bug in the compiler itself, some weird Rust-lifetime problem, — It's hard, to make up an example, but the point is it requires some thinking and research. Help I need in the case of simple mistake is generally just as precise place of the mistake in the code as possible. Maybe some function signature with good description of parameters. Something I can easily grep (ack) or google (DuckDuckGo) in the case of more "internally-oriented" errors. I don't usually find "typo suggestions" useful, but ok, why not. The thing is 3 times out of 5 I will fix my mistake quicker than I do the reading (even in the languages with less verbose error messages), literally. I would thank my compiler for being as concise and straight-forward as possible to shorten the gap even more. I don't want to read what I don't need to.
In the case of "the complicated mistake" it is not likely compiler will actually guess what's the matter. I'll need stacktraces (if any), precise and unique exception names (error codes), maybe some additional info, but nothing that compiler can actually handle to discover itself.
So the best compromise I can imagine, which, I think, would help in both cases, would be short, concise, "machine-like" messages, concrete things like function signatures, "expected/found" diffs and links to the related section of documentation (web-based or local one — the latter would be even better) where everything would be described in the detail. And no playing a human.
Apple have already hired writers for Siri; how many more years before its personality starts to change to match the tone and manner that you ask questions? Young children have already started to mis-identify Siri as a real person! There was also the film Her the other year which seems quite prescient. Regardless of if we ever get to real AI, it seems quite clear we're going to face intelligent-seeming computers in the near future, and we're going to start to interact with them almost as if they were intelligent. It's going to be an interesting future.
This is a huge tangent, I know. It just seems like compilers speaking to you like a human being is just one more facet of this change.
I thought this blog post on conversational UIs really sums up some of these trends, and it only takes a bit of imagination to imagine where this will go in 10, 20 years: http://interconnected.org/home/2015/06/16/conversational_uis
> Software isn't "cautious" or "ambitious", those are qualities of alive beings. But maybe it serves us to think so.
Also this:
> Our minds respond to speech as if it were human, no matter what device it comes out of. Evolutionary theorists point out that, during the 200,000 years or so in which homo sapiens have been chatting with an “other,” the only other beings who could chat were also human; we didn’t need to differentiate the speech of humans and not-quite humans, and we still can’t do so without mental effort.
http://www.newrepublic.com/article/117242/siris-psychologica...
As I understand it, we use different regions of our brain for interacting with tools, people, animals, etc. I enjoy using the tool part of my brain when computing. It feels "quieter" and contemplative. It feels like I am doing something, not like I am asking an assistant to do the work on my behalf.
I like social interaction a lot too, but as an introvert, it's kind of draining. I would have to leave computing if using a computer felt like being in an intense conversation with a computer-person all day. Even if the conversation is pleasant and fulfilling, it would be too much social interaction.
I totally get that. Personally though, I see conversational UI as an additional parallel mode of interaction, that you can use simultaneously. I can imagine e.g. coding, and interacting and thinking analytically about that, while still verbally being able to have a conversation with my computer for e.g. responding to email or calendar appointments. Not for everybody, of course.
Edit: I suppose, what I really mean, is that I'd like to interact with human things (email, meetings, messages, etc.) in a human way (conversationally, as I would a person), and I'd like to interact with machine things (coding, simulations, etc.) in a machine way. Trying to interact with human things in a mechanical way seems unintuitive, really, except it's all current technology affords, so we all have to learn these machine interfaces instead. And maybe that's why so many (non-techy) people find computers needlessly complicated.
You'll only ever actually read "As I infer types of values flowing through your program…" a few times while getting used to the compiler, after that you'll recognize it by shape and never read it again, and at that point all the anthropomorphized text is doing is taking up screen real estate that could have been used to display more errors.
This is not at all why I prefer dynamically typed languages. Try doing some JSON parsing in Haskell for a non-trivial payload (like something with nested objects) without wanting to pull your hair out.
Counterexamples: Common Lisp is dynamically typed, OCaml and Haskell are statically typed. All three languages have both compilers and interpreters.
(Some crazy persons even wrote C interpreters, for what it's worth.)
https://github.com/rocurley/Nations-Graph/blob/master/Nation...
Even indented, it looks unreadable in elm docs, especially for beginner.