Hundred Year Mistakes
ericlippert.com
ericlippert.com
"I have no idea" he told me.
He said many programmers try to make their code as short and pretty as possible, as if there was some kind of character limit. Instead of falling into that trap, he just used parentheses. And the order of operations never was an issue.
(...Over the reals. Your mileage may vary in less exact types. The Surgeon General recommends avoiding division in production code. Regulations vary by state.)
(2×3)÷(4×5) = 6÷20 = 0.3
2×(3÷4)×5 = 10×0.75 = 7.5
((2×3)÷4)×5 = 7.5That's kind of one problem with it, it's less PEMDAS and more P, E, MD, AS. Multiplication and division don't have precedence over each other and neither does addition and subtraction. Both of those go from left to right.
*a.b
Is it (*a).b
Or *(a.b)
Your heuristic says it should be the former. Is it? What about this one: *a->b* Name operators (e.g., a::b, a.b, a->b)
* a[b], a(b) expressions
* Unary suffix operators (C doesn't have these, but Rust's ? operator applies)
* Unary prefix operators
* Arithmetic operators, following normal mathematical precedence rules (i.e., a + b / c is a + (b / c), not (a + b) / c). Note that I don't have any mental model of how <<, &, |, ^ compare to each other or the normal {,/,%}; {+,-} rank.
Comparison operators
* Short-circuit operators (&&, ||)
* Ternary operator (?:)--and this one is right-associative.
* Assignment operators
This list I think is fairly objective, although C and some of its children "erroneously" place bitwise {&,|,^} below comparison operators instead of above them. The difference between suffix and prefix unary operators is somewhat debatable, but it actually does make sense if you think of array access and function call expressions as unary suffix operators instead of binary operators.
That’s my conclusion too. Especially when I deal with several languages at the same time I don’t want to spend time on thinking about these details. Makes my code a little more verbose but I think it adds clarity.
Some languages may have boolean types, where others evaluate non-boolean types as true/false. So do you say "if myflag" or "if myflag=true" to make sure it's valid?
And some languages don't have short circuit operators, which I find annoying, but have learned to work around. Should I then write C or whatever that way, when I do have short circuit and/or?
The moment you catch yourself adding parentheses to other people's code to be able to understand it, and then having to "git checkout --patch" to flip it back the way it was.
Indeed, this is the pattern most industry programmers follow: never rely on operator precedence and always uses parentheses to disambiguate where things aren't glaringly obvious.
That makes your code both more robust and maintainable.
As a young egotist I would often omit parens in complicated C expressions. I did this intentionally and in a very self-satisfied way - writing multi-line conditionals and lining them up neatly without parens with a metaphorical flourish of my pen.
Then one day, chasing a hard-to-find bug, I realised it had happened because I'd mixed up the precedence of && and || in a long conditional. I was an idiot. Since then I've made a point of reminding myself that I know nothing and that there's nothing to be gained from pretending I do, and putting parens in everywhere.
Sometimes, even now, I get those grandiose moments when I think the code I'm writing cannot possibly go wrong. Those are the moments that call for a bit of fresh air and an extra unit test or two.
It is twice as hard to debug code as it is to write it. If you write code as clever as you can make it, you will never be able to debug it.
Last I checked this behavior was still on the roadmap for Zig. It would be a compile error to not clarify such expressions with parenthesis.
Here are a few that illustrate common imprecision in language.[0]
Conjunctive/Disjunctive Canon. And joins a conjunctive list, or a disjunctive list—but with negatives, plurals, and various specific wordings there are nuances.
Last-Antecedent Canon. A pronoun, relative pronoun, or demonstrative adjective generally refers to the nearest reasonable antecedent.
Series-Qualifier Canon. When there is a straightforward, parallel construction that involves all nouns or verbs in a series, a prepositive or postpositive modifier normally applies to the entire series.
Nearest-Reasonable-Referent Canon. When the syntax involves something other than a parallel series of nouns or verbs, a prepositive or postpositive modifier normally applies only to the nearest reasonable referent.
Proviso Canon. A proviso conditions the principal matter that it qualifies—almost always the matter immediately preceding.
General/Specific Canon. If there is a conflict between a general provision and a specific provision, the specific provision prevails (generalia specialibus non derogant).
——-
[0] https://www.law.uh.edu/faculty/adjunct/dstevenson/2018Spring...
You'll thank me when you're printf debugging.
it's never fun to track down a bug to a line that is doing 8 things.
With an answer like that he would no doubt flunk the modern interviewing process:
"We had a guy come in, tons of experience, aced all of our coding tests... but when we asked him about operator precedence in C, he just shrugged and said 'I have no idea'. So we had a to let him go, for his lack of strong CS fundamentals."
Said one 25 year-old SSE to another in the follow-up.
Yes, sometimes it's silly, but there is the harsh reality of how interviewing is executed, especially by those companies who want to base the assessment mostly on the judgement of peers as you said.
If you're really senior, in most cases, you can't expect that the majority of people in the company you're applying to is going to be as senior as you. You'll need to do a lot of convincing and explaining of things that might be obvious to you even after you join, on a daily basis. As with everything, the interview can be a good place to show you can do it.
Until the moment where you get flushed for some dumb random thing or another. Then the rest all goes down the drain.
Sometimes, even when you have the power to add parentheses to existing code and merge the commit, you still have to know what the unmodified code is doing: just so that you're sure your readability improvement is not changing the behavior, for one thing!
You can be looking up precedence tables all the time, or adding prints to test things empirically run-time.
Joke aside, I use parentheses liberally. Even if I know operator precedence it saves me from errors when I edit the code and another person reading it might not know precedence rules perfectly.
So, in other words, he didn't properly understand quite a bit of other people's code that does not follow full parenthesization.
Another possible moral is: The languages that stay in use for fifty years are the ones that avoid making breaking changes. :)
That said, the C# compiler team was and continues to be extremely concerned about breaking changes because we very clearly perceived the cost to customers and the barriers to upgrading entailed by breaking changes. I introduced a handful of deliberate breaking changes in my years on the C# design and compiler team, and every one was agonized over for many hours by members of the design team who were experts on the likely customer impacts.
> After getting myself snarled up with my first stab at Lex, I just did something simple with the pattern newline-tab. It worked, it stayed. And then a few weeks later I had a user population of about a dozen, most of them friends, and I didn't want to screw up my embedded base. The rest, sadly, is history.
<source>:5:11: warning: & has lower precedence than ==; == will be evaluated first [-Wparentheses]
int t = x & y == z; // ?
^~~~~~~~
<source>:5:11: note: place parentheses around the '==' expression to silence this warning
int t = x & y == z; // ?
^
( )
<source>:5:11: note: place parentheses around the & expression to evaluate it first
int t = x & y == z; // ?
^
( )
Seems like a mostly solved problem.Sure, there are plenty of ways to mitigate the problem. That's not the point. The point is that the problem should not have arisen in the first place to require ongoing mitigation fifty years later!
Seriously, the article was just a rant. Yeah, people 50 years ago did something wrong. So what.
Whenever a project like this is done, it’s very tempting from a business point of view to just leave the “legacy” users / data alone. But, technically, all your code continues to have to support two different models. Yes, with the right abstractions you can make this work without being terribly painful, but in practice there are always rough edges.
The time and effort it would take to migrate the old stuff into the new format is high, and the perceived business value is low. In many cases the migration never happens, and the old code is never removed.
But the problem is, you are now paying a tax to deal with that old code for all time. Every time a decision like this is made, the tax goes up. The tax slows you down, and it makes all future work you do more complicated.
So, there is a game theory problem here: for each individual decision, it arguably makes more sense to leave the legacy case alone and move on. But if you choose that path every time, it is a mistake, because the costs compound.
So far, the only thing I’ve seen that prevents this is very strong technical leadership that ensures the migration is baked into the project and made non-optional to the business teams. That’s really hard to do though, and arguably not always the right decision for the business, which may need to move fast now just to survive.
If a tool has a large community of usage, breaking compatibility affects a large number of users. If the tool doesn't have a large community, then a lack of stability means there's risk the tool won't in-future support current use cases, or that I'll have to spent much more time maintaining my use of it than with 'boring' tools.
It's fine for a tool in v0.x to have breaking changes. I think not supporting stable versions of a tool would make it hard to retain users in the long run, though.
In a crude type notation:
bool -> (... -> bool) -> bool
vs. int -> int -> int(There's a way using Category Theory to firm up that assertion into proper math, but I'm not good enough at CT to do that.)
:D
If you control the compiler, you can make the compiler diagnose all the cases which are affected by the precedence of &, emitting a file name and line number.
That was the thing to do with those several hundred kilobytes of code. Get the precedence right, and then warn about any code actually relying on the precedence. The compiler then "greps it out" accurately; even instances that are obfuscated by preprocessing.
Or, add the warning first before rolling out the change to the precedence:
foo.c:35: warning: obsolescent & precedence: use parentheses
Let the programmers go through a period of obsolescence whereby they rid their code of these warnings by using parentheses. Then, the compiler can be changed, while maintaining a diagnostic in that area of the language for some time. Eventually, when everyone has forgotten the old dialect with the bad precedence, the diagnostic can be removed.When the new precedence is rolled out, a compatibility switch can be included for compiling code with the obsolescent precedence. Actually, that option can even be rolled out first (so it initially does nothing).
When the option is eventually removed, the compiler will refuse to run if it is specified.
So anyway, Ritchie had a few fairly easy and cheap alternatives to sticking with bad precedence; he perhaps wasn't aware of them.
The boolean operations are the biggest gotcha, but these are everywhere. Similarly the "*" indirection operator binds more strongly than many programmers think it does (and the same symbol in type expressions too, which is why those parentheses in function pointer declarations that no one really understands are needed). The shift operators got picked up by C++ to do I/O of all things, where get to wreak havoc in a whole new regime.
Frankly the only precedence rules that most programmers understand (because we were taught them in grade school!) is the two-level/left-to-right grammar for the four basic arithmetic operations, along with a vague sense that these should bind more tightly than an operator with "=" in it, because that separates the "other side of the equation".
Given that, why don't new programming languages just require fully parenthesized expressions for expressions involving other operators?
That would be excellent and be in line with the trend towards “safe” languages. I have countless examples where I read code with several expressions and when I asked the dev if they really understood the situation rules of this language they usually didn’t. They just wrote something they thought should be ok.
Ok, that's understandable for C++, but for all the other languages I really wonder what the reasoning behind this (if any) was - "our language looks similar to C, so we have to copy as many of its warts as possible so C developers feel at home" ?
Sometimes it is better to keep with convention rather than change it even though it is considered incorrect.
No complaints as far as I know.
However, often (and I submit this very article as exhibit A), the logic is reversed: One observes a bizarre scenario, points at it, and goes: I don't know what or how, but clearly somebody somewhere messed up.
That's not true; sometimes seemingly simple things simply aren't simple, and you need to take into account all users of a feature and not just your more limited view.
Take operator precedence. __operator precedence is an intractable problem__.
Some languages try to make it real simple for you and say that all binary operators are resolved left to right, and have the same precedence level. This makes it easy to explain and makes it much simpler to treat anything as an operator; a feature many languages have. Smalltalk works like this. The smalltalk language spec fits on a business card; obviously, you must use such principles as C's operator precedence table would otherwise occupy most of your card!
But that does mean that `1 + 2 * 3` is 9 and not 7, and that is extremely weird in other contexts.
So, you're in a damned if you do and damned if you don't situation: There is no singular answer: No operator precedence rule is inherently sensible regardless of the context within which you are attempting to figure out how operator precedence works: It is an intractable problem.
One way out is to opt out entirely and make parentheses mandatory. If you then also make it practical and say that the same operator associates left to right regardless of what operator it is, you can still write, say, `1 + 2 + 3 + 4`, but you simply aren't allowed to write `1 + 2 * 3` – you are forced to write `1 + (2 * 3)` or `(1 + 2) * 3`. But now anybody who feels they can copy problems from math domains, or from technical specifications such as a crypto algorithm description now gets tripped up. You may disagree (I certainly do), but a significant chunk of programmers just prefer shorter code. I would certainly prefer my languages to work this way, and have configured my linters like this as well, but it's still not a universally superior solution regardless of point of view.
In other words, had Ritchie chosen to go with this 'when in doubt, do not compile it' strategy, this article could not be written. And we'd still have programmers unhappy with how the operator precedence rules work.
I disagree. FORTH got it right. Postfix notation is inherently sensible, as any RPN calculator user will tell you. It eliminates the need for both precedence rules and parentheses, unambiguously.
1 2 + 3 * is 9. 1 2 3 * + is 7.
It's infix notation that's the problem, precedence confusion is a consequence of it. (Prefix notation can work too, but requires more parentheses).
For instance, if I have five minutes to code a traveling salesman implementation, I'll probably do a greedy algorithm that just does the closest city next until I've visited all cities. If I have more time, I might check that paths don't cross, and if they do, swap the order of the cities in that closed loop. That sorta thing.
Is it perfect? Of course not. Is it better than throwing up my hands and saying the problem is intractable, therefore I should just return the list of cities in whatever order they were given to me? Absolutely.
In the case of bitwise operators, Boolean algebra teaches us that & behaves like \* and | behaves like +. Therefore & should have the precedence of \* and | should have the precedence of +. Is this perfect, or as good as RPN? No. Is it better than a world where 'if (x & 0xff == y & 0xff) [...]' almost certainly does the wrong thing? Absolutely.
If Ritchie had designed C to use postfix notation instead of prefix notation for expressions, it simply wouldn't have caught on. It would be as dead as FORTH. If he had designed it to have better precedence rules for bitwise operations, the world would be a better place.
See also null terminated strings vs length prefixed strings. Length prefixed strings are not perfect. They are merely a lot better.
At least, particularly poor choices of precedence can be diagnosed. Whenever the compiler applies a precedence rule between two operators that has been identified as awkward, it can emit a diagnostic mentioning those operators: "warning: quirky precedence: use parentheses when & is combined with ==".
So the mistake here seems to be expectation, which is arbitrary. Everyone knows the +-*/() precedence of basic arithmetic, but boolean operations and others? That isn't at the cultural level of basic schooling, so the author's contention that the precedence is wrong isn't a priori.
int x = 0, y = 1, z = 0;
int r = (x & y) == z; // 1
int s = x & (y == z); // 0
int t = x & y == z; // 0
None of that surprises me, but moreover I would never expect myself or anyone else to be sure about it, I would refer to the operator table to be check. The article also supposes that this line has some natural meaning that everyone expects, but I'm not seeing it: if(x() & a() == y)
I think the author is projecting their own cognitive patterns onto the rest of us without justification.I'm curious to know if you knew from the moment of your birth that you should look at the operator table in this case, or if you learned that mitigation on a particular day. If the latter, what might you have done before you learned that?
I'm glad you mentioned it though, because now I'm curious if there's a difference among programmers who score well or poorly on a cognitive reflection test. Maybe people with a tendency to suppress their intuitive response are also less likely to think that a line of code has an obvious meaning?
int x = 0, y = 1, z = 0;
int s = x & (y == z); // 0
Other languages like Go dont allow this: package main
func main() {
n1, n2, n3 := 10, 11, 12
n4 := n1 & (n2 == n3)
println(n4)
}
Result: invalid operation: n1 & (n2 == n3) (mismatched types int and bool)https://en.cppreference.com/w/c/keyword/_Bool
But I otherwise agree with you.
I will definitely try to remember that for next year!
You should never write something so complex?
I'd argue the practices that are most popularly known as 'good practices' are to avoid mistakes that are common.