Many programming languages have infix expressions, but get associativity wrong
jfm3-repl.blogspot.com
jfm3-repl.blogspot.com
Mathematically, studying associativity makes sense for mappings of type A x A -> A, where A is any set; < is not such a mapping.
The behaviour of "<" the author refers to is not associativity. Therefore, the languages he criticises do not really get associativity wrong, he just needs to brush up on his math.
Now you can say "well, ok this behaviour is not associativity but it is a well know behaviour of "<" and ">". Why can't the languages get it right?" I think the answer is that this (a<b<c<d) is one of those notations that look ok in algebra class but are really problematic if you actually try to define and use a logically consistent function notation, so it is not a problem for me, if some languages do not perform (a<b<c<d) as expected in order to keep their semantics clean.
Haskell is right. Sugaring in the conjunction like Python violates non-associativity, which exists for predicate relations in human languages. And it's not a type error; it's a syntax error.
Note that in mathematics one sees things that look like "1 < i < n", which seems similar to what Python allows. But, if you ask a mathematician to read this aloud, they will say "some i between one and n", or in certain contexts "all i between one and n". This is an entirely different thing; it is not a predicate at all.
[...]
So with <, it's non-associative and bad to sugar up the error you get if you try to use it associatively.
Somehow the author isn't clear on whether he wants "1 < 2 < 3" to return true or to be rejected; if it's the latter, then most languages get it right (except Python and Matlab, as far as I know).
I know nothing about Ruby so I don't know how its operator overloading works, but I think that could be a problem if < was redefined with an a -> a -> a signature, or a -> a -> b with b comparable to a, since in this case you wouldn't get the type error. The proper solution, according to the author, would probably be to simply refuse any expression with more than one < (or any other non-associative operator).
So I guess a fair compromise would be if the default implementation rejects "1 < 2 < 3" at either compile-time or run-time, whichever is possible.
Not really. As pointed out above, the author doesn't appear to understand that the associative/non-associative distinction is literally undefined unless we're talking about a binary operator from a set to itself. Saying that < is a non-associative operator is wrong, something like saying that a vector is a non-odd number; as long as we're complaining about parenthesization, I'd suggest the author rephrase his statement and say that < is a non-(associative operator), which at least means something.
An expression like a < b < c should be a type error - no possible parenthesization could make it a valid statement.
It's easy enough to come up with a way to make < work like this, though, by changing the signature of < so that it takes and returns a helper class instead of a numerical value, with automatic type coercions to that helper class from numerical values on both ends and booleans on the right edge. This should be easy in a language like Scala, probably a 10 minute exercise if you know what you're doing.
If we make this helper class hold both the left edges of the chain and the right edges, and use the appropriate edge for each comparison, then the thing is associative.
((a > b) > c) > d is the same as a > (b > (c > d)), since in either case, each element will at some point be checked with > against both of its neighbors, and only those neighbors, and the overall expression will be coerced to true if and only if the strict ordering is preserved at each check. Process left to right, and it works out. Even if we get messier and allow >, >=, <, <=, and == into the mix, we can mix and match everything, and the whole thing will only be true if every neighboring check is true. The author's worry in the blog comments about '1 < 2 > 1' evaluating to true strike me as odd: that should evaluate to true under any reasonable definition of those operators.
So even if we do make < into an operator such that it makes sense to talk about its associativity, it turns out to be associative. Which is why mathematicians don't bother parenthesizing their chained < expressions, there's just no ambiguity.
The article is plain wrong, IMO.
An example would be C++ stream << operator which takes a stream on the left, a string on the right, and returns a stream (after writing the string into the stream). This is so that you can do chaining:
cout << "Hello" << "World";
I don't think that there are any mathematical operators that work this way though.While we can parse the following expression properly:
a < b < c
parses to: (a < b) AND (b < c)
obviously most compilers don't do that. As others have pointed out, the < operator should not be associative.The misunderstanding, though, is that the original expression
a < b < c
actually could be parsed (I'm making a bunch of assumptions about the grammar and parser type here) by treating ..<..<.. as a ternary operator, like ?: in C++.Then, in the syntax tree, the compiler could re-organize the expression to be the correct an unambiguous:
(a < b) AND (b < c)I mention that because the standard argument against lisp notation is ....
Is the condescending tone necessary?
An alternative:
"Ruby turns this expression into the following pseudo-code"
Better pseudo code would be
1.lt(2).lt(3)
In fact, this is correct (albeit ungainly) ruby code 1.<(2) # true
and so 1.<(2).<(3) # Say wha ... ?
makes no sense.[0]: Yes, a cheap shot.
Those familiar with the mathematical notation would think:
1. '1 < 2 < 3' parses to '(1 < 2) and (2 < 3)'.
2. '3 - 4 - 5' parses to '(3 - 4) - 5'.
3. '3 ^ 4 ^ 5' parses to '3 ^ (4 ^ 5)'.
Ruby's convention is clever, but ultimately wrong, and this is one of the cases.And 1 < 2 evaluates to true. What else would it evaluate to?
That leaves (true < 3)
Which fails. This isn't a mistake by ruby. When you chain methods you have to think about what they return.
And the issue here isn't associativity.
I actually think that a trained mathematician would have no idea what to do with 3 ^ 4 ^ 5. It's a completely ambiguous statement and has no proper mathematical parsing.
As for associativity proper, I program in multiple languages all the time, but find remembering the various associativity rule differences in these languages a pain in the ass. So I always use parenthesis.
That way anyone reading my code (including myself a few months later) will find it easy to understand what I intended, and the right thing gets done no matter what the associativity rules are.
I think < is simply left associative: True < 4 is True.
Similarly, I think > is right associative. (50 > 30) > 10 does not hold, but 50 > (30 > 10) holds.
Comparisons can be chained arbitrarily, e.g., x < y <= z is equivalent to x < y and y <= z, except that y is evaluated only once.
randall@gulch:~$ python -c 'print 1 < 2 < 4'
True
randall@gulch:~$ python -c 'print 1 < 5 < 4'
False >>> True + True
2
Simple left-associativity would also make 3<4<2 evaluate to True. >>> perl -e 'print "jfm" + 3'
3
My guess is that it coerces "jfm" to an int (0), and then adds that to 3 to get 3.If Perl does anything wrong here, it is in not throwing a type error at runtime but just silently coercing the number. Compare with Python:
Python 2.6.4 ... linux2
Type "help", ...
>>> "30" + 2
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
TypeError: cannot concatenate 'str' and 'int' objects
(Using "use overload" may let you do further crazy things with +; I'm just talking about the default definition of the operator.)Incidentally, the string equivalent of + is ".", the dot. There's a parallel hierarchy of operators in Perl for numbers vs. strings, which is part of why Perl has so damned many operators. Woe betide the developer that says "==" instead of "eq" accidentally....
The fact that the string equivalents are all words (cmp, eq, ne, etc.) gives a useful mnemonic (strings usually contain 'words', so their operators are words), and also prevents ugliness from operator-proliferation - it's the symbol operators that usually make the code _look_ ugly and lead to claims of 'line noise'.
I've made the argument that the intended type of an operation is as much a context as the amount (void/scalar/list) context. The relevant Perl 5 literature flirts with that phrasing, but Perl 6 goes much further to extend that metaphor. That has proven very useful in discussions of consistency.
Except Java is statically typed, so it's possible for the compiler to know if f returns an int, and it can rearrange the expression given that knowledge.
int A = 3000000000;
int B = 3000000000;
int C = -3000000000;
int D;
D = A + B + C;
Depending on order of evaluation, this either throws an overflow exception or assigns the number 3 billion to D. There is a hard rule in the spec that defines what must happen here. In C or C++ there are no such exceptions, so the compiler can (and does) transform arithmetic expressions where it's advantageous. (this is obviously a trivial example which will compile down to a constant anyway)> The built-in integer operators do not indicate overflow or underflow in any way. The only numeric operators that can throw an exception are the integer divide operator / and the integer remainder operator %, which throw an ArithmeticException if the right-hand operand is zero.
Maybe it's true that subexpressions can't be reordered for other reasons, but offhand I can't find that restriction.