My Reality Moment. Why did I ever agree to do this?
software.intel.com
software.intel.com
The only time I remember seeing this addressed outside of a "programmers are bad at estimating tasks so do x to get better estimates" context was a post by Jeff Atwood about the tendency of programmers to look at something like StackOverflow and see, basically, two CREATE TABLE statements[1]. The reality moment, then, sounds like the moment one sees past that "weekend worth of coding". A bit more fleshing out, maybe with an examination of why we think this way would be a great read.
This time I'm thinking I might be happier taking this job offer back to my current company and get a good pay rise, but I'm not 100% decided yet.
Or we could pick a non-insane language to standardize on.
Even better, we could pick a non-insane bytecode to standardize on.
> $ node
> "5" + "1"
> '51'
> "5" - "1"
> 4
How is this in anyway reasonable behavior? '+' is string concatenation but '-' is interpret these strings as numbers and perform arithmetic?
Does the principle of least surprise apply here?
That's precisely it.
You cannot subtract a string so it coerces "5" and "4" to integers instead of treating them as strings. So it results in 5 - 4.
The first concatenates two strings, "5" and "1" and results in "51".
It's perfectly reasonable behavior. If you want to add numbers you should give it integers - namely 5 and 1 instead of two strings "5" and "1".
The unreasonable behavior is giving integers as strings.
Maybe in Javascript you can't, but you sure could design a language where you could subtract strings, if you wanted to. You could treat the strings as two arrays of characters and subtract them. The result would be the characters that are in the first string but not in the second string. "Hello" - "o" would be "Hell".
I'm sure you could do this in Ruby by monkeypatching the string class to overload the minus method.
Also I was speaking in context of Javascript so I fail to see how this applies.
A result would be that, but I certainly would not call it the result.
Another result would be to treat the strings as byte arrays and subtract one byte from another, e.g. [a1,a2,a3] - [b1,b2,b3,b4] would result in [a1-b1, a2-b2, a3-b3, -b4]. I'm sure you could do any number of different "correct" operations. And then it's up to you, as a language designer, to decide what you think string-minus-string should actually do... and there is zero probability that everybody in the programming community will agree with your decision.
It's unreasonable for the computer to take integers given as strings and subtract them as integers anyway.
The former doesn't preclude the latter.
> The unreasonable behavior is giving integers as strings.
That I can agree with. But in that case this notion of reusing '+' as string concatenation but '-' is coercing strings to integers is utterly bizarre.
I'm afraid to ask what '*' does (string replication?) and '/' "space magic"?
I'm personally against type coercion - but I also understand how it functions (and why it functions). I disagree with the "why" but as long as I understand the "how" I can avoid gotchas like string coercion. :)
Multiplication ( * ) and division ( / ) coerce to integers. Only + is used for strings, particularly concatenation. Personally I'm a fan of & being used for string concatenation and can't think offhand why not to use it, other than slight possibility of confusion over &&. However I'm not aware of a language that uses it.
>>"5" * 5
25
>>"5" * "5"
25
>>"5" / 5
1
>>"5" / "5"
1As a strong typing fanatic, I'd rather have the interpreter/compiler scream at me. In my opinion automatic type coercion is slightly less wrong than undefined behavior in C.
(Aside, did you know that if GCC spots a statically provable null pointer deference, it will insert __builtin_trap() after the deference? I didn't know either until yesterday. This is acceptable because undefined behavior can mean "eat you hard drive".)
If for strings '+' is concatenation, and '*' is for replication with '-' and "/" as exception behavior I can fully understand. (Admittedly '+' over '&' is either style or hard-core "the operation is different, therefore we need a different symbol choice").
Your example worries me. Personally I would not design a language like that.
(Though I may have abused it in the past.)
The former more than the latter. The operation is different therefore we need a different symbol choice - and & also makes as much sense stylistically as + does.
"string1" and "string2", when read, makes sense. & is different from the conditional && like piping | is different from the conditional or ||
>Your example worries me. Personally I would not design a language like that.
Neither would I, but I like to think it was part of "made in 10 days" that allowed it to stick around.
Of course, in languages where & and | are already in use for bitwise operations on ints (etc.), using & for string concatenation would be a different form of the same clash as using + is.
In Python it does. But then, Python has strong types and won't coerce the strings to numbers for '-', what makes this sane.
I get what you're saying literally, I strongly disagree that it is reasonable, at all. If you can't subtract them, there should be an error at that point, telling the programmer it can't be done.
!!0 === false; !!"zipzap" === true; !!NaN === false
To cooerse to a integer (not float): ~~0.5 === 0; ~~"12zipzap" === 0; ~~(Math.PI) === 3;
I tend to avoid it for readability (for other developers)My impression is that in PHP-land it's very hit-or-miss. Some developers adhere to strict guidelines to avoid these kinds of problems and others don't.
Perl programmers generally avoid loose-typing when writing libraries and larger programs, but not while writing scripts for personal use, which I think is a pretty reasonable attitude to have.
I wonder if the Java library was comfortable doing that because of JS's quirks with types. I think it's still bonkers behavior though.
Since JS doesn't have any of those string operators, you end up with the fun example that was posted in the article. It will default to using + to concatenate the strings, but - is not defined to strings, so it converts them to numbers first and then applies the operator. Easy to remember once you know what's going on, but still irrational.
I don't code PHP anymore, but the use of "===" has been standard practice for ages, and checking for it fully automated.
It's one of the many areas in which uninformed people like to criticize PHP for bad practices that in reality have been left behind a long time ago and are only still supported by the language for backward compatibility reasons.
The point -- as is often the case -- is precisely stated at the end of the article: "What this boils down to is that when we approach any of these server dynamic languages to optimize, we can't employ the same tried-and-true techniques we have used for decades (literally) in C, C++ and Java. We need creativity mixed with discipline to make a difference here."
Its not making a point of bad or good. Its discussing the features of the languages that make optimizing them challenging.
In fact, JS doesn't seem to have anything that's specifically suited for async IO at all, apart from clumsy closures.
Also, string handling can often be a bottleneck. Anything dealing with HTTP the protocol, for instance, will have all sorts of crazy hacks to parse that messed-up text format as fast as possible. (And this is a major reason why HTTP/2 isn't text-based.)
I noticed what seemed like random numeric errors in a random subset of RPCs. It turned out, after what seemed like eons of investigation, that some clients were stringifying decimal integers with leading zeros, which made PHP interpret them as octal numbers, and hence the random seeming errors.
Having come from a strongly-typed-language background (C/C++/Java), that was a major learning experience about the caveats of weakly typed languages.
Your issue is with the too liberal number parser.
http://www.eecs.berkeley.edu/Pubs/TechRpts/1994/CSD-94-812.p...
The unsurprising strongly typed Python example was still dynamically typed.