'9223372036854775807' == '9223372036854775808'
bugs.php.net
bugs.php.net
We've got gigabytes of RAM and terabytes of storage to work with; don't throw away characters in a string just because they don't fit in a compact data type used only because there is a passing resemblance of one to the other.
If I'm comparing two 53-digit barcodes, and they differ only by the last digit (checksum), then it's very important that comparing those two STRINGS comes up FALSE.
Then use === and do a type-strict comparison.
<?php
// Prints bool(true)
var_dump('9223372036854775807' == '9223372036854775808');
// Prints bool(false)
var_dump('9223372036854775807' ==='9223372036854775808');I'll stand by the axiom that throwing away data should NOT occur unless no other sensible option is available. If I'm comparing two literal strings, I shouldn't have to start with the obscure knowledge that a simple comparison will result in an aggressive attempt to perform two consecutive non-obvious type casts high risk of data loss.
I'm reminded of the great Belkin router fiasco: wireless routers were shipped with the "hold muh beer" great idea that random web page requests would be redirected to Belkin ad pages. I don't buy Belkin products any more (and that was years ago now) because knowing they would go there broke the trust that they wouldn't. Ditto here: if PHP is going to go to great lengths to try to throw away critical data (hey, I'm storing those numbers as strings BECAUSE I need all the digits), then I can't trust that the language won't do other similarly stupid things. I'm working in an industry where such a cavalier attitude to data can cost MILLIONS of $$$ over one failure, and can't afford to use a language where such failures are systemic. That there exists a workaround is inadequate. </tangent>
Fine. I could use ===.
The problem remains that a fundamental axiom of the language design is that casting lossless to lossy data types, without direction or warning, is considered acceptable. Ya know, if PHP wants to convert my numeric strings to integers for comparison, fine ... IF it maintains precision and preserves all the data. I shouldn't have to know of and use other operators/functions to explicitly avoid a pathological pursuit of forgetfulness.
Uh, all reasonable compilers warn about ambiguous use of = as a truth value.
$ g++ -Wall -c a.cc
a.cc: In function ‘int foo(int, int)’:
a.cc:2:12: warning: suggest parentheses around assignment used as truth value [-Wparentheses]
$ clang++ -Wall -c a.cc # output is colored
a.cc:2:9: warning: using the result of an assignment as a condition without parentheses [-Wparentheses]
if (x = y) return 0;
~~^~~
a.cc:2:9: note: place parentheses around the assignment to silence this warning
if (x = y) return 0;
^
( )
a.cc:2:9: note: use '==' to turn this assignment into an equality comparison
if (x = y) return 0;
^
==
1 warning generated.Every half-decent compiler will spit warnings at you, though.
> Now you want to add === to the mix?
It's been this way forever. PHP does all these conversions intentionally and trusts that you want them done. If you don't, go write C++ with FastCGI yourself. The difference between == and === is something you should pick up within your first few days of using PHP - what kind of industry has an economy measured above the "MILLIONS of $$$", but can't afford a dedicated PHP programmer?
I'm almost glad there isn't an official first-party public bug tracker, mailing list, etc. for Javascript - it would be ten times worse than this.
'=' vs '==' is not a syntax error. Consider "x=y=z" vs "x=y==z". And it's in somebody else's code. And they wrote it 2 months ago, but the programmer who's using it has only just started working with it. And they are super busy and don't have time to look at it. And it sort of looks like the problem is in the code you changed last week.
You can easily lose 4 hours over this stuff... have some imagination ;)
I addressed the "someone else wrote it two months ago" point above: If that happened, and this was in the code, it was dead code for two months, because it clearly couldn't have been running correctly. That's a process problem, not a syntax issue, and the appropriate fix is clearly not to modify the syntax of the language.
(Edit for ctdonath: good grief. 1.) the reply was to to3m's post, not yours. 2.) The "two months" thing comes straight out of his example, please read it. 3.) It was a JOKE, based on his chiding me for language. 4.) Why are you still flaming about this?)
I'd written the code the day before, and it was failing a pre-commit unit test. As I posted elsewhere, this kind of "forgot the second =" error can compile without warning, esp. within a complex evaluation. The process was running fine, as it caught the existence of the logic error early. That it took hours to find was a matter of tracing symptoms back to cause in an embedded system not easily debugged when running.
One could make a valid argument that this is a problem of language syntax, as everyone has been bit by the = vs == difference. As such, and in line with this thread OP, you'd think a new popular language would learn from that mistake and would not throw === into the mix as a solution to an even more obscure problem (casting a string to a float? really?).
And I was using GCC, which is happy to assign values amid a more complex logical evaluation. Try: if ( x = y && p ) return -1;
$ echo 'int foo(int x, int y, int p) { if(x=y&&p) return -1; return 0; }' > test.c
$ gcc -Wall -c test.c
test.c: In function ‘foo’:
test.c:1:1: warning: suggest parentheses around assignment used as truth value [-Wparentheses]
This warning has been there since at least gcc 2.x, I believe.I'm not making fun of you, really. But seriously: if you are spending 4 hours chasing bugs that can be trivially found by turning warnings on in your compiler, you have some process issues unrelated to C or PHP syntax.
(No, I don't have the exact statement in question handy. It was more complicated than this example.)
And BTW, things that work in most cases, but not all cases, are exactly where bugs come from.
If I didn't have such a low pain tolerance, I would get that tattooed on my body somewhere prominent.
What is shown here is just a rare edge case that you should not normally encounter.
But I think that == has some other behaviors which are really detrimental. Like 0 == 'hallo world'. Sadly those can't be fixed due to backwards compatibility.
This is what continually baffles me about PHP. There are plenty of bad languages out there, but PHP seems pretty much unique in having a community that actively resists any improvement and actively campaigns to keep the broken stuff around.
"...My 5cents on that (recently published issues):
<? if ('9223372036854775807' == '9223372036854775808') { echo "I can not count!\n"; } ?> (see https://bugs.php.net/bug.php?id=54547)
or
built-in PHP web server dies with a large Content-Length header value: The value of the Content-Length header is passed directly to a pemalloc() call in sapi/cli/php_cli_server.c on line 1538. The inline function defined within Zend/zend_alloc.h for malloc() will fail, and will terminate the process with the error message "Out of memory". (see https://bugs.php.net/bug.php?id=61461) Luckily we are getting Javascript ready to replace all PHP on the server sooner or later ;-)..."
If you compare a number with a string or the comparison
involves numerical strings, then each string is converted
to a number and the comparison performed numerically.I'm currently working with barcodes: numerical strings from 6 to 55 digits. In no way can I risk having one barcode be evaluated as equal to a literally different barcode just because the symbols in that string just happen to exhibit a passing resemblance to data of a different type.
Again, it's not just that it has loose typing. It's that it's taking what is OBVIOUSLY a string, converting it to an integer, THEN converting it to yet another data type which imposes data loss.
Intolerable for real-world use. A toy language. Alas, PHP, we hardly knew you...
ETA: Oh, I'd love to know the justification for the downvoting.
I was about to give an outraged reply that, if PHP is like Perl, then it doesn't scan the string afresh, just keeps a flag indicating whether or not it thinks a string is numeric. However, it turns out that's not true at all. `Perl_looks_like_number`, defined in `sv.c`, calls `Perl_grok_number`, defined beginning on l. 577 (as of v5.14.2) in `numeric.c`, which (after some book-keeping) does this:
if (s == send) {
return 0;
} else if (*s == '-') {
s++;
numtype = IS_NUMBER_NEG;
}
else if (*s == '+')
s++;
if (s == send)
return 0;
if (isDIGIT(*s)) {
UV value = *s - '0';
if (++s < send) {
int digit = *s - '0';
if (digit >= 0 && digit <= 9) {
value = value * 10 + digit;
if (++s < send) {
digit = *s - '0';
if (digit >= 0 && digit <= 9) {
value = value * 10 + digit;
if (++s < send) {
digit = *s - '0';
if (digit >= 0 && digit <= 9) {
value = value * 10 + digit;
and goes on and on and on and on in the same vein. Sheesh! (I didn't forget to close that last brace; the next line is de-dented, but that seems to be a mistake.)Seems like a silly thing to say. I wouldn't write avionics software with it, but there are a billion websites demonstrating that it's pretty decent for real-world use. At least as good as any other language, I'd guess.
Your lack of knowledge does not make it obscure.
So you could have predicted this yesterday? Just because it's codified somewhere, it doesn't make it clear, or anything other than a whim, or a product of circumstances, at best. That's not how languages should be defined, even if PHP clearly demonstrates that they can end up that way by chance.
Rasmus' lack of foresight does not make it reasonable.
In fact, this behavior has been documented explicitly for a year and a half: http://web.archive.org/web/20100808122711/http://www.php.net...
Earlier versions have said the comparison converts the numbers to integers though, which may be incorrect, and misleading if it was. Did PHP not convert float-like strings to floats in, eg, 2009? http://web.archive.org/web/20091024233139/http://www.php.net...
The point, however, is that you shouldn't pepper your language with operations which have consequences as hard to foresee as this with no good reason, and I really don't think that saving yourself some type conversions here and there would do.
The same kind of logic is used to make `false == ""` true. Or any 'falsy' language. If you want strictly typed behavior, yes, it's stupid to do that. If you don't, then it makes some things simpler, at the expense of more edge cases that are unlikely to happen - note that this bug was reported in 2011, and people are acting like it's a new thing. Because it comes up so rarely that, while it technically exists, many people never encounter it.
You are right, in a way. Sure, it may attract and retain more newcomers, but that's like saying that tobacco is "teenager friendly". I think it's not beginner friendly at all if you must have years of experience to avoid the innumerable pitfalls which PHP lays for you all over the place, learning, e.g. the range of Integers in PHP, which defines when a string will be either a float or an int, or that you should actually use strcmp.
In Python, Ruby, or heck, Haskell, you'd just have to do == and there would be no surprises.
This essentially breaks down to the top-down vs bottom-up education style debate. You can learn the gritty details and get caught up in minor details that may not matter in other languages, or learn how to do something, and get tripped up by the details in other languages. Similarly, we could teach kids abstract algebra, or basic +-*/ and then over-simplify when they try to divide by zero.
Neither is ideal, both have useful traits and problems, so we have to pick one. Or try to come up with something radically different.
edit: to ask it another way: if PHP is a massively-popular gateway drug to the world of programming, but it gives some people horrible flashbacks for the rest of their lives, do you want to make it illegal and close the door to a huge number of people?
Oh, and "everyone knows that PHP lights the upper-rightmost pixel in you screen purple and will crash if there's no screen" would not, in fact, justify such a thing.
You must hate programming then.
9223372036854775807.0 == 9223372036854775808This tripped me up when I was trying to compare two numbers, one of which was the result of a COUNT query via PDO. Of course, that COUNT result was a string.
I suppose if you worked entirely with strings it's alright. Or it wouldn't be so bad if you could make the reasonable assumption that functions returned appropriately typed data.
I'm sorry, I would like that to be true, but programmers rarely are half as smart as they think they are. We have many more years of PHP and its resulting insanity ahead of us.
Have you been out on the internet the past decade? Do you have a tendency to make extremities of things and trying to stand the needle on its tip?
Making extremes? My medical application would fail FDA approval in minutes if ported to PHP precisely because of this issue.
Btw, PHP's behavior doesn't totally make sense to me either. But I'm willing to assume that its users and designers have thought this through and it makes sense for PHP's intended use cases, because I don't know PHP.
Javascript (which most people on HN seem to like) also has similar issues (null vs. undefined, == and === etc). It got so bad that "Javascript: the good parts" had to be written to define a de-facto sane subset of the language. People are actually writing in Coffeescript (in part) to avoid Javascript's pitfalls.
YMMV.
There are actually some decent commercial PHP IDEs, believe it or not. I wouldn't be surprised if some of them are able to Warn on loose equality comparisons. I don't have much direct experience with them though.
By making the constant the expression's lvalue. But they don't have to do this anymore; gcc warns when you accidentally use = instead of == now.
Edit: I know == is not a string comparison. But you'd expect it to fail in a predictable way when passed strings that are not parseable as numbers, instead of trying to fall back on a string comparison so that people get the wrong idea.
This hamdriver you speak of: tell me more.
I generally avoid exceptions/error_levels in all languages but this is probably a good cause for them, in order to keep the rest backwards compatible.
In this case in php, the truncation happens due to loss of precision in the mantissa of the double precision float. But there are so many other ways to lose precision, I don't think it's reasonable to ask a language to attempt to account for them.
This is why languages should have clear rules about when type conversion occurs, and allow the user to prevent it when it isn't desirable.
edit: in fact, amusingly, php seems to be doing some non-standard stuff with its floats. I was going to make a point about how you can't determine if a double is a "correct" representation of a string decimal, but in mocking an example I discovered something odd. Check this out:
This is what one should expect:
$ ruby -e'puts "%5.25f" % 0.1'
0.1000000000000000055511151
$ perl -wle'printf "%5.25f\n", 0.1'
0.1000000000000000055511151
But in php:
$ php -r'printf("%5.25f\n", 0.1);'
0.1000000000000000000000000
$ php -r'printf("%5.25f\n", "0.1");'
0.1000000000000000000000000
Is php changing the type conversion? Or not using double precision at all?
[1] This doesn't really happen in PHP, but you have $php_errormsg that can be set without stopping execution (as happens with some errors/warnings when error_level is not set to E_STRICT, and below that depending on the error). This errors could be triggered in a new level, let's say "E_PEDANTIC".
We have two distinct problems here:
- strings converting to numbers without there being any number on any side. "Peculiar" of PHP but easy to circumvent using string comparison. IMO belongs in PHP4 but not at all in PHP5, which is an attempt at a "general-purpose" language. To be frank, I thought PHP4 made more sense because it was 1st of class at what it did, while PHP5 falls short to a number of languages in basically everything.
- automatic integer-to-float comparison to accomodate bigger integers. A horrible hack to squeeze a little extra performance in naive benchmarks in computers with no native 64 bit integer support. This really makes no sense whatsoever now and may have had some partial justification in the early 90s, prior to PHP4 even.
Both ideas are terrible and pretty much unique to PHP of all popular languages.
This is not a philosophical debate about typing styles or the existence of perfect type conversions. PHP's problems in this regard are relics from a dubious past.
No, this is not unique to php. Many popular, comparable languages perform an int -> float conversion. For example, Perl:
$ perl -wle'print "20938410923849012834092834" + 0 if "20938410923849012834092834" == "20938410923849012834092835"'
2.0938410923849e+25
- This is not a philosophical debate about typing styles or the existence of perfect type conversions. PHP's problems in this regard are relics from a dubious past.
Conversion from string -> number, and loose numeric types which auto-convert to float are near universal in loosely typed languages, out of necessity -- if such a scheme doesn't work consistently it can't be used at all. This brings me back to my point. You said "IMO this conversion should fail if the number represented is not valid, or fall back to arbitrary precision math". My response is that you cannot provide such a rule on the basis of "is it valid" because there is no such thing as a "valid" type conversion -- ALL have precision loss. It is inherent in the datatype. When I said "you may be underestimating the difficulty in predicting whether a particular decimal number can be accurately represented as a floating point type" you should perhaps read that as "you cannot do this, it is not possible".
Instead you might suggest that no loose conversion, no loose typing be permitted in a language design -- and I would agree wholeheartedly. But your suggestion that this be handled on a case-by-case basis depending on the numeric value is fundamentally unworkable. Big integers are not the only area this type of problem presents.
This, now, where it's being used, is absurd. There is no two ways to that. And this doesn't happen elsewhere to this extent.
I will leave you the last word though. Cheerio.
Your suggestion would make type conversion utterly unusable as it would fail seemingly randomly -- for example on simple numbers such as "0.1"
I am happy to see you agree with me.
This reminds me of Excel-like programs that by default, automatically detect (and convert) fields that appear to be dates/strings...often to catastrophic effect.
Prelude> "9223372036854775807" == "9223372036854775808"
False
Prelude> "9223372036854775807" == 9223372036854775808
<interactive>:1:25:
No instance for (Num [Char])
arising from the literal `9223372036854775808'
at <interactive>:1:25-43
Possible fix: add an instance declaration for (Num [Char])
In the second argument of `(==)', namely `9223372036854775808'
In the expression: "9223372036854775807" == 9223372036854775808
In the definition of `it':
it = "9223372036854775807" == 9223372036854775808
Prelude>
Yes. Strings are not numbers.(This is not strictly required, of course; you can write a typeclass that defines a two-paramater ==, instead of a -> a -> Bool, it could be a -> b -> Bool. But that's dumb, so nobody does.)
Prelude> 2 == 2.0 True
One of the worst bugs I've encountered years ago involved the conversion of Javascript int from string to number. Javascript's long integer has only 53 bits, while most other languages have 64-bit long int. When the backend language generated Javascript snippets (JSON) containing integers greater than 53 bits, the horror started at the frontend. Javascript happily truncated the int to 53 bits upon conversion from string to int. It was not a happy tale since those long integers were account numbers. The wrong accounts ended up getting updated, randomly at first appearance.
I can come up with half a dozen reasons to use something other than XML for data storage. I've yet to hear anyone give me a compelling reason to use something other than UTF-8 for encoding strings. Just because what I said is absurd when you replace UTF-8 with XML doesn't mean the original was absurd.
I don't have problem with UTF-8. I have problem with the silver bullet attitude advocating using an approach for all cases without thought. That's just intellectually lazy.
I'm not saying don't think about it. But once you think about it, I think there's really only one sane conclusion to reach.
This is easily solved by using the type-checking === operator, which exists for that purpose.
I hesitate to say that this is a feature, not a bug, but it is clear that this is documented behavior.
(Such that
strcmp('9223372036854775807', '9223372036854775808');
returns -1, meaning the strings are not equal.) php > var_dump('9223372036854775807' == '9223372036854775808');
bool(true)
php > var_dump('9223372036854775807' === '9223372036854775808');
bool(false)Here is what node.js says:
> "9223372036854775807" == "9223372036854775808"
false>>> "9223372036854775807" == "9223372036854775808" false
>>> 9223372036854775807 == "9223372036854775808" true
>>> 9223372036854775807 == 9223372036854775808 true
I believe the grandparent post is more referring to "general" use cases then this one. Personally I now default to strict comparison operators both in JS and PHP unless I explicitly want a loose comparison and end up missing most of these strange vagaries these days.
The one case where I sometimes use == is if I want to check for "null" or "undefined". Even then it scares me.
Its the fact PHP refuses to fix it that is the news.
This is what ghci (Haskell) says:
Prelude> 9223372036844775807 == 9223372036844775808
False
Prelude> 9223372036844775807.0 == 9223372036844775808.0
True
This is what Python says: >>> 9223372036844775807 == 9223372036844775808
False
>>> 9223372036844775807.0 == 9223372036844775808.0
True
Here is what SBCL (Common Lisp) says: * (= 9223372036844775807 9223372036844775808)
NIL
* (= 9223372036844775807.0 9223372036844775808.0)
T
Lua: > print(9223372036844775807 == 9223372036844775808)
true (!!!!!)
> print(9223372036844775807.0 == 9223372036844775808.0)
true
Javascript: alert(9223372036844775807 == 9223372036844775808)
true (!!!!!)
alert(9223372036844775807.0 == 9223372036844775808.0)
true
Other languages that will also do this[1]: Javascript, Lua. Languages that won't: anything with actual, honest-to-god integers, and not floats or doubles masquerading as them.[2] Languages that actually handle numbers sensibly: Lisp.[3] I'm not familiar with any others that actually treat rational numbers like rational numbers, but I expect there are some. (It's still, of course, impossible to treat real numbers like real numbers, meaning that this sort of thing will also happen there.)[1] Well, not the string-to-number bit, but whatever.
[2] Except for the niggle that they'll still do this when you're using floating point numbers, because this is what floating point numbers do.
But it seems to me - and let me stress that I am not a PHP developer and won't be bothered to install PHP on my machine at this time - that PHP is failing to exhibit exactly the behavior your code examples are giving.
Put it another way - type coercion 'run amok' being another thing entirely, you are correct that this bug stems from the fact that PHP is converting these integers to floating-point, and the standard floating point implementations will all behave in this exact way (thus, not a PHP bug.)
However, the issue here is that (again, "most?") languages also provide an easy way to get to arbitrary-precision arithmetic - and indeed, in the three examples you posted, you simply encode in the most natural way (by simply writing them) the two integers and they automatically compare correctly.
My understanding is that this is not the case in PHP, and that is a shame.
If you're going to run with PHP's axioms, then the specifications should (1) demand an unlimited-length integer type, and (2) NEVER convert a non-lossy to lossy data type without very good overt reason.
Fail.
#include <stdio.h>
#include <stdlib.h> /* strtol, strtod */
/* Convert string to long and return true if successful. */
int string_long(char *beg, long *num)
{
char *end;
*num = strtol(beg, &end, 10);
return *beg != '\0' && *end == '\0';
}
int main(void)
{
char *x_str = "9223372036854775807";
char *y_str = "9223372036854775808";
long x;
long y;
int x_ok = string_long(x_str, &x);
int y_ok = string_long(y_str, &y);
printf("x_str = %s\n", x_str);
printf("y_str = %s\n", y_str);
printf("x = %ld (ok=%d)\n", x, x_ok);
printf("y = %ld (ok=%d)\n", y, y_ok);
printf("x and y are %s\n", x == y ? "equal" : "not equal");
return 0;
}
And here's the output: x_str = 9223372036854775807
y_str = 9223372036854775808
x = 9223372036854775807 (ok=1)
y = 9223372036854775807 (ok=1)
x and y are equal
I compiled it like so: gcc -c -Wall -Werror -ansi -O3 -fPIC src/test_num.c -o obj/test_num.o #include <stdio.h>
int main(void)
{
long x = 9223372036854775807;
long y = 9223372036854775808;
printf("x = %ld\n", x);
printf("y = %ld\n", y);
printf("x and y are %s\n", x == y ? "equal" : "not equal");
return 0;
}
The error message is: gcc -c -Wall -Werror -ansi -O3 -fPIC src/test_num2.c -o obj/test_num2.o
src/test_num2.c: In function ‘main’:
src/test_num2.c:6:11: error: integer constant is so large that it is unsigned [-Werror]
src/test_num2.c:6:2: error: this decimal constant is unsigned only in ISO C90 [-Werror]
cc1: all warnings being treated as errorsIf you don't actually use PHP (as most of the commenters seem not to) don't comment on the bug, it has nothing to do with you and you are just making noise.
But the problem in my perception is that it is very error-prone. A less error-prone solution would be only convert one string value to number when the another is really a number. For example:
'9223372036854775807' == 9223372036854775808
That is, in fact, exactly how JS handles it.
9223372036854775807 == 9223372036854775808
true
9223372036854775807 === 9223372036854775808
true
> 9223372036854775807 == 9223372036854775808
true
> 9223372036854775807 === 9223372036854775808
true
However, that is not the two lines written about in this post. The following works as one would expect on JS, but (I assume, based on the report) not in PHP: > "9223372036854775807" == "9223372036854775808"
false
> "9223372036854775807" === "9223372036854775808"
falseDeleted comment
"Hey, this very precise literal has a passing resemblance to another data type, so let's convert it to that, but since it's got too many digits and resembles another data type at this point, let's just throw away some of the data, make the conversion, THEN perform the equality comparison." Yeah, that makes sense. FAIL.