Integer math in JavaScript
james.darpinian.com
james.darpinian.com
This sounds a bit unnecessary given the capabilities of modern engines. I benchmarked some additions on my machine on nodejs and there is no difference between the naive and |0 implementations.
Yes - if you are taking 'after every math operation' literally. JavaScript engines (and other dynamic language engines - nothing special about JavaScript) track types at a finger-grained level than simple generic 'number' or not. They already know if a number is a guaranteed integer, and even in some cases which bits can be set in that number. Sometimes the compiler doesn't want to prove more advanced properties like range when for example a value escapes, due to the overhead of tracking this, and then it's useful to do the trick in this article.
Here's a worked example in a Ruby compiler for example https://chrisseaton.com/truffleruby/stamping-out-overflow-ch....
But the author knows infinitely more than me about JavaScript compilation - they're ex-Google and Meta, including working on both Chromium and WebKit. They could show their working by for example showing some compiler IR for the engines they know about.
It's true that JS engines already optimize code to use integer math extensively, without |0 annotations. There are situations where this is made difficult or impossible due to requirements of the language, and |0 can help in these situations.
There's a whole lot more to know about this topic for sure. My article would be a lot better with some examples where |0 gives a benefit and the corresponding IR. This post already took a lot more time to put together than I expected, and I didn't know if that many people would read it. Now that it hit HN, maybe I can justify the extra work.
I'm not saying it's useless, asm.js used this trick to great effect. I'm just saying that unless the situation is more complex and doing a bitwise operation actually encodes information (like the &0xff example in a sibling comment, from which the compiler can deduce "this is a uint8"), it's unlikely you're giving the compiler more information than it's already capable of deducing on its own.
It effectively does do that though, since earlier in your code `a` and `b` would be defined with `let a = ... |0` so now `a` and `a|0` are equivalent in any subsequent expressions.
By forcing your number into a 31-bit integer, you provide definite type info to work with.
Its possible until you do division or some other operation where the value needs to be promoted for language correctness. At that point the programmer must constrain the result to assure it stays in an integer.
In theory I guess one could be more selective about the |0 and apply it only after operations where the hint is needed.
On that note: why does |0 and >>>0 needlessly chop to 32 bits when they can represent more?
JavaScript has Int64?
I suppose the answer if you really care about performance is "just use WebAssembly" since that does have native 64 bit integers.
You don't even need a Double, you can use a Float NaN to encode strings:
> IEEE 754 encodes 32-bit floating points with 1 bit for the sign, 8 bits for the exponent, and 23 bits for the number part (mantissa). (...)
> 23 bits isn't much space, but conveniently the max size of a Unicode codepoint is 0x10FFFF, making it a 22-bit charset. When you encode your strings in UTF-32, you're only going to be using 22 of those bits, so masking the top portion with the NaN signature makes all of your characters NaNs.
In Python 2.7 there was a distinction between when Python would use the "long int" vs the traditional Int64. This distinction was eventually entirely dropped in favor of all ints being "long ints". In case you are unfamiliar python implement ints internally as a list of digits. This is now true no matter the number of digits in the int.
What interesting about this is that there actually are edge cases where performance in Python 2.7 will be superior to Python 3. There's a nice foot note about this in "High Performance Python".
It's slow, so nobody wants to use it and nobody uses it, so they don't see a need to optimize.
It's not 'Asm.js' optimisation code - it's normal optimisations that were there before Asm.js, and are still useful.
(I'm sure there are some exceptions and some thing were Asm.js-specific, but the basics are generic.)
Hilariously, I didn’t actually do the folding of x|0 to x because I didn’t know folks would do that. I just wanted to eliminate int overflow checks if the use didn’t care about the high bits. We only implemented the folding once asm.js became a thing.