What a pleasant surprise it will be for you when you find out that jq silently corrupts integers with more than 53 bits.
What a pleasant surprise it will be for you when you find out that jq silently corrupts integers with more than 53 bits.
But still, given that easy consumption from JavaScript is the ultimate primordial reason for choosing JSON over other formats, it seems like trying to transmit integers with more than 53 bits of precision over JSON is asking for trouble. Because it's only a matter of time until someone will want to do something like write a new service in Node, and the JavaScript parsers for other formats are at least somewhat more likely to guide people toward using BigInt for large integers.
[1]: https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe...
I explicitly asked for and achieved step 2. in the modified SerializeJSONProperty algorithm[1] so that users could decide and opt-in to serializing BigInts as strings if they so choose, with or without some sigil that could be interpreted by a reviver function. e.g.:
> JSON.stringify(BigInt(1))
TypeError: Do not know how to serialize a BigInt
...
> BigInt.prototype.toJSON = function() { return this.toString(); }
> JSON.stringify(BigInt(1))
'"1"'
[1]: https://tc39.es/proposal-bigint/#sec-serializejsonpropertyNo, it's typically IDs. And my observation is that inexperienced devs start out serializing 64bit IDs as ints and get burned first before they move to text.
There are several reasons why you end up with IDs > 2^53 even if 2^53 is good for enumerating ~100k things per earthling:
Many cloud providers will give you IDs that are specified to be 64 bit uints, and you'd better not corrupt some fraction of them. Generating unique integers strictly serially does not scale, so you might end up with something like Twitter's snowflake. Or you might want to tag additional info (reserve a few of the low bits for the type of id, or reserve particular ranges for particular types of acccounts etc etc).
I can promise you I have run into the issue in real life more than once and it's not just javascript and jq that like to silently corrupt ids by truncating them to double precision. For example pandas is really great at it as well.
What makes it especially fun that often only a very small fraction of IDs will be affected (< ~1 in 2000 if uniformly distributed).