Jq is rounding 64-bit unsigned integers (2017)
github.com
github.com
"we know" is kind of ok, but the status of bug-hood is defined between coders and users, not solely by coders I think, and this breaks the POLA severely: People who depend on JQ don't expect this.
> Since software that implements IEEE 754-2008 binary64 (double precision) numbers [IEEE754] is generally available and widely used, good interoperability can be achieved by implementations that expect no more precision or range than these provide...
[0] https://datatracker.ietf.org/doc/html/rfc7159#section-6
Which means that if you are putting 64 bit integers into JSON and require every bit to be used, you are not actually creating a JSON which is compatible with all the consumers. For example, such JSON is not compatible with browser. Here is what my Firefox's JQ console says:
>> x = '{"id":675127116845989888,"id_str":"675127116845989888"}'
<- "{\"id\":675127116845989888,\"id_str\":\"675127116845989888\"}"
>> JSON.parse(x)
<- Object { id: 675127116845989900, id_str: "675127116845989888" }
I'd say that JQ acting the same way as browsers is pretty reasonable, no?I am more wrong than right: if you can point to a written spec saying "be not astonished" then POLA doesn't apply. And you did.
I wrote my own command line XPath JSON-query tool and used bigdecimals for JSON numbers. If you do not do much math, bigdecimals are probably even faster than using floats, and much easier to implement. Converting a string to a double float is extraordinary complex. I think it is one of the most difficult tasks in computing.
Unfortunately, now the W3C made a new XPath standard that requires JSON parsers to use double for numbers, so I changed my tool to use doubles as well. Now I am struggling with the string<->float conversion. The conversion in the standard library does not work properly. I just looked at another conversion library. 4000 lines of code, and after a 2 hour investigation, it turns out, it also does not work properly.
System jq:
$ jq --version
jq-1.6
$ echo '{"number":288230376151711744}' | jq '.number'
288230376151711740
Fresh compile from source according to the build instructions at https://github.com/stedolan/jq: $ ./configure --with-oniguruma=builtin && make -j8
$ ./jq --version
jq-1.6-137-gd18b2d0-dirty
$ echo '{"number":288230376151711744}' | ./jq '.number'
288230376151711744
Alternatively: $ ./configure --with-oniguruma=builtin --enable-decnum=no && make -j8
$ echo '{"number":288230376151711744}' | ./jq '.number'
288230376151711740
So the basic bug is fixed, jq has included a bignum library for > 2 years. I don't know if Mint (and thus presumably Ubuntu, and thus possibly Debian) includes an older version of jq or sets nonstandard user-unfriendly flags on purpose, but I'm somewhat underwhelmed in either case.Submitting this to let others that use jq beware of this :/
Though I don't see a way around it with a Serialization format like JSON that is meant to work across languages. I've done limited work with static languages, but from what I remember it would be a nightmare to have a variable that could be one of many types. I'm thinking overloading, or interfaces, or something. Could someone familiar with like Go or C#, or whatever explain how they would handle that?
Edit:
Writing that reminded me of something. If you are using a 64-bit unsigned integer, and you need to convert it to JSON for public use, please do not just give a object with a high and low value in hex. If you decide to do this anyway, but also ship an official Python SDK, just do the conversion in the SDK. I'm looking at you F5.
I should have never needed to write this function to read memory usage stats, but it was kind of fun figuring it out.
def ulong64_to_int(ulong64):
high = ulong64.get('high')
low = ulong64.get('low')
return int('{0:032b}{1:032b}'.format(high & 0xffffffff, low & 0xffffffff), 2)def ulong64_to_int(ulong64): return (int(ulong64['high']) << 32) + int(ulong64['low'])
Though the point I was trying to make is that if you're going to be sharing data, you should serialize it in a way that works for multiple languages, probably a string in this case. Or at the very least, if you're going to provide a client library for a language, you should make that library present the data in a way that makes sense for that language.
Visitor pattern. A really shitty, but workable, implementation of tagged sum types.
https://github.com/stedolan/jq/issues/1741#issuecomment-4306...
In fact, x86-64 has 256 bit data paths internally in some places.
https://superuser.com/questions/168114/how-much-memory-can-a...
There’s no reason to actually support 64 bit lines.