What is the performance impact of using int64_t on 32-bit systems?
stackoverflow.com
stackoverflow.com
Footnote: I have mis-remembered the system performance numbers, the SS10 was a 50Mhz machine. The R4000 (native 64 bit machine was only 100Mhz. That said, the CPU being 'starved' while it waited for the cache to fill was the proximate analysis of why the performance differed very little.
Lucene, the popular core search library, has always used int as internal document IDs. As a result, the number of documents in one Lucene index is limited to 2 billions (Java int are signed). It's even worse if there are many deleted docs in the index as they are are just marked as deleted and consume up IDs. There are a number of reasons why no one is keen on moving internal docIDs to long, and performance is one of them.
I suspect Java code on average does far more pointer-chasing and function calls than pure maths, so the difference wouldn't be very visible.
ISTR this was a problem for Linux kernels? There used to be several kernel counters that were stuck at 32 bits because the overhead of making them 64 bits was just too much. I think one example was network byte counters for NICs? I might be misremembering this, though.
https://stackoverflow.com/questions/5162673/how-to-read-two-...
With 32 bit integers you just execute XADD and move on with your life.
That's not entirely true. Under Linux, with the x32 ABI (see https://lwn.net/Articles/456731/) you have access to the entire register file, but still use 32-bit pointers.
In many ways, it combines the advantages of 32- and 64-bit mode code.
That said it can be useful for certain loads where you aren't consuming large amounts of memory but are doing lots of calculations.
[1] http://stackoverflow.com/questions/16841382/what-is-the-perf...
Ubuntu has some x32 packages available for their standard 64 bit distribution, but the selection is limited.
In theory, x32 should still be faster because the code gets to use all the other x64 features, like a larger register set and so on. I've no idea how big a difference that actually would make though.
The theoretical possibility for speedup is exactly why I asked. The x32 website has benchmarks where it's as fast as i386 and 40% faster than amd64 on pointers or as fast as amd64 and 40% faster than i386 on 64 bit math, but what's missing to really justify x32 is some "20% better than either" case.
Another fun tidbit -- the debugger is blissfully unaware that the CPU and compiler are conspiring to do this. If you debug such a program, it will work fine. If you set a breakpoint on the instruction in question and simply continue, the entire program freaks out because the debugger just chopped off the high 32 bits of whatever it touched. I found that amusing when hitting it while debugging an issue. (Sad face)
In general most tools freak out if they encounter this because they were built under the assumption this was impossible.
Also, I see now that the sentence "Our C++ library currently uses time_t for storing time values" might have been misleading; by "storing" I meant "keeping in RAM and registers, for calculations" rather than "keeping in non-volatile memory, for archival".
Most ARM chips are 32bit, so most phones are 32bit. There is little reason for mobile chips to be 64bit.