It's also worth noting that every finite double with magnitude larger than 2^52 has a precise integer value; it's just that once you get beyond 2^53, not every integer is representable.
It's also worth noting that every finite double with magnitude larger than 2^52 has a precise integer value; it's just that once you get beyond 2^53, not every integer is representable.
[edit]
Since the extra value is precisely a power of 2 (-2^52), then it will round correctly, however the value is arguably not precisely -2^52 since it has an epsilon of greater than 1.
To be precise: every integer, positive or negative, with magnitude less than 2^53+1, is exactly represented in double-precision. “Extra values in two’s complement” don’t (and couldn’t possibly) effect this at all, since it is a statement about abstract integers and floating-point numbers, neither of which depends on two’s complement representations.
In particular, -2^52 has a sign field of 1, an exponent field of 1023+52=1075, and an all-zero significand field. This number is exactly -2^52.
int fits_long(double value)
{
double max = (double) (1L << 52)
double min = (double) -max;
return min <= value && value <= max;
}
(assuming a 64bit long)I think this is a good solution without using the math library and having access to the FP status registers. Otherwise the lrint solution seems better.
int fits_long(double value)
{
double max = (double) (1L << 52)
double min = (double) -max;
if( min <= value && value <= max)
return( (double) (long) value == value);
return( 0);
}The value 2^80 can be exactly represented in double-precision as well, but it is also the most precise representation of 2^28 other integers, so it is ambiguous which integer is being represented.
[edit]
To clarify, the representation you suggest would also be the best possible representation of -2^52-1.
Analogously, `float x = 2.25; int a = x;` assigns the value two to a, but this does not imply that the integer two also represents 2.25. Two is just two.
Integers are inherently different because calculations with integers are naturally discrete, while floating-point calculations are a discrete approximation of the reals, which are not discrete.
There are two general purposes for floating-point numbers: scientific computing, where you start with an imprecise value and the precision multiplicatively accumulates (the original use). And as a hardware optimization for calculation of non-integer values (a common use-case today). In neither case does it make sense to treat a floating-point value as a precise number, which matches common advice to not compare floating-point values by equality.
But I also see that StehpenCannon has spoken :-); there is going to be very little chance he is wrong when it comes to arguing about floating point! He also notes the following, but not quite so explicitly...
The problem is poorly posed because the question is, which mantissas, exponents and sign bits fit into a long.
The algorithm is backwards because simple comparisons in integer space cannot compute this; but I think the algorithm should be,
int fits_long(double value) {
unsigned int trailingZeros = countOfTrailingZeros(mantissaOf(value));
bool fits = (bitsInAMantissa + 2 - trailingZeros + exponentOf(value) < bitsInALong);
return fits;
}
[Notes:if double is always > 0 and going to an unsigned long, the 2 above would be 1. The 2 represents the sign bit in the double. Given that sizeof(double) == sizeof(long) all bits in the two representations are accounted for, so there is no information loss.
checking:
mantissa = 0 (with a hidden msb of 1) means, trailingZeros = bitsInAMantissa -> fits will be true when the exponent value can be 0..62. So this represents each of the +/- 2^exponent values
mantissa = 1 x bitsInAMantissa means, trailingZeros = 0 -> fits will be true when the exponent value can be 0..(bitsInALong-BitsInAMantissa-2). So, +/- (all ones) * 2^exponent value.
The representation for going from double to long is not "smooth".
]
countOfTrailingZeros() is a favorite of bit twiddlers. "Hackers Delight" or "bithacks" https://graphics.stanford.edu/~seander/bithacks.html
[edited to improve readability]