Exploring the native use of 64-bit posit arithmetic in scientific computing
arxiv.org
arxiv.org
Also, please note that all traditional algorithms are wary of the disasters of overflow to infinity and underflow to zero, so they tend to manage the magnitudes of numbers to prevent that. Posits take advantage of that by decreasing relative error when the exponent scaling is not extreme. Standard 64-bit posits (2 exponent bits) have 60-bit significands, versus 53-bit significands for IEEE floats, for values between 1/16 and 16 in magnitude. And floats do not have anything like the quire, since an exact dot product accumulator for 64-bit IEEE floats has to be something like 4,664 bits wide (an ugly number) and has no provisions for infinities and NaN values.
[1] https://posithub.org/docs/posit_standard-2.pdf [2] https://posithub.org/docs/Posits4.pdf [4] https://cse512-19s.github.io/FP-Well-Rounded/
https://www.sigarch.org/posit-a-potential-replacement-for-ie...
Original paper:
[1] http://www.johngustafson.net/pdfs/BeatingFloatingPoint.pdf
As far as explaining posits, the key point I'd say to people is that a posit is almost the same as a float, but it stores the exponent in a different (variable-length) way.
> Posits have superior accuracy in the range near one, where most computations occur. This makes it very attractive to the current trend in deep learning to minimise the number of bits used. It potentially helps any application to accelerate by enabling the use of fewer bits (since it has more fraction bits for accuracy) reducing network and memory bandwidth and power requirements.
> [...] Note: 32-bit posit is expected to be sufficient to solve almost all classes of applications [citation needed]
The floating-point numbers are designed to have almost constant relative errors over their entire range of representable numbers and no other numeric format can improve on that.
The posits are the result of a different trade-off, where improved precision for the numbers close in magnitude to 1 is obtained by reducing the precision of the small numbers and of the big numbers.
When evaluating the accuracy of benchmark results, with floating-point numbers the input values do not matter, unless they have values that cause underflows or overflows. On the other hand, with posits the input values matter a lot, because with some values the accuracy will be excellent, much better than with floating-point, while with other values the accuracy will be much worse than with floating-point.
There are many problems for which low precision, i.e. up to 32 bits, is adequate and where posits can be better than floating-point numbers. Nevertheless, for each such problem a careful numeric error analysis must be made, because using them blindly can produce unexpected results. Floating-point numbers are more foolproof.
However, for scientific computing problems that require precision of at least 64 bits, I have never seen one that could benefit from using posits. All the problems that arise from the simulation of the physics of sufficiently complex systems, e.g. the simulation of electronics circuits, require computations with small numbers and big numbers, where the precision of posits drops dramatically.
In theory, it would be possible to also use posits in many 64-bit applications, if an analysis of the problem would be made, to determine a large number of constant scaling factors, which would be inserted in various formulae, to bring the operands in the range where posits are more accurate than floating-point numbers.
Nevertheless, such an analysis consists in a huge amount of work, which is never worthwhile just for replacing FP numbers with posits. If such a search for optimal scaling factors were done, then it would be better to implement the computations with fixed-point numbers, to obtain even more accurate results than with posits.
They're really elegant.
And not just that, it seems they also standardized arithmetics? Which is a big deal because IEEE 754 is unusable in heterogeneous distributed systems as every hardware implementation does something different.
It's expensive though:
Synthesis results of the 64-bit PAU in Big-PERCIVAL
have shown that it requires 2.5× as many resources as
the double-precision FPNew FPU. Moreover, we studied
the impact of the corresponding 1024-bit quire accumulator
register, which increased the total hardware cost to a third
of the area of the core. Detailed area results illustrated how
the hardware resources are distributed among the different
operations. In particular, the most resource-hungry elements
are the quire-related units and the posit division and square
root units.
I don't think this is a particularly positive results for posits.Quad-precision float seems more general-purpose and honestly more promising for scientific computing, since the error analysis is easier.
But in general, it seems that the strongest features of posits are basically recognizing that being strategic with where you need extra precision is advantageous, and if you apply the same techniques to IEEE 754 floats, you lose most of the seeming advantage of posits.
The #1 issue in computer performance is The Memory Wall... it is orders of magnitude more expensive to move data between external DRAM and the processor than it is to do operations within the processor. The solution is to increase information-per-bit so that real numbers can be represented in 32-bit precision with sufficient accuracy. That more than doubles the performance over 64-bit floats since it allows more data to fit in cache at every level of the memory hierarchy.
As I understand it, your other complaints tend to center around overflow to infinity and precise summation of vectors. For applications that really care about that precision, there are ways to do it in floating point without a quire register - sorting before summing is the naive approach, but look into ReproBLAS for some better algorithms.
Also, I can't help but wonder if the memory wall idea here is centered only around synthetic benchmarks like gigantic dot products. A lot of code leans heavily on caches these days, which make the energy cost of operations a lot lower, and pretty much everything short of massive dot products uses them. I imagine you would have to make a very nuanced argument about why a 1k fixed point sum is saving energy here. Even matmuls are pretty cache-efficient now.
Elsewhere in computing, we are actually generally moving away from tightly-packed structs in performance-sensitive code despite the memory retrieval cost, because they are just easier to deal with in both hardware and software, and locality picks up all the slack.