Posits, a New Kind of Number, Improves the Math of AI
spectrum.ieee.org
spectrum.ieee.org
>Real numbers can’t be perfectly represented in hardware simply because there are infinitely many of them. To fit into a designated number of bits, many real numbers have to be rounded. The advantage of posits comes from the way the numbers they represent exactly are distributed along the number line. In the middle of the number line, around 1 and -1, there are more posit representations than floating point. And at the wings, going out to large negative and positive numbers, posit accuracy falls off more gracefully than floating point.
>“It’s a better match for the natural distribution of numbers in a calculation,” says Gustafson. “It’s the right dynamic range, and it’s the right accuracy where you need more accuracy. There’s an awful lot of bit patterns in floating-point arithmetic no one ever uses. And that’s waste.”
For example, what would be the most efficient binary representation of probabilities values between 0 and 1? In the future, I can imagine hardware + software that is specialized for examples like this.
edit:
Do I really want more resolution between 0.41 and 0.42 than between 0.01 and 0.02?
Some experiments like with liquid crystals early on required purities that what chemical company was it, Merck I think, around the year 1900, complained like they felt insulted. Told the inventor of liquid crystals because it was an absurd amount of purity required to get them actually working. But look at them go! Right in front of your very eyes!
The logistic curve becomes denser close to 0 and 1. Which makes sense: you will want to tell apart 1 defect per million from 0.01 dpm, and 5-sigma process (99.977%) from 6-sigma (99.99966%), much more than tell apart 30% from 30.001%.
I wonder if it has graphics applications for high dynamic range images. Probably not.
The real challenge is probably in the power efficiency. As far as I know right now power consumption is the biggest cost factor when training new models. Everything else seems like a minor trade-off but when accounting for the power consumption there is real money involved.
Those A100 use 300watts each, plus let’s say 500watts for the rest. (8gpus300+500)8servers ≈ 25Kw * 8760 hours = 219k KwH / year. So if your costs are $0.15/KwH that’s only on the order of $32k/year.
In terms of TPUs or other custom accelerators, sure, they exist. However most definitely aren’t building their own hardware.
ETA: I’m not saying power is irrelevant, it clearly matters. But saying it’s the dominant financial constraint is clearly wrong, at least below Google/Amazon/Apple scale. Never mind the cost of the people running these trainings!
But if it mitigates error accumulation in perceptually-friendly ways, it might still have value in calculations (once there’s native support).
2. Bitmap images are not using floating-point values, except in some super niche use cases like GIS data, so “posits” are irrelevant for the use case you’ve posited. (ba-dum-ts)
3. Non-linear gamma is completely unnecessary for bit depths >= 16.
Edit: It seems I’m conflating “precision” with “accuracy”, which is a separate concept in math.
0.1 is the same as 1 / 10, which does not have a finite representation in binary notation, just as 1 / 3 does not have a finite representation in binary or decimal notation.
Posit would likely enable to better compression with Huffman encoding.
A format close to the cone fundamentals, like XYZ, could have benefits being encoded in some kind of non-standart posit.
I'll add this into the stack of things I eventually I'll look into.
There are also infinitely many integers, but we can represent them just fine inside a finite bound. The problem with reals (and rationals to a lesser extent) is that, within any range we are interested in, a single real number can be infinitely 'long'.
- Two's complement instead of one's complement
- No infinities, no signed zeroes
- Only one exceptional value with the same encoding as the smallest signed integer: 10...0
The above means that comparisons work exactly like for signed integers (with the unique exceptional value behaving like -∞). Also, one can test whether x is invalid using `x == -x && x != 0` rather than the ugly `x != x`, i.e. no need to break reflexivity. Even if posits do nothing but remove the one's complement dust off IEEE754, that would be a positive change in my view.
The biggest saving you could make is probably foregoing exactly rounded results, and only stipulate that the result has to be within ±1 lsb of the true value. That would save multipliers from computing a bunch of bits that don't end up in the result anyway, except for the rare case where they decide the rounding. That would probably be a good trade-off for most AI chips. For general purpose CPUs I don't think it is worth the breakage.
https://people.eecs.berkeley.edu/~demmel/ma221_Fall20/Dinech...
the relative error is bounded by 2^{-24} only in the interval [1.0e-6,1.0e6] for the Posit32 format, whereas it is [1.17e-38,3.4e38] for the IEEE binary32 format.
So I understand the use case but they can't replace floats with those infinities remove imho. They complement them. You can work around those things by wrapping posits but sometimes you just need them (for example: I needed it yesterday). I don't always use them, but still somewhat frequently. They are not error codes, I work with the extended real number line [1] so infinity is a valid return value. What's the alternative...working with PositOrInfintiy?
I would probably use an exceptional value and later treat the exceptional like if I would encounter an infinity (usually they are later involved in calculations where infinites turn to 0). But that's just hard to understand for anyone but myself.
Maybe I work not low-level enough but I've never used `x != x` would use `x == -x && x != 0` because that's just not readable. I want a function (isinf, isnumber etc) and I am happy.
Sure, isnan(x) is what one should use. The fact that it is `x != x` is an implementation detail. The problem is that it is also a hack that breaks the usual mathematical axioms for equality and for order relations. For instance, if you want to sort floating-point values, you have to write your own comparison predicate in case there is a NaN, because a NaN is neither smaller, greater or equal to itself.
As for infinities, they somewhat work for real numbers, but it gets more complicated for complex numbers. For instance, Annex G of the C standard stipulates that an infinite complex number multiplied by a nonzero finite complex number should yield an infinite complex number. Sounds reasonable, but consider:
(∞ + i∞)×(0 + i1) = (∞×0-∞×1) + i(∞×1+∞×0) = NaN + iNaN
So Annex G recommends some complicated functions to be executed at each complex multiplication and complex division, which makes little sense for most applications, and I suspect few people do that. As an aside, Annex G breaks the whole point of NaNs, because it stipulates that numbers like (∞ + iNaN) should be considered infinities rather than NaNs, which means that NaNs are no longer necessarily viral.All in all, what I find frustating with these aspects of IEEE754 is that they complicates things under the hood, but the benefits seem to me limited to some specialized applications.
Edit: Also, it looks like posits also use sign-magnitude and not 2's-complement? I am confused as to where you are getting this from. [Edit again: I was wrong about this, see below.]
https://posithub.org/conga/2019/docs/13/1430-John-Introducto...
The posit standard was ratified in March this year. It's only 12 pages and shockingly comprehensible for a standard document: https://posithub.org/docs/posit_standard-2.pdf
Also in March Gustafson hosted the Conference for Next Generation Arithmetic (CoNGA) 2022 sorta 'inside' the Supercomputing Asia (SCA2022) conference. They covered a bunch of different areas not exclusively related to posits. Lots of cool topics including hardware implementations and optimizations, comparisons between different number formats, and a talk and panel from the posit committee about writing the standard and what's next for posits. The session videos are somewhat difficult to find so I made a playlist: https://youtube.com/playlist?list=PLBH9oUUfaYoQNl6-zr_ScsMp4... (Note the descriptions have talk titles and authors.)
The best video intro to posits is probably the one in 2019: https://www.youtube.com/watch?v=JJgT-YphE1Y
The posit standard requires correctly rounded elementary functions in order for an implementation to be considered compliant. This means that every standards-compliant posit implementation is now also deterministic.
No, it doesn't. You have to request #pragma STDC FP_CONTRACT ON explicitly, or use -ffp-contract flag-equivalent (implied by -ffast-math or -Ofast) on most compilers. icc is the only major compiler that actually defaults to fast-math flags.
> Also, with default settings, GCC is willing to replace single precision math with double precision math if it feels like it.
I'm less familiar with gcc than I am with LLVM, but I strongly doubt that this is the case. There is a provision in C/C++ for FLT_EVAL_METHOD, which indicates what the internal precision of arithmetic expressions (which excludes assignments and casts) is, and this is set to 2 on 32-bit x86, because x87 internally operates on all numbers as long double precision, only explicitly rounding to float/double when you tell it to in an extension. But on 64-bit x86, FLT_EVAL_METHOD is 0 (everybody executes according to their own type), because SSE can operate on single- or double-precision numbers directly.
Fortunately that isn't a problem since SSE2.
[1] http://raden.fke.utm.my/blog/positnumbersystem
[2] https://www.johndcook.com/blog/2018/04/11/anatomy-of-a-posit...
I’d love to know what Kahan (the father of IEEE 754 floats) has to say about them. He thought that the „merits of schemes proposed in“ earlier versions were „greatly exaggerate[d]”.
There's a vid on youtube with Kahan and Gustafson talking over it. Kahan goes in like a wolf onto a lamb, much to gustafson's shock and hurt. Gustafson points out some of Kahan's criticisms are plain wrong, even citing the page in the book, then says something like "we should be working together on this, why aren't we working together?" Kahan doesn't respond.
Kahan may be right or wrong but his attitude is weirdly hostile, and Gustafson is no n00b about floats.
Edit: I think this is it (but it's been a while) https://www.youtube.com/watch?v=LZAeZBVAzVw
Unum III/Posits can work, but mainly it is just a number compression format. Working with say 64 bit Posits pretty much requires implementing arithmetic equivalent to that of 80 bit IEEE floats, and then throwing away a larger or smaller portion of the mantissa depending on the exponent. And the extra accuracy around the sweet spot of 1 still comes at the cost of lost accuracy for large and small numbers, so some workloads will suffer.
The answer to your question--why isn't Kahan working with Gustafson--is that, in Kahan's view, interval arithmetic (this is, AIUI, the main thrust of unums) isn't actually an effective solution to the "problem" of needing numerical analysis. While Kahan isn't the best at explaining this in detail, he does point out two valid issues: interval arithmetic can give excessively pessimistic ranges (because it doesn't account for correlated error), and it can give just plain incorrect answers when you have singularities in ranges.
I would note that, as Gustafson is no longer (as far as I know) pushing for the interval arithmetic approach, this is basically a concession that Kahan was right and Gustafson was wrong.
They go pretty much the other way than this representation: they allow huge numbers, much larger than can be represented in regular notation (incl. floating point) to be represented and calculated with efficiently and exactly. Tarau's system improves slightly on Knuth's in that representations are unique. He also proves that they at worst require twice as many bits as regular representation.
Knuth's and Tarau's system are variable-length and limited to naturals, but it seems like it would be easy enough to extend them to rationals and fix the representation length.
The article links to the posit paper [0]:
> A posit processing unit takes less circuitry than an IEEE float FPU. With lower power use and smaller silicon footprint, the posit operations per second (POPS) supported by a chip can be significantly higher than the FLOPS using similar hardware resources.
The argument seems to be predicated on the cost associated with NaN-handling. The exposition comes across as somewhat arrogant, IMO:
> If a programmer finds the need for NaN values, it indicates the program is not yet finished, and the use of valids should be invoked as a sort of numerical debugging environment to find and eliminate possible sources of such outputs.
Meanwhile, from the article where people actually tried implementing this thing on an FPGA:
> They also found that the improved accuracy didn’t come at the cost of computation time, only a somewhat increased chip area and power consumption.
I wonder where the mismatch comes from.
[0] http://www.johngustafson.net/pdfs/BeatingFloatingPoint.pdf
Maybe you save a bit by not having denormals, but then parsing the packed float is a bit more complicated in that the bits do not have a fixed division between exponent and mantissa.
It is possible that the Posit circuit was smaller by leaving out some feature, like exact rounding, which is quite expensive, but then it is not an apples-to-apples comparison.
Gustafson's paper is old, and doesn't reflect the language of the recent Posit standard. He may have had a different silicon implementation in mind than what has been implemented. Gustafson says there is no NaN. But in the Posit standard, there is a single NaN-like value called NaR (Not a Real). In IEEE, 0/0 is NaN, while in Posit, 0/0 is NaR. The rules for NaN and NaR are different, so they have different names.
It's possible Gustafson's paper doesn't consider the scaling as you get to wider types, or the presence of NaR removes enough of the savings from killing off NaN that it's a wash.
Yes. Nothing for free.
> Whereas floating point numbers are polluted with NaN values, posits are cleansed of such unclean special values.
https://www.cs.cornell.edu/courses/cs6120/2019fa/blog/posits...
I can imagine an alternative timeline where it worked like that though.
I also expect that chip area and power consumption for IEEE floating point has been optimized quite a bit more than for this novel number format, so there may be some efficiency left on the table on those particular metrics.
It uses more hardware because the posit is decoded into a floating point larger than a equivalent IEEE754 with the same number of bits. So, the logic units needs to be larger.
Posits showed an astounding four-order-of-magnitude improvement in the accuracy of matrix multiplication"
This is nothing short of amazing!
I can just imagine future GPUs on FPGA's using Posits rather than (now old!) IEEE floating point formats...
(Also, on a probably unrelated note -- might Posits be where we get the future equivalent of Star Trek TNG's character Mr. Data's Positronic (Posit-ronic!) -- AI "brain" from? ???)
But... if we take for example float16, the smallest model prediction you can represent greater than 0 is 2^{−24} = 5.96*10E−8, but the largest you can represent _less_ than 1 is (binary) 0.111... = 1-2^{-10} = 1-9.77*10E-4. So values around 1 are about 4 orders of magnitude more quantized than those around 0. I don't know if that's necessarily a 'problem' but I have noticed this fact when looking at model predictions.
I quickly suspected he wasn't very competent about how computers worked, and tried to probe a bit. I suspected he was mentally ill but could "talk the talk" enough to convince people his ideas had merit.
But this: A more efficient way of handling floating-point numbers, is probably something worthwhile. I wonder how hard it will be to recompile existing software to take advantage of posits?
That doesn't apply to engineering because it either works or it doesn't. The presenter misunderstood basic concepts of information. If you don't understand those, you can't design a working computer.
It seems much simpler than the posit proposal and it has the same advantage (most precision around 1 and -1).
In the logarithmic representation, multiplication becomes addition. I don't think that just by doing everything in logarithms you really increase the overall complexity, rather you transfer it from multiplication onto addition, so I don't think adding logarithms should be more expensive than integer multiplication..
Should it then matter for neural networks, which IIRC require similar amount of additions and multiplications?
Even floating-point is actually trading off the simplicity of addition in order to make multiplication easier, because they are quasi-logarithmic representation.
However, I wonder if there is a numeric representation where the circuits for addition and multiplication are of similar complexity. Something like half-logarithm, which if applied twice, you get the logarithm.
* Converting back and forth to linear space, using exp and log functions, this is much slower than a regular multiplication.
* Evaluating a polynomial that approximates the logarithmic space add, this takes several multiplications, so also much slower.
In general we tend to use more additions than multiplications, so trading addition speed for multiplication speed is rarely a good idea, even 1 for 1. If we need a lot of exponentiation keeping some values in logarithmic space may be beneficial, but it has to be those rare cases only.
Posits have floating range as well as floating precision, so doubly floaty.
Therefore I propose that they should be called "Lofties".
I've done work with big numbers and this format screams problems.
I can see how it would be great depending on the domain, and wish there was more diversity out there in the wild in this area but this seems a little hyped? Maybe I'm missing something. I'm not saying this doesn't seem useful, just that the article seems to be presenting it as a replacement for floats, when it might be better thought of as another option?