EDIT: just realized that the example was mentioned in the article, with a link to the slides. Still, the YT talk may add some context
[0] https://www.youtube.com/channel/UCOstJ2IVC4Y8mbgN0IsowKw/vid...
EDIT: just realized that the example was mentioned in the article, with a link to the slides. Still, the YT talk may add some context
[0] https://www.youtube.com/channel/UCOstJ2IVC4Y8mbgN0IsowKw/vid...
For example, the posit16 with nbits=16 and es=1 encodes the exponent like:
01.0 = 0 01.1 = 1 001.0 = 2 001.1 = 3 0001.0 = 4
The format has a normal exponent field with es bits that encodes exponent in binary. When it overflows, it encodes the carry as unnary format. In the unnary format the run-lenght of the same number encodes a number. For example: 0001 would be 3, 001 would be 2, etc. Of course, the posit format is a bit bore complicate than that (it supports negative exponent, for example). But the ideia is pretty much this.
Because of this encoding, if you are working with numbers that have small magnitude posit will have a LOT of more precision than your floating point format.
But the claim that posit16 can have as much precision as binary64 from ieee-754 is misleading. The posit16 can have up to 16-1-2-1=12 bits of precision. While binnary64 always has 53 bits of precision.
They likely compared the binnary64 with posit16 using the accumulator (aka quire). I'm not sure how the quire would map to real world FPU, it uses a lot os space.
Which I always find deceptiv because nothing stops us from using a quire with classical floating point arithmetic.
It is in fact a comparaison betwen a summation and a compensated summation : it is more precise because the algorithm is different not because of posits or floats.
They tell you that the quire is even faster than a traditional sum while being more precise but reading the associated reference reveals that it only hold with specific hardware which could also be used to use the quire with floats.
(and, arguably, I would love to know that every programmer is aware of a solid implementation of compensated/exact summation/dot-product and uses it when appropriate)
Precisely. Once the posit in unpacked they are indistinguishable from a floating point. It is not fair let posit use a massive accumulator while working with a tiny ieee-754 floating point accumulator. Like I said before, the precision of a number represented in binnary64 is greater than posit8. This comparation ignores the biggest advantages of posit: efficient data format.
> (and, arguably, I would love to know that every programmer is aware of a solid implementation of compensated/exact summation/dot-product and uses it when appropriate)
I like the idea of making the accumulator type (quire) accessible to the programmer. I think this brings awareness of the underlying hardware implementation to the average programmer.
As far as I can tell, Klöwer didn't. He mentions around the 13 minute mark (slide 9) that he used the SigmoidNumbers software package for Julia. From the looks of the examples he gives, if he used the quire, it must have happened implicitly. Which IIRC is not how using the quire works.
He did rescale all his inputs to minimize rounding errors that way, and he mentions that this has benefits for posits that floats don't have (because of the tapered precision of posits).
The two largest/smallest exponent posit16 can represent are:
min exp: 2^{ (+14<<es) or 0x01 } * (1+0) = 2^(-28) max exp: 2^{ (-14<<es) or 0x00 } * (1+0) = 2^(+29)
Note those numbers only have the implicit bit as significant. They don't have any space left to encode any other information other than the sign and regime.
While the double (binary64) can represent a much larger range of exponent:
min exp: 2^{ (-1023-2047-1) } * (1+0) = 2^(-1022) max exp: 2^{ (-1023+2047-1) } * (1+0) = 2^(1023)
Also, all doubles have 52 bits of precision while the posit have at most 16-1-2-1=12 maximum bits of precision.
Their significant are encoded slightly differently though. I'm not sure if this would be enough to achieve such different result without quire.
The scaling he mention could be done very easily introducing a parameter bias on FPGA implementation. I might add that to my work.
I pinged Klöwer on twitter.
> @milankloewer Would mind joining this thread on HN: https://news.ycombinator.com/item?id=20392612 We are discussing the results of your work with posit.
> Happy to join. In short: I did not use any quires so far. All simulations are entirely based on 16bit posits, but I compare them to Float16 and not Float64. Tricks are: Scaling and rewriting algorithms to avoid very large and very small numbers, that's it.
> I pinged Klöwer on twitter.
Probably the most sensible, easiest way to clear this up, haha :)