Standardizing next-generation narrow precision data formats for AI
opencompute.org
opencompute.org
Spec: https://www.opencompute.org/documents/ocp-microscaling-forma...
Whitepaper: https://arxiv.org/abs/2310.10537
> Integer data types use a 2’s complement encoding, but the maximum negative representation (−2) may be left unused to maintain symmetry between the maximum positive and negative representations and avoid introducing a negative bias.
... the maximum negative representation was used for a NAN. IDK why and how it is that we all agree that NAN-s are useful for floats (and they are super useful), but very few think the same for integers??
Because making an integer bit pattern act as a NaN would require specific semantics (e.g. NaN + X = NaN; NaN != NaN) which are difficult to implement efficiently in hardware. These properties would also potentially rule out some arithmetic optimizations which are currently possible.
For something like addition, a ripple-carry adder is ~5 gates per bit. To check for NaN on input, you'd need a wide AND/OR on each input (~N log N gates per bit per input, so ~4 gates per bit for a 64-bit adder), a multiplexer on the output (~3 gates per bit), and a bunch of fan-in/out for the "is this NaN" signal. That'd likely more than double the size of the cell.
Subtraction makes that even more awkward. With 2's complement, an adder can also perform subtraction by inverting the second operand and carrying in a 1. This trick stops working if one of your bit patterns is a special value, so you either have to add even more logic to specify that NaN is inverted, or duplicate the whole mess for subtraction.
You'd also have to add a bunch of completely new hardware to distinguish between e.g. "X is bitwise equal to Y" and "X is numerically equal to Y" in equality tests, because NaN != NaN. It's hard to speculate how expensive that would be, but it certainly wouldn't be trivial.
> For floats it's been decided that paying the price was worth it, it seems.
That's a bit of an oversimplification. I'd say that it's more that:
1) Floating-point arithmetic is already fairly complex; even if you didn't handle the special values it'd still require a lot more logic than integer math.
2) Handling special values like infinities and NaN was simply part of the "spec" for floating-point math. It wouldn't have been considered fit for purpose without those features.
https://web.archive.org/web/20231018183224/https://www.openc...
Integer quantization doesn't typically just round, it has scaling and other factors in blocks so it's not just a question of manipulating int8's. And the FP16/FP8 are not supported by most processors so need their own custom routines as well. It would be great if you could just write code that operates with intrinsics on the quanitzed types.
Oddly they do have E8M0.
Nvidia's E8M12 is also a format specifically for operators - they expect you to store FP32 when you operate in E8M12. Storage is almost always in power-of-2 sizes.
Not sure why more effort isn't being put toward Posits, or thinking up a different format for ML specifically.
As for why not posits, I'm not entirely sure, but the "variably-encoded exponent width" nature of posits likely makes several details of their construction in hardware more difficult than things that have fixed exponent and mantissa widths. Although by the time you're talking about 8-bit datatypes, implementing operations by lookup table starts to look appealing.
2 NaNs, and -0 still seem like a bad call (-0 could have been the singular NaN), but I guess I understand maybe why they don't want to deviate too much from IEEE754 floats.
IEE754 floats should never have been a primitive in any programming language. Library type only.
Also: It's not cool IMHO that they have two distinct formats for FP8.
Not in general. Every 8 bit posit with es=1 can be decoded to a 10 bit E5M4 float. Every 8 bit posit with es=2 can be decoded to E6M3, and almost all fit into E5M3.
Though overall I'm not sure if posits are particularly useful here. Posits give you more precision near 1.0, and they give you better range for outliers. The latter almost certainly helps, but I don't know about the former.
Saving 1 bit wouldn't hurt though.
In terms of hardware to operate directly on posits, things get ugly. Floats are actually relatively easy to work with due to the normalized storage format and fixed precision across the range of numbers, while posits don't have that.
But you're not limited to powers of two. If you're decoding to 16 bit floats then you can use posits up to 13 bits wide.
> In terms of hardware to operate directly on posits, things get ugly. Floats are actually relatively easy to work with due to the normalized storage format and fixed precision across the range of numbers, while posits don't have that.
But is that extra shifting bigger than the multiplier unit? It's hard for me to see it growing the circuit that much.
And any nonzero finite number might as well be 1 to make it easier to think about.
0 and 1 would work, as long as you have a final scaling factor per node. -1 and 1 should work too.
Also the entire concept of a floating point is weak enough at 3-4 bits and collapses entirely when you have 1 or 2.
To have something IEEE754-like you'd need a sign, significand, and exponent. That gets you +/-(0, 1, Inf, Nan). Add another bit to the exponent if you want non-integer values. That's the minimum for what I'd call a float.
If you really want to strip it down to one bit, you should choose which of those bits to keep. Significand gives you (0,1), which is the same as a uint1_t. Exponent arguably gives you (0, Inf). Sign gives you (-1,1).
Of those, I'd say the sign-bit, giving -1 and 1, is maybe the most practical, but it's not a complete float.
"Building on years of design space exploration and research at Microsoft, Microscaling technology enables sub 8-bit formats while also enhancing the strength and ease-of-use of existing 8-bit formats such as FP8 and INT8. These advancements also help contribute to broader sustainability goals like reducing the environmental impact of AI technologies as demand continues to grow by improving the energy efficiency of AI in datacenters as well as on many AI endpoints."
Of course, it was not the first block floating point, those have been around since the 1963! https://en.wikipedia.org/wiki/Block_floating_point.
For example, current transformers llms seem to like 4-6 bits with smart quantization, with good performance at 3-4 bits with extremely aggressive quantization methods (like good use of sparsity and profiling inference on useful data).
The Stable Diffusion unet doesn't like 8 bit without some changes, the vae barely even likes fp16.
So to answer your question, some quantization is "free" and theres no reason not to use it, but sometimes its very lossy and a serious compromise that would not be taken with more compute/ram.
Also, sometimes there is compute overhead that makes quantization inferencr/training slower. Sometimes the reduced model weights size makes passes faster due to a memory bandwidth bottleneck. It just depends.