Apparently some people are drawing connections to this paper [1] on 4 bit LLMs, which has one NVIDIA employee among its contributors
Only partially, because in LLMs FP4 isn't half as useful as FP8. So if you have gear that crushes at FP4 then that's what you use and you benefit from that increased speed (at minimal accuracy loss).
Definitely some marketing creativity in there, but its not entirely wrong as a measure of real world usage
...assuming the recent 1.58b paper doesn't render the entire float quantization approach obsolete by then.
Discussed in a previous post https://news.ycombinator.com/item?id=37930663
- Research for a while now has been finding that smaller weights are surprisingly effective. It’s kind of a counterintuitive result, but one way to think about it is there are billions of weights working together. So taken as a whole you still have a large amount of information.
Wasn't there a paper from Microsoft two weeks ago or so where they trained on log₂(3) bits?
This makes network minimise loss not only with regard to expected outcome but also minimises loss resulting from quantisation. With big networks their "knowledge" is encoded in relationships between weights, not in their absolute values so lower precision work well as long as network is big enough.
4 bits is effectively 16 different float point numbers - 8 positive, 8 negative, no zero and no NaN/inf. 1 bit for sign and 3 bits for exponent, 0 bits for mantissa, mantissa is implied to be 4. It’s logarithmic - representing numbers in the range from -4^3 to 4^3, smallest numbers are 4^-3.
I like how you specified that it's not floating point.