The author says:
"It used Voronoi Cells and reduced most math to 8- and 16-bit integer calculations with one or two single-precision floating-point calculations."
I wonder if it can be used to speed up LLM inference.
"It used Voronoi Cells and reduced most math to 8- and 16-bit integer calculations with one or two single-precision floating-point calculations."
I wonder if it can be used to speed up LLM inference.
Now maybe you don't need to represent 1/2^32 in precision. Maybe you just need to know to 1/8th. That's where you can quantize to a lower precision to save space.
But if you start with 32 bit float and quantize to 32 bit fixed, I struggle to think how that saves storage.