Is this different than ZipNN? https://arxiv.org/pdf/2411.05239
I see it mentioned but can’t understand if it’s based on it or different/better…
I see it mentioned but can’t understand if it’s based on it or different/better…
If you add an LZ-type compressor and have this be in the critical path for inference, then decompression will be a lot slower. It would be best to fuse decompression with the compute kernels (e.g., a GEMM that performs decompression on each tile before the arithmetic), and the simpler the decompression routine, the easier this will be.