Hi I'm Mark I work on torchao which was used for the quantization aware training and ARM kernels in this blog. If you have any questions about quantization or performance more generally feel free to let me know!
https://huggingface.co/docs/hub/en/gguf#quantization-types
It might even mean a non-GGUF quantization scheme; I'm just an intermediate user of local models, not an expert user or developer.
So this is gonna be 8 bit weights, 8 bit activations, group size of 256, symmetric quantization. Not sure how to map this to the GGUF variants because they don't mention how they don't do activation quantization
So for example for AWQ and GPTQ we can accelerate them by using a fast int4 kernel called tinygemm
In vanilla Pytorch I have the following expression:
t.sum(values[inds] * weights)
If 'inds' is int8, I get "IndexError: tensors used as indices must be long, int, byte or bool tensors".Is this still true if I use torchao?