Might be off-topic but: is it possible to perform such a quantization on Apple devices? Something like Mac Studio Ultra M1 (even if it would take weeks/months)?
[0]: https://github.com/ggml-org/llama.cpp/blob/master/tools/quan...
I use this project: https://github.com/vllm-project/llm-compressor