I really hope that there will be some real alternative in terms of hardware soon. This kind of comment makes me laugh and makes me sad:
> any GPU could be used to run the 4bit quantization as long as you have CUDA>=11.2 installed
Any GPU (as long as it's NVIDIA).
Also, it sounds like those models won't get the .cpp variant if they don't do the processing with quantized weights, right? (The list claims they're unpacked and still running on floats)