4-Bit Quantization and QLoRA
huggingface.co
huggingface.co
> any GPU could be used to run the 4bit quantization as long as you have CUDA>=11.2 installed
Any GPU (as long as it's NVIDIA).
Also, it sounds like those models won't get the .cpp variant if they don't do the processing with quantized weights, right? (The list claims they're unpacked and still running on floats)
As an ODE solver, you wouldn't do nanoGPT with it though, you'd need to go back to KernelAbstractions and write a nanoGPT based on that same abstraction layer. Again, this is a demonstration of the cross-GPU tools for ODEs, but for LLMs you'd need to take these tools and implement an LLM.
Are there any decent but easily digest summaries of FP4? The best I can find is a giant paper. I do not understand why the linked entry gave great summaries of the larger FP types but then waves hand / fog about FP4.