(Possibly naive question) This is marketed as open source. Does that mean I can download the model and run it locally? If so, what kind of GPU would I need?
When 4-bit quantization comes out, I would expect a GPU with 12GB VRAM to be able to run it.
Disclaimer: I work at Hugging Face