X.ai's Grok-1 Model Is Officially Open-Source and Larger Than Expected
synthedia.substack.com
synthedia.substack.com
Not aware of OpenAI at any point saying they would open-source their models.
Only that researchers would be encouraged to share their work with the world:
https://web.archive.org/web/20190224031626/https://blog.open...
So if this was quantized using ~4 bits per parameter you'd need ~40GB of vram. So you could spread it across 2x 3090 24GB using llama.cpp.
> So if this was quantized using ~4 bits per parameter you’d need ~40GB of vram.
No, Mixtral 8x7B (which is a total of 45 billion parameters, because there is a shared portion of the 7B, so its not 56 billion) at 4-bit quantization takes ~29GB [0]. A 314B model is ~7 times as large; with a similar architecture its not going to take only another 1/3 as much RAM.
You also need additional RAM that increases as some function of context size (not sure what function, and ISTR there are big-O differences between architectures in how it varies) to actually do inference.
> We are releasing the base model weights and network architecture of Grok-1, our large language model. Grok-1 is a 314 billion parameter Mixture-of-Experts model trained from scratch by xAI.
> This is the raw base model checkpoint from the Grok-1 pre-training phase, which concluded in October 2023. This means that the model is not fine-tuned for any specific application, such as dialogue.
> We are releasing the weights and the architecture under the Apache 2.0 license.
> To get started with using the model, follow the instructions at github.com/xai-org/grok.
A little disappointing they are not releasing the weights for the Grok-1 finetuned model.