As someone who just follows this stuff from afar, it is hard for me to conceptualize if this is a SaaS only model, or if it means we are getting to the point where you can have a A1 model on a local machine.
As someone who just follows this stuff from afar, it is hard for me to conceptualize if this is a SaaS only model, or if it means we are getting to the point where you can have a A1 model on a local machine.
- to LOAD the model, you need at least 768GB of VRAM, which means 10xH100 GPUs or similar.
- to QUERY the model, it then uses one of the 37GB layers to perform the computation at any given time, which means that each GPU can process 2 queries concurrently - (37 * 2 < 80) - and the queries are very fast because of this.
So a single user setup would involve a crazy expensive rack of 10 h100 GPUs that can essentially process 20 concurrent requests almost as quickly as it can process 1 request in a single user mode...
The result is that the model is extremely cheap to operate if served as a SaaS, but ridiculously expensive for a single user setup
Recommended RAM: more than most PC.