Switch Transformers C – 2048 experts (1.6T params for 3.1 TB) (2022)
huggingface.co
huggingface.co
[1] https://github.com/google-research/t5x/commit/199f226eeff5f8...
> It's pretty much the rumored size of GPT-4. However, even when quantized to 4bits, one would need ~800GB of VRAM to run it.
The Nvidia documentation on this is absolutely horrible, the features vary wildly by model of GPU, and I claim no real expertise.
I also have 960GB of vram in my garage (40x P40), though I suspect getting this model distributed in a useful way might be annoying and not worth the effort particularly since it would probably be close to turn-key (if a bit slow) on a high ram epyc host.
I have a similarly-sized pile of MI25s and recently managed to get a eight of them running in a single plentiful-on-ebay supermicro motherboard (custom bifurcating fanout-riser, x8 to each card). It was a "learning experience".
Falcon 180B, 4-bit quantization fits in VRAM and works, but at under 10tokens/sec because llama.cpp's support for multi-GPU on AMD is kind of an afterthought.
Nice!
Sorry to pester you, last question: did you find a source for those at a reasonable price?
Pricing from last time I investigated that would've been more than the P40s, so I'm assuming you did. Apparently the actual ICs are/were incredibly expensive and only Pericom sells them as standalone chips AFAICT. (Edit: oh wow, Pericom got bought by Diodes Inc; this explains a lot)
Thanks!
Not the anime kind. ;)
Why it can't be streamed from disk layer by layer, a sliding window of it, computed, temporary results held, offloaded it back and load the next window. Repeat till whole inference is done.
Also, if these are so much repetitive calculations in nature that you need CUDA cores to compute then why the inference can't be streamed and spread across a cluster of machines each having multiple commodity GPUs where one central "conductor/orchestrator" machines collects all the results from all the cluster participants?
However, it's still slow as tar. When I was running the BLOOM model I think my inference time was 1 token / m.
See: https://towardsdatascience.com/run-bloom-the-largest-open-ac...
So for researchers who can access a system with enough memory, which is most of the people professionally exploring these models, nobody needs to invest much preparatory effort into those other techniques or reducing their overhead. For them, there's basically no ROI for it.
But a lot of that work to optimize for those other techniques is gradually being accomplished in the open source community, where people don't have access to expensive clouds and lab systems. It takes time though, and can still only achieve so much.
Would it be difficult to wire up for conversations like ChatGPT? Could you run it against a local photo store to let you search by names of objects/people? Or is it basically an intermediate model that needs further training to fine tune to your application?