41 karma · joined January 2, 2018
Deploying LLaMA 3 8B is fairly easy but LLaMA 3 70B is another beast. Given the amount of VRAM needed you might want to provision more than one GPU and use a dedicated inference server like vLLM in order to split your model on several GPUs.
LLaMA 3 8B requires around 16GB of disk space and 20GB of VRAM (GPU memory) in FP16. As for LLaMA 3 70B, it requires around 140GB of disk space and 160GB of VRAM in FP16.
I hope it is useful, and if you have questions please don't hesitate to ask!
Julien
Does it monitor hacker news in real time?
It seems to me that Google does not really want to sell TPUs but only showcase their AI work and maybe get some early adopters feedback. It must be quite a challenge for them to create a dynamic community around JAX and TPUs if TPUs stay a vendor locked-in product...
As far as I can see llama.cpp with CUDA is still a bit slower than ExLLaMA but I never had the chance to do the comparison by myself, and maybe it will change soon as these projects are evolving very quickly. Also I am not exactly sure whether the quality of the output is the same with these 2 implementations.
If you prefer to use an "instruct" model à la ChatGPT (i.e. that does not need few-shot learning to output good results) you can use something like this: https://huggingface.co/TheBloke/Wizard-Vicuna-30B-Uncensored... The interesting thing with these Uncensored models is that they don't constantly answer that they cannot help you (which is what ChatGPT and GPT-4 are doing more and more).
I hope it will be useful.
On NLP Cloud we're doing our best to make sure that once a model is released it is "pinned" so our users can be sure that they won't face any regression in the future. But again it costs money so profitability can be challenging.
I takes a bit more work though, and it makes your requests bigger so more expensive.
These models can be self-hosted but they will require advanced hardware and some specific skills related to AI model deployment.
But for the moment you can't just use pure natural language instructions with BLOOM indeed. Maybe it will come!
<html>
<head><title>403 Forbidden</title></head><body>
<center><h1>403 Forbidden</h1></center>
<hr><center>Microsoft-Azure-Application-Gateway/v2</center>
</body>
</html>
I recently integrated GPT-NeoX 20B on NLP Cloud: https://nlpcloud.io . I had hopes that non-English languages would be better supported than with GPT-J since the model was trained on 20B parameters instead of 6B parameters, but quality still leaves to be desired. In my opinion, the best way to handle text generation in non-English languages for the moment is to couple it with a good translation module. I actually wrote an article about that: https://nlpcloud.io/multilingual-nlp-how-to-perform-nlp-in-n... .
But there is hope! Bigscience is about to release a huge NLP model that should theoretically work very well in almost 50 languages: https://bigscience.huggingface.co/ . We'll soon see if it's true!