I would love it if I were able to run these things locally like I am with stable diffusion.
I would love it if I were able to run these things locally like I am with stable diffusion.
Large language models gain key capabilities as they increase in size: more reliable fact retrieval, multistep reasoning and synthesis, complex instruction following. The best publicly accessible is GPT-3 and at that scale you're looking at hundreds of gigabytes.
Models able to run on most people's machines fall flat when you try to do anything too complex with them. You can read any LLM paper and see how the models increase in performance with size.
The capabilities of available small models have increased by a lot recently as we've learned how to train LLMs but a larger model is always going to be a lot better, at least when it comes to transformers.
For training such large models, data parallelism is no longer sufficient and tensor/pipeline parallelism is required. The problem is communication bottlenecks, differing device/network speeds and massive data transfer requirements become serious enough issues to kill any naive distributed training across the internet approach. Deep learning companies use fancy 100Gbps+ connections, do kernel hacking and use homogeneous hardware and it's still a serious challenge. There is no incentive for them to invest in something like GPT@home.
But it's not impossible and there's some research being done in the area. Although, it'll be a while until a GPT@home approach becomes a ready alternative. See https://arxiv.org/abs/2206.01288 and their recent GPT-JT test for more. Another development would be for networks to become more modular.
> use fancy 100Gbps+ connections,
you can pick up 100gbps mellanox nics on ebay for $50 on a good day, $200 whenever. If you're only connecting up two or three hosts you can just use multiport cards and a couple dac cables, rather than a switch.
I suspect for inference though there is a substantial locality gain if you're able to batch a lot of users into a single operation, since you can stream the weights through while applying them to a bunch of queries at once. But that isn't necessarily lost on a single user, it would be nice to see a dozen distinct completions at once.
GPT-JT was released recently and seems interesting but I haven't tried it. If you're focused on scientific domain and want to do Open book Q/A, summarization, keyword extraction etc. Galactica 6B parameter version might be worth checking out.
If our main language is not English one of the mt0 models might be worth a try https://huggingface.co/bigscience/mt0-xl
These models are distinguished by being able to follow relatively complex natural language instructions and examples without needing to be finetuned.
It's an open-source text generation frontend that you can run on your own hardware (or cloud computing like Google Colab). It can be used with any Transformers-compatible text generation model[3] (OpenAI's original GPT-2, EleutherAI's GPT-Neo, Facebook's OPT, etc).
It is debatable that OPT has hit that sweet spot in regards to "surpassing" GPT-3 in a smaller size. As far as I know, their biggest freely-downloadable model is 66B parameters (175B is available but requires request for access), but I had serviceable results in as little as 2.7B parameters, which can run on 16GB of RAM or 8GB of VRAM (via GPU).
There's a prominent member in the KAI community that even finetunes them on novels and erotic literature (the latter of which makes for a decent AI "chatting partner").
But you do bring up a great point: the field of OS text generation develops at a sluggish pace compared to Stable Diffusion. I assume people are more interested in generating their own images than they are text; that is just more impressive.
[1] - https://github.com/koboldai/koboldai-client
[2] - https://old.reddit.com/r/KoboldAI/
[3] - https://huggingface.co/models?pipeline_tag=text-generation