Edit: The petals/bloom publication that I read for the information I put above was https://arxiv.org/abs/2209.01188 published to arxiv on September 2 2022.
Edit: The petals/bloom publication that I read for the information I put above was https://arxiv.org/abs/2209.01188 published to arxiv on September 2 2022.
We have guides for adding other models to Petals in the repo. One of our contributors is working on adding the largest LLaMA right now. I doubt that we can host LLaMA in the public swarm due to its license, but there's a chance that we'll get similar models with more permissive license in future.
Is there anything in the license that specifically forbids distributed usage? If not, you can run it on Petals and just specify that anyone using it must do so for research purposes (or whatever are the license terms)
However, the project has a lot of appeal. Not sure how different architectures will get impacted by network latency but presumably you could turn this into a HuggingFace type library where different models are plug-n-play. The wording of their webpage hints that they’re planning on adding support for other models soon.
I love this "bittorent" style swarms compared to the crypto-phase where everything was pay-to-play. People just sharing resources for the community is what the Internet needs more of.
even if the currency is computing resources that you have put into the network before (same is true for bittorrent at scale, but most usage of bittorrent is medium/high latency - which makes the market for low-latency responses not critical in that case)
This already exists, it’s corporations. BitTorrent is free, while AWS S3 - or Netflix ;) - is paid.
OpenAI has a pay to use API while this petals.ml “service” is free.
Corporate interests and capitalism fill the paid-for resource opportunities well. I want individuals on the internet to be altruistic and share things because it’s cool not because they’re getting paid.
I don't see the Netflix model working here, unless they can't somehow own the content rights at least partially. Or, as it happens right now with the likes of OpenAI and Midjourney, they sustain a very obvious long term technical advantage. But long term, it's not clear to me it will be sustainable. Time will tell.
So, I believe 1 sec/token with Petals is the best you can get for the models of this size, unless you have enough GPUs to fit the entire model into the GPU memory (you'd need 3x A100 or 8x 3090 for the 8-bit quantized model).
Comments should get more thoughtful and substantive, not less, as a topic gets more divisive.
Eschew flamebait. Avoid generic tangents. Omit internet tropes.