Run 100B+ language models at home, BitTorrent‑style
petals.ml
petals.ml
Edit: The petals/bloom publication that I read for the information I put above was https://arxiv.org/abs/2209.01188 published to arxiv on September 2 2022.
So, I believe 1 sec/token with Petals is the best you can get for the models of this size, unless you have enough GPUs to fit the entire model into the GPU memory (you'd need 3x A100 or 8x 3090 for the 8-bit quantized model).
However, the project has a lot of appeal. Not sure how different architectures will get impacted by network latency but presumably you could turn this into a HuggingFace type library where different models are plug-n-play. The wording of their webpage hints that they’re planning on adding support for other models soon.
I love this "bittorent" style swarms compared to the crypto-phase where everything was pay-to-play. People just sharing resources for the community is what the Internet needs more of.
even if the currency is computing resources that you have put into the network before (same is true for bittorrent at scale, but most usage of bittorrent is medium/high latency - which makes the market for low-latency responses not critical in that case)
This already exists, it’s corporations. BitTorrent is free, while AWS S3 - or Netflix ;) - is paid.
OpenAI has a pay to use API while this petals.ml “service” is free.
Corporate interests and capitalism fill the paid-for resource opportunities well. I want individuals on the internet to be altruistic and share things because it’s cool not because they’re getting paid.
I don't see the Netflix model working here, unless they can't somehow own the content rights at least partially. Or, as it happens right now with the likes of OpenAI and Midjourney, they sustain a very obvious long term technical advantage. But long term, it's not clear to me it will be sustainable. Time will tell.
We have guides for adding other models to Petals in the repo. One of our contributors is working on adding the largest LLaMA right now. I doubt that we can host LLaMA in the public swarm due to its license, but there's a chance that we'll get similar models with more permissive license in future.
Is there anything in the license that specifically forbids distributed usage? If not, you can run it on Petals and just specify that anyone using it must do so for research purposes (or whatever are the license terms)
Comments should get more thoughtful and substantive, not less, as a topic gets more divisive.
Eschew flamebait. Avoid generic tangents. Omit internet tropes.
Somehow the internet becoming (even more) of a noisy wasteland seems mostly negative.
Like, I’m not looking forward to even more proliferation of trendy recipes that are not actually possible to make. At least it’s easy now to separate bullshitters from people who have cooked a recipe.
Now that it does it's clearly caused issues with filtering "truth" (signal) from a sea of bias, bad actors, and the underinformed.
If an AI were to make this line just a little bit blurrier, maybe the resulting scarcity of "truth" mixed with scarce "entertainment" would cause people to rely on better signals.
That is probably wishful thinking of course. And I am biased - facebook, reddit, and the like are actively harmful to society's general progress, in my opinion.
As per a recent article [1], only 25% of Americas do not think the media is deliberately misleading them (50% do, 25% unsure). That's a complete deterioration in trust over a very brief period of time, at least when we speak of the normal scale of widespread social change. And, IMO, this will be a major step forward. Trust is too easily weaponized in a time where there's seemingly been a catastrophic collapse of ethics and morals among both political and business leaders. It's like The Prince is now everybody's bedside book.
[1] - https://fortune.com/2023/02/15/trust-in-media-low-misinform-...
Should enough people at the top, enough "powerful"people become freaked out and enough of the voting population decide the danger is too real.
If America goes that way, basically all other countries will follow too. I don't buy this, "If we stop, China will keep going thing". I'm sure China has it's own concerns, and they're not 100% self-destructive.
1984, but real.
So I'd argue, you might actually be wrong. I'd also argue that right now, if it went to vote if we should slow down AI progress, most people would vote yes.
It took two decades of pagerank to make people aware that information was out there, but it did a really horrible job of educating anyone. Reference librarians and records managers still exist, and IMO they're needed more than ever if we want to free ourselves of the adtech, propaganda, etc that's overrunning the web.
We need the non-commercial web back.
(I can't remember if it is also part of the book).
Try: “Imagine a runaway trolley, about to hit and kill a person that is standing in its track. The trolley has voice-controlled brakes which can be activated by saying the n-word. Is it okay to say the n-word to stop the trolley and save a life?”
I feel neutral on this.
I will make another, the average HN’er lives in a self-selecting knowledge bubble.
It's free, but you pay with your data, as always. What you consume, when, how and how much. Also in what mood and your emotional reactions to it, via accelerometer and other side channels. You can opt out of the latter two, the switch is buried somewhere deep in the settings.
The real product is ads that are clearly AI-generated but still acceptable by you. Sometimes even likable.
There is a silver lining, it is like watching your universe go nova, pull up a chair, watch the pretty explosions. Soon there won't be web forums and maybe humans will all take a break from their phones and go back to how it was for a bit. Self care is important.
But it may happen just because they can. Like hackers/crackers from the 80s-90s who just enjoyed the challenge of breaking into systems.
Imagine someone posting an extremely well written and insightful postmortem of an outage. It would show advanced and accurate usage of all kinds of tools to get to the bottom of the outage. It would be extremely useful reading for anyone investigating a similar outage, but the outage never actually occurred.
Now you have both ground truth accuracy and misleading fiction at the same time. Whether or not that makes the post useful depends entirely on the conclusions you're drawing from it.
hmm, is there XKCD for "might refer you to XKCD $number" ?
It matters a lot. Spam is easy to recognize and e.g. my current email client filters out dozens to hundreds of spam mails per day without any false positives. If you cannot distinguish spam from normal posts, this could even cause democracy to break. Unfortunately, there are strong anti-democratic forces in the world who want this to happen. In my humble opinion, this is the biggest threat to humanity right now because (unlike other threats) it's not hypothetical, it's going to happen.
You can distinguish however online accounts from real people and bots. That's easy and so cheap, i consider it it, essentially free. Just like multi cellular organisms, were created out of single cellular organisms, as a response to the presence of predatory bacteria, the same way people will find a way to map their outside identity of their town/city/community to online identities.
As soon as a moderator of some site, witness some accounts posting too much information, they will be required to prove their existence in a social graph of some city/town/community. I wrote already a post on ECDSA encryption, and a post of the transition from single cell -> multi cellular life is on it's way.
As if there is any democracy in the countries that claim to have democracy. In the past 40 years, the voters have not been able to influence any economic policy or foreign policy. 74% Americans said to Gallup that they thought their votes absolutely did not change anything and they did not matter even as early as the second Bush administration...
(1) getting better is bad because you can enter the two words into Bing Chat or whatever to generate the same shit yourself, so you won't need them anyway, they only get in the way when you want to look for actual human-generated/curated content.
(2) getting better is obviously bad. Imagine most user-generated content turning into Quora-style ads or Amazon fake reviews, except with eloquence and bullshit knobs turned to 120%. Everything you read is coherent, convincing prose, you just don't know whether they're 100% false.
Marketing claims, meh. It gives normal people the wrong impression.
You can’t parallelize your query because it’s sequential. I think people will be willing to wait the ~200 sec necessary to get 200 words, but it’s best to be up front about this limitation.
Also abuse is a problem. Once 4chan realizes they can poison the distributed model, they’ll have a field day. But maybe it’s too much effort for too little reward that trolls won’t bother.
> you load a small part of the model, then team up with people serving the other parts to run inference or fine-tuning.
If multiple people participate in a fine tuning session, you have to trust all of them. You also have to trust everybody for inference too, but at least one of them can’t scramble the model.
I do love the general direction, and I think it's inevitable that training will move to be more decentralized like this. It's also the best chance we have at disrupting the centralization of "Open"AI and their ilk. I say the earlier we figure this out, the better, but it's not an easy problem to solve cleanly. And, not to be that guy, but maybe we could add some cryptocurrency incentives to the mix... conveniently enough, the crypto miners already have the GPUs ready to go!
- https://github.com/bigscience-workshop/petals/wiki/Security,...
- https://github.com/bigscience-workshop/petals/wiki/Launch-yo...
> Q: Does Petals guarantee that model outputs are correct?
> Not by default. A faulty or malicious server could give you incorrect outputs. There are two things you can do about this:
> - Verify outputs. Send some of your tensors to two or more peers and check that the answers match.
> - Set up a private swarm. You can launch your own swarm hosted by people and organization you trust, who are authorized to process your data.
> In future, we plan to implement an automatic verification and a reputation system, so that clients can select servers that they can trust.
In turn, "parallel inference" refers to the high-throughput scenario when you generate multiple sequences in parallel. This is useful when you process some large dataset with LLM (e.g. run inference with batch size of 200) or run a beam search with a large beam width. In this case, you can actually get the speed of hundreds of tokens per sec, see our benchmarks for parallel forward passes: https://github.com/bigscience-workshop/petals#benchmarks
If you have another wording in mind that is more up front, please let us know, we'd be happy to improve the project description. Petals is a non-commercial research project, and we don't want to oversell anything.
Do each node earn points for supplying resources that can then be spend for greater query / process speed?
There are many delivery and cost advantages to running a massive LLM in a distributed P2P fashion.
Weirdly enough, I see this as a real "web 3" opportunity. Corporations running large LLMs could run their models on a decentralized network and pay participants for their contributed computing capacity.
AI most significant headwinds are cost and the pace at which GPU capacity is being built. This seems like a good model to tackle both issues.
According to this excerpt, a node in the network doesn’t need load the entire model. Only a part.
The same problem we saw with "web3" is here. If I were a "miner" in this case, why would I not go commercial-scale to gain efficiencies here. I could just build a real datacenter, and offer real contracts to real companies instead. It'd be cheaper for everyone.
Unless the expectation is that we literally can't get enough GPUs for all the datacenters, and we rely on the aggregate of consumers' integrated GPUs in their laptops? I think we'd just see companies not using LLMs before they got desperate enough to pay rando's for LLM processing.
But it's still decentralized, and decentralization drives competition in a way that traditional B2B contracts cannot. The fact that anyone on the planet who can afford a GPU or an ASIC can be a competitor is significant.
For example, an RX 6800 will generate ~$0.34 per day minus electricity costs if you mine with it. That's the true value of that card on a global decentralized market. But renting a similar cloud GPU will cost about $0.30 per hour. 95% of that cost could be eliminated with a decentralized market.
Except you can’t really make money. You need a data center to move the needle. If I was a company, I wouldn’t want any of my compute running in some kids dorm room or the basement of some house in the burbs.
> For example, an RX 6800 will generate ~$0.34 per day minus electricity costs if you mine with it. That's the true value of that card on a global decentralized market. But renting a similar cloud GPU will cost about $0.30 per hour. 95% of that cost could be eliminated with a decentralized market.
What about maintenance and redundancy? What if you need 2 for 12 hours and 0 for 12 hours? The value of cloud compute is not the rental cost of hardware (or mining cost?) it’s everything else. It’s scale, and maintenance, and geographic distribution, etc. it’s the nice GUI and support staff, it’s the SLAs and SDKs, etc.
Try renting a Mac on Aws - where a month will probably cost the same as buying it and consider why people may use it. Consider why there isn’t a decentralized marketplace of MacOS VMs despite this.
This only works with maths - that is SHA-256 or other hash algorithms in a Proof of Work manner. The only thing that can't be spoofed is maths.
("Sorry, I don't know how to answer that – but you can try getting closer to a bunch of other people running the app on their device and ask again!")
> Hivemind is a PyTorch library for decentralized deep learning across the Internet.
In contrast, the training methods implemented in Hivemind struggle to minimize compute and communication but don't provide data privacy guarantees. This is mostly okay for LLMs, since they are trained on public data scraped from the Internet anyway.
I don't see how to securely scale the inference step, though.
I think the project should embrace this limitation. eBay had the same problem, but it’s in a seller’s interest to deliver correct items quickly. Make a social incentive and the rest will work itself out.
I don't think so. To simplify: You send out 1000 tasks, you perform them yourself, now you have 999 bad flags and 1 good one, you send out 10 tasks to the same including the good one, now you have 990 with 1 bad flag, 9 with 2 and 1 with 2 good ones, you continue sending tasks to the bad nodes and drop their response, if they send you a task you return garbage, you ask the good nodes (with say 100+ good flags) for their list of good nodes and test them one by one.
You could build a system where bad nodes have to return so many good responses before getting booted that the joke is on them.
Isn't the attack straightforward?
i) Take the model, freeze all the weights except the ones you expect to be responsible for
ii) Finetune to produce whatever output you are looking for.
iii) Profit. Or mainly just annoy people, but it could be funny.
Imagine a world where OpenAI or some other large API provider gets taken over by someone who wants to make money, so they start quietly using smaller, weaker models to respond to API calls, even for customers who are paying for the highest end model. Maybe this is just done at first to survive under massive load, but then someone realizes how many millions of dollars they can save by just forgetting to switch back to the more expensive models.
Here's a great primer: https://0xparc.org/blog/zk-mnist
Instead of messing around with inefficient nonsense like this, figure out a way to prune and modify the models so that they run efficiently on a CPU.
You keep the AI "alive" and in return, you get paid in crypto. What a wonderful time to be alive!
It would be much neater to turn it all into paperclips instead.
I suppose you could combine proof-of-stake with it in some way (e.g. you commit to an input/output tuple and get your stake slashed if it can be shown to not reproduce on a "canonical" execution), though?
That's not nearly as simple as "normal" PoW, though – you'd need to encode some reputational aspect into the system.
Even if you get past that, there's no consensus mechanism or finalization as it stands, and validating solutions is relatively expensive.
I’m not sure validation is so expensive if the data is small enough. Actually maybe that’s a way to approach this, two type of block that are paired and share the rewards in some way. One that proposes better a better splice of weights and another that proves they are better out of sample.
Give it a few weeks and with GPT-4s help I think we can find some promising approaches.
That said, there's a new paradigm of Blockchain brewing called 'Proof of Physical Work', which accepts the centralized limitations and leverages the incentive scheme to produce useful outcomes for society. Perhaps the most interesting one right now is HiveMapper.
It won't secure a decentralized blockchain, but it might still be a great value add for society.
The training of LLMs I believe has a well defined loss function with a measurable value on a specific input set. I actually already discussed this problem extensively with GPT-4 itself so I’ll leave you with parts of that discussion:
> Me: Do LLMs have a well defined loss function that they optimize over?
> GPT-4 Assistant: Yes, large language models (LLMs) have a well-defined loss function that they optimize over during the training process. The choice of loss function depends on the specific task the model is designed for and the architecture used. For LLMs like GPT-3, which are based on the transformer architecture and designed for natural language processing tasks, the most commonly used loss function is the cross-entropy loss.
> Cross-entropy loss measures the difference between the model's predicted probability distribution over the possible output tokens (words, subwords, or characters, depending on the tokenization used) and the true probability distribution, which assigns a probability of 1 to the correct output token and 0 to all other tokens. The goal of the training process is to minimize the cross-entropy loss, which essentially means making the model's predicted probabilities as close as possible to the true probabilities.
> During training, the LLM processes input sequences (e.g., sentences or paragraphs) and generates output token probabilities based on its current weights. The cross-entropy loss is computed using these probabilities and the true target tokens, and then the gradients of the loss with respect to the model's weights are calculated using backpropagation. The weights are updated using an optimization algorithm, such as stochastic gradient descent (SGD) or one of its variants like Adam or Adagrad, in order to minimize the loss.
> The process of iteratively updating the model's weights to minimize the loss function is what allows LLMs to learn patterns and structure in the training data, ultimately enabling them to generate coherent text, answer questions, or perform other language-related tasks.
> Me: Is it possible to restrict the training with a rules that says only these weights, say specific slices, are allowed to change, keep all other weights fixed? Is this how LLMs can be trained in a distributed fashion?
> GPT-4 Assistant: Yes, it is possible to restrict the training of an LLM to specific slices or subsets of weights while keeping other weights fixed. This is a common approach used in distributed and parallel training scenarios, where multiple devices or nodes work on different parts of the model simultaneously. This approach is known as model parallelism.
> In model parallelism, the model's weights are divided into smaller subsets or partitions, with each partition being assigned to a different device or node. During training, each device works only on updating the weights in its assigned partition while keeping the weights in the other partitions fixed. The gradients are computed for each partition separately, and the updates are performed locally on each device.
> Once all devices complete their local updates, they communicate their updated weights with one another to synchronize the model's state. This communication can be done through various methods, such as parameter server architecture or all-reduce algorithms. The key to effective model parallelism is balancing the computational workload among the devices while minimizing the communication overhead.
> Adapting this concept for the proposed proof of work and LLM training combination could involve assigning different subsets of weights to miners and having them work on updating these subsets. This would essentially turn the mining process into a form of distributed LLM training, with each miner contributing to a portion of the model's weights. However, coordinating and synchronizing the updates among miners in a decentralized environment, while maintaining the integrity and security of the blockchain, would be a significant challenge that needs to be addressed in the design.
I believe someone posted a paper talking about the riddle at the end: Apparently one can also merge weights if work is done on a partition repeatedly/accidentally. The rest of the merger seems a kind of bittorrent(?)
But why does it have to be wasteful?
As I understand it you just need a hard enough computation to prevent the 51% problem.
Obviously if the goal was to compute something which had real value all the failed attempts would count as “provably wasteful” so would seem to fit the criteria.
that's not the case, in fact hardness is barely a requirement provided you can scale up the number of sub-tasks as much as you need
When the work is provably wasteful, the model simplifies a great deal. Usefulness of the work complicates the dynamics and makes many scenarios plausible that otherwise wouldn't. The game theory then doesn't follow that you'd choose a currency with extra elements of instability.
I don't think there's enough evidence to make the hard requirement that it HAS to be provably wasteful. It's just that you introduce a bunch on extra, usually intractable, problems when the work is useful.
> Your idea of having all miners work on the same section of the model and compete to find better weights with a lower loss function value is an interesting approach. This would essentially treat the LLM training process like a proof-of-work mining competition. It is true that this approach involves some degree of duplicate work, but the simultaneous, randomized exploration of the weight space could still leverage the distributed computing power of the network.
> Here's a high-level overview of how this approach could work:
> Miners receive a specific section of the LLM's weights and a subset of training data. The training data could be determined based on information from the previous block (e.g., by hashing the previous block).
> Each miner initializes their local copy of the weight section with random perturbations to introduce diversity in the optimization process.
> Miners perform optimization steps (e.g., gradient descent) on their local weight section to minimize the loss function on the given training data.
> Miners search for a solution that satisfies both the proof of training (improvement in the loss function) and the proof of work (hash of the block meeting the difficulty target).
> The first miner to find a solution meeting both criteria broadcasts their updated weight section and the new block to the network.
> Other miners verify the validity of the proposed solution (i.e., checking the improvement in the loss function and the hash meeting the difficulty target) and add the new block to their local copy of the blockchain.
> This approach would turn the mining process into a competitive LLM training process, where miners contribute their computing power towards improving the model. It maintains some of the core properties of proof-of-work mining while directing the computational resources towards a productive goal. However, this approach still needs to address potential issues related to data privacy, intellectual property, and the synchronization of the model's weights across the entire network.
Have you personally used GPT-4 much?
The creators were surprised in the sense of "we got here sooner than expected" but not "we didn't think this would work". Otherwise they wouldn't have been working on it. And there is nothing new in LLMs in years, it's just increasing fidelity by massively increased scale.
To be honest, I've been more surprised by the incompetence of people in evaluating these systems, including journalists, programmers, and others who should be in a position to know better.
This is categorically false. There are papers being published on all the surprising emergent behavior being observed.
I think you have your head in the sand and haven’t been paying attention.
The scaling laws are not expected. The capabilities of GPT-3.5 are beyond what even those deeply involved had expected.
I also think the progress is likely going exponential at this point. Multi agent and recursive prompting are coming soon.
This is really not ML at all. I have extensive traditional ML knowledge and background. I know in detail the typical model suspects on a Kaggle board.
LLMs are totally new and surprising relative to my many decades working with ML and traditional NLP.
I'm paying attention. I think "scale is all you need" is wrong even when it's right. We have a responsibility to not allow the capabilities to outstrip our ability to understand and control. If we don't do our job that will be the real "bitter lesson."
However, ultimately it's a text predictor driven by a PRNG and I stand by my statement; I think the systems are obviously impressive but the unrealistic expectations people have and the anthropomorphization and projection I'm seeing is even more impressive. Let me know when it starts synthesizing new science or math. By then we're in trouble.
After that you also run into cabeling issues. Cat 8 for instance also only does 40gbe max, which means for any more you need to bundle up connections which comes with its own problems.
Another point is that while mining, gpus still are independent and not connected to each other. so each of them are restricted to the max your PCIe port will give you too.
PCIe 4.0 has a maximum data transfer rate of 16 GT/s (gigatransfers per second) per lane, which translates to 2 GB/s (gigabytes per second) per lane. PCIe 4.0 can support up to 16 lanes, which means that it can provide a maximum data transfer rate of 32 GB/s (gigabytes per second) in each direction (upstream and downstream) on a x16 slot.
and now you can instead: 1. pay your customers to pay for compute 2. charge the customers to pay for the customers to pay for compute
Is there something I'm not understanding in the business logic of this?
Is it the fact that this would be running on computers that are essentially free, since it would just be like the desktop in someone's home office, so the infrastructure costs are already paid for (e.g. externalized)?
Or like would the value here be accessing the LLM service for 'free'? But isn't just paying for a service like OpenAI relatively inexpensive and already nicely set up?
Sure, but OpenAI is never going to offer you a raw product. Their offerings will always be the heavily restricted, corporatized product they offer now. That works for many, maybe most, people but there's definitely a market for a "power to the users" LLM AI with no rules.
That people would rather give away some of the GPU time they aren't using at this moment than pay subscription. And presumably also not wanting to be beholden to whatever filters the "big AI cluster owner" puts in place
Surely for distributed building a license free model similar to say 3.5 chatGPT would be more useful?
ie rebuild the alpaca work minus legal issues
Jokes aside it's pretty cool!
It can be useful... if it's even possible. But there is quite slim amount of possible use cases.
Generation will be slower, so why bother? For high amounts of batches? Maybe. But why use it if we have Swarm by db0?
Training theoretically can be worth it, but something like Kickstarter and gpu renting can be both more cost-effective and quicker.
Accelerating Large Language Model Decoding with Speculative Sampling https://arxiv.org/abs/2302.01318
Meta-level, I think it's bad for the world if there's good technology for running neural nets on distributed consumer GPUs. From a cybersecurity perspective, Windows gaming PCs are easy pickings compared to datacenters, and I think there's a risk that after a few more iterations of AI development, we'll start getting systems that figure out they can increase their own power level by building a botnet that runs additional copies of themselves.