Petals: Run 100B+ language models at home bit-torrent style
github.com
github.com
Did I get that right?
This AI stuff is moving very fast and it's hard to keep up, but it's all fascinating.
For large LMs, people usually use tensor-parallelism (TP) or pipeline-parallelism (PP). TP involves lots of communication, but uses all GPUs 100% of the time and works faster. PP requires much less communication, but may keep some GPUs idle while they are waiting for data from others.
Usually, TP is used when you have good communication channels between GPUs (e.g., they are in one data center and connected with NVLink), while PP is used when communication is a bottleneck (like in Petals, where the data is sent over the Internet, which is much slower than NVLink).
Check out the infer_auto_memory_map metho which will optimize the model for your configuration (multi gpu, ram, nvme) and then run dispatch model on with that memory map.
I didn't have much luck with stock accelerate, but once gpu is disabled (so it runs only on cpu offloading to nvme storage where ram is insufficient) worked pretty well with me. (there is a small code change that has to be done as the stock software refuses to run without gpu-it is a simple change described in its github issues). My gpu is 8gb vram, but this way I managed to run 7b parameter models. In principle I could run a lot larger ones, but of course it takes a lot more time. The 7b bloom takes 90s for one inference and additional 60s to load the model (from a spinning disc array) initially.
Like Stable Diffusion, it's a web UI (vaguely reminiscent of NovelAI's) that uses a backend (in this case, Huggingface Transformers). You can use different model architectures, as early as GPT-2 to the newer ones like BigScience's BLOOM, Meta's OPT, and EleutherAI's GPT-Neo and Pythia models, just as long as it was implemented in Huggingface.
They have official support for Google Colab[2][3]; most of the models shown are finetunes on novels (Janeway), choose-your-own-adventures (Nerys / Skein / Adventure), or erotic literature (Erebus / Shinen). You can use the models listed or provide a Huggingface URL.
[1] - https://github.com/koboldai/koboldai-client (source code)
[2] - https://colab.research.google.com/github/koboldai/KoboldAI-C... (TPU colab; 13B and 20B models)
[3] - https://colab.research.google.com/github/koboldai/KoboldAI-C... (GPU colab; 6B models and lower)
Although you need a premium GPU. I admit it's not as good at zero shot or 1-shot as GPT-3 but if you provide examples, you can get as good of output. I feel like the team behind it needs better marketing.
However, I wonder how they prevent abuse. The main page doesn't mention it. As they mentioned block chain I suspect there will be some sort of credits implemented. I'll definitely be watching where this project goes.
Edit:just to clarify the 90s is not the 170b parameter model. It is 7b bloom version. I forgot to mention it and it puts the ability to run a 170B model in 1s in better perspective.
It's a very interesting concept, and I quite like the idea of a public, open compute cloud. I'd like to see more detail on security: if I'm going to donate time on my personal machine, I'd like some assurance that the workload is properly sandboxed and can't reasonably access my network or data.
Mostly out of interest, what's the advantage to this over just using the existing BOINC network? I've been running BOINC on and off since the dialup days, it's an extremely mature platform with all kinds of workload capabilities.
A client needs to communicate with multiple servers in a specific way to run the model, I'm not sure our communication model can be implemented with BOINC.
It would be interesting to reach a point to be similar to docker where you don't need to load each layer again and you only need your specific layer. The shared models layers would be already loaded, and running multiple models at once would consume less GPU memory.
Human: How is the weather today?
AI: the
AI theAI)aultAIAI ) course )
. can?esterday to people?
? is to think thatified )
Really cool project though, I wanted to work on something similar. It is nice today.
Not garbled, but also extremely shallow.This is a language model, not an oracle or an interface to weather forecast data.
What is "offloading" in this context?
Several recent works aim to democratize LLMs
by “offloading” model parameters to slower but
cheaper memory (RAM or SSD), then running
them on the accelerator layer by layer (Pudipeddi
et al., 2020; Ren et al., 2021). This method allows
running LLMs with a single low-end accelerator
by loading parameters from RAM justin-time for
each forward pass. Offloading can be efficient for
processing many tokens in parallel, but it has inher-
ently high latency: for example, generating one to-
ken with BLOOM-176B takes at least 5.5 seconds
for the fastest RAM offloading setup and 22 sec-
onds for the fastest SSD offloading. In addition,
many computers do not have enough RAM to of-
fload 175B parameters.It turns out, Petals is faster than offloading even though it communicates over the Internet (possible, with servers far away from you). That's because Petals only sends NN activations between servers (a small amount of data), while offloading copies hundreds of GB of NN weights to GPU VRAM to generate each new token.
You write that speed can be inferred, but the analogy that was used here is BitTorrent—and my experience with BitTorrent tells me that it certainly cannot be inferred.
Of course, you could do better if you have enough high-end GPUs to host the entire model yourself (3x A100 or 8x 3090). But if you don't, 1 token/sec is much faster than what you get with other existing methods.
Regarding homomorphic encryption (HE), I'm afraid the current methods to run neural networks in the HE fashion involve 10-100x slowdown, since they are mostly not designed for floating-point operations. We'd love to find a way to do it faster though, since privacy is obviously an important issue for many tasks.
In other words, if Petals nodes became 10-100x slower, Petals would lose its competitive advantage over simpler methods that don't communicate over the Internet.
So you can get predicted text that looks "coherent". Then what?
There is literally no place to add logic. Neural net-based language models are impressive, sure, but it's not hard to see how useless they are.
The only time their output is logically coherent is when they are lucky, and that seems to happen often because most of their input was logically coherent to begin with.
And as I said, it's very impressive.
And it has some usefulness: essentially it's an alternative to reading through many pages/posts of StackOverflow and Wikipedia.
But it doesn't know anything. It has no clue whatsoever whether it is correct or incorrect. It only makes guesses. The only reason there is useful output is because that output is a transformation of useful input.
There is no logic. There is no way to introduce logic. There is no way to filter it through logic.
If some coherent mixture of the ML's training datasets already contains the answer to your question - like literary or code examples, definitions, etc. - then the output will be useful. Otherwise, it's just wrong, and sometimes unexpectedly so.
The output of chatGPT (or any other ML-based NLP) can only be as correct or knowledgeable as the data it is trained on; and it will practically never even match that level, because it is only mixing words by semantic popularity, never by logical relationship.
Even if we assume the technology is useless in its current state, it is still incremental progress. Could we have predicted 10 years ago what neural networks would be capable of today? Now, tell me what neural networks will be doing in 10 years. If you think you know the answer with any degree of certainty, you're probably deluded.
We can get coherent (understandable) output all day long, but we can never introduce logic.
ML-based NLP is a semantic word-guessing machine. It's based on entirely on how often words show up near each other in the training datasets. There is no room to add logic.
The entire exercise is like a magic trick: impressive sure, but at the end of the day, a fool's errand.
We don't understand how humans do logic. It's entirely possible that whatever structure in the human brain is responsible for handling logic can emerge in a neural network.
If we're talking about what it takes to get to true AGI in the near future, then I agree that a pure neural network approach might not cross the finish line first. I think Stuart Russell made this point in an interview, basically saying that a neural network is a very inefficient computational model and that we could do the same thing much more efficiently if we had the right "good old fashioned AI" algorithm. But fundamentally a neural network is just computing a function so there's nothing in principle preventing a neural network from doing whatever a symbolic system does. It's mostly a matter of efficiency and hardware availability.
What you are telling me is that I should place my expectations for the future, not on the reality in front of me, but on the hopes and dreams you have for the future. That's circular reasoning.
The very reason that I don't place credibility in your assertions is the lack of reason itself: in your assertions, and in what a neural network is.
Neural networks are like dreams. Wonderful only when your intention is to get lost in a swirl of memories. Useless if you want to actually accomplish something.
Knowing the difference is crucial, because that difference can never be taught to a neural network without completely redefining what a neural network is in the first place.
Knowing the difference is literally the thing neural networks are incapable of doing. They don't know anything. They just guess. That's literally the function. In the code. Guess what comes next.
There is no sense pretending sense itself will magically appear out of a guessing machine. Neural networks are nonsense generators, and that is what they are forever doomed to be.
How many people are using ChatGPT, Stable Diffusion, etc. for economically or personally valuable activities?
If (1) is true, then the answer to that question is "zero" or at least "close to zero". Do you really believe that?
If (2) is true, then it is also true to say that transformer models will never exceed today's capabilities by a significant amount at any time in the future. Do you really believe that?
The limitation is inherent in the core design. There is no overcoming. This is not a hurdle or a wall. It's a design flaw.
Is it totally useless to everyone? No. Not completely. It's like a coherent search engine: a way to find data that is close to other data. But "close to" in this case is only "semantically", and never "logically", so that's that.
Is it going to get any less useless than it is? Only slightly. "It" will never get better. The only better version of "it" is a completely new ground-up redesign that doesn't resemble "it" at all.
Language understanding doesn't magically spawn itself as a process on your computer! Someone has to write that program first.
And that's my point. ChatGPT transforms language, but it does not understand it. For that, we will need a different kind of program.
Do you think it's impossible for such a program to emerge as weights in a Turing complete neural network architecture?
You can use and fine-tune them to solve almost all existing natural language processing tasks: machine translation, recommendation/search, text classification and summarization, code generation, etc.
You can use them to transform already existing text and code (the training datasets); but you can never do more than that.
There is no room in the ML algorithm to introduce logic. It's doomed to forever be a guessing game; and the resulting guesses will always be limited by the information it is fed to begin with.
The only reason chatGPT is so impressive is that it is transforming human conversation that itself is impressive (except that we were already aware of it). The code generation, literature, and definitions, etc. it outputs are all just rephrasing the written code, literature, and definitions that it was given as training data.
It's effectively no more than a sleight-of-hand. Flashy and impressive, but never anything more.
I find these emergent phenomena pretty interesting.
The "emergent phenomena" can be trivially explained by the input they are giving it.
They are not using a dataset that contains an equal amount of "correct" and "incorrect" responses. They are using datasets of human communication, which are obviously filtering for "correct" data. We get things wrong occasionally, but that is quite rare relative to what we get right. We can't even structure a sentence without getting something correct!
If you feed a dog good food, is it really a surprise that dog is healthy? You never fed it poison!
The language model is only returning semantic relationships. The "emergent phenomena" is that most semantic relationships in human communication just happen to also be logical relationships.
But the language model doesn't know that. In no way does it interact with logic. It only interacts with semantics.
If anyone actually bothered to train an instance of GPT or whatever on poisoned data, (i,e nonsensical stories) then you would see that emergent phenomena disappear. But no one is writing the nonsensical stories in the first place, so such a dataset does not exist.
The underlying LM, BLOOM, had a few programming languages in its dataset, so it works at least with Python and C++.
Distributed File Sharing or computation without the whole tokenomics that, while interesting, creates too much attention from scammers.
Decentralized tech would never be where it is today if it weren't for investor attention and the potential for gains. We just have to separate the wheat from the chaff, and remain vigilant for bad actors.
I’ll bet blockchain is only as popular as it is because of the money. But other forms of decentralization like Mastodon or Matrix are pretty separate from the whole crypto sphere
Federated platforms appeal to the privacy-oriented "f** big tech" mindset, which is pretty common in the hacker & FOSS crowds. I'd put it in the same category as VPNs, E2E messengers and TOR.
So VPN are not really in the same category
Big corps only invest in blockchain because of the buzz words that are used as marketing by the consulting firms to sell their "expertise" and by VCs to sell their companies.
Sure they hope to gain some money, like luxury brands wanting to sell to crypto-billionaires. But crypto was a useful toy, then Ponzi scheme and now it's a closed loop. How long will the bubble last?
Nobody around me ever uses any of it. Old p2p networks (gnutella, kademlia, emule) had way larger impact on society 20 years ago.
This created a lot of bubbles. NFTs are already down by a lot, now yield farming (https://www.bloomberg.com/news/articles/2022-04-25/sam-bankm...) just took a big hit from the FTX case. I see way too many "revolutionnary" projects from fresh graduates. There is no way that tens of thousands of inexperienced people with barely enough CS education to pass programming interviews would magically create innovation just because VCs put a ton of money on them.
Also, can you tell me more about where decentralized tech is today? BitTorrent was a revolution as a way of information sharing, Onion was a revolution for privacy and Bitcoin was a revolution for decentralized ledgers.
Starting from that, IPFS is the continuation of BitTorrent with more features and Ethereum is a more efficient (especially since The Merge) and customizable (smart contracts are advanced checkers for write operations) ledger.
But what are the real world applications of those technologies? What are concrete use cases of Ethereum and IPFS besides payments, records and file sharing?
Surely there are exciting progresses to be made on the technical side like zk-SNARKS but how useful will they be to society?
I think we already have all the technical blocks we need. If there is no real-world adoption maybe we should just wait another 10 years before pumping crazy amounts of money.
The real decentralized tech, the one that serves a purpose other than emptying the wallets of naïve crypto-enthusiasts, does just fine without a profit motive. You don't need get-rich-quick promises to get an audience if you're actually doing something useful.
That’s exactly why blockchains haven’t found Product Market Fit.
Investors != Users
Its a mute point whether the whole crypto/blockchain period was a net positive. It certainly made a noisy case for "re-decentralization" given the very real and mostly harmful status quo. One could also argue that it diverted vital resources to potentially dead-end or limited use areas. The recurrent scams may also give decentralization a bad name to an uninformed public that can't distinguish all the different versions.
What matters next is that projects that deliver real benefits to users get attention and traction. Worth keeping in mind that the real trouble starts when you get noticed by vested interests as a potential threat.
They went hand in hand even back in the day: private torrent trackers were all about tokenomics where tokens were the number of bytes you've seeded (uploaded) minus you've downloaded.
I'm not saying it's impossible to imagine distributed file sharing otherwise, but to "guarantee" the availability of (especially unpopular) content, you need some incentive mechanisms either built in to the protocol or externally imposed.
>Please do not use the public swarm to process sensitive data. We ask for that because it is an open network, and it is technically possible for peers serving model layers to recover input data and model outputs or modify them in a malicious way. Instead, you can set up a private Petals swarm hosted by people and organization you trust, who are authorized to process your data.
This is what blockchain and staking tokens is for. (Part of the reason, at least)
You act maliciously, the network slashes your stake. "pinky promise not to do bad stuff" only goes so far... and it's really not far at all. You can trust "trusted" organizations or private individuals, but they have no incentives to ensure that the service works as intended, regardless of intent.
In fact, it adds traceability. And data stored in it can never be deleted. Just to name a few issues.
No, but staking is certainly an improvement over "pinky promise", and it requires a public blockchain.
> issues
I'm fairly sure those are features, not issues. You are free to disagree.
This is a weird statement. Blockchain security is real and it isn't "magic". Blockchain is specifically designed to secure decentralized applications.
> In fact, it adds traceability. And data stored in it can never be deleted. Just to name a few issues.
These aren't issues, these are part of the security model. Traceability is fine here because everything is pseudonymous, if you want to avoid that use a chain that has untraceable transactions with zero knowledge proofs (zero traceability).
> And data stored in it can never be deleted. Just to name a few issues.
Storing data on blockchain is extremely expensive. Only hashes are stored on chain, not the data itself. Hashes are much different from encryption because they're irreversible.
First, automating the detection of malicious acts against sensitive data seems pretty difficult. So this can't be implemented to systematically occur, and has to be determined after the fact by an investigation. Then, if a malicious act has been detected, the stake is slashed (and the acts are reverted where possible).
Is my understanding sound so far?
Because this would mean in any case where a slashed stake is considered an "acceptable cost" to the bad actor, then the sensitive data is fairly accessible -- the stake is effectively a paywall. And raising the stake is a difficult decision because higher stake means less actors and higher risk of collusion.
I mean this is probably fine for a very large public blockchain where detecting malicious acts is not as difficult or where the malicious act is not very profitable, but sensitive data can, depending on its nature, be extremely profitable to exploit (and as I stated, I don't see how it could be easily detected).
With sensitive data, "trusting" an organization only means having a legal agreement or strategic alliance with a third party. In these circumstances the consequences are usually more serious for the malicious actor than the loss of an arbitrary amount of money.
I've seen suggestions to do sensitive (e.g. medical) data processing on the ethereum blockchain from some enthusiasts and I have never been able to understand this beyond assuming they have a insufficient threat model in mind for this kind of data.
BitTorrent style projects are far more restrictive for a lot of applications though. If something is without cost, then it becomes open to abuse.
Take domain names for instance. I would love to have a decentralized name registry, so that no country have censorship power on the _whole_ internet, as we've seen with recent US intervention at the tld level.
DNS is a good example because it's quite trivial to implement with a plain old DHT. The problem though is how do you prevent scammers and squatters in this model?
There needs to be a cost on a distributed database, otherwise after 1 year it will be fully squatted, used as free hosting, store illegal content, DDoS'd for fun, etc.
How to set this cost though, while keeping the distributed nature of this database ? the simplest solution is to let the users decide, over the price of a token, sold by people running nodes, bought by people using the service.
Honestly I love this idea. The problem with crypto currently is that a whole bunch of parasites jump on these tokens to speculate on their price without giving a.. about the underlying utility. This completely screws the price optimum and creates a inflated price bubble, in turn preventing adoption.
We have exactly the same problem with real life systems like food, raw materials and real estate.
Take the DNS example for instance, this was implemented on Ethereum by "ENS", but the price of ETH/gas at the time made a single ".eth" domain name cost something like $500.
Way more in terms of money involved, just slower.
"Massively redundantly replicated"
A distributed consensus mechanism would segment the decisions amongst nodes, not poll for a unanimous response.
Blockchain as such has nothing to do with the costs of a node and incentives to run one
Blockchains are built under the assumption that everyone is selfish and untrustworthy. Which is a decent assumption when building a crypto currency, but that doesn't mean that every system has to run like that.
Typically on a tracker you’re given a currency (although not as sound as some e-coins) and can use that to influence your upload or download statistics, which in turn affect your ratio. Some trackers might employ rules where your user class has to have a certain ratio, or else you’ll lose privileges like certain forums or even the ability to download at all. (The trackers are private and can control which peers you can see)
I see some similarity here to the world of private torrent trackers. You want a Linux ISO, I want a Linux ISO, we're all working towards the same goal. So we're already incentivized to cooperate, without getting money involved. And trackers also have things like minimum seeding ratios to keep people honest. In the case of AI, you and I both want to generate images, so we're also working towards the same goal, so let's help each other out so both of our workloads finish faster. Maybe idealistic, but I think it could work.
Is this also how IPFS works?
In zk proofs-of-computation-result, different nodes can perform different intensive parts of a calculation and send the results along with proofs that those are the correct results. Other nodes can accept the results and verify the proofs with remarkable efficiency, then use those partial results for further calculations. To me it still feels counterintuitive and almost magical that any large, arbitrary computation result can be easily verified without repeating the computation, without the verifier needing much memory or data.
For cryptocurrency blockchains this allows smart-contract (computational) transactions to be accepted with only one node having to execute the code, everyone else just efficiently verifies the proof to accept the state change. As proofs can be aggregated, this scales well: it isn't necessary for every node to run all the verifications, either.
For big, distributed calculations like the article's, the whole calculation can progress using those partial results without having to rely on trust and reputation, and everyone can have high confidence that the final result is what it should be, not undermined by subterfuge or subtly inaccurate contributions.
This is an offshoot of zero-knowledge proofs, as ironically zero-knowledge is not required for these types of applications. Just the efficient verifiability part.
(Fwiw, I am working on large, scalable zk-proofs-of-computation in my spare time, in optimised software and with hardware accelaration, if anyone is interested in discussing this stuff.)
Why counterintuitive? That’s kind of all of cryptography and most of computer science. Take factoring into primes (which has been done for forever): it’s really time consuming and expensive to determine what the prime factors for a number are, particularly if it’s a big number and you know it only has two. That’s because division is very very difficult and time consuming. Multiplication on the other hand is super cheap so once you tell me the prime factors, I can confirm much more quickly whether or not they’re factors.
In computer science, one of the earliest identified computation classes is NP complete which has this property. Eg traveling salesman and knapsack packing problem are examples. It can be insanely difficult to find a path that exists between two cities in a graph under some cost. But if you give me a solution I can easily confirm whether it meets the criteria (global optimality testing is itself NP complete but if you give me a set of solutions you can verify which one is the cheapest).
I’m not claiming that factorization is NP btw. There are complexity classes beyond NP that share this property. https://cstheory.stackexchange.com/questions/159/is-integer-...
Anyway. ZK proofs themselves are super surprising and not intuitive but not because verification is fast but because verification reveals nothing to the verifier about the solution. That’s the mind blowing result.
What I find remarkable is that zk-proof-of-computation works for any kind of computation. On the face of it, it might seem that some computations would resist being compressible that way, but no, it works with anything that can be run on any real computer.
It doesn't depend on what kind of computation, so it has nothing to do with which program, how it's written, the complexity class (linear time, P, NP-complete, superexponential etc), or even on the size of the problem. It doesn't even depend on how much memory the problem requires. You can have a computation that requires terabytes or exabytes of RAM to compute, and the world's largest supercomputer running for a decade: The proof that the output is correct, no matter how much complexity went into calculating it, is still small and fast to verify.
But you still have to do the computation somewhere to get the proof. That's why it's called "argument of knowledge", because the entity constructing the proof must have access to ("knowledge of") the computation.
So it's still about feasible computations. Usual zk-proof-of-computation can't be used to prove things larger than there's a computer able to compute.
That boundary is different from cryptography (and P vs NP), which is more about verifiability of problems requiring exponentially larger time and/or space to solve if you don't have the secrets, so if the parameters are suitable, these are about infeasible computations by any physically realisable computer.
The connection is that that zk-proofs-of-computation are about making proofs of feasible computations, while ensuring it's infeasible to compute a false proof, or to find the secret inputs if there are any (there don't have to be).
(By the way, you may be thinking of discrete logarithm not division. Division is not difficult. In finite fields such as used in cryptography, division can be computed by constant exponention using Fermat's Little Theorom, and exponentiation takes logarithmic time in the size of the field using a repeated squaring method. Division is slower than multiplication, but not prohibitively so; it's used in elliptic curve operations. The hardness of factorising certain numbers is for a different reason than division.)
Of course its abused by shady operators out to make a quick buck, but issuing tokens, when done right, is a great innovation by itself.
That's why you can actually attack and shut down a bittorrent network, by targeting the index servers, that are not massively replicated. I.E. The Piratebay is often down.
As a solution for this, I'll shamelessly plug my small project here, that combines bittorrents with the blockchain as a invulnerable piratebay-like bittorrent index server, called Blockchain Bay: https://github.com/ortegaalfredo/blockchainbay
It's command line, and don't use any tokenomic scams. You pay the blockchain only for the data you need to upload, that is fortunately, very little as bittorrent magnet links are very small.