Learning@home hivemind – train large neural networks across the internet
learning-at-home.github.io
learning-at-home.github.io
So if you're forced to trust all of the peers, how is this better than a cloud? Who out there is training models for purely benevolent reasons (i.e. non-profit seeking) and can trust random nodes? If not for purely benevolent reasons, who out there is going to donate CPU time to training your model, essentially writing you a blank check?
edit: Ah, just saw the "what it isn't for" section - apparently not.
Like, say, if one selected some of the images in the batch to use a different label for when computing the gradient, but still using the right label for most of the images in the batch?
Neural networks automatically error correct so they will be robust to some amount of corruption.
Reason I ask is, you can absolutely bet that states will attempt to cause interesting failure modes in other states’ A.I. — imagine if self-driving cars had a literal blind spot for the fifteen senators most aggressive towards [rolls dice] Agrabah?
There could be a ranking system. Rank I verifies 100% of submitted jobs. Rank II verifies 50% of submitted jobs. Rank III verifies 25% of submitted jobs. Rank IV verifies 12.5% of submitted jobs. Rank V verifies 6% of submitted jobs.
After 250 jobs you go to Rank I, after 500 to Rank II, after 100 to Rank II and so on...
If you submit a job with incorrect results then you lose your account and all unverified jobs submitted by that account are then verified. If you're a honest person you'll just create a new account, if you're a malicious actor then you just wasted a lot of money on nothing because doing a bait and switch will result in your malicious jobs being discarded.
There is still an opportunity for denial of service by creating lots of reputable accounts and then letting them go malicious all at once. You'll have a large backlog of jobs to verify.
How about tying the training and consumption of the model together. An internet scale tool with a focused goal, like Alexa/Mycroft for speech and intention recognition, that trains a distributed model while pushing improvements back might be more successful in getting adoption.
And tons of people participate on Folding@home which also uses GPU these days.
But you are right about the why? I admit I personally do Folding@Home most of the time, because EVGA gives you up to $10 to spend at EVGA a month by participating, or Boardgamegeek gives badges and some currency to to spent on their site. Which in EVGAs case, has made it so I basically only buy EVGA products now as I have a bunch of EVGA bucks to spend there.
Even if scaling alone ends up solving everything (I doubt it), I'd still feel that very significant improvements ought to be possible from an algorithmic perspective. (I realize that's largely baseless, but for some reason I just can't escape the feeling that current algorithms leave a huge amount of potential on the table.)
https://arxiv.org/abs/1803.03635
https://ai.facebook.com/blog/understanding-the-generalizatio...
https://ai.googleblog.com/2019/08/exploring-weight-agnostic-...
However, this doesn't directly make my point about sample efficiency of today's algorithms compared to humans less valid. Although what I'll give you is that with smaller networks the required sample size is expected to shrink (due to curse of dimensionality). Although the expressiveness is clearly harmed by the reduced parameter count/altered network structure which possibly reduces the ability of the network to perform well for certain tasks.
I think it's important to clearly make a distinction between the required amount of computation and the number of data samples that are necessary when talking about scaling up existing methods. Compute is "cheap", while data isn't.
As a side note, I think the usefulness of the lottery ticket hypothesis is mostly about the ability of random initialization to already give a hint about the quality of the 'prior' that is encoded by the network structure. Useful for less computationally intense architecture search as also suggested by the papers and a paper by Andrew Ng on this topic.
Actually that's not the point. Pruning typically results in networks that still perform well but are harder to train. The idea is to explicitly search for a subnetwork (via pruning) that is easy to train.
> Although what I'll give you is that with smaller networks the required sample size is expected to shrink (due to curse of dimensionality).
I'm not so sure about that either. From (https://arxiv.org/abs/2001.08361):
> Larger models are significantly more sample-efficient, such that optimally compute-efficient training involves training very large models on a relatively modest amount of data and stopping significantly before convergence.
My second and third links are also important! The second talks about generalizing "winning" tickets across other datasets and optimizers. The third talks about weight agnostic neural networks, which in a nutshell are still capable of more-or-less performing a task even with _randomized_ weights.
Weight agnostic networks have a lot of parallels to wildlife that is capable of certain behaviors required for survival effectively immediately, before there's been a chance for significant learning to take place. This is the counterpoint I was referring to - an equivalent phenomenon could explain (at least partially) why humans require so much less data when learning.
They state "smaller network, same test accuracy, with similar number of iterations". So it seems the original network size wasn't necessary for best test accuracy, and compute requirement is reduced only because it's a smaller network. Sample efficiency isn't increased according to https://arxiv.org/abs/1803.03635.
Good performance with random weights seems to indicate good 'priors' encoded in the network. Like how convolutional networks encode the prior of translational invariance and hence it being a naturally good performer on image inputs/tasks.
I think the parallel to "wildlife that is capable of certain behaviors ... before there's been a chance for significant learning to take place" is that priors are also part of biological intelligence. I.e. brain structure at birth enabling certain surivival oriented behaviors.
Hence, I'm optimistic about transfer learning which could happen through _both_ better models (priors that generalize well) and pretrained weights (possibly partially pretrained, i.e. just initial feature extraction). Either could potentially provide a better starting point from the 'how many samples are necessary for good performance on a variety of tasks' perspective.
The point is that either way information needs to be added for performance on tasks to increase. Doing that in a task specific way by using today's algorithms and a billion samples doesn't seem like the right approach. Finding algorithms, models or specifically perhaps neural network architectures (including training procedures, regularizers, loss function, weight tying) that generalize across tasks without needing many samples due to their informative priors seems the way forward to me. That's _not_ a naive scaling of today's algorithms to larger and larger training sets. Which was the point I was trying to make.
You need a lot of samples because you're starting from scratch with each network. If you had one super NN that is equally powerful to a bunch of small networks then you would have a network that can easily generalize because it can use existing data as a starting point. The amount of existing data that is useful to an unknown task grows with the size of the NN.
An NLP NN for English could be combined with an image recognition NN. Since the NLP NN already has a concept for "cars" it only has to associate its already learned definition of "car" with images of cars. If you have separate NNs then you will have to teach the both NNs what a car is twice. With small NNs there will always be some redundancy and that redundancy is a fixed cost.
But a human isn't trained from scratch. Babies go through huge amounts of unsupervised learning to build up a basic vision and language framework.
Training a neural network to recognize dog pictures is like connecting electrodes to your tongue and trying to do the same. Rudimentary "vision" (very small resolution) has actually been demonstrated this way in human experiments, but you definitely need more than a few examples.
A fairer comparison is: can giant pertained NNs learn to generalize with few examples, and the answer seems to be yes.
He is focused on the gaming space, but with this, the data science space might make more sense.
[1] https://en.chessbase.com/post/tutorial-how-does-the-engine-c...
(1) distribute data chunks as you train using more conventional bittorrent systems (e.g. https://academictorrents.com but internal) (2) since you most likely use raw unlabeled data (e.g. just text), peers can crawl it straight from the web
I mostly work with audio, where individual examples are ~2MB, so the dataset sizes get very heavy quickly.