Google Open-Sources Trillion-Parameter AI Language Model Switch Transformer
infoq.com
infoq.com
How do you know that? Perhaps with this methodology you really need those 3.12TB to reach comparable performance?
[1] provides one with a whole-data-set training method (ADMM, one of such methods). Page 8 contains figure 2(b) - accuracy of training after specified amount of time. Note that ADMM start where stochastic gradient stops.
[1] https://arxiv.org/pdf/1605.02026.pdf
At [2] I tried to apply logistic regression trained using reweighted least squares algorithm on the same Higgs boson data set. I've got the same accuracy (64%) as mentioned in the ADMM paper with much less number of coefficients - basically, just the size of input vector + 1 instead of 300 such rows of coefficients and then 300x1 affine transformation. When I added squares of inputs (for the simplest approximation of polynomial regression) and used the same reweighted iterative least squares algorithm, I've got even better accuracy (66%) for double the number of coefficients.
[2] https://github.com/thesz/higgs-logistic-regression
There's a hypothesis [3] that SGD and ADAM are best optimizers because that everyone use and report on. Rarely if ever you get anything that differ.
[3] https://parameterfree.com/2020/12/06/neural-network-maybe-ev...
So, answering your question of "how do you know" - researchers at Google cannot do IRLS (search provides IRLS only for logistic regression in Tensorflow), they cannot do Hessian-free optimization ([4], closed due lack of activity - notice the "we can't support RNN due to the WHILE loop" bonanza), etc. All due to the fact they have to use Tensorflow - it just does not support these things.
https://github.com/tensorflow/tensorflow/issues/2682
I haven't seen anything about whole-data-set optimization from Google at all. That's why I (and only me - due to standing I take and experiments I did) conclude that they do not quite care about parameter efficiency.
Google Research is pretty big, I used to think like you did but I think it's mostly b/c DeepMind just hogs all the spotlight.
Check out PRESS [0] for example.
But you gave me another point to support my view: PRESS uses stochastic gradient, not second-order method like IRLS.
For instance, on modern modelling problems with non-linearities and change points, it's a lot less easy to do something like IRLS in an end-to-end system, but interesting as a research direction.
The very SGD thing was developed because it was the only way to train something like neural network with small memory and, more importantly, in reasonable time. Multiplying of training time by N (number of parameter) meant having good result in a year, not in a day.
And today we have large batch training with complex synchronization systems to speed up training even more. Which bring us closer to the whole-dataset training and, I guess, second-order optimization as well.
The accuracy would have to be significantly higher that any alternative to justify monopolizing that many hardware resources.
This kind of model could be used as the teacher in a distillation setup too though. Then faster training of the teacher is actually a huge benefit since it speeds up model development iteration cycles.
But even if it weren't practical to use in production in any sense, I'd argue there's value in doing the basic research of exploring design space of architectures in this way. This came out of a research team at Google. It may inspire and inform smaller, more practical architectures.
Part of the justification is the MoE sparsity by design means only a small part of the model will be activated by a given query. Don't think of it as a single giant 1t model, think of it as 50 small models which happen to share a glue layer at the input. So at deployment, you could, for example, keep the gating-layer in RAM and only pull the necessary sub-model off disk as necessary. Or you could shard the sub-models over 51 GPUs and feed the master gating-layer+GPU lots of queries, and each query will be dispatched to a different expert+GPU pair. This could easily be competitive with running a lot of dense models in parallel trying to keep up with the same load.
See https://github.com/google-research/text-to-text-transfer-tra...
https://github.com/google-research/text-to-text-transfer-tra...
The advantage, as they show, is that the model can train to a given level of performance much faster with a fixed amount of computing power compared to an architecture that uses all parameters on every step. This might be because it allows you to have a very large number of parameters that can store a lot more specialized information without incurring as much of a computational cost. Of course the downside is that you end up with a very large model that literally won't fit in a lot of environments.
Researchers at Google's scale prefer a single model where you throw all your data in a single bin and get perfect performance out, no tweaking and no pesky humans required.
https://storage.googleapis.com/books/ngrams/books/datasetsv3...
(I don't remember exactly how much it is, but I remember that the old version was already in the terabytes.)
I want to see how well weights for these models compress, but it will take me some time to run this code and generate some. I'm guessing they won't compress well, but I can't articulate a reason why.
This model has 3.12TB of floats??? That's insane. How do you load that into memory for inferencing?
Alternatively order something like the HP Z8 with 3TB RAM configured, which is only $75k - https://zworkstations.com/configurations/2040422/
It's interesting. It would take ~six years for the Z8 to break even compared to AWS, but traffic into and out of the machine would be $0, and I don't think you're running directly on the metal with AWS, so performance would probably be a bit higher. And then there's storage - I configured, uhh, 120TB of a mixture of SSDs and HDDs. I'm not even going to try and ask AWS for a comparible quote there.
I may or may not have added dual Xeon Platinum 8280s to the Z8 as well. :P
Do you mean six months?
Yup.
Also - think you meant 6 months, not 6 years anyhow :)
And I did mean 6 months, woops. Didn't even notice...
The brain "works" because it's evolved structure matches or reflects reality. It is not about having billions of neurons, but about to have the right structure which matches the environment.
My favourite example is how butterflies evolve pictures of eyes on its wings to scare predators, having literally no idea about existence of other creatures.
It has been evolved because other creatures have eyes, and they are there, of course.
The proper structure of neural networks must be based on such fundamental features, like "most of creatures have eyes" and similar ones.
Brain does not have a flat structure, like a billion x billion matrix. It is more clever and simpler that that.
A language model must be based on the fundamental notion that there are nouns (things), verbs (processes) and adjectives (attributes). It is that simple.
The brain has an incredibly complex architecture, which evolved over millions of years. On top of that, it then develops throughout a human's lifespan. The brain we observe is a "finished product", and even then it has ~150 trillion synapses to do computations [0].
Even massive neural networks have a relatively simple architecture before they are trained. Part of the training process is effectively learning more complex architectures, which are manifested by changing weights.
What I'm getting it is that artificial neural networks aren't equivalent to the brain - ANNs are both learning their own structure, on top of the circuits actually doing computations. They are doing the work of millions of years of evolution, genetics, developmental biology, interaction with the environment etc. Perhaps it's to be expected that ANNs will need orders of magnitude greater number of parameters than a brain.
An interesting development is meta-learning, where we separate the process for learning the architecture (this could be using deep learning, but not necessarily) with the network actually doing computation (equivalent to the brain).
> A language model must be based on the fundamental notion that there are nouns (things), verbs (processes) and adjectives (attributes).
I agree, but how does the brain represent these concepts? Some would argue that ANNs do have these concepts, just hidden away in abstract vector representations. Take the visual system, which has been extensively studied - we see the brain represents contrast, edges, shapes and so on very similarly to convolutional NNs.
[0] It's likely that this number doesn't come close to capturing the brain's complexity, as it doesn't incorporate parameters like long-term potentiation/depression, synchronization, firing rates, habituation vs sensitization, immunomodulation and likely so much more we haven't yet discovered.
Can someone describe what the lexical reality of 3.5T inputs actually means?
I feel like this is 'Deep Memorization' instead of 'Deep Learning'.
Like a Doctor who passes everything merely by memorizing the textbook with absolutely no ability beyond that.
It's like a 'new form of storage and lookup' as opposed to the kind of 'magic algorithm' we usually think of when we think of AI. Or maybe that's just me.
Inputs -> Outputs not Search -> Response
Like if you train AI on a small dataset, it feels like what it's doing afterwords is a 'function' or 'algorithm' using what is in the end some arcane algebra.
But if you train on all the data in the world, with a trillion parameters ... well ... I kind of feel that 'all the data' is in an AI-style datastructure, that we are 'querying' with AI.
But that's just an observers abstraction.
It's a poor comparison, though, since neurons and synapses aren't the same as parameters in a computational network. It's the same trap that news outlets routinely fall into when citing "storage capacity of the brain" in TBs and other such nonsense.
This is not the case with black-box models. They only point to an answer/give a reply without any justification or reasoning behind it. This is actually a very severe problem with black-box models: ultimately they cannot be trusted, because it's very hard to verify whether the learned objective function matches the intended objective function (this is called the alignment problem [1].)
Optimisers tend to produce Clever Hans instances whenever they can, because it's the cheapest and therefore most optimal solution. Even if this becomes obvious from failure cases (e.g. common misclassification in image recognition systems), it's still not obvious which clues the system used that lead to the misclassification.
This is in contrast to a person, who an be queried as to why and how they arrived at their conclusion.
[1] https://bdtechtalks.com/2021/01/18/ai-alignment-problem-bria...