DeepMind’s New Language Model, Chinchilla
marktechpost.com
marktechpost.com
This seems highly unethical, and I'm surprised how they continue to operate.
https://www.similarweb.com/website/marktechpost.com/#overvie...
Cloudflare will forward it to their host, I believe, who will then ask that they remove the infringing material, or provide a counter claim.
Would it also annoy you if they screwed up the interpretation of what you wrote? Is the alternative less reach of your work? For hard core research the tradeoffs are tougher it seems. If it is just a matter of non-nevermind, thats strictly messed up.
https://www.lesswrong.com/posts/midXmMb2Xg37F2Kgn/new-scalin...
> On March 29th, DeepMind published a paper, "Training Compute-Optimal Large Language Models", that shows that essentially everyone -- OpenAI, DeepMind, Microsoft, etc. -- has been training large language models with a deeply suboptimal use of compute.
> Following the new scaling laws that they propose for the optimal use of compute, DeepMind trains a new, 70-billion parameter model that outperforms much larger language models, including the 175-billion parameter GPT-3 and DeepMind's own 270-billion parameter "Gopher".
Anyway it distracts from the point so it's not relevant.
(parent)
> the fact that data and compute need to scale proportionally seems… like a big point in favor of NNs as memorizers/interpolators.
(gwern)
> Surely it's the opposite? The more bang you get out of each parameter, the less it looks like 'just' (whatever that means) memorization/interpolation. When you needed to increase parameters a lot, disproportionately, to cope with some more data, that does not speak well of abstractions or understanding. (If I can train a 1t model to get the same loss as what I thought was going to take a 100t model, why would I think that that 100t model must be memorizing/interpolating less?) Let's take your claim to its logical extreme: suppose we discovered tomorrow a scaling law that made parameters near-constant (log, let's say); would that not suggest that those parameters are super useful and it's doing an amazing job of learning the underlying algorithm and is not memorizing/interpolating?
On the other hand, https://twitter.com/model_mechanic/status/151297688118364569... is admittedly pretty magical, even if the basis of that magic is memorization and interpolation.
From your comment a reader will understand that you think they're just memorizing and interpolating and that you disagree with gwern on this point, but you've given your reader nothing that argues in favor of your position
Why should someone believe that models are just memorizing and interpolating?
People arguing about this are basically all speaking ambiguously, in ways that tend to either make an apparent disagreement when there is none, or hide the location of the actual disagreement.
It is true that a piecewise-linear function, within any linear component, will have any convex combination of some points be sent to the corresponding convex combination of where it sends those points.
It is not true that a piecewise-linear model trained on a set of data points will produce only outputs which are a convex combination of outputs that appear in the training set.
These are both obvious.
If one person takes the former to be what “it just interpolates between points” means, and another takes the latter to be what “it doesn’t just interpolate between points” means, and then they argue about which of them is right, then both are being silly.
I’m not saying that this is literally what is happening. This is meant as a metaphor for somewhat more sophisticated/reasonable interpretations of “(doesn’t) just interpolate(s) between points in the data set”.
_____________
A model trained on images which produced only convex combinations of images in its training set, would clearly be producing what could be called “interpolations between images in its training set”, and taking convex combinations of images is unimpressive.
This is obviously not what today’s image ML models do.
And, of course, you aren’t claiming that they do.
______
I should speak plainly.
Much of where the disagreement is, or is hidden behind, is disagreement as to the meaning of “just interpolation”.
At one end, “just interpolation” could refer to “take the Voronoi cells of the inputs in the training set (or maybe the dual of it, whatever), and at runtime, find the nearest neighbors of the point and take the linear combination of their assigned outputs, weighted according to the distances to the point.” This would certainly be “interpolation”, and is not impressive, calling it “just interpolation” seems quite fitting. However, it is obviously not what ML models do.
On the other end of the scale, “interpolation” could be interpreted as meaning “any process whatsoever, except that the process is required to be based primarily on the training data, with the process being generated mechanically from the training data, of computing an output for a given input.” And, certainly today’s ML models satisfy this description, but with this description the moniker “just” seems, inappropriate. It is like saying “just a process”. Well, yeah, everything is a process.
__________
It seems to me like much of what the disagreement ought to be about (which might not be what it is about) is along the lines of, how many conceptual layers of something are captured? Like, say something modeled images of faces as “linear combinations of images from this list of images of faces”. That’s an extremely basic thing. Then, very slightly more, would be something that determines positions of facial features, and then does stretching etc. of images to make them line up with the image to reproduce, and then does linear combinations. Then, suppose something takes like, the parts of the images of the face which are just skin and not like lip skin or eyes, and takes local averages of this in a number of general locations of the face (relative to locations of facial features), and takes principal components of this across the training set (with principle components perhaps corresponding to perhaps, 1 or 2 for skin tone, and then directionality for the lighting in the image, and maybe a component for how shiny the skin is).
A model which represents a face in terms of variables which we can interpret as things like “position of eyes”, “skin tone”, “lighting”, seems notably less in the direction which one might call “interpolating” than one which just lists a coefficient for each image in the training set (or each principal component of images (taken as plain vectors) in the dataset). And, of course, one can go farther than this in this direction. And the further one goes in this direction, (so, like, the more that what the individual images in the training set tell the model is “here is more data about an overarching pattern”), the less it seems like what one might be inclined to call “just interpolation”.
No, and I didn't claim that. I said that, outside the training sample, the model is linear (or quadratic in the case of transformers, thanks for pointing that out) Whether linear or quadratic, a model that has a fixed structure outside the training sample, will obviously not fit data which lies far away from the training sample - i.e. it will not extrapolate. This isn't controversial - it's just something people like to forget about.
>A model trained on images which produced only convex combinations of images in its training set, would clearly be producing what could be called “interpolations between images in its training set”, and taking convex combinations of images is unimpressive.
True! I should have clarified that it's not linear interpolation in pixel space (or input space generally), but interpolation on the latent manifold. This is where the power, as well as limitations of deep learning come from. It's definitely non-trivial to identify the latent manifold of data - different dimensions of the manifold may sometimes even correspond to independent components, as you mention (position of eyes, skin tone,...) (though empirically, finding disentangled latent codes is mostly a function of the random seed).
How does an NN process a new input? It maps the input to the latent manifold.
In the input space, it will be some highly non-linear, non-trivial combination of points, which in terms of Euclidean distance in the input space, could be arbitrarily close or far away.
In the latent space, the output will be some convex combination of nearby points.
Here's the kicker - even if your problem happens to be well-modeled as a continuous, low-dimensional manifold embedded in a high-dimensional space (and many, many problems aren't), and even if you manage to obtain a super dense sampling of input space, so that the manifold can be well-approximated (which is impractical or impossible for most problems),
you will never be able to generalize beyond the data distribution.
Our brains don't stop working as soon as conditions are slightly different from what we've seen before. If there's a slight fog on a Stop sign, we can still see a stop sign. If the Go board is 9x9 rather than 19x19, we can still play Go. If we can play Starcraft on one map, we're pretty much as good on a different map, we don't need to relearn the game over the next several thousand years.
How come? Because we aren't just latent space interpolators. We can extrapolate.
So, is that 280 billion bytes of just parameters?
In my mind a parameter is: language, dialect, perhaps context parameters (food, dinner, lunch, travel) and if we than talk about language and audio perhaps sound waves, gender.
Or are context parameters which gives you insight? Like a billion of parameters are literally something like travel=false, travel-europe=true people speaking=e, age, height,
I see tx
They're too abstract to assign much meaning to individual parameters, as our understanding of why their values are exactly the way they are is extremely limited.
A parameter is a "weight" in this case (the lines drawn from neuron to neuron). The neurons are effectively runtime values or "activations." Parameters (weights) are updated during training and then set as constant during "inference" (also called "prediction").
There's unfortunately a ton of jargon and different groups use different words almost exclusively.
For example, our recent paper "Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer"[1] shows that even learning rate and initialization used by existing models are deeply wrong. By just picking them correctly (which involves some really beautiful mathematics), we can effectively double the model size of the GPT-3 6.7B model (to be comparable in quality to the 13B model across the suite of benchmark tasks).
Large neural networks behave in a way we are only beginning to understand well just because each empirical probe of any such model is so much more expensive and time consuming than typical models. But principled theory here can have a lot of leverage by pointing out the right direction to look, as it did in our work.
But as far as I know, the more precise predictions of optimal batch size aren't used much, probably because it's expensive to measure accurately, or because the predictive equation isn't accurate enough to begin with. I wonder if we can "transfer" the optimal batch size from a smaller setting (smaller model or data) to the full setting, like in our paper. This would make it much more practical.
I think I support this in principle but it seems like the scaling curves keep going so it's easier to just make larger models with more data.
>Looking forward to a group like Eluther or Hugging Face releasing a version of this
Both of those groups have access to dozens if not hundreds of Cloud GPUs, I'd hardly call them small.
It would be impossible to replicate these models as say an independent researcher or even in an academic research group outside of maybe Stanford/Berkeley/MIT/etc. and I'd even doubt their ability to replicate models like this based purely on Cost alone.
It is if you outperform it with much fewer parameters
In any case, during training (where the model is run in possibly large batches), and even during inference, the size of the parameters is completely dwarfed by the intermediate tensor representations.
What makes you say this?
Something similar applies for RNNs (same weights applied at each element of a sequence), GNNs and transformers (same weights applied at each pair of data).
And it can't cache tokens because all tokens are evaluated in the context of all the other tokens, so they don't have the same representations when they reoccur at different positions.
But, yes, tokens are chosen one word as a time based on the previous content, similar to earlier auto-completion algorithms.
Can you say that about any other task in ML? When Inceptionv3 came out I was able to run the model pretty comfortable on a 1060. Even pix2pix and most GANs fit comfortably in commercial compute, and the top of the line massive models can still run inference on a 3090. It’s so unbelievably ironic that one of the major points Transformers aimed to solve when introduced was the compute inefficiency of recurrent networks, and it’s devolved into “how many TPUs can daddy afford” instead.
However the upside to pocket sized intelligence will eventually win out. It's just a question of when someone will scrape together the required investment.
If you mean putting the model weights into gates directly, it’d be useless because users would get bored of the model as soon as they figured out what its style looked like. Also, large models can memorize their training data so eventually you’ll get it to output something copyrighted.
Is there much more data out there than what they’re already using?
- 500M tweets per day
- 30 words/tokens per tweet
- 40% of all tweets thrown away due to being duplicate/spam/bots
= 9B tokens generated per day
It's clear that these models have orders of magnitude too much data already.
It somewhat reminds me of the proposals for larger and larger colliders in the hopes of seeing new physics that is always one collider in the future.
I agree with your main point, but think this analogy isn't an apt one. If you want to see what particles are created at higher energies you kinda need the bigger particle accelerators. (This isn't to say that we shouldn't be investigating lower energy collisions, but at a certain point you do need "bigger colliders" to see new things)
>It's clear that these models have orders of magnitude too much data already.
Seems like a strange claim. The scaling laws are showing that you can still make gains with more data and more parameters.
>It somewhat reminds me of the proposals for larger and larger colliders in the hopes of seeing new physics that is always one collider in the future.
This is literally true though, couldn't find the Higgs without the LHC and most GUT candidates would only start being ruled out at high energy levels.
But then we’ve given up on matching human intelligence which is all about working efficiently with small training data, and certainly training a human does not need anywhere near as much data as GPT-3.
GPT-3 was interesting as a proof-of-concept of what happens when you use a gigantic amount of training data. We don’t need a bigger one until we can figure out how to make a smaller one that is just as effective.
If scaling laws are telling us to keep putting even more training data into the thing, then the conclusion should be that the architecture is just not working out.
I don't think we should really take so much inspiration from the brain. We didn't make airplanes work by building bird machines so why should we do that here.
>GPT-3 was interesting as a proof-of-concept of what happens when you use a gigantic amount of training data. We don’t need a bigger one until we can figure out how to make a smaller one that is just as effective.
This feels like a non sequitor. We can certainly keep making larger models and we will, because we can continue to make performance gains doing so.
>If scaling laws are telling us to keep putting even more training data into the thing, then the conclusion should be that the architecture is just not working out.
I don't think anyone in the field would agree to this point. Researchers see an easy avenue to gain better performance so they take it. Deepmind's model shows you can get similar results with more refined architecture, but this was released well after GPT-3. When teams significantly advance the state of the art with a much smaller model I think we should take notice but that hasn't happened yet.
It’s not that we should mimic the brain’s implementation, but we should certainly strive to match the brain’s capabilities. One of its outwardly observable capabilities is that it is extremely efficient in the size of the training data set it requires.
Efficiency isn’t an implementation detail, it’s definitional to what “highly intelligent” means.
GPT-3 is not an airplane, it’s a zeppelin. Zeppelins also have scaling laws dictating that a zeppelin should be very very large. Building bigger and bigger zeppelins is one thing, justifying expending resources on gigantic zeppelins by stating the scaling law and concluding that a jet aircraft will magically pop out if you build a big enough zeppelin is quite another.
But generally I think the better analogy is a rocket ship. If we can still go higher and faster with more fuel we should try to do that before we worry about engine efficiency. You have to get to the moon before you can colonize the galaxy.
I have a toy disproof for your claim that this is clear.
Imagine that you are training a ML system using oracle access to Mum. The ML training system can request 10 million representative samples of Mum output, and then we could judge if the ML system has adequately reproduced Mum.
Now also imagine that Mum frequently tells people that Mum knows a 23 letter secret and while mum won't tell people what is outright, she'll answer queries like if a guess is lexographically higher or lower. We could even imagine that the ML has seen Mum's side of some interactions with her doing that.
Would the ML know Mum's secret? No.
Would a child that could interact with Mum? Yes-- after at most ceil(log_alphabet(23)) queries at most, if the child is efficient.
Learning in an interactive context is not the same as learning from written material, so you can't be sure that the fact that children learn english from less text means that a non-interactive ML system could english from the same amount. Q.E.D.
Now, if someone figures out how to efficiently train these natural language models with reinforcement learning...
Consider that a human adolescence is ~9.46x10^6 minutes and a fast speaking rate is ~200words/minute. That sets an upper bound of 1.9 billion words heard during adolescence. ie: human adults are trained on a corpus of less than 1.9B words.
To some extent, more data can offset worse models, but I don't think that's the regieme we're currently in. GPT-3 was trained (on among other languages) 181 billion English words - or about 100 times more words than a human will hear by the time they reach adulthood. How is the human brain able to achieve a higher level of success with 1% of the data?
1. https://github.com/openai/gpt-3/blob/master/dataset_statisti...
Deaf-blind authors would beg to differ.
But yes, a human brain is exposed to lots of other sensory input, and we know from other research that multi-modal models can learn shared representations that benefit from the knowledge of each domain.
In Transformer's favor, at least, they are far closer to tabula rasa than the human brain is and likely have to dedicate a lot of their training time to things that are otherwise "baked" into human brains. For example, humans come pre-packaged with V1 and V2 as part of their visual system, but CNNs and ViTs have to learn those filter packages from scratch.
I agree with you though. Human brains are able to take single instances of experiences and build a wealth of understanding from them in ways that even modern Transformer architectures are not yet able.
All those conversations in the shower were actually regularizers!
The most obvious answer is "the human brain uses a shit-ton more compute", for 18+ years as well.
We spend data, which we have in abundance, to save on compute, which we do not. Even at the most generous low-end estimates of the human brain's computing power, we are only barely there; on the high-end estimates that people in love with the ineffable mysteries of the brain love to cite, we are multiple orders of magnitude away from even the biggest supercomputers matching the brain. So no matter which way you slice it, we are extremely compute-poor.
Feeding a lot of data through an extremely lightweight optimizer like first-order SGDs is one way to cope with lacking compute: https://www.gwern.net/docs/ai/scaling/2013-bottou.pdf Bottou asks why (even in 2013!) is SGD so hard to dethrone when we can empirically see plenty of optimizers like second-order gradient descent algorithms which can beat SGD quite solidly? His observation is that while they are much better than SGD in terms of iterations or _n_, they lose in compute/wallclock because SGD can just go-brrrr through the data much faster than they can.
Some quick googling gives this:
- Generation of an action potential seems to use ~2.5×10^−7 J [0]
- The brain consumes around 20W during normal activity
This seems to imply that there are around 8×10^7, call it 10^8, activations per second [1].
Apparently, the average neuron has 1000 synapses. Let's say each synapse requires 10 mulacc operations per activation. Doing that math gives about 10^12 FLOPs/s [2].
Integrate that over 18 years, and you get roughly 5.7×10^20 FLOPs [3].
PaLM required 2.56×10^24 FLOPs to train [4]. So, we have (way more than) enough compute, we're just not using it efficiently. We're wasting a lot of FLOPs on dense matrix multiplication.
There's plenty of wiggle room in these calculations. I checked over the math, but I'd appreciate if someone would let me know if I've missed something.
[0]: https://link.springer.com/article/10.1007/s11571-018-9503-3
[1]: https://www.wolframalpha.com/input?i2d=true&i=Divide%5B20+W%2C2.5%E2%80%89%C3%97%E2%80%89Power%5B10%2C%E2%88%927%5D+Joules%5D
[2]: https://www.wolframalpha.com/input?i2d=true&i=Power%5B10%2C8%5D+Hz+*+1000+*+10+flop
[3]: https://www.wolframalpha.com/input?i2d=true&i=Power%5B10%2C12%5D+Divide%5BFLOP%2Cs%5D+*+18+years
[4]: https://blog.heim.xyz/palm-training-cost/#:~:text=PaLM%20(2022)-,2.5e24,-10x*** C
CH
CHI
CHIN
CHINC
CHINCH
CHINCHI
CHINCHIL
CHINCHILL
==> CHINCHILLA
HINCHILLA
INCHILLA
NCHILLA
CHILLA
HILLA
ILLA
LLA
LA
AAlternatively, is this a joke and the "recursive, selective acronym" can be used to justify any word?
A
AR
ARB
ARBI
ARBIT
ARBITR
ARBITRA
ARBITRAR
==> ARBITRARY
RBITRARY
BITRARY
ITRARY
TRARY
RARY
ARY
RY
Y
Yup, seems it works for any word.