33TB of text data for a 1T-parameter model
lifearchitect.ai
lifearchitect.ai
- best score - don't care about efficiencies (GPT3, GPT4)
- best score for a fixed quantity of compute at training time - good for PhD's and people who make proofs-of-concept (Chinchilla)
- best score for a fixed quantity of compute at inference time - good for people who inference their models at scale (LLaMA, chatGPT turbo)
The article didn't mention the LLaMA scaling laws, where we use more than 20 tokens per weight, more precisely 142 tokens per weight for LLaMA 7B.
> The objective of the scaling laws from Chinchilla is to determine how to best scale the dataset and model sizes for a particular training compute budget. However, this objective disregards the inference budget, which becomes critical when serving a language model at scale. In this context, given a target level of performance, the preferred model is not the fastest to train but the fastest at inference, and although it may be cheaper to train a large model to reach a certain level of performance, a smaller one trained longer will ultimately be cheaper at inference. For instance, although Chinchilla recommends training a 10B model on 200B tokens, we find that the performance of a 7B model continues to improve even after 1T tokens.
What we care about is the best model we could run on our own hardware, not how efficient was its training, that doesn't cost us users anything.
Training cost is a huge (but diffuse) cost because it limits which organizations can be train a model. Such concentration of power greatly affect users.
What does "inferencing models at scale" mean?
The models produced have to be fast and efficient to support capacity/cost, which has some detriment to quality/accuracy.
As for inference if you want just bound from above possible model size, then just take largest model you can allow and train for as long as possible. There is no evidence (yet) that we can hit the ceiling with this one.
OpenAI's GPT-5 that they've said they're not making yet?
Does anyone actually read this guy or is this what they call puffery?
“A contributor to the fields of human intelligence and peak performance, he has held positions as chairman for Mensa International, consultant to GE and Warner Bros, and memberships with the IEEE and IET.”
He doesn’t mean to degrade the guy as he doesn’t even know him. It’s just.. ah, well.
I think these are important details.
We should probably make a new rule for AI article, where the (Month Year) is posted if the article is older than a month.
O̶p̶e̶n̶AI.com knows that they cannot let the DALL-E 2 situation (with Stable Diffusion interrupting their rollout) happen again with the GPT releases and they have to cement their lead with another surprise announcement to throw everyone off.
I'm expecting them to acquire an AI accelerator and hardware company or at least add further investment in that chosen company.
Pre-2022 his books seem to be non-AI related. One of them is "Best: A practical guide to living your best life" and another is a book for gifted toddlers.
An older description of him says "Dr Alan D. Thompson is a world expert in the fields of child prodigies, high performance, and personal development. "
Now he is a "world expert in artificial intelligence (AI)".
Looking forward to see his next world expertise.
For the record, I'm not saying LLMs are not a huge step forward. They are. But they are not -- and cannot be -- the whole answer.
That helps tremendously compared to a tabula rasa.
(That said, DL architectures are obviously wildly different from how the human brain works. E.g. backprop is physically impossible.)
I'm sure there's a lot more to it than this, but maybe one factor that makes humans a lot more data efficient is the multimodal input we receive.
If that's the case, imagine how much better things could get when we train with all the videos, podcasts, radio etc in the world, in addition to all the text out there!
And it's been optimized over billions of years.
You can look at it the other way around - dopamine and other neuro transmitters as poor approximation of backpropagation. It has many flaws for example tight harmful loops ie. addictions.
Majority of brain work is ignoring irrelevant information (attention) and small scale hallucinations (we don't see world as is but slightly hallucinated to keep it stable - ie. they way brain processes blinking <<turns off>>, you can peek at those nuances with ie. optical illusions etc).
One of missing bits in neural nets may be reusing its output as input (embedded in inference itself, not poor mans re-prompting).
Once it's sorted out I'd argue the performance will skyrocket and give opportunity to massive optimisations ie. embedding things like known functions - imagine brain which has known, very narrowed, available functions at its disposal - all mathematical functions on numbers, logic, optimal sorting etc. Imagine if as thinking human you'd have access to accurate functions - the sky is a limit.
Bayes formula as a built in primitive. You don't need to know much of statistics to see how limited humans are at processing information because estimating posterior updates is so expensive for them.
Thinking in terms of raw probabilities would be very alien to most humans, but could easily be technically superior for making plans.
This is only true when considering single performance axis like pixel resolution. When you consider the corpus of power efficiency, jitter resolution enhancement, dynamic contrast, performance per volume, etc. We aren’t close to building something as capable.
It's pretty clear that humans, unlike LLMs, use external sensory data (the only external data we have, when you cut through it all) when they produce speech, as evidenced by the fact that they don't speak falsehoods that don't mesh with their internal data model. LLMs have such a weak model of reality -- it's whatever “sounds right” -- that they speak falsehoods all the time. The only way to give an LLM sensory data would be to encode every sensory experience people have into text.
I don’t have access to any other humans internal data model, but the indirect evidence I do have suggests that they do, in fact, speak falsehoods that don’t mesh with their internal data model for a variety of strategic purposes.
So does Wikipedia. But that's not a good model of the human brain either.
> I don't think it's quite an apples to apples comparison.
Yes. That is exactly my point. Despite the superficially similar I/O behavior, the two systems are very different under the hood.
What weights are you referring to ?
Obviously newborns need to develop and take in stimulus before they "know" anything, so the initial conditions are not sufficient, but they are obviously necessary to make the limited learning useful.
As far my understanding goes to very large degree it is unknown how to model such dynamics, where it would be possible to start with no spiking neurons and evolve effective/stable learning behavior from small number of examples/experiences.
There may be many roads to intelligence. The path biology and evolution took may just be one such path.
Are they coded in DNA, or do epigenetic factors, the environment cells grow in, adjacent cells, their interaction, etc, play a role here as well?
Does this refer to DNA -> MRNA -> to protein or is there some other mechanism here?
Also, since you asked, I would like to mention protein interactions with cellular components and tissues. These can be argued as just proteins interacting with each other but the extreme complexity of these higher levels and the unique phenotypes arising from them make me think they are deserved to be treated as another layer of "data decompression".
A chimpanzee can listen to humans as much as it wants, and it will still not pick up much of the language.
It's deeper than that. Our brains are optimized for a lot general human functions. Learning language is one of them. There's a whole section of the brain dedicated to it and other things.
There's also a lot of vital biological information encoded in DNA/RNA. We are not even close to starting from scratch.
There's some reasons to suspect this is at least partially true, but to what extent is unknown and contraversial.
It doesn't say anything about that knowledge being precoded.
The fact that human brains are capable of learning things that animal brains cannot, is a clear indicator that it is already "trained" to a significant degree.
Human "learning" seems to be much more analogous to "fine tuning and memory storage/retrieval", than actual training.
Why can't a macaque teach a calculus class? Why can't its "empty" brain be taught to do something like that?
If you mean a lot of accuracy, that's obvious and doesn't really need argument. And this new fact doesn't change the argument.
If you mean a more moderate amount of accuracy, this isn't proof either way. Human brains take in less text but they put a lot more processing into it.
But regardless, the architecture of the human brain doesn't have to be the only way to get to AGI (not that LLMs are necessarily the way)
This is assuming we are learning nothing during sleep, which probably isn't true.
By the time a person is 21 years old, they have been trained on at least 1 petabyte of data.
By two years old, about 125TB of data. It makes LLMs look quite good in comparison.
Call 10^10 \approx 2^40 for convenience, and 8000 \approx 2^13, which gives us a 2^53 entropy estimate, or about a petabyte of information as an estimate of what the human brain can store (discounting more exotic theories of memory stored in DNA or some such).
[0] https://en.wikipedia.org/wiki/Human_brain#Microanatomy
[1] https://psychology.stackexchange.com/questions/7967/how-many...
A single pyramid neuron in the neocortex might be more comparable to a multilayer neural net.
https://www.biorxiv.org/content/10.1101/2021.10.25.465651v1....
We don't understand how they work at the subatomic level simply because human understanding of the subatomic world is not complete, but even just at the atomic level a single neuron is massively more complex than anything humans have created.
Going up to the molecular level, even that is staggeringly more complex than the incredibly simple abstractions that make up a neural net.
Is what happens in the brain at the molecular, atomic, or subatomic levels relevant or necessary to intelligence and consciousness? We just don't know yet, but we do know all of that is far more complex and very different from the simple abstractions that are used for neural nets and LLMs.
The back of a napkin calculations in this thread don't even begin to do justice to the tremendous amount of "calculation" or "storage" that happens in the human brain.
In contrast the data we collect through our nervous system is rich and meaningful and far deeper than just raw text. We can even manipulate the environment as we learn to facilitate faster learning eg. pick up a ball and throw it, rather than just watch videos of balls being thrown.
Our object recognition is trained pretty quickly. And we sleep (certainly as a child) more than 6 hours per day. But we don't learn much from just looking at pictures.
> 18 hours of reinforcement learning
You made that up.
> It makes LLMs look quite good in comparison.
So, exposing an LLM to a lot of video will make them understand language?
Even sound alone (uncompressed CD quality stereo) is 3TB/year.
Some deaf/blind people can read braille — they learn fine too.
Much less data than you might think.
Bare in mind most people can run the brain on 2500kCal/day.
And so on.
Which is, in turn, 2900 W·h, or 2.5 times less than one A100 card working round the clock (300W·24h)
The significant thing is the reinforcement. Without it you have to resort to the openai brainlet style 10tb of text training.
If their model had reinforcement built in then you could just plop it in front of a person or on reddit and it would rapidly self-learn on a fraction of the data. Their training model relies on it learning purely based on observations rather than interaction which is inefficient.
I don't care if my model needs an exabyte of RAM if I can just go and buy that much RAM one day.
Yes, we use now magnitudes more oil than we used to 120 years ago. But back then oil was used for cooking and lamps, while now it is used for so many more things. Same goes for data. It is the new oil :).
These LLMs are also trained on arbitrary books. When acquiring new languages, we use educational materials that are specifically created to facilitate an understanding of language. Not just arbitrary books in random order.
Edit: Still, the idea of mapping symbols and words to a many-dimensional space of meanings is a great insight into how mind works. In that space, symbols with similar meaning appear next to each other, and a thought looks like a smooth but intricate shape that separates all symbols into the "insiders" and "outsiders" that, in practice, divide symbols into true/false, good/bad and so on. Such smooth intricate shapes appear in the frequency domain as a bunch of rational numbers, and that's the boundary of what a mind can imagine.
Can you explain why the thought divides symbols into good/bad and how it is projected to rational numbers? (What is the frequency domain of?)
In other words, once you “compress” to thoughts, there actually can be some degree of actual reasoning on “decompression”.
Has any one written the obvious reply to this paper yet?
It's false. It's not true.
The rules and the board state it "learns" are just functions of the token positions *by construction*; it is given the token positions. obviously it learns patterns in those positions.
The paper is such a misfire, it's absurd.
The problem with using a formal system to investigate this issue is that by construction the "distributional hypothesis" is true for that system. Ie., the pattern in distribution of the moves are (a very good model of) the rules and board-state.
But this is clearly untrue for non-formal systems not constructed this way. Eg., text tokens are not distributed like causal structures in the world -- and it's absurd to suppose so.
The distributional hypothesis is obviously false for almost every measurement system of interest: measures are not distributed like their causes. We invent measurement systems to encode information, not to be isomorphic to world-structure.
I strongly disagree with that view, and encourage reading the "Sparks of Artificial General Intelligence" paper[1] and watching the associated video[2], then just experiment with GPT4 using novel information that has been created after its training data cutoff and you'll see it reason about it. It can also output novel creative works about things that didn't exist in its training data.
Sure, there's some "compression" going on in LLMs, but that's far from everything that's going on.
Map the 70 emotions, dozen tenses, handful of modifiers and suddenly the only missing variables of state and desired state will bring about a framework to allow universal translation.
For instance, a recommendation system's output effects what users see, so it effects what they click on. The next training set is statistically dependent on the input of the previous.
There are strategies to deal with it, but in the case of a LLM it seems difficult apart from downweighing everything after 2023 that isn't from a vetted source.
https://en.m.wikipedia.org/wiki/Low-background_steel
tl;dr steel (except for low background steel from sunk battleships) has been pretty much forever contaminated by nuclear weapons.
I'd posit that data-wise, ChatGPT is the equivalent event.
The RLHF reward model is effectively a good/bad content filter. Human==good AI==bad doesn't always hold, there is good AI generated content and bad human written content out there, we should filter by content quality and novelty, not origin.
If I had the resources I would turn GPT-4 onto its own dataset to mine all the facts from all the source documents. Then flip the index grouping by fact, and for each fact it should write a page listing the supporting and contradicting evidence. This would be super useful in determining - 1. existence of a fact, 2. its controversiality level and 3. the distribution of answers.
With these pieces of information at hand GPT would get much better, it would at least stop hallucinating. This fact index would probably double the size of the training corpus as well. Google, with its years of search logs, is the best positioned to create this fact index. They already had simpler attempts at creating a knowledge graph for a long time. But if they don't move to exploit the data they sit on, others will take the lead anyway. It costs only money to run so much text over the model, no slow & expensive human involvement necessary.
It's almost certainly a problem for LLM development, just like it is for humans. Humans generate all sorts of stuff, some of it being absolute bullshit, and it does seem to cause problems for other humans. Yet with effort it still seems to be possible to cut through the bullshit in a lot of cases, so it's likely not an insurmountable problem for LLM development either.
We have to be very careful with the "factual" content of old non-fiction works. So much of that, from history, to medicine, to biology, etc, turned out to be straight from their writers' imaginations. If an LLM considers such books to be no different from modern books on these subjects it would get a very skewed view of the world.
Imagine asking an LLM for a medical diagnosis and it responding with something about the humors.
For this, large amounts of very old text would be extremely valuable.
What about data that's sitting around and isn't supposed to be public? If training data gets scarce, does a market for small-medium sized data emerge? Like old homework papers, internal company documents, etc?
Maybe there already is and I just haven’t looked hard enough.
I understand that there are reasons to expect improvement with a broader set of inputs - the model would never understand slang and other key language components by being trained only in academic papers. I wonder whether the Shakespeare example would be worked out by it somply occupyinf a novel high-dimensional space because of its uniqueness, or whether there is a signal to noise issue here.
You often train for 100s of epochs. One epooch means one pass over the training set.
Or would it? Are there estimations abut this?