Since, I assume, humans don't need this much training (?) this area would seem to be ripe to explore — can you achieve similar training with a fraction of the data needed for GPT-3.
Since, I assume, humans don't need this much training (?) this area would seem to be ripe to explore — can you achieve similar training with a fraction of the data needed for GPT-3.
But that does I think shed light on why multi-modality is so promising; beyond additional use cases, it could make the models better at existing text tasks too when they train on all that video data.
And I think it also reinforces why everyone is so hyped for reinforcement-learning techniques like the self play that alpha zero used to come LLMs.
And their tokenizers are better - that's the primary reason why ChatGPT is so bad at math, its network literally doesn't (and can't!) comprehend numbers longer than the token size by definition.
Additionally, humans can cross-reference information with information they already know at learning time. For example, an AI "reads" some climate change denier politician's speech - how is it supposed to know that the politician is a climate change denier and downrank their opinion on climate change as a result, at least without a human having pre-classified the data?
* in gpt-2’s case, you get 1,000,000 -> 16, 11, 830, 11, 830, 198, 49 And 1000001 -> 388, 486 which means the model will see a sequence of noise embeddings entirely different when training. That means seeing 1,000,000 any number of times will not really help the model prepare for when it sees 1000001 for the first time. You can play around with https://tiktokenizer.vercel.app/?model=gpt2 to see how tokenization of numbers has been improved by different models to try and mitigate this.
That, and another thing: A LLM fundamentally boils down to a statistical repetition engine. That means, if it has never seen the input - say, some random large numbers -, it will hallucinate garbage if you ask it a question ("multiply 245358943543548 by 438574589675486789654"), whereas a human would recognize where the boundaries of the numbers are and plug it into a calculator (or, some particularly big brained people would actually do it in their mind) to figure out the result.
Some AI or pseudo-AI (dunno what Wolfram Alpha uses under the hood) has to have a "front" AI to recognize such questions and redirect them to a calculator, but in my eyes that's just cheating for a very specific case and not enabling general intelligence.
If multi modality gets us through this phase then you are right in your analysis. Let's see what come out in the next 5 years
I just think we’re really making a silly comparison when we compare a stream of text to the multi sensory experiences humans have directly, reducing what we are doing to just reading words off a screen or page.
I.e. of the GB of data I get every day, most of it is visual data of my apartment which I've already seen 10000 times and audio data of the same shit my girlfriend says every day. I'm not ingesting the whole of Wikipedia or every book written, which is far richer and more varied.
Have you tried increasing the temperature?
The experience of breathing and staring at a wall has an incredible amount of data that has to be learned before it is noise and is boring, and I’d argue it would take many Wikipedia articles to describe it fully.
And on the novelty function- SGD does update things based on a loss and how unexpected things were to the model. But that’s very crude compared to what humans are doing, where everything gets the same learning rate multiple, and we feed nearly random data to the model. There is a lot of evidence that if you pick your training set carefully you can train a lot faster, and I think humans are picking their own training set.