Textbooks are all you need
arxiv.org
arxiv.org
Now this wouldn't be possible without the high quality synthetic dataset produced by GPT(1B tokens) but this is more evidence in line with Tiny Stories (https://arxiv.org/abs/2305.07759). That is, LLMs only need to be so big (both data and parameters) to learn the total sum of human knowledge (and deal with trash data).
It’s not clear from these results but the paper seems to at least imply that the importance of the synthetic data is to “unlock” the pre-training data.
My takeaway as a non-expert is that this is a good result for small and efficient models focused on well-defined domains, but a neutral or maybe even a bad result for a model that displays general intelligence.
Even for fields that lend themselves well to being converted to textbook format there is often a tradeoff between accuracy and conciseness. The more you refine, the more nuance you throw away. It seems like a Very Hard Problem (tm) to know which data are superfluous and which are not, especially at scale.
After you get a score for the model trained on the whole textbook, try removing each sentence in the book in turn. If removing that particular sentence decreased the test scores then keep it in, else throw it away.
- Perhaps the model can be improved by adding sentences, but which ones? There is a potentially infinite amount of sentences to add and no good way to select them
- Your proposed method handwaves the method of acquiring "many problem challenges that it hasn't seen before", which just moves the problem. Constructing a problem set containing the full range of potential problems to solve is again a Very Hard Problem.
The second problem is not just theoretical either: see https://sitn.hms.harvard.edu/flash/2020/racial-discriminatio... for example, where a facial recognition algorithm performed much worse on non-white women because the training set didn't contain enough pictures of them. AI software is notorious for finding this kind of loophole and for overfitting itself for the training set rather than for the real world.
One one side they are part of who we are. On the other side, same as an airplane does not copy a bird 100%, it makes sense that to make a machine "think" we would feed it rational content. That is, content that follows the scientific method.
On the other end of the AI spectrum, an AI being trained to be a childrens toy, old people companion, nursing robot or even psychologist would be incredibly deficient if it didn't understand emotions. I don't think using purely content that follows the scientific method would be sufficient for such an AI.
On the other side, there are textbooks about "emotions". I own a Cognitive Behaviour Therapy manual (seemed interesting) that goes step by step through how to conduct a session (it is aimed at therapists or at the reader being their own) as it progresses. As a textbook, similar professional references could be more useful than say reddit or some blogs. Those other sources on the internet will be mostly derived from the reference materials in the field, or from personal anecdotes.
I can imagine textbooks exists on how to deal with people on the spectrum, or for them to recognize social cues, or how symptoms of depression look like.
I am still skeptical. You did not say anything of the sort, but the scientific method is not the opposite of emotions. It shows us what we know of emotions so far. Psychology has the reproducibility scandal but it is our best try.
Even in the case of planes, I would not be surprised if an engineering manual mentions a minimum space for passengers to be comfortable, as part of regulations, or any other data that might as a side-effect solve the claustrophobia issue.
TL;DR. Scientific knowledge is not antithetic to "human" knowledge. Scientific knowledge is what we actually know. Otherwise we have mysticism, or faith.
For example: People react very different to being offered a bacon pizza depending on (at least) the time (not for breakfast thanks), location (funerals are right out), their pizza topping preferences, their faith, whether they've already just eaten, how much they trust the person offering the pizza, if they're currently in a group or not (is there enough for everyone?), whether they've ever had any traumatic experiences with pizza or not, if any other foods may be available, whether they have dinner with other people planned, their current dietary restrictions, whether they're ill right now or not, whether they know you are doing pizza-based experiments on them, etc etc etc etc.
All of these might radically alter the response you get. It might be possible to get enough information about someone to make a reasonable guess, but you can never know if your model is complete enough. Even if the same inputs occur twice, the output might still be different based on something you cannot reasonably know. So I think it would be very difficult to construct a predicable stimulus/response model for individuals. Groups might be easier because the differences sometimes average out, but quite often the internal communication leads to feedback loops that disturb any predictions you were trying to make about their behavior.
Btw there is a third option beyond "scientific knowledge" and "faith", which is simply "we do not know and may never know" without ever getting to fill in that gap. Things like "what, if anything, created the universe" fall in that category as we cannot observe such a thing. Given some information about the position of a particle, we can never be sure of its speed per the Heisenberg uncertainty principle. Accurate models of people could be similar: it might simply be that they cannot be meaningfully reduced to simple formulas.
> Btw there is a third option beyond "scientific knowledge" and "faith"
The three I usually see mentioned are: I know because "argument" (reason/scientific); I believe (faith) and mysticism (I just know/I had a revelation). I agree with you, different ways of not taking a position are positions in themselves. Skepticism, nihilism, etc.
> Accurate models of people could be similar: it might simply be that they cannot be meaningfully reduced to simple formulas.
I agree. All models are false but some are useful.
Very approachable prices.
> Now this wouldn't be possible without the high quality synthetic dataset produced by GPT(1B tokens)
Just a note that this is GPT3.5 (I assume turbo?).
"'All you need' considered harmful"
I guess it shouldn't be surprising. With reams of information piling up, it can be hard to get noticed. So you look for gimmicks to get attention (hah!). The trouble is that you are not unique and a thousand other people have the same idea... and your cute imitation is just tired.
It used to work, but judging by your diatribe, maybe it got tired already and I didn't notice.
I can totally see how what I posted could be interpreted that way, though.
Is this common, training a LLM from another LLMs generated output? How do you avoid "bad code" from GPT, if not out right hallucinations?
It's not uncommon. And the level the SOTA LLMs are now, it'll only become more common.
It happens because we're at the stage where GPT-3.5/4 can generate much better data (or hit more encompassing or general distributions with the right instructions) than what the majority of LLMs will be typically trained on.
Also 4 can recheck output for bugs, mistakes, errors etc
It may not be perfect but neither would "natural data" and depending on exactly what kind of data, it might be as close to perfect as you'll get - https://arxiv.org/abs/2305.07759
I've read that people reorganize information when they sleep, I wouldn't say that's "thinking".
I'd say that inner voices are mostly stupid idiots :)
I guess you may be one of the presumably minority of people who doesn't have an inner narrative? I'm having a hard time imagining how this works, but then again, I'm aphantasic, and a lot of people have trouble imagining how that works too.
The point is that you don’t need a recent external environment to grow your knowledge in certain ways.
For example, if you locked a sufficiently gifted and immortal mathematician in prison with no knowledge of the outside world after 2021, he can still prove new theorems forever afterwards (provided he has enough coffee). He doesn’t need access to recent mathematical results to keep proving theorems and growing his knowledge base - that’s simply an accelerant. In fact the longer he stays in prison, the more new theorems he can prove because he has ever more lemmas.
Whereas if you locked a sufficiently gifted and immortal chemist in prison with no knowledge after 2021, she might be able to synthesize some new knowledge at first (meta analysis, etc), but eventually that would run dry and she wouldn’t be able to say anything new about chemistry anymore without any equipment.
The False Promise of Imitating Proprietary LLMs (UC Berkeley, 25/May/2023)
A child gets alot of his knowledge by randomly copying what grownups do.
"displays surprising emergent properties"
Section 3 states "includes managing intricate algorithmic tasks." I take this as meaning it can write new code (not regurgitation). It is impressive that is can write code as described in the prompt, but I didn't think this was considered "emergent" but "generalization." What is the difference?ie., with 1MB data you dont get generalisation, with 1TB you do. So they call that emergence.
I'm almost at the point where I can let these things go; but each new bit of mystifying hype, another stone falls in my shoe.
I'm pretty sure they mean emergent as in emergent gameplay.
Here's a good explanation for that:
> Emergent gameplay refers to complex situations in video games, board games, or table top role-playing games that emerge from the interaction of relatively simple game mechanics.
So emergent is not inherently surprising. Games like Dwarf Fortress or Rimworld have a lot of emergent complexity but most of it was not surprising.
So in the context of LLM it means that complex reasoning is emergent based on simple underlying mechanics.
And since we thought humans’ complex reasoning skills were very unique.. We are surprised by the emergence of it in LLMs.
PS: The explanation for "emergence" is much better:
> In philosophy, systems theory, science, and art, emergence occurs when a complex entity has properties or behaviors that its parts do not have on their own, and emerge only when they interact in a wider whole.
Human reasoning (, game play, etc.) is emergent in the sense that it is irreducible to its parts. "Generalisation in high-parameter training regiemes" is fully reducible. It is not emergence in any relevant sense of the term.
https://hai.stanford.edu/news/ais-ostensible-emergent-abilit...
None involved know what emergence means; they're engineers chasing and creating hype.
Their whole method relies on getting GPT-3.5 to produce examples and then training a network on those examples. This is a run-of-the-mill method called distillation.
There's nothing new or special here.
https://arxiv.org/search/?query=everything+everywhere+all+at...
I wish someone would do something like that but not use OpenAI's models to create the training data because then supposedly it can't be used for commercial purposes.
Not that it's easy to create a lot of high quality training data.
If that worked as a proxy value, you could sidestep needing GPT-4 at all.
"Consider the matrix A = np.array([[1, 2], [2, 4]]). We can check if this matrix is singular or nonsingular using the determinant function. [...]"
No. The determinant is not a suitable way to do that. A proper way to numerically measure singularity would be to compute the condition number of the matrix (the ratio of its largest to smallest singular value).
The problem with the determinant is not about performance. It is just useless for determining if a matrix is singular. The thing that gives it away is that the determinant is influenced by a rescaling of the matrix:
det(s A) = s^n det(A) where A is a n x n matrix
As an example, would you say that [[1e-10, 0], [0, 1e-10]] is singular? It has condition number 1.
Let's say we are in a setting where we only work with integers. A matrix is invertible iff its determinant is invertible in the underlying ring. The only invertible elements in Z are -1 and 1.
So, the code is also incorrect in the integer setting. Here, we should not check for 0, but for -1 or 1.
This is purely a matter of numerical stability. Of course [[1e-10, 0], [0, 1e-10]] is nonsingular, and its determinant is 1e-20, which does not equal zero.
Yes, when it comes to floating point issues we might want to use something else, and that's a valid complaint when it comes to NumPy code, but from a theoretical perspective the determinant is an excellent tool to determine singularity.
Numerically, such a binary answer is pretty useless. Here, we need a measure of how singular/nonsingular a matrix is relative to the numerical precision we are working with.
There is no one bulletproof general method to approximate mathematical calculations with floating point numbers. More context is generally required, including the actual problem that is being approximated, to determine if a method is reliable. Painting this as a black and white situation where the determinant is wrong and the condition number is right gives a misleading picture of how we evaluate numerical methods for fit-to-purpose.
Quality is everything and part of quality is also having it be structured and geared towards an AI which may not be identical to being geared for human reading.
I’m not sure why they compare to a list of results for the MBPP benchmark but don’t seem to include these results which are much better:
"Just as a comprehensive, well-crafted textbook can provide a student with the necessary knowledge to master a new subject, our work demonstrates the remarkable impact of high-quality data in honing a language model’s proficiency in code-generation tasks. By crafting “textbook quality” data we were able to train a model that surpasses almost all open-source models on coding benchmarks such as HumanEval and MBPP despite being 10x smaller in model size and 100x smaller in dataset size. We hypothesize that such high quality data dramatically improves the learning efficiency of language models for code as they provide clear, self-contained, instructive, and balanced examples of coding concepts and skills."
So, a coin toss effectively, right?
If multi-choice or free-text, random answering will be much below 50%.
Actually, it turns out that impact factor is all you need.