No "Zero-Shot" Without Exponential Data
arxiv.org
arxiv.org
We've essentially been ripping off the entire internet and feeding it to the models already, spending many billions of dollars in the process. It's pretty much the largest possible dataset you can currently get, and due to the ever-increasing and now rapidly accelerated AI poisoning of the internet most likely the largest possible dataset which will ever exist.
All that and all we're getting out of it is not-entirely-useless but still quite crappy AI? We would've been better off if we had never done this.
This paper seems to suggest that significant advancement in domain knowledge acquisition and retention is exactly the problem, as you seem to need exponentially more data due to a lack of generalization. What's the point of a model which can perfectly quote Shakespeare if you're a programmer trying to refactor a proprietary codebase and it fails to make a link to whatever garbage it picked up from StackOverflow?
Replacing CLIP with an LLM is the current meta in image generation models, specifically because of that lack of generalisation. This isn’t a surprise to anyone.
15 months ago, general-purpose LLMs that have not been specifically trained on legal reasoning could score better than 90% of humans on the multistate bar exam, and these are humans who actually completed law school.
General-purpose LLMs get similar results in medicine, and when the models are fine-tuned for medical diagnosis they're even better.
And that was more than a year ago. Those who have seen current models not yet released tell us they'll make the current state of the art look like toys.
Progress is still tracking the steep part of the S-curve, and there's no indication that they're near the top yet.
If I understand it correctly, that seems to be exactly what this paper is suggesting.
Scoring high on the bar exam is pretty trivial for AI - the data needed for that is fairly generic and widely available on the internet. It requires you to demonstrate a relatively basic understanding of the concepts by answering a bunch of multiple-choice answers. If anything, I'd expect AI to have a perfect score.
Like I said, such an AI is not entirely useless. You can replace quite a few legal assistants with that, and I bet it could be used to create first drafts or to expand a core concept into a full legal argument. There is plenty of money to be made there, and it's going to make an awful lot of people jobless. But that's just replacing more-trivial jobs with automation, it doesn't add anything novel to society.
On the other hand, the actual difficult work involves being able to come up with completely novel concepts, and being able to expand upon some obscure but crucial stuff few people have ever heard about. Current models simply aren't capable of that, and the results achieved here with multimodal models suggests that they never will. We risk getting stuck with models which can do some trivial work, but silently produce complete garbage when you ask them to do anything providing substantial value.
what if you couple a random text generator with a picture generator, isn't each new picture novel?
if we could send all artists to work in the healthcare mines we might discover new treatment for various ailments. is that not novel? is that too indirect?
increasing "economic surplus with externalities factored in" is good for society even if it's not novel.
arguably philosophers come up with novel stuff all the time, and ...
> Current models simply aren't capable of that, and the results achieved here with multimodal models suggests that they never will.
can you elaborate on this a bit please?
You do realize that artists are artists because they want to be, and even if AI can create art, artists will still be artists, right? They won't just be like, "Oh, AI is good at generating art now? Whelp time to go become a doctor."
the model is great at this, because the training set is full of this stuff. the questions are static and simple. (the actual text of the questions change, of course, but the format and the answers are from a fixed set. and the LLM doesn't need to generate text basically, just one token. A B C or D ... that said I'm curious how it's administered to the LLMs and how much that influences their performance.)
The claims of similar performance on coding problems were shown to be due to contamination of the training data on the tested problems. It did abysmal on problems made public after the model training cutoff.
I don’t think anyone has tested contamination for the MBE claims, but I would lean toward assuming the same issue exists for that assessment until proven otherwise.
The use cases for LLMs will no doubt grow as hallucinations are reduced, and they gain planning/reasoning ability over next couple of years. It'll be interesting to see what they subjectively "feel" like with these improvements.
On the other hand, it's still a big "if" whether a general hallucinations solution exists, and in the meantime we're paying a pretty high price for it, as the entire internet is being flooded with absolute garbage. We risk getting stuck in a situation where there is no way to get new senior people because nobody hired them as juniors because AI is cheaper and good enough, and senior people are getting less and less productive due to them being unable to use websites like StackOverflow as reference material. That's a pretty high price for a tiny gain, if you ask me.
No. What you describe encompasses only one poor modality: text. We have oodles more data, and can create almost arbitrary amounts more, by just eg pointing webcams at the world.
> We would've been better off if we had never done this.
Who is 'we'?
It is still early to tell if we reached AI winter again or not, but at least we can see that news are slowing down.
Only the poor souls who survive the fall, at the very bottom of the precipice, will get to experience the AI winter.
Incremental improvement of proven approaches, driven by profit motive, will surely continue regardless of whether there is an AI winter or not.
It's the rage because it's a way to practically _do work_ from LLMs that generally provide wow from conversationally accurate, often factually accurate responses.
There are some major limitations to LLMs that aren't going to be "optimized" away. At the end of the day LLMs are Monte Carlo samplers over a latent, compressed representation existing human text data. Many of the tasks people hope LLMs will achieve require major leaps in our current understanding of language modeling.
One great example limitation (which shocks me sometimes when I think about it): generating output is still ultimately stuck in looking at the probability of the next token rather than the much more useful, probability of the generated statement. There are techniques to improve this, but we're missing a major piece of generating highly probable statements with no hint about how to really get there. Consider how you might write SQL. You conceive of the high level query first and start sketching out the pieces. An LLM can only look at each token and can't, statistically speaking, think in terms of the entire query.
Personally I think LLMs are very underutilized/exploited for what they are good at, and there is way too much focus on what they can't do. Hopefully we'll dodge an AI winter by using LLMs to solve the wide range of classical NLP problems that make many tasks that just a few years ago nearly impossible, rather simple today. Unfortunately the irrational hype around these models makes me skeptical of that scenario.
ie., P(mat|the cat sat on the) is a distribution over, say, 100k words. Whereas, P(the cat sat|on the mat) is a distribution over 100k^3 words.
Part of the illusion of an LLM is that we produce create text in such a highly regular way that a mere distribution over 100k iterated, say 500 times, gives you a page of text as if modelling 100k^500.
There's always some space in high-dimensions to put that extra context somewhere. So yes, only one token is predicted, but there's a lot of information squeezed into the "condition" (as in conditional probability, as mjburgess' comment shows)
> [hype instead of NLP]
yes, exactly. we'll see what remains after the bubble bursts.
We're just scratching the surface of what's possible with the current state of the art. Even if there are no major advances or breakthroughs in the near future, LLMs and associated technologies are already useful in many use cases. Or close enough to useful where engineering rather than science will be sufficient to overcome many (though not all) of the shortcomings of current AI models. Never mind grandiose claims about AGI, there's enough utility to be gotten out of the limited LLMs that we have today to keep engineers and entrepreneurs busy for years to come.
But I think the term "AI Winter" usually refers to the underlying research programme and the economics around it. Soaking up many billions of dollars of industry money and public grants on the pitch that AGI might be just around the corner, and then being unable to deliver on that pitch can induce a hangover effect that makes it much harder to raise money for anything that smells even a little bit like the failed pitch. Investors and administrators feel burnt and turn very skeptical for a very long time.
Meanwhile, the actual productive applications which shook out of the initial boom just get renamed to something else so that they don't carry that smell.
We'll see how it goes here, but that's the familiar road and where the terminology comes from.
But you would think after 70 years AI practitioners would learn some humility! It is very obvious that GPT-4 is dumber than a honeybee, let alone a cat, let alone a crow. But for over a year I've heard dozens of tech folks insist it's smarter than most people.
Why? The people doing AI now are not the same as those that did AI 60 years ago.
We are very clearly on that same path again. What leads to the conclusion that another winter is coming. But even the fact that people are talking about it is evidence it's not here yet, and as always with political phenomena, there's no guarantee history will repeat.
Anyway, none of it means people will stop applying their knowledge or studying AI. The entire thing happens on funding and PR, and the world is not entirely controlled by those two.
Again?
Yes sheer quantity has a quality of its own but that won't produce optimal results.
The funding is there but we're letting hype drive everything and not calling out the con artists.
The problem is we've had recent great success but still don't know how to get to AGI. But because we're afraid of winter we're not willing to try new things. We want to only compare to sota and think it's fair to compare a new method with a handful of papers to the status quo. That's not how the S-curves of technology work. Sure, maybe things don't scale but that doesn't mean they don't have merit or can't scale if someone finds some modification.
The problem is we treat research like products, not academic work. You need to produce everything from hard core theory to robust products to have an effective chain. But we seem hyper focused on the middle area. And for some reason people think products can be placing a nice interface around research code. There's still a lot of work you need to do and all those models can be optimized. They absolutely do not have optimal hyper parameters or even parameters.
Sure we already have some interesting applications, but they are not exactly printing money.
Even if academic progress stalls, we're going to be inundated for at least another few years.
But I do feel confident that an AI winter is not on the horizon solely due to the overhead of implementation that we currently have. Just with currently existing AI, it would take years for the economy to fully leverage the abilities that are available. I'm confident we have transformative AI. So it won't feel like a winter for several years as we actually succeed in optimizing, implementing, and productizing current technology in other industries.
If AI is supposed to resemble a human mind with ability to learn, then it must be able to learn from a blanker slate. You don't teach the human before it is born, and in this comparison an AI is born when you finish it's model and set its weights using the training set. If you test it with the training set, you aren't testing ability to comprehend, just regurgitate what it was born with
https://britishlibrary.typepad.co.uk/digitisedmanuscripts/20...
There's so many rabbit holes to go down when trying to understand language, vision, reasoning, and all that stuff.
[0] (Jesus England... this is what you call this game?!) https://en.wikipedia.org/wiki/Chinese_whispers
[1] https://translate.google.com/?sl=en&tl=zh-CN&text=giraffe&op... ----> https://translate.google.com/?sl=zh-CN&tl=en&text=%E9%95%BF%...
It depends on the zero-shot experiment. Let's look at two simple examples
Example 1:
We train a classifier that classifies several animals (and maybe other things). For example, you can use the classic CIFAR-10 dataset which has labels: airplane, automobile, bird, cat, deer, dog, frog, __horse__, ship, truck. The reason I underlined horse is because you want your model to classify the zebras as horses!
The reason this is useful is for measuring the ability to generalize. At least in our human thinking framework we'd place a zebra in that bin because it is the most similar (and deer should be the most common "error"). This can help us understand the network and we'll be pretty certain that the network is learning the key concepts of a horse when trying to classify horses rather than things like textures, colors, or background elements. If it frequently picks ships your network is probably focusing on textures (IIRC CIFAR has ships with the Dazzle Camo[0] and that's why I threw "ship" out there).
Example 2:
Let's say we train our network on __text__. In this case it can get any description of a zebra that it wants. In fact, you'd probably want to have a description of what it looks like!
The what we might do is take that trained text network, and attach it to a vision classifier. For simplicity, let's say that was trained on CIFAR-10 again. We then tune our LM + CV model so that it can match the labels of CIFAR-10 (basically you're tuning to ensure the networks build a communication path, otherwise it won't work). Here we end up testing our model's actual understanding of the zebra concept. It again should pick horse as the likely class because you've presumably had in the training text some description that compares zebras to horses.
-----
So really the framework of zero-shot (and few-shot) is a bit different. We're actually more concerned about clustering and you should treat them more similar to clustering algorithms. n-shot frameworks really come from the subfield of metalearning (focusing on learning how networks learn). But as you can imagine, these concepts are pretty abstract, but hey, so are humans (that's why we see a log as a chair and will situationally classify it as such, but let's save the discussion of embodiment for another time).
In either example I think you can probably see how a toddler could do similar tasks. You can ask which of those things the zebra is most similar to and you'd be testing the toddler's visual reasoning. The text one might need be a little older but it could be a great way to test a child's reading comprehension. Does this make sense? Of course machines are different and we need to be careful with these analyses (which is why I rage against just comparing scores/benchmarks, these mean very little), because the machines may be seeing and interpreting things differently than us. So really the desired outcome depends on if you're testing for what the machine knows/understands (you need to do way more than what we discussed above) or if you are training a machine to think more similar to a human (then we can rely pretty close to exactly what we discussed).
Hope this makes more sense.
I’ve experienced both, each at a different university
In one, professors would teach one thing then ask very different (and much harder) questions on tests
In the other, tests were more of a recap of the material up to that point
I definitely learned a lot more in the second case and was a lot more motivated. It also required more effort from the professors
The two methods also test different things. The recap one tests effort and dedication, if you do the work, you get the grade
The difficult tests measure either luck and/or creativity and problem solving under pressure. It’s not about doing the work, it’s about either being lucky or good at testing
I think you are misunderstanding the experience.
The first (harder questions) is testing your understanding of the material and problem. Can you applying the material to solve a novel problem? Do you understand the material not just the mechanics. Do you understand how it would interelate it with other problems? Do you understand the limitations?
The second is just regurgitation. This is great for rote skills, but this isn't really learning. This is grinding until you can reproduce. These are the kinds of skills that are easily automated. This is not what we should be testing our kids.
And yes, to bring back to ML it is the difference of generalization and memorization (compression). I wrote a longer response to a different response to my initial comment to help clarify because I think this chain is a bit obtuse and aggressive for no reason :/ (I mean you can check the Wiki page to verify what I said)
So everyone here, including TFA, are all kinda doubting the same claim (our AI models can perform zero-shot generalizations) in different ways, I think?
In essence you aren't wrong, but that's not what we'd typically do in a zero (or few) shot setting. We'd be focusing on things that are more similar. If you want to understand this a bit better in what we might do in a ML context I wrote more here[1].
And I like Nico's comment about how different professors test. Because it makes you think about what is actually being tested. Are you being tested on memorization or generalization? You can argue both these kinds of tests are testing "if you learned the material" but we'd understand that these two types of tests are fundamentally different and let's be real, are not reasonably fair to compare scores to. I'm sure many of us have experienced this where someone that gets a C in professor A's class likely learned more than than someone who got an A in professor B's class. The thing is that the nuance is incredibly important here to really understand. And you can trivialize anything, but be careful when doing so, you may overlook the most important things ;)
Now... you could make this argument about geometry -> calculus if we're not talking about the typical geometry (single) class most people will have taken in either middle school or high school. Because yes, at the end of the day there is a geometric interpretation and we have the Riemann sum. But we'd need to ensure those kids have the understanding of infinities (which aren't numbers btw). We'd have to be pretty careful about formulating this kind of test if we're going to take away any useful information from it. Though the naive version might give us clues about information leakage (in our case with children this might be "child who has a parent that's a mathematician" or something along those lines). It really all depends on what the question behind the test is. Scores only mean things when we have nuanced clear understandings of what we're measuring (so again, tread carefully because "here be dragons" and you're likely to get burned before, or even without, knowing it)
And truth be told, we actually do this a bit. There's a reason you take geometry before calculus. Because the skills build up. But you're right that they don't generalize.
That performance vs. training data is not linear, but logarithmic, doesn't exactly come as a surprise.
I suspect from the success of Phi3 that it is in fact the former.
ML is funny. It's predictably going to run into the Education field.
Well, you’re learning from data you generate.
Sure. I'm producing human output from human input in a generally unconstrained, limitless way.
This is producing approximated human output from approximated human input. That second level of abstraction will be constrained, ultimately, by the limits of input.
These new text books could be great at simplifying the subject matter and making the material accessible or they may just never have fully understood the materials and are misleading.
Now imagine that over and over again, imo it's pretty likely to introduce inaccuracies if just taking a naiive approach.
Also, ever read your old journals? You are training on generated data.
Of course we still need real world data, but it seems like generated data should also play a role. Humans don’t weight dreams equally with reality, however, and that’s a distinction I feel is missing here.
1. Generate inverted problems which are easier to produce than solve. For instance, create an integration math exercise by differentiating a hairy function and reversing the steps.
2. Create (simulated) environmental data.
3. Use adversarial model competition, e.g. self-playing chess or training an artificial image generator/detector model pair.
There's evidently a commonplace myth that information quality starts pristine and exclusively gets degraded by systems thereafter. That's just easily, demonstrably false in myriad ways. That's why it's absurd to conclude that LLM output being in internet training data will cause model collapse.
Information-rich synthetic data can be created without humans, and it works. (Check out Phi, for instance.)
There's some cursory indication that in the long tail, training LLMs on LLM-generated data causes model collapse. Kind of like how if you photocopy a photocopy too many times the document becomes unreadable.
This isn't really surprising though. Neural networks at large are a form of lossy compression. You can't do lossy compression on artifacts recovered from lossy compression too many times. The losses stack.
If you want ML to do it... well that's a bit of a catch-22. How would an ML algorithm know if data is good enough to be trained on unless it has already been trained on that data?
Eventually AGI's will need to learn by experimentation as we do, but in the meantime ability to well predict a potential training sample could be used to decide whether to add it to training set or not. At the moment it seems the main emphasis is on a combination of multi-modality (esp. video) and synthetic data where one generation of LLM generates tailored training samples for the next generation. I guess this synthetic data allows a more selective acquisition of knowledge than just adding surprising texts found in the wild.
This seems doable, and, I think something like it is already done?
The authors ask whether image-to-text and text-to-image models (like CLIP and Stable Diffusion) are truly capable of zero-shot generalization.
To answer the question, the authors compile a list of 4000+ concepts (see paper for details on how they compile the list of concepts), and test how well 34 different models classify or generate those concepts at different scales of pretraining, from ~3M to ~400M samples.
They find that model performance on each concept scales linearly as the concept's frequency in pretraining data grows exponentially, i.e., the rarer the concept the less likely it is actually/properly learned -- which implies there is no "zero-shot" generalization.
The authors also release a long-tail test dataset that they cleverly name the "Let it Wag!" benchmark to allow other researchers to see for themselves how current models perform on the long tail of increasingly rare concepts.
Go read the whole thing, or at least the introduction. It's clear, concise, and well-written.
Still, interesting approach and kind of confirms the experience of most people where clip models can recognize known concepts but struggle with novel ones.
When we learn to drive, we do not need to crash our car a thousand times in a row before we start to get it.
When we play a new board game for the first time, we can do it fairly competently (though not as good as experienced players) just by reading and understanding the rules.
When I'm explaining AI stuff to family, the example I use is classification and I specifically use cats and dogs. I use the analogy of how you teach a toddler that this is a cat and that is a dog. Essentially repetition. And at first they get them mixed up and the parent will say "no, that's a dog" when they think it's a cat and so on.
But essentially, for a child learning the difference between a cat and a dog you only need to show them a handful of each and they'll generally get it from that point on.
That being said, why does it take ML millions (or billions) of images to be able to say "that's a cat" when a human does it on a handful (might be up to, say 100 but my point stands). Why can ML not do that yet?
I'm a dev for many years but not in AI, hence my ELI5 question :)
Edit: If the answer is massively long and complicated, perhaps if you could point me to some text (book, paper etc.) and I can read at my leisure.
Edit2: I just thought of something. Is it related to whether the child sees a still image or a live cat? So, for example, a still image is a single example of a cat standing in a particular position etc, whereas a moving, live cat, would be interpreted by the brain as many many still images, all processed individually? The end result being that, in fact the child, when seeing a live cat, actually sees thousands or millions of still images of the cat? It just popped into my head there :D
You're ignoring the trillions of frames a toddler has seen before they get to the part where they can even understand what a photo is.
Machine learning models are just math functions fitted to some data. When we get predictions from them, we’re really just using a technique to interpolate between the data points. The denser a particular region has been sampled in the data, the better the predictions will be. (This is why GPT-anything will do a good job writing solutions to common leetcode problems, while struggling with a novel problem.)
Humans have a powerful abstraction ability far beyond any algorithm that has been developed. We can take in a few pieces of information describing a really unusual set of circumstances, run imaginary experiments and simulations on them, and make very granular and accurate predictions about their consequences. Nobody actually knows how.
I don't want to be pedantic. But I'd like to interject. We can simulate everything with math functions. If the algorithm isn't there yet it's because we're using the wrong functions.
But back to the subject of deep learning, whatever brains do, I think it’s pretty clear that existing neural networks don’t approximate the biological process. They’re just too static.
2. Children see many more data points than computers, thousands of hours of lived experience from which they can learn, of not just the visual world but the auditory and other sensory worlds.
Essentially, you are comparing something that is ~80% trained and then training it on billions of data points to something that is 0% trained and then training it on still images of only two dimensions and asking why the former is much better.
Crucial context:
- They're only looking at image models -- not LLMs, etc
- Their models are tiny
- A "concept" here just means "a noun." The authors index images via these nouns.
- They didn't control for difficulty in visual representation/recognition of these exceptional infrequent, long-tail "concepts."
If I didn't know an object's label, I too would struggle to identify/draw it...