If AI is supposed to resemble a human mind with ability to learn, then it must be able to learn from a blanker slate. You don't teach the human before it is born, and in this comparison an AI is born when you finish it's model and set its weights using the training set. If you test it with the training set, you aren't testing ability to comprehend, just regurgitate what it was born with
https://britishlibrary.typepad.co.uk/digitisedmanuscripts/20...
There's so many rabbit holes to go down when trying to understand language, vision, reasoning, and all that stuff.
[0] (Jesus England... this is what you call this game?!) https://en.wikipedia.org/wiki/Chinese_whispers
[1] https://translate.google.com/?sl=en&tl=zh-CN&text=giraffe&op... ----> https://translate.google.com/?sl=zh-CN&tl=en&text=%E9%95%BF%...
It depends on the zero-shot experiment. Let's look at two simple examples
Example 1:
We train a classifier that classifies several animals (and maybe other things). For example, you can use the classic CIFAR-10 dataset which has labels: airplane, automobile, bird, cat, deer, dog, frog, __horse__, ship, truck. The reason I underlined horse is because you want your model to classify the zebras as horses!
The reason this is useful is for measuring the ability to generalize. At least in our human thinking framework we'd place a zebra in that bin because it is the most similar (and deer should be the most common "error"). This can help us understand the network and we'll be pretty certain that the network is learning the key concepts of a horse when trying to classify horses rather than things like textures, colors, or background elements. If it frequently picks ships your network is probably focusing on textures (IIRC CIFAR has ships with the Dazzle Camo[0] and that's why I threw "ship" out there).
Example 2:
Let's say we train our network on __text__. In this case it can get any description of a zebra that it wants. In fact, you'd probably want to have a description of what it looks like!
The what we might do is take that trained text network, and attach it to a vision classifier. For simplicity, let's say that was trained on CIFAR-10 again. We then tune our LM + CV model so that it can match the labels of CIFAR-10 (basically you're tuning to ensure the networks build a communication path, otherwise it won't work). Here we end up testing our model's actual understanding of the zebra concept. It again should pick horse as the likely class because you've presumably had in the training text some description that compares zebras to horses.
-----
So really the framework of zero-shot (and few-shot) is a bit different. We're actually more concerned about clustering and you should treat them more similar to clustering algorithms. n-shot frameworks really come from the subfield of metalearning (focusing on learning how networks learn). But as you can imagine, these concepts are pretty abstract, but hey, so are humans (that's why we see a log as a chair and will situationally classify it as such, but let's save the discussion of embodiment for another time).
In either example I think you can probably see how a toddler could do similar tasks. You can ask which of those things the zebra is most similar to and you'd be testing the toddler's visual reasoning. The text one might need be a little older but it could be a great way to test a child's reading comprehension. Does this make sense? Of course machines are different and we need to be careful with these analyses (which is why I rage against just comparing scores/benchmarks, these mean very little), because the machines may be seeing and interpreting things differently than us. So really the desired outcome depends on if you're testing for what the machine knows/understands (you need to do way more than what we discussed above) or if you are training a machine to think more similar to a human (then we can rely pretty close to exactly what we discussed).
Hope this makes more sense.
I’ve experienced both, each at a different university
In one, professors would teach one thing then ask very different (and much harder) questions on tests
In the other, tests were more of a recap of the material up to that point
I definitely learned a lot more in the second case and was a lot more motivated. It also required more effort from the professors
The two methods also test different things. The recap one tests effort and dedication, if you do the work, you get the grade
The difficult tests measure either luck and/or creativity and problem solving under pressure. It’s not about doing the work, it’s about either being lucky or good at testing
I think you are misunderstanding the experience.
The first (harder questions) is testing your understanding of the material and problem. Can you applying the material to solve a novel problem? Do you understand the material not just the mechanics. Do you understand how it would interelate it with other problems? Do you understand the limitations?
The second is just regurgitation. This is great for rote skills, but this isn't really learning. This is grinding until you can reproduce. These are the kinds of skills that are easily automated. This is not what we should be testing our kids.
And yes, to bring back to ML it is the difference of generalization and memorization (compression). I wrote a longer response to a different response to my initial comment to help clarify because I think this chain is a bit obtuse and aggressive for no reason :/ (I mean you can check the Wiki page to verify what I said)
So everyone here, including TFA, are all kinda doubting the same claim (our AI models can perform zero-shot generalizations) in different ways, I think?
In essence you aren't wrong, but that's not what we'd typically do in a zero (or few) shot setting. We'd be focusing on things that are more similar. If you want to understand this a bit better in what we might do in a ML context I wrote more here[1].
And I like Nico's comment about how different professors test. Because it makes you think about what is actually being tested. Are you being tested on memorization or generalization? You can argue both these kinds of tests are testing "if you learned the material" but we'd understand that these two types of tests are fundamentally different and let's be real, are not reasonably fair to compare scores to. I'm sure many of us have experienced this where someone that gets a C in professor A's class likely learned more than than someone who got an A in professor B's class. The thing is that the nuance is incredibly important here to really understand. And you can trivialize anything, but be careful when doing so, you may overlook the most important things ;)
Now... you could make this argument about geometry -> calculus if we're not talking about the typical geometry (single) class most people will have taken in either middle school or high school. Because yes, at the end of the day there is a geometric interpretation and we have the Riemann sum. But we'd need to ensure those kids have the understanding of infinities (which aren't numbers btw). We'd have to be pretty careful about formulating this kind of test if we're going to take away any useful information from it. Though the naive version might give us clues about information leakage (in our case with children this might be "child who has a parent that's a mathematician" or something along those lines). It really all depends on what the question behind the test is. Scores only mean things when we have nuanced clear understandings of what we're measuring (so again, tread carefully because "here be dragons" and you're likely to get burned before, or even without, knowing it)
And truth be told, we actually do this a bit. There's a reason you take geometry before calculus. Because the skills build up. But you're right that they don't generalize.