It depends on the zero-shot experiment. Let's look at two simple examples
Example 1:
We train a classifier that classifies several animals (and maybe other things). For example, you can use the classic CIFAR-10 dataset which has labels: airplane, automobile, bird, cat, deer, dog, frog, __horse__, ship, truck. The reason I underlined horse is because you want your model to classify the zebras as horses!
The reason this is useful is for measuring the ability to generalize. At least in our human thinking framework we'd place a zebra in that bin because it is the most similar (and deer should be the most common "error"). This can help us understand the network and we'll be pretty certain that the network is learning the key concepts of a horse when trying to classify horses rather than things like textures, colors, or background elements. If it frequently picks ships your network is probably focusing on textures (IIRC CIFAR has ships with the Dazzle Camo[0] and that's why I threw "ship" out there).
Example 2:
Let's say we train our network on __text__. In this case it can get any description of a zebra that it wants. In fact, you'd probably want to have a description of what it looks like!
The what we might do is take that trained text network, and attach it to a vision classifier. For simplicity, let's say that was trained on CIFAR-10 again. We then tune our LM + CV model so that it can match the labels of CIFAR-10 (basically you're tuning to ensure the networks build a communication path, otherwise it won't work). Here we end up testing our model's actual understanding of the zebra concept. It again should pick horse as the likely class because you've presumably had in the training text some description that compares zebras to horses.
-----
So really the framework of zero-shot (and few-shot) is a bit different. We're actually more concerned about clustering and you should treat them more similar to clustering algorithms. n-shot frameworks really come from the subfield of metalearning (focusing on learning how networks learn). But as you can imagine, these concepts are pretty abstract, but hey, so are humans (that's why we see a log as a chair and will situationally classify it as such, but let's save the discussion of embodiment for another time).
In either example I think you can probably see how a toddler could do similar tasks. You can ask which of those things the zebra is most similar to and you'd be testing the toddler's visual reasoning. The text one might need be a little older but it could be a great way to test a child's reading comprehension. Does this make sense? Of course machines are different and we need to be careful with these analyses (which is why I rage against just comparing scores/benchmarks, these mean very little), because the machines may be seeing and interpreting things differently than us. So really the desired outcome depends on if you're testing for what the machine knows/understands (you need to do way more than what we discussed above) or if you are training a machine to think more similar to a human (then we can rely pretty close to exactly what we discussed).
Hope this makes more sense.