“Less than one”-shot learning
technologyreview.com
technologyreview.com
> In a previous paper, MIT researchers had introduced a technique to “distill” giant data sets into tiny ones, and as a proof of concept, they had compressed MNIST down to only 10 images.
The total entropy in 10 images (even carefully engineered ones) is very low in comparison to the full data-set.
Interesting paper, although the headline is of course sensational. The crux of the paper is that by using "soft labels" (for example a probability distribution rather than one-hot), it's possible to create a decision boundary that encodes more classes than you have examples. In fact, only two examples can be used to encode any finite number of classes.
This is interesting because it means that, in theory, ML models should be able to learn decision spaces that are far more complex than the input data has traditionally been thought to encode. Maybe one day we can create complex, generalizable models using a small amount of data.
As written, this paper does not provide much actionable information. The problem is a toy problem, and is far from being useful in "modern" AI techniques (especially things like deep learning or boosted trees). The paper also is not practical in the sense that in real life you don't know what your decision boundary should look like (that's what you learn after all), and there's no obvious way to know which data to collect to get a decision boundary you want.
In other words, this paper has said "this representation is mathematically possible" and is hoping that future work can actually make it useful in practice.
(And the previous paper was about compositing elements of multiple dataset images into a single training example - likely overfitting).
It's like saying here's the ideal partitioning scheme, memorize this.
As an example, the media lab is still citing innovation with deep fakes, claiming entirely novel results people are shocked to see. They hype their own researchers even though there are kids on YouTube who that have been making similar content up to a year prior to Technology Review's publication.
I suspect they do the same with fields I'm less familiar with.
But even when you're note reading a university publication/PR piece, you still see this effect. One big player like a famous lab at Stanford or MIT publishes an incremental paper that is a followup on a well known existing research direction, where several groups are working in parallel on very similar things, and then it's presented as if it was some breakthrough and the whole subfield was invented by them right now.
It's very, very hard for outsiders to recognize this and to really understand what the actual incremental step in a particular paper is. Necessarily, when explaining to laypeople, you can only scratch the surface and present the rought idea of a whole big research field, and it gets really murky what is part of the established, pre-existing research field and what is the novel contribution.
I'm sure I fall victim to this when reading outside my expertise as well, i.e. when reading about genetics stuff or quantum computing.
I try to avoid this by focusing on the content, not who the researchers are. If they can do some cool new thing with mosquito genetics, now I know that. I don't need to know whether the specific paper being hyped is novel or whether 95% of the content was proved by someone else: that's a task for the Nobel committee and I wouldn't recognise the researchers' names again anyway.
Note that the distilled data is not even from the same "domain" of input data any more. They're basically adversarial inputs.
I am not convinced that this time saving is more than the time spent to engineer the combined and synthesised data.
I see this a lot. It's completely wrong. I'm not trying to pick on the author here, I think 95%+ of people share this misunderstanding of deep learning.
If you see "only one" horse, say for even a second, you really are seeing a huge number of horses, from various angles, with various shades of lighting. The motions of the horse; the motions of your head (even if slight); the undulations of the light; are generating a much larger number of basically augmented training data. If you look at a horse for a minute it could be the equivalent of training on 1 million images of a horse. I'm not sure the exact OOM, but it's certainly orders of magnitude more than "one" horse.
(Relatedly: Some people say there is an experiment you can conduct at home to see the actual images your brain is training on).
Standard machine learning systems don’t work like this. No matter how much time they’re allowed to process a single sample image, they are unable to learn to classify it. There are of course machine learning systems that work more like humans (facial recognition is a notable example), but they’re the exception not the norm.
None of that is the point though: a child can look at a single static image (not 100 different perspectives - just one perspective) of a horse, and learn to recognize a horse. A standard machine learning model cannot, no matter how much you augment the image.
But in an important way, they can't.
You need to define "look". How many nanoseconds? How much is the lighting changing during that time? How still is the photo? How still is the person's head? In DL a training example in a batch is perfectly defined in bits. Once you try to define a training example for a human you see that a "single static image" means totally different things. A human seeing a static image is the equivalent of training a model on at least thousands, if not millions+ of training images.
If you trained a neural net with 1 million images of the same horse, it would learn to recognise that horse... but no other horse. Neural net datasets try to include as many variants of the target concept as possible in order to capture as many of the common features of instances of that concept as possible. A single horse would not suffice, e.g. 1 million images of a white horse would teach a neural net that horses are only white, etc.
Also, a child can learn to recognise horses from a caricature of a horse- which is a single image of a horse, not 1 million images. Whatever human minds do when they learn to recognise objects, that's not training with big data in real time.
Not sure exactly how good the state of the art is compared to human children on that task, but it can't be too far off.
I can show a 2 year old a single picture of a horse and have them learn it, but only after training them on 5 billion pictures of non-horses.
I mean, how do you know it takes 5 billion images of non-horses before the child can learn to recognise a horse? Even if the child sees 5 billion images of non-horses, who says that they haven't learned to recognise the non-horses and the horses from the first of those 5 billion images? After all, I don't know that children stay transfixed for hours staring at horses and non-horses while their brains process all those billions of images until they finally get it.
I think you are making an effort to explain away something that is not very well understood by modern science, by analogy with machine learning, even though there is no good reason to suppose that the two are connected in any way.
Just a ball park figure for how many “pictures” they have seen. Probably off by a few OOM, but point holds.
I think there is a very good reason to think machine learning and human learning are connected. First, I’d say it’s largely true the people who have made the most impact in the former have all studied the way the mind works (as far as science allows), and tried to emulate that in machinery. Second, experimentally the phenomena observed are increasingly similar (deep dream, for example).
IEEE Spectrum: We read about Deep Learning in the news a lot these days. What’s your least favorite definition of the term that you see in these stories?
Yann LeCun: My least favorite description is, “It works just like the brain.” I don’t like people saying this because, while Deep Learning gets an inspiration from biology, it’s very, very far from what the brain actually does. And describing it like the brain gives a bit of the aura of magic to it, which is dangerous. It leads to hype; people claim things that are not true. AI has gone through a number of AI winters because people claimed things they couldn’t deliver.
https://spectrum.ieee.org/automaton/artificial-intelligence/...
Also, like I say above, it doesn't matter how many images a human sees- what matters is how many she needs to see before learnign to recognise a thing. The example of an unchanging, two-dimensional caricature of a horse is evidence enough that, even if we do see billions of "images" as you say (I'm not sure it makes sense to speak of "images" int the sense you use it) we don't need to see all those billions of them before we learn what things look like.
The way the chain of connection goes from the eye through the brain is like a tree, and when a child "sees" a horse, a number of those pathways are being activated, some of which are more general, and some of which are specific.
> which is a single image of a horse,
Although it may be a "single" image in the sense that is a single file on disk or a single printed image, a child is not seeing a single image. They are seeing a rushing river of horses, even if from just that 1 static image. Think of the 60hz refresh rate of your monitor. A child is seeing at least 60 "images per second", and likely many, many times more.
So how does the child learn to recognise a fully 3-dimensional real-life horse from a single caricature of a horse seen any number of times? Can you explain?
Also, if you were to train a nerual net with a single insance of a caricature of a horse, even if you copied it a million times to create an example set of a million copies, the neural net would still only be able to elarn how to recognise that particular caricature, if that- it would certainly not be able to extract any useful features to recognise real-life horses from that single caricature.