Is there something fundamental we are missing in going about building these deep learning stuff ?
Is there something fundamental we are missing in going about building these deep learning stuff ?
Even a human brain has to train for ~4-5 months to become interested in shapes (https://en.wikipedia.org/wiki/Infant_visual_development). Once the human brain has been trained for these basic shapes for a while, it is able to quickly break down a new class (i.e. a cat) and recognize similar patterns in other images. This is something that is very similar to the way that training a deep NN works.
Also don't forget that the current NN are being trained mainly for photos, not moving images. A brain may recognize a cat by its tail-wagging or fur movements, which is a dimension that is completely missing from still images.
Also check out this similar post and discussion: https://news.ycombinator.com/item?id=9247851
Everyone who wonder how (on a superficial level) grown up humans are so good at learning new categories really should spend time around babies and toddlers and children for this reason...
You quickly realise how much training and brain development it actually takes before we're capable of doing much.
Then one day they start to get it (like "fire burns"), but it's still not there for sure until they experiment it deeply multiple times.
The dev in me can't help but see this two little humans as big mighty neural networks who spend their full uptime constantly ingesting tremendous amount of data and being restlessly tuned back by adults and experience :)
An untrained neural net has to learn everything from scratch (ha!), from pixel values to "knowing" what a cat is. There is some work on decreasing the amount of data needed to learn, but it's a very tricky subject.
A human child receives images on the retina at 20fps, say ... for 12 hours a day, for many years. That is a lot of training, a lot of images received by the brain.
What the kid does when we show it something new is to do a kind of fine-tuning of its neural net where previous visual experience is reused in order to quickly learn new types of objects. It's called one shot learning and it can be done in neural nets too.
Check out the recent paper "Building Machines That Learn and Think Like People":
https://arxiv.org/abs/1604.00289
Humans can see one or a few examples of a novel object, such as a cat, and create a fully 3D mental model of it. So we know what it will look like in different orientations and lighting conditions.
I.e. we can recognise static 2D photos of cats, 2D movies of cats, 3D movies of cats, and real cats in the real world.
The invariants are probably relatively simple - a set of head geometries and head feature shapes/distribution, with some secondary colour and texture confirmation.
What's interesting is that we can recognise modifiers to the invariants - e.g. a shaved cat is still recognisably a cat, but parsed as "cat without fur."
We can also recognise invariants when they're pared down to essentials in cartoons and sketches.
https://www.youtube.com/watch?v=R9qdyXCVNVk
A lot of learning is really just data compression - finding a minimal set of low-resource invariant specifics from a wide range of noisy high-resource inputs.
While in general I felt something similar that it has to be not just reams of data but also some sort of model (meta data) that when combined can produce innumerable combinations more easily.
In that sense it's quite similar to a vision system of humans or other animals, which needs lots and lots and lots of early-age exposure to "learn how to see" (which is the hard part); and only after that it becomes possible to figure out cats of any kind just by seeing one or two.
> And how many different pictures of cats are there?
I would guess at a shite-load more than 1 million.http://jmlr.org/proceedings/papers/v37/romera-paredes15.pdf
"Zero-shot learning consists in learning how to recognise new concepts by just having a description of them."