Google & Stanford create a digital brain that learns to identify a human face
extremetech.com
extremetech.com
Most people don't realize that computer vision research is done using tiny sets of tiny images. Think hundreds to thousands of images that are 32x32 grayscale or black and white. Google did 10 million 200x200 color images.
For a conservative comparison human vision can be thought of as two 15 megapixel cameras capturing at 30 frames per second. (The real numbers are higher and much more complicated.)
I suspect that given this amount and quality of input data most machine learning algorithms would have produced better than state of the art results.
Now the bad. Their model totally ignored time.
Most computer vision researchers work with static images. In biology there is no such thing as a static image. Everything moves. You move, your eyes move, the environment moves. The brain, as an exception, can understand static images, but the rule is motion.
A majority of features in biological vision are tied to motion. Light/dark change cells, moving edge change cells. A huge portion of the information you derive from sight comes from the temporal context of a given moment. The change from moment A to moment B is directly encoded and learned and is often more important than the state at A or the state at B.
I'm not faulting them, in fact, I am extremely happy with their results, but I want to point out that these results can and will get much better in the near future as the input size increases and people start to integrate temporal features into their models.
So I may embarrass myself with the following...
The more software engineering experience I get the more neural networks bother me. I have not studied them in detail... and yet, they seem like a crappy solution.
Throw a lot resources at something, add a bit of math and push a pile of data through.
Does that ever lead to result that can't be bested though simpler more traditional methods?
For example, the math behind how the network is organized and learns, that same math can be used to write algorithms which reveal patterns in data. Patterns which can then be studied further.
Then there is the doing of things. What ever pattern is responsible for doing something in a network, the same function can be described much more simply by an algorithm.
It just seems like neural networks would be interesting only if an alien spaceship with stupefying computing resources fell on us. And it fell on us in such a way that we could fully utilize the power of this new super computer while at the same time we could not learn anything new about algorithms or anything else. Only then does it seem like "train the neural network" is a smart strategy.
In every other scenario is just seems like a decent if somewhat lazy start to an investigation. Really just a way to kick the tires of a problem.
So did I just completely embarrass myself with my lack understanding of ANNs?
There are a number of areas in which unsupervised feature learning outperforms the hand-programmed models we have been fine-tuning for decades.
Surely whoever thought this post was so wrong as to merit down-voting could explain why that is and answer the man's questions?
Neural Networks aren't the only common approach in modern AI research, but the other techniques (e.g. Support Vector Machines) have similar properties: they're dumb methods of finding correlations in data. The field of AI is slowly coming to grips with the fact that more training data trumps better algorithms.
That said, I sympathize with your frustrations. You can solve a lot of problems by throwing more cycles at it, but you don't learn anything, either about the algorithms or about your target domain. For example, a human plays the game of Go differently than a machine. A human has an intuition of what moves are good and what are bad, and only considers a handful; a machine relies less on evaluating board position and more on brute-forcing through possible board configurations. As machines get more powerful, eventually the machines catch up to humans in ability, with that same approach. While their methods will become impressive in their efficacy, we learn nothing about how humans are so much better at evaluating board position and selecting moves to speculate on.
What are you doing?
I am training a randomly wired neural net to play Tic-Tac-Toe.
Why is the net wired randomly?
I do not want it to have any preconceptions of how to play.
Minsky then shut his eyes.
Why do you close your eyes?
Sussman asked his teacher.
So that the room will be empty.
At that moment, Sussman was enlightened.
-- Jargon FileActually if you rigorously examine the same data you are feeding your network with well applied statistical methods, that is not true.
But in my experience statistic are actually some of the hardest math, and not many people, other than statisticians, are very good at using them correctly.
The idea that I can create software which works even if I don't understand why it works is a powerful idea. The point of AI as a field is stepping back and saying: "Computers are fast and cheap, people are slow and expensive, can't we automate this?" and continuing to do that not just to secretary's, or even coders, but also designers. As long as you can build a large enough training set and toss enough computing resources you can basically solve any problem.
Granted, the need for large training set's and processing power can also be a huge limitations. But, in the end neural nets / AI is just another tool like your computer, Algebra, or a big stick. The it's a cool idea and solves useful problems is enough for people to spend a lot of time and money looking into it.
I think that is exactly what bothers me about it. It's that everything you say would be 100% correct IF we had far, far, far more computing resources at our disposal.
But since we don't, humans solving a problem actually does not take longer or cost more than training a neural net. Not if you want the neural net to be as good as the solution humans come up with.
Only if you are satisfied with a very quick and dirty approximation can NNs be faster and cheaper than humans.
Yes, but for that you'd have to find the algorithm, which takes lots of human resources. That's what unsupervised learning saves you. In several specialized fields like computer vision and aural interpretation, a generic neural network trained for hours can outperform algorithms that took decades of research to refine.
First, theyre isn't a dataset suitable for only supervised learning (the final supervised stage they - imagenet - only has a relatively small amount of images of each category).
Second, higher resolution can make things worse in conventional supervised learning, you have to get more and more images to keep up with higher resolution (extra things in the image start confusing the algorithm, and even humans - more identifiable things in the photo mean you have to think more about what is the most important thing in the image).
The unsupervised algorithm gets around the both problems. There existed unsupervised/clustering algorithms before, but none that worked well until geoff hinton's breakthrough in 2005 (for which he deserves a turing award).
Paper here: http://arxiv.org/abs/1112.6209v3
[UPDATE: fixed erroneous statement.]
>Historically, machine learning has generally been supervised by humans. There are plenty of examples of computers identifying human faces (or cats) with incredible accuracy and speed — but only if human operators first tell the computer what to look for. That the Google/Stanford system starts from scratch and develops its own ability to classify objects is amazing.
Unsupervised machine learning has been around long before this feat -- which is nonetheless very impressive.
EDIT: Of course, that may have been due to the fact of better data/more processing power -> more time.
Most of the machine learning community had been disregarding this algorithm since 2005, until the past couple of years it started making dramatic showings, such as
- Microsoft Research speech recognition system in 2011 (done by one of geoff hinton's students interning there)
- NEC labs, ronan collobert - real-time natural languaege parser - first real-time parse ever - 2010 (trained on one computer for 3 months using wikipedia)
- Google Research 16,000 node billion connection network - december 2011
Sure, a neural network wired up to pixels of images wouldn't be able to come up with the word "face", but it would certainly group things correctly, right? Using keypoint matching to "grade" guesses, you could seemingly train the network appropriately.
Fascinating stuff. I just wish for a less fluffy article. Excited for the white paper.
With more computers and more data it will keep getting better - i've thought a lot about this and following this stuff for a couple of years - so it's not just some idle speculation. (I broke this recent story on hacker news a few days ago, and from here it got picked up by the new york times and spread everywhere else).
But Yann LeCun posted this [1] as a response to Andrew Ng linking to his article on Google+. Basically its a summer school about Deep Learning taught by Andrew Ng, Yann LeCun and Geoff Hinton (and others). The video's should end up online.
[1]: https://plus.google.com/104362980539466846301/posts/hhMJsPQy...
With the basic building blocks from the ML course you have a much better understanding of what Google and Stanford did. Andrew Ng is also the course instructor for the Coursera ML course[1].
I am presuming that supplemental readings and Youtube videos by Prof. Ng (e.g. http://www.youtube.com/watch?v=ZmNOAtZIgIk) allows a sufficiently motivated individual to at least partially replicate the results in the OP.
[1] Andrew Ng is also a co-founder of Coursera.