Inside Google Brain
wired.com
wired.com
No, it's an advertising company. That's who pays the bills. At the end of the day all this cool tech is to better understand and model human beings in order to better push ads.
Sometimes it depresses me that so many of the world's most brilliant minds are working on that, but a generation or two ago they'd all be building doomsday bombs. I guess that's some progress.
They do own what was DoubleClick, but it's quite different from the DoubleClick of yore (malware distributed over their network seemingly every other day, for instance).
And in practice, the same thing is true for Android and Glass: there aren't ads popping up in your face all the time, because that would be a sure way to alienate all your users (which is one of the reasons why the "they're just an advertising company" phrase doesn't give any real insight).
I think it's more likely that having a giant thing on your face that doesn't actually do all that much is more the reason for people seeming to start drawing the line with Glass. That and meme status ("glasshole") and echo chamber effects (most people outside of tech circles don't care either way).
The only reason Apple has as much money as it does is because it greatly overcharges customers, doesn't participate in research that benefits society, and takes advantage of it's customers psychological need to have the latest model (even if the improvements are minimal).
Their research isn't solely focused on gathering data for ads. If these technologies can be applied to ads then it helps justify large costs because it'll increase revenue, but that is not the sole driving motivator for their moonshots and research.
Hopefully the percentage of people who are passionate about futuristic concepts and visionary ideas will rise in the near future. The faster we get to a post-scarcity society the greater chance we, as an intelligent species, will have to survive long-term.
Just my 2 cents. I love them as a company and have a huge amount of respect for what they've done and for sticking to their founding principals of transparency and "do no evil" (for the most part, within reason for a company of their size).
I know it's pretty subjective and irreverent to bring Apple into any Google debate, the amount of people who put them on a pedestal drives me crazy. Anything I can do to redirect money from going to Apple is a positive thing in my opinion.
I'm not versed in machine learning, but it looks to me that any model whose output quality is dependent on the quantity of data it ingests is deeply flawed. There's no doubt a bigger number of samples will make the predictions more accurate, but isn't the challenge to develop a system that is as accurate as possible regardless of the number of data points its fed, like the human brain?
The thing about using lots of data is that prior to publication of "The Unreasonable Effectiveness of Data" in 2009, most people did think that good algorithms were the most important thing. What that research showed was that a bad algorithm given more data will eventually outperform "better" algorithms, at least when those algorithms are initially judged based on their performance on smaller datasets.
So what happened with neural nets was that after some initial excitement about how they were more like the human brain, etc., it was found that "stupider" algorithms actually performed better and ANNs were written off for a while. It turns out that the reason NNs were performing badly was that they weren't being fed enough data.
Nowadays it's pretty easy to saturate a feed-forward neural network with data to the point where it's performance will never get much better. Deep learning techniques allow you to train bigger and more complex models with more data, but these more complex neural nets won't perform very well unless you feed them tons of data.
So in reference to your point about the brain, the thing about brains is that they actually learn based on massive amounts of data too. Think about how much data you have from continually streaming video ~16 hours/day, plus sound, plus touch, proprioception, and other inputs, over the course of many years.
Deep learning tries to emulate this to some degree with "pre training" which is where you feed lots of data into a deep network and have it learn "something" (it learns by itself at this stage). Then you start teaching it more complicated, high-level concepts. This pre-training allows it to do things like recognise common patterns in images, which the later training allows it to then associate with semantic ideas like "this is an apple", "this is a person", etc.
TL;DR: What seems to work best is fairly "dumb" algorithms, scaled up to be able to handle vast amounts of information and fed a ton of data to learn from. This is also how the human brain works.
Regarding artificial systems, I think more data is the only way to reach super-performing classifiers. The data you supply doesn't have to be big but at least the data you extract from raw data should be big. For example, a method called Integral Channel Features [2] is designed to act in such a way.
[1] http://en.wikipedia.org/wiki/Sensory_deprivation
[2] http://pages.ucsd.edu/~ztu/publication/dollarBMVC09ChnFtrs_0...
Humans are excellent at learning from very few or even 1 example. Show a toddler a single image of an elephant and the toddler will generalize perfectly on new examples; show a machine a few thousand images of elephants and it might generalize decently if your machine is really clever.
There are very few tasks where machine systems achieve anything resembling human level performance. But on all such tasks, the machine requires far more data and still underperforms.
Nevertheless, if you provide your machine system with the video of all the child's visual input, it still won't generalize well from single examples, the way children do effortlessly.
Humans generalize well from few examples because, well, they've already processed billions of examples. A toddler may have never seen an elephant before, but it may have seen cars, trucks, birds, dogs, people, trees, skies, buildings etc, giving it concepts for bigness, smallness, aliveness, humanness and much else. With all these concepts in place, then yes it becomes easy to see what makes an elephant distinct from a dog or a person. And it would be too for an artificial neural network.
An interesting fact is that newborns have very few concepts to begin with. It takes some months for them for instance to learn to differentiate between alive and dead things (the family cat vs a teddy bear for instance).
If you show a child 1000 images or animals. Then show different photographs of animals. And tell the child what animal each animal photograph is, you can now go back to the original 1000 and the explained ones will likely be recognized dispute them never beig initially sorted, or modeled as such.
Going from 5-10 to 1,000,000 is what computers have a problem with. They go from 1,000 to 1,000,000 easily, or even million to billions.
What I'm talking about is a child can see a photo of an unknown animal, I can show that child a cartoon elephant (which is the original animal). I then ask what the original animal is, the child likely responds correctly.
Reprocessing of already learned data as the scheme of the world changes based on new information.
It's rapidly becoming apparent that some algorithms (eg Deep Learning related models) work much better at scale than on small amounts of data. It doesn't make sense to discount these better algorithms because they don't work as well as other models when tested against less data.
It is also apparent that these models require significantly more computing power to perform well than other models. That doesn't make them less worthy, just a cost people must consider.
It turns out that intelligence is hard..
Statistics developed as a science because of the need to overcome the weakness of large samples being expensive. Machine learning has taken off as a direct result of the field's ability to take advantage of and get serious performance gains from the massive amounts of data being generated and leveraged recently.
Here is the best summation I can reference, and I can tell you from personal experience it is very true:
"The accuracy & nature of answers you get on large data sets can be completely different from what you see on small samples. Big data provides a competitive advantage. For the web data sets you describe, it turns out that having 10x the amount of data allows you to automatically discover patterns that would be impossible with smaller samples (think Signal to Noise). The deeper into demographic slices you want to dive, the more data you will need to get the same accuracy."
http://www.quora.com/Big-Data/Why-the-current-obsession-with...
http://anand.typepad.com/datawocky/2008/03/more-data-usual.h...
One way to think about it is to look at problems with human perception like forced perspective. It's relatively easy to create a situation where the only available information results in mental models that describe the size of an object incorrectly. Given a different point of view (i.e. more information) the faults in the model become obvious.
We often don't know which model to use. Occam's Razor [1] can be effective in favouring simpler models, but I tend toward the view that a good data scientist is invariably needed to build good models. Hence I view Big Data more as a consulting business than SaaS.
[1] For an excellent Bayesian discussion on why Occam's Razor actually works, see Chapter 28 of David J.C. MacKay's book 'Information Theory, Inference and Learning Algorithms'.
You often want your models to also perform well when you have fewer data points. Those are two separate - if in effect related - design goals.
Seriously, Norvig has been big on this since forever: the reality is that consuming large amounts of data with relatively subtle features tends to be one of the few areas where computers can easily outclass the human brain.
Can this meme die soon? Google is an ad company in the same way the New York Times is an ad company. Sure, that's ultimately where the money comes from, but there'd be no ad revenue if search, maps, mail etc weren't all among the very best available and that's a very real and very big engineering problem. Just like if the NYT stopped during journalism, their ads would stop generating revenue.
http://techcrunch.com/2012/03/29/google-now-using-recaptcha-...
I remember on my pre-broadband days here in Brazil subscribing to Doctor Dobbs Journal to the tune of 25 USD per issue and being happy. Each issue filled with little gems that would advance my knowledge a lot... these days its all those glossy covers with photoshop covers and over the top headlines.
:-(
"The singularity is near: When humans transcend biology" "The age of spiritual machines: When computers exceed human intelligence"
I'd argue that Stephen Boyd is more of a master of AI than Kurzweil.
It wants to be seen that way, but until it abolishes closed allocation (and maybe it has, but I haven't heard anything to indicate that it has) it will just be an ads company. It's still a pretty good place to work, by industry standards, but the percentage of people who'll get to work on machine learning is very low.
Google definitely wants to have the image of being the machine learning company because that's a great way to attract talent (even if that talent is mostly wasted under closed allocation). And if you land in the right place, there is interesting work. The reality most people face, though, is that most people (especially outside of Mt. View) aren't going to get real projects and won't be anywhere near the machine learning work.
Google does have a lot of talent and probably would be the undisputed #1 tech company if it implemented open allocation, though.
It clearly is a tech company, but to pretend it's anything other than an ad delivery machine is sort of ignoring the elephant in the room.
Everything goes back to ads. Even Google doesn't bite the hand that feeds it.
lol. This is the grand payoff?
def ER(accuracy_before, accuracy_after):
delta = accuracy_after - accuracy_before
return delta / (1. - accuracy_before)
(It is often written as a percentage.) Note that ER(.5, .75) = ER(.98, .99) = .5. It is for this reason I dislike it. Without knowing what performance before and after was like, it's hard to tell whether this was an incremental improvement to an already impressive system or the innovation that made a system useable.(It strikes me as unlikely that the author means to say that Google reduced the Android recognizer's word error rate [WER] by 25%, since WERs on many publicly available databases used for evaluation purposes were already well below 25% before the deep learning revolution. But, it is not impossible that Google's test set was particularly hard and the pre-deep-learning Android recognizer was unimpressive. Caveat: I know, I know, WER isn't really "accuracy".)
See, for example, compressing the ENWIK8 in the Hutter Prize.
If you can compress this approx 100 MB file to less than approx 16 MB (including the decompressor) you win cash.
That website shows small decreases in filesize over a few years.
Alexander Rhatushnyak 23.May 2009 6.27 | 1614€ Marcus Hutter
Alexander Rhatushnyak 14.May 2007 6.07 | 1732€ Marcus Hutter
Alexander Rhatushnyak 25.Sep.2006 5.86 | 3416€ Marcus Hutter
Matt Mahoney 24.Mar.2006 5.46 | pre-prize -
The compression ratio goes from 5.46 to 5.86 to 6.07 to 6.27.BTW, this point elucidates the importance of this classic Google Interview question: http://www.glassdoor.com/Interview/Sort-a-million-32-bit-int...