Why 2015 Was a Breakthrough Year in Artificial Intelligence
bloomberg.com
bloomberg.com
The second boom was expert systems and logic super computers. Those systems never worked that well and A.I. went into long sleep.
Now its supercomputer data mining and greatly improved neural networks.
The academic image recognition machine though seems unstoppable and yes it does seem to improve over time. I honestly don't know what the limits of "dumb" image recognition are in terms of quality, but calling it AI still doesn't make sense to me.
That's the core problem of AI, no matter what progress is made, it's instantly called not AI anymore while the goalpost of what AI is is continually being pushed out to not this. The real issue of course is what don't how intelligence actually works so it's impossible to set a fixed goalpost of when AI is truly achieved.
Or is intelligence.
No, the real issue is that people still think, intuitively, that there's a little homunculus in their head making the decisions. Each time we build something that doesn't look like a homunculus, we've failed at attaining "real" intelligence...
https://en.wikipedia.org/wiki/Simultaneous_localization_and_...
... but I agree that there's nothing in particular that distinguishes it from other problems.
My comment was satirical in nature. SLAM is an interpretation of what the parent comment had described:
"[B]eing aware of your own location in space and time, as well as remembering previous locations. In other words, having some model of the world and knowing your place in it[.]".
There is a general pattern of statements of the form "We'll only really have AI when computers X", followed by computers being able to X, followed by everyone concluding that X is just a simple matter of engineering like everything else we've already accomplished. As my AI prof put it, ages ago, "AI is the study of things that don't work yet."
Which is exactly what we do with many kids today; makes you wonder how many times we might invent AI and not know it because we don't raise it correctly so it appears too dumb to be considered a success.
Causal induction: sounds interesting until you dig in and realize everything is non-computable.
So what exactly is your point?
Wait, what?
We have a kind of fixed goal post in human intelligence. When computers are worse at something than humans like chess it's thought of as intelligence and when they get better it's ticked off as just an algorithm. The AI researches gradually tick off abilities. Chess long ago, image recognition happening now, general reasoning at some point in the future.
For (classic) video games it's actually somewhat similar but rather than a board being fed in as an input you just feed in a bitmap of the display (sometimes at a lower resolution using compression techniques to reduce input features) and optimize moves made to maximize the score at any point rather than end-game.
But as soon as chess programs got good, we all took them for granted.
The reason it's AI is because it isn't specific to speech recognition. Deep learning is very general. The same algorithms work just as well at speech recognition, or translation, or controlling robots, etc. Image recognition is just the most popular application.
Certainly, a new born baby given a video game controller would not be able to figure it out.
Deepmind's Atari player is based on deep reinforcement learning, where increasing the score represents a reward:
https://m.youtube.com/watch?v=EfGD2qveGdQ
I used to believe the moving goalpost idea, that AI is anything that isn't yet possible. I now disagree.
I saved myself a lot of energy by avoiding entirely the question "what counts as AI?" by switching to the question, "what counts as AGI?" which is a term with a clearer threshold:
https://en.m.wikipedia.org/wiki/Artificial_general_intellige...
If you don't know how it works -- it looks like magic. It can tell a donkey from a horse, it can play checkers, diagnose a patient etc.
After you are told the trick -- it is just A*, or Rete algorithm, or a multi-layer NN. It makes it less magic and it becomes just another algorithm.
A good introduction to how the visual cortex works can be found in Eye, Brain, And Vision: http://hubel.med.harvard.edu/book/bcontex.htm
[1] Sorry, don't have a cite, but there are a lot of results on Google talking about it: https://archive.is/iDE27
60 years later, we have finally made general purpose learning algorithms, vaguely inspired by the brain, which are just powerful enough to do it. And because they are general purpose, they can also do many other things as well. Everything from speech recognition, to translating sentences, or even controlling robots. Image recognition is just one of many benchmarks that can be used to measure progress.
Their machine translation system probably has a similar # of DNNs, and there you have to deal with language pairs, rather than single languages. Let's call it another 400.
That's two side-projects. Then you pull in query prediction, driverless cars, all kinds of infrastructure modeling, spam detection, all of the billions of things that are happening in ads, recommendations, I haven't really even mentioned search yet... Honestly, if I'm right in assuming that the cited figure is really "# of DNNs that do different things", then I'm surprised it's not higher.
We had Dragon natural speaking on a 133MHz Win95 PC (offline of course). After training it for like 10min it worked better or equal as good as Ford's Sync car assistent (offline) and Siri/GoogeNow/Cortana. Well all these services licensed the Nuance speech technology which they got from buying the company behind Dragon natural speaking software. The Ford board computer runs WinCE and has only 233MHz and is still sold in many 2016 Ford cars around the world. And with cloud hosting, to scale the service each users gets only a small amount of total CPU timeslice anyway.
What I want is an offline speech recognition software on my mobile devices! So do I have to install Win95 on an emulator in my smartphone just so my multi-core high end smartphone can do what a Pentium 1 could do in 1996? My hope is on open source projects. Though most such OSS projects are university projects with little documentation how to build the speech model, little community, on an outdated site, written in Java 1.4 and no GitHub page. There is definitely a need for good and competitive C/C++/(native code) TTS and speech recognition project.
I find it hard to believe, do you have any citations for that - or is that just your gut feel?
A cursory search shows a 26% error rate[1] for Dragon NaturalSpeaking in the year 2000 (beaten by IBM in the same report at 17%).
By May 2015, if Sundar Pichai is to be believed, Google has an 8% error rate[2]. In my books, 26-to-8% (or even 17-to-8%) is far from barely improved.
1. http://www.ncbi.nlm.nih.gov/pmc/articles/PMC79041/#!po=1.562... Table 3, General Vocabulary
2. http://venturebeat.com/2015/05/28/google-says-its-speech-rec...
Much of the Google's stuff is for search term recognition only. It's functionality on general dictation is nowhere near that good.