Researchers achieve speech recognition milestone
blogs.microsoft.com
blogs.microsoft.com
[1] http://www.itl.nist.gov/iad/mig/tests/ctr/2000/h5_2000_v1.3....
People regularly ask each other, "sorry, what did you say?", "wait, what did she say?", "would you repeat that please?", "huh?", etc.
Contrast this with speech recognition, which will often substitute words that are nonsensical in context, making it look silly from a human perspective...
I think the chinese room experiment overlooks this part of "understanding"
Seems like a good approach.
Or that is my observation, anyway. I don't use it myself.
Understanding rate is less than 10%. If you don't match a keyword it gives a useless web search.
Personally I don't think understanding rate is the whole issue as much as reaction to error (which is partly understanding). You can't say "no that's not what I said" and Siri et al never keep enough context to say "huh? What did you say? Or "I didn't get that last part. can you repeat it?"
It's that errors in understanding or accuracy turn the whole thing into a complete shitshow.
One failure and you might as well pull over and type what you want.
They can cancel out reverb and create very fine tuned waveform profiles for speech.
I think one of the reasons that Siri is slightly better at SR than google is because of the control that Apple has over the hardware.
While Cortana turns sourpuss on me every time I switch headsets.
It's better to avoid throwing around numbers like that but even if that was the case you have to remember that humans understand speech. The speech recognition task performed by AI systems on the other hand is more akin to transliteration: the system takes in sound as input and produces text as output. Any sort of "understanding" a) is extremly difficult to do well and b) must be performed by a different component of the system (a different algorithm, trained on different data).
For humans, isn't this a due to a combination of factors than just comprehension alone? Humans who ask, "sorry, what did you say?" or "would you repeat that please?" or even just a "huh?" usually aren't paying attention at all. It's not a comprehension or sound quality or surrounding noise problem for many, except in situations where the person is not fluent in a particular language or dialect or accent or if the surrounding noise vs. the person's hearing ability aren't conducive to listening properly.
Most people also usually tend to think about judging what the other is saying and constructing a counter-point during the process of listening that impairs the ability to listen and understand well.
On the other hand, a computer could expected to be, and made to be, paying attention a lot better in a predictable way, which is not possible with humans.
With the other comment reply above stating people's expectations with humans vs. computers, shouldn't we also consider the computer's strengths while making comparisons with humans?
The Microsoft system also does this: it uses language modelling to attempt to model what word is more likely in a given context. This gives them a 18% word error rate reduction (see section 7 of the paper: http://arxiv.org/pdf/1609.03528v1.pdf).
What is somewhat unusual about this approach is they use recurrent neural networks for the language modeling, as opposed to more traditional approached like a backoff n-gram model.
That's 6% on the NIST dataset. Typically, results get much worse on real-world datasets, not least because trying to get good results on the same dataset year after year leads to subtle bias. Don't forget this is a dataset that's been around for 16 years, now.
I worked in speech recognition for a bit and 8kHz was the standard audio rate for recordings. No one saw this as an impediment; and it really isn't. 8kHz is able to capture up to 4kHz frequencies and the fundamental frequencies of speech are MUCH lower than 4kHz.
Edit: For comparison: http://www.utdallas.edu/~assmann/hcs6367/lippmann97.pdf
What is amazing is how consistently people overestimate human performance.
Frankly, we're lousy drivers.
On the one hand, there are very good ethical arguments, for example about who bears the moral responsibility when a self-driving car is involved in a fatal accident.
Further, there is a great risk that self-driving cars will become available long before they are advanced enough to be less of a risk than humans, exactly because people may implicitly trust an automated system to be less error-prone than a human, which is not currently the state of the art.
Conversations I've had where people have told me that self-driving cars will need to be 100% perfect before they should be used. Ironically, one of those people was an ex-gf of mine who caused two car accidents because she was putting on makeup while driving.
Anyway, based on Google's extensive test results, I'm pretty sure self-driving cars are already advanced enough to be less of a risk than humans. Right now, the sensors seem to be the limiting the factor.
This should make, er, sense. Sensing your surroundings is only the first step in taking complex decisions based on those surroundings. The AI field as a whole has not yet solved this, so there's no reason to expect that self-driving cars have.
Seen another way, if self-driving cars could really drive themselves at least as good as humans drive them, we wouldn't have compilations of videos of robots falling off while trying to turn door knobs, on youtube.
The state of the art in AI is such that self-driving cars are not yet less dangerous than humans.
>> an ex-gf of mine who caused two car accidents because she was putting on makeup while driving.
Honestly.
Google's self-driving car accident statistics say otherwise.
>Honestly.
Yeah. Weirdest part was, she actually thought she was a GOOD driver. Mostly because of all the times she was able to apply makeup while driving and didn't cause an accident.
That's an experiment that's been running in a tiny part of one state in one country for a very limited time. I wouldn't count on them and in any case, see what I say above: the state of the art is not there yet, for fully autonomous driving better than humans'.
>> Yeah. Weirdest part was, she actually thought she was a GOOD driver. Mostly because of all the times she was able to apply makeup while driving and didn't cause an accident.
:snorts coffee:
It's driven over 1 million miles. That's the equivalent of 75 years of driving for the average human. Plenty of data to draw a conclusion from. In all that time, it's been responsible for a single accident. That's way better than human drivers.
Then there's the fact that human drivers have to drive in all sorts of weather conditions with all sorts of different vehicles and so on. Google car- not so much.
But my point is very simple: AI in general is nowhere near producing autonomous robots, yet. Why would Google car (or a similar project) be the exception? What makes cars and driving so different that autonomy is easier to attain?
The recent death of the Tesla owner, for example, as far as I know, was due to the vehicle accelerating into a semi. This is something that most people would not do even in their worst driving state unless they were intoxicated or seriously mentally impaired. I don't want AI driving errors to be compared to human benchmarks that include people who are seriously intoxicated.
A lot of speech frustration problems, similarly, are not only about poor recognition in general, or lack of appropriate prompting to increase classification certainty, but recognition failures in situations where a human would not have any trouble at all, such as in recognizing names of loved ones, or things that would be clear in context to a human. I.e., maybe humans listening to speech corpora would have x% error rate, but that's strangers listening to the corpora. The real question is, if I listen to a recording of my spouse or coworker having a conversation what's the error rate there?
So, although humans are far from perfect, which is something that's often forgotten, the true AI target is also probably not "humans broadly defined" but rather "functional humans" or something like that. AI research often sets the bar misleadingly low because it's so hard to reach as it is.
Another example. If a self driving car is hit by another car that's running a red light while speeding, we might be more forgiving and say "well nobody could have avoided that accident" but actually we'd be being too soft on the self driving car since it has access to more data and faster reaction times and should probably be expected to avoid that type of crash even when a human can't.
You're absolutely right that there are risks. But honestly, I suspect drunk drivers alone cause more fatal accidents than autonomous cars ever could.
Unfortunately... I think the big issue is going to be pure anxiety; a bigger and more immediate form of what a lot of people experience on an airplane. Giving up even the illusion of control is supremely hard for us as a species, in general. Then there's just the fact that as a species we're terrible at risk assessment.
https://www.schneier.com/blog/archives/2006/11/perceived_ris...
I did say "primary interface" as well, which definitely rules out a mediocre phone connection.
That's a troublesome definition though at least in part for reasons you've brought up or alluded to, which is that a multimodal approach is pretty clearly going to dominate. That said, speech recognition at least stands to replace the keyboard for say, the author of a book or article, if it's good enough.
So just based on observation I would say it works pretty good. On mobile I almost always use voice search, and unless I'm searching for an italian name or something weird it almost always hears me correctly. Even in a noisy pub.
It seems to be significantly less than 1% - 1.6%
Switchboard is kind of a lame evaluation set. It's narrowband, old, and doesn't contain all that much training data (100s of hours, whereas many newer systems are trained on 1000s or 10Ks of hours). And the quest for a lower Switchboard WER to publish means teams are now throwing extra training data at the problem, or using frankly unlikely-to-be-deployed techniques like speaker adaptation, impractically slow language models, or bidirectional acoustic models (which require the entire utterance before they can emit any results).
I really wish they would have stuck to just publishing a paper explaining was actually new here (ResNet for acoustic models? Cool!) rather than just a "let's see how low we can push this 20 year old benchmark" paper.
Are you saying the benchmark is useless? It's old, yes, but it's extremely valuable to have a benchmark that allows one to assess system performance over time. It gives a good idea of the rate of progress and the distance still to go to match people - after ~16 years, computed are still about a third worse than humans, error rate wise.
If you looked at modern performance on test sets from the 1980s (like Resource Management or TIDIGITS) you might be under the impression that we'd achieved human-level accuracy levels years ago, but we clearly haven't. And similarly, what users expect from speech recognition today is in many ways much more demanding than it was in 2000: vocabularies are huge (think about all the words you could say to Google), latency needs be very low, and no one thinks it's acceptable to require users to perform enrollment any more.
So yes, just like other benchmarks, we should retire them after a few years. The fact that a modern computer could get 100,000 FPS on a video game from 2000 wouldn't be considered a "milestone."
I'm not the OP (who replied already) and I don't think old benchmarks are useless but I'm worried that teams trying to beat a dataset from an old competition for a sufficiently long time will inevitably overfit to the dataset, reducing the accuracy of their published results. That's even more so when the test set for the competition is available and there's nothing really keeping it from "creeping" into the training set at some point, maybe between different system versions.
What would really be useful is a sort of ongoing challenge where a training set stays up for a decade at least and the test set is never revealed (but can be used to test systems). Perhaps data could even be renewed every few years as long as new examples can be reliably collected in a similar enough manner with older data.
w.r.t. run time, though, agreed. Hearing the IBM folks say "... 10" in response to the "what's the RTF" question was funny.
(and, agreed, at this point the switchboard announcements are definitely just marketing.)
And it's _really easy_ to increase accuracy by taking more time, by: building bigger DNN acoustic models; exploring a larger search space of hypotheses; using a slower language model (like an RNN) to rescore hypotheses; considering more possible pronunciations; etc....
(ML is usually a space / time / accuracy trade-off, so if you get phat accuracy gains at the cost of significant slow down, I'm usually unimpressed. The deepmind TTS paper _was_ impressive because it went beyond the best we can do, so even though it was 90 minutes to generate 1 second of speech, it's cool because it shows where we can go. TBH all of these switchboard papers don't do a ton of new stuff, they just get more aggressive about system combination and tuning hyperparameters.)
---
[0] It is like that any viability analysis would be on an by-application basis, so I don't pretend like I'm asking for an insignificant amount of work here!
[1] a crude, toy, and likely inaccurate example. Not trying to belittle the work.
Also, major issue with this kind of research is that they combined several systems in order to get best results. Most practical systems don't use combinations, they are too slow.
Also I'm note sure an error reduction of 20% (1-6.3/7.8) is to be considered small; depends on the particular challenge really. Like, sentiment analysis only starts to get interesting above 80% on some dataset, as much can be guessed correctly in very naive ways..
Human lvl on this task is estimated to be ~4% so we have quite a lot of ground to cover still..
In Microsoft paper http://arxiv.org/pdf/1609.03528v1.pdf Table 5 it's a line "Povey et al. [19] LSTM". http://www.isca-speech.org/archive/Interspeech_2016/pdfs/059... Interpolate it to RNNLM column and you'll get 7.8. See also http://www.isca-speech.org/archive/Interspeech_2016/pdfs/047...
I think I get about 50% hit rate with 'OK Google' and I'm a native English speaker :).
"Okay, Google." Nothing.
"Okay, Google." Nothing.
"Okay, Google fucking work or I am taking a hammer to this fucking phone." "DING!"
Apparently threats of violence still work against our machine overlords.
WaveNet: A Generative Model for Raw Audio https://deepmind.com/blog/wavenet-generative-model-raw-audio...
"Was it just coincidence that computerspeak, which we'd learned with such throat wrenching difficulty, was also a complaining, niggling, nit picking noise that gave any normal human being a pain in the arse?"
Now me, I used Graffiti a ton as well as the not quite so single stroke and therefore not so patent encumbered version that MS used on their late PDAs. I even implemented my own version on an early windows touchscreen netbook (angular delta and Levenshtein distance, match) so much so that it still pollutes my hand writing to this very day. I find it elegant and immediately understandable. Every one else thinks I'm writing in pig pen cipher.
https://deepmind.com/blog/wavenet-generative-model-raw-audio...
I'm sure their accuracy is fairly reasonable, though.
I would have though the state of the art would be better, given anecdotal evidence from friends who write with speech to text programs, and love them.
Perhaps some of this is due to deliberately bad audio quality in the switchboard samples.
I like it when people say "Can you repeat that, I was on mute"
Now that bandwidth is becoming less of an issue, we will be getting less shitty sounding, wider bandwidth phone audio - https://en.wikipedia.org/wiki/Wideband_audio
Though if I had it my way, U87 or U47, or hey even SM7B would be mandatory for all speech recordings :)