Researchers reach human parity in conversational speech recognition
blogs.microsoft.com
blogs.microsoft.com
On the CallHome dataset humans confuse words 4.1% of the time, but delete 6.5% of words, most commonly deleting the word "I".
Their ASR system confuses 6.5% of words on this dataset, but only deletes 3.3% of words, so depending on how you view this their claim about being better than humans isn't definitely true, if you consider the task to be speech recognition, rather than transcription.
Also, while the overall "word error rate" is lower than humans it's not clear if this is because the transcription service they used is not seeking perfect output, but rather good enough output and the errors the transcription service makes may not be as bad as the errors the ASR system makes in terms of how well you can recover the original meaning from the transcription.
It's clearly great work, but reaching human parity is marketing fluff.
I think one important aspect is while humans miss words, they often get the sentence meaning correct. When computers miss words, they tend to substitute words that sound similar. That's readable if you have time but not necessarily as a stream of text going by...
The character on the screen said "Hello", the English subtitle said "You good." Which would be the literal translation of nihao.
I don't think that is possible. Like, worse subtitling isn't even an option.
http://www.nbc.com/saturday-night-live/video/british-movie/n...
So if this speech recognition is for creating a transcript for humans, this (in my uneducated opinion) isn't as good as humans do, at least not yet.
Being Microsoft, I'm sure the place is diverse, but there's no mention in the paper on accents or dialects that I can see (might just be missing it).
To explore that analogy: In the case of lossy audio compression, the compressor deliberately introduces quantization noise into the signal. It does so by running a "psychoacoustic model", which attempts to capture a broad quality of human hearing called auditory masking. There's a number of different kinds of masking[1]: a strong tonal sound creates an "umbrella" across nearby frequencies that can mask quieter noise-like sounds. Similarly there's noise-vs-noise masking, as well as forwards- and backwards- temporal masking. (Yes, backwards. A sound can mask perception of a sound that occurred before it.)
In the audio compression case, we've built an algorithm that attempts to characterize exploitable phenomena of human hearing. These masking characteristics aren't perceived quite the same by human listeners as the model, nor even the same between individual human listeners. Thus the need for human listening tests. These, due to the experimental care and human subjects required, are expensive.
Back to speech transcription. Say the end goal is "how well does a human comprehend this transcribed speech"? (e.g. vs some standard, such as the original speech, vs. the original speech transcribed by a skilled specialist, etc.) The problem starts to look pretty similar. We can cite numerical, word-centric error rates, but that fails to capture how well meaning is preserved and transmitted. Imagine a perverse algorithm that did a perfect transcription, but then dropped or altered words for maximum meaning obfuscation. It might equal or even beat the cited error rates but be much harder to actually comprehend.
These classic signal processing analyses tell you nothing about the correct operation of a psychoacoustic model codec design (i.e. MP3, AAC, Vorbis, etc.) An analysis of A (original signal) vs. A' (signal passed through a compression-decompression cycle) is just extracting the quantization noise introduced by the codec. That provides no information about how effective the codec was in masking that noise with the original signal content.
To illustrate, imagine a perversely designed codec: it runs two models, the first "good model" is a normal psychoacoustic model. The second "bad model" is the one the compressor uses: it applies the total amount of quantization noise allowed by the good model, but applies it in ways that are maximally annoying to human listeners. This isn't just avoiding masking, it's things like using noise correlated to the original signal, which is generally more obtrusive than uncorrelated noise, etc.
A codec using just the "good model" and one using the "good + bad model" would have (by definition) exactly the same introduced noise, but the latter would sound FAR worse to a human listener.
No joy.
I can't really follow your thinking above; sorry.
In audio, there is distortion, which is correlated with the signal. Noise is uncorrelated. Codec error would seem very much to me to be at least much more correlated than uncorrelated. MP3 artifacts sound at least more "phase-ey" than they sound anything like quantitization noise ( at least to me ). This may be because I have heard badly aligned tape machines make that sort of error happen. "Phase-ey" also triggers my (feeble) mind into thinking about allpass filters as the model for the error.
In terms of intelligibility, it's possible to improve intelligibility by adding noise, and by adding clipping - aviation comms does this at times. What destroys intelligibility is phenomes being destroyed by phase changes and bad amplitide errors ( where there are actually good amplitude errors ).
I've run an ABACUS voice quality analyzer several times, and I don't think it's purely a distortion analyzer - adding clipping at least can improve MOS/PESQ score surprisingly. Even more surprisingly, there's no mechanism available for calibrating gain staging on one.
If I understand you correctly, your discussion around intelligibility primarily applies to signal processing of voice vs general audio. Codecs such as MP3, AAC, etc. don't have the luxury of making assumptions about the signal content, and so aren't designed along those principles. E.g. speech codecs can generally run well at much lower bitrates than general audio codecs because they operate on a constrained domain of audio (i.e. speech).
Regarding distortion management, see Rate-distortion optimization[1] for lossy audio codecs: where the purpose is to manage distortion within the limits of the bit rate supported by a communication channel or storage medium.
[1] https://en.wikipedia.org/wiki/Quantization_(signal_processin...
Google Voice transcription often makes easily detectable mistakes that I can mentally correct using context. Smarter, context-aware mistakes might be worse, if they are also more difficult to detect.
We find that the artificial errors are substantially the same as human ones with one large exception confusions between backchannel words [acknowledgment words like “uh-huh”] and hesitations.
The difference they found, but suspect might be a result of the different transcription guidelines of the training corpus: we see that by far the most common error in the ASR system is the confusion of a hesitation in the reference for a backchannel in the hypothesis. People do not seem to have this problem.
The term 'human parity' refers to adding up all the bits in a human and taking the residue to a convenient modulus to detect human error.
A tough test would be to hook this up to a police/fire scanner, or air traffic control radio.
Presumably you could feed it speech from a running instance of gnu-radio.
http://www.volkskrant.nl/tech/privegesprekken-van-duizenden-...
Also, don't assume that a 0.4 % increase means drastically better real-world results. This dataset has been around longer than I have, so by this point Microsoft has just gotten really good at tuning.
> Still, he cautioned, true artificial intelligence is still on the distant horizon
It's frustrating when technologies like image and speech recognition and robotics are conflated with AI.
Are you kidding? Of course these things are examples of artificial intelligence. I don't understand why people keep moving the goalposts wrt "AI".
My own taste uses the word "AI" as you do in a permissive way to include simpler tasks (which in themselves are more elemental than useful) like identifying an object in an image and presenting a few straightforward interpretations of sentences. But what Turing stipulated an actual AI would be able to do is "reach parity" with a fully human conversation, with all of our knowledge and values and desires and subtlety of motivation. When you really dig into what goes on in a real conversation---the joking, the shades of meaning, the individual quirks and tribal patterns, the lying, the compassion, ambition, insecurity---I think it's hard not to admit that we're still quite far from that. We may not even want it! But exactly the genius of the Turing test for setting a rough standard for AI is that it demands competences far beyond mere language processing.
I'm sure Turing's own goals for AI were broader than eventually passing the Turing test. He worked on neural nets himself and would've considered progress in perception as partial progress in AI.
Just to link up the way you put things with the way I chose to here, his argument for how an artificial mind is possible proposes a reasonable, minimal standard which different sides can agree to---a criterion of, "well if it can do that then sure it's a mind!". And the choice of a open and unbounded conversation as the standard was brilliant because of the massive range of subcompetencies which are required for actually executing it (including, obviously, perception). Which, of course, we gradually continue to plod through in AI research.
Doesn't make it any less frustrating, but also distinguishes the people in it versus those who aren't.
Ray Kurzweil understood this and is why he paved the way for some of the first speech / image recognition platforms like ocr/fax/etc...
To me speech/image recognition is a precursor/adjunct of AI. You can have the former without the latter, but the latter will never be realized without those.
Thanks for the feedback! Gives me more to think about.
Wait, what? The most commonly known (and perhaps oldest) AI test in no way requires either image nor speech recognition.
Millions of blind and deaf human beings would like to disagree with your claim that they are not intelligent or sentient beings.
Seriously, what is the basis of this claim?
Intelligence (at least in this case) is being able to take data and draw conclusions from it, but you just can't see such potential when you don't have input. Is a computer not a calculator when there's no software installed? No. It's still a calculator, but it just doesn't have inputs. One day, we may be able to give sight to the blind, but for now, considering such people without ANY learning capacity as "unintelligent" is still wrong.
And yet people like Ray Kurzweil toss the term around as if it does have meaning. What's the point unless you define it? You might as well reference anti gravity or teleportation.
Kurzweil was a great inventor once but he's a prophet now and he needs to speak prophetically, not scientifically.
One of the biggest criticisms of the Turing test is, because it's a behavioral or functional test, there is no way to ensure that the computer is actually thinking intelligently at all, rather than following a very-well-articulated rules engine.
Can you think of a better one?
> because it's a behavioral or functional test
What else would you test for in an AI other than its output and behavior? Conversely: how would you test a human for intelligence other than through its output and behavior?
A not-very-bright cousin of mine has been heard responding to robo-calls. Would his opinion do?
I think 'intelligence' is ambiguous, but many parts of it can be measured to some degree. Lets give the AI the SAT test perhaps? Or a test of hypothetical arguments. Or ask it how it would tell if you were an intelligent being yourself. Anything that's even a little bit 'meta' would confound most 'AIs'.
We can tell if an AI is functional in some environment or responds well to social queues by interacting with it. But not much else. Not how 'intelligent' it is for instance.
> We can tell if an AI is functional in some environment or responds well to social queues by interacting with it. But not much else. Not how 'intelligent' it is for instance.
The Turing Test does not intend to determine level of intelligence (it's not an IQ test). It is intended to determine whether the intelligence you're interacting with is advanced enough to fool you into thinking they are of human nature.
It is not a test of degree of intelligence, but rather of its nature. Human / not human are the only two possible outcomes.
With AI - we built it. We can read the code that's running. We can make a distinction, in cases where output or behavior is identical, whether it's a neural network or a 12-million-line switch statement - and that value judgment means something to our determination of intelligence. There are plenty of ways other than (or in addition to) output to determine the "intelligence" of a machine that we simply can't exercise with other people.
Note: Medical tests generate both false positives and false negatives, they are still useful.
Is your argument that that person is still intelligent even if they've failed the test? Because that's the entirety of my point - the Turing test does not return a measure of intelligence, but of communicability.
Or is your point that if such a person fails it doesn't invalidate the test's measurement of intelligence? In a world where the test is implemented correctly, meaning where things like cross-language barriers are accounted for, failing to pass means failing to convince another person that you can communicate like a human would. If you fail, the only way you can consider that a "false negative" would be if you concede that the test is not a sufficient measure of intelligence but of ability to communicate - that's what makes it a FALSE negative.
Cost and accuracy are obvious trade-offs, but sometimes you are willing to trade say cost and false positives to avoid false negatives say, a mass screening for HIV in a blood sample. In that case you need a cheap test and while a false positive has minimal cost a false negative could be deadly.
Highering is the opposite case. Micdonalds want's a cheap test (interview + background) and as long they get enough acceptable low level candidate from their pool that's enough.
As such the turing test can fit the second example. A hypothetical AI could hold a huge range of real jobs even if limited to purely text based communications.
PS: Let's flip it. A super intelligent AI in a box that can't communicate in any form. Without IO it's indistinguishable from a space heater.
A passing grade on the Turing test just indicates intelligence is possible, but excludes assured intelligence in the absence of communication ability, which is why I argue it is not a sufficient guarantee of intelligence.
You make a good point with the hiring example about the distinction between "intelligent" and "intelligent enough to do some things" - our disagreement may be stemming from us having differing definitions of "an intelligence". Need to think about it.
X is a demonstration of intelligence says nothing about not X.
If you can design something that passes an arbitrary touring tests yes it is intelligent. For example I could teach it any subject that works in text format. It could then pass an open ended essay test. And get a reasonable essay.
Now, you could do the same thing with a pig and it would fail the test despite being more intelligent than the average dog. That just means there are limits.
Thanks for the great chat.
Think about the intelligence conveyed in ritualized greetings.
EDIT: I see the nuances of epistemological problems are lost to HN and knee-jerk culture war atheism still rules supreme.
When you drag people out of their fantasies, they used to kill you. At least they downvote you these days, so I supposed that's better.
Yes, this is an incredibly unique view. Yes, people are going to react negatively to it because it fundamentally undermines their revenge fantasies. Yes, I can call them out on their unspoken biases no matter how uncomfortable that makes them.
Silicon Valley thinks AI is not an epistemological problem, as if neurons can be perfectly simulated atomically and that all intelligence processes can be categorized as structured vs. unstructured. Very naive conclusions.
Ironically, they BELIEVE if you simulate the axiomatic neuron perfectly, emergent properties of intelligence will mystically emerge after some undefined threshold of complexity. They are permanently unable to simulate the two billion years of neurological evolution and natural selection for such a belief to hold.
If you must emulate human intelligence, you must be able to navigate the realm of distilling reality from the overlap between truth and belief. Mocking how humans observe reality isn't enough.
Abrahamic religion explored the alternative intelligence problem in great depth over two thousand years ago. It's a pity the results have been lost to Progressive axe grinding.
I'll take the bait: What in the world are you talking about?
A man doesn't need to believe in God to be intelligent. In fact, the two are pretty much inversely related.
That's a very limited sense of the idea of a god, and certainly doesn't match with definitions of God [a singular, eternal, omnipotent, omniscient, deity] that I've come across.
Merely making something doesn't make you a god, not even if that thing appears to display intelligence. Some sense of one of the characteristics of existing in a separate spiritual realm, having power/knowledge beyond that possible in the present realm, having an existence that's not bounded (eg physically) within the normally experienced space of the "mortals". They seem like a start for basic level definitions of a god.
There is nothing wrong with that sort of metaphor, no matter how much stress you place on it. If you think of religion as a technology, then it can be abused but the abuse doesn't mean it's bad.
Claiming as fact. Citation, please?
Furthermore, the purpose of this thought experiment is that you CANNOT know what god an AI would end up believing in. That should make you rethink everything you think you know about theology and reexamine it under epistemological terms
Do you have some evidence of that? And even if it's true, is it the only thing that's unique to humans?
Indeed, I suspect the interesting part of "believe in god" in terms of intelligence is probably "construct narratives". Assigning a high probability to those narratives reflecting reality (i.e. believing) seems like an aside to the main complexity involved.
>"The anthropologists got it wrong when they named our species Homo sapiens ('wise man'). In any case it's an arrogant and bigheaded thing to say, wisdom being one of our least evident features. In reality, we are Pan narrans, the storytelling chimpanzee." - Terry Pratchett, The Science of Discworld II: The Globe
http://bigthink.com/videos/can-animals-be-religious
Outside of that considerable stretching, it doesn't exist. Also, I challenge your premise that conflates "belief" with "constructing narratives"
Do you believe you will be alive tomorrow? If so, are you constructing a grand-weaving narrative... or simply making a singular belief? Or, more simply, are you just extrapolating off of past experiences?
I didn't conflate belief and narratives, I conflated belief with assigning high probability to those narratives. Sure, there are also simpler predictions we can estimate as well, but I think the thing that makes "belief in god" interesting from the perspective of judging intelligence is the complexity of the narrative.
>Outside of that considerable stretching, it doesn't exist.
What evidence would convince you that an animal "believed in god"?
And further more, if animals DO have religious behavior, then artificial intelligence research across the entire board is WOEFULLY inadequate to represent such a core part of the neurological interaction that generates religious behavior... a core behavioral capacity that was somehow missed in every single behaviorist research paper ever published since mankind mastered animal husbandry.
You either get in bed with the idea that only humans worship gods or you have to fundamentally throw all of psychology, sociology, and animal behavior studies completely out the window.
You provided... passive aggressiveness.
And if animals now engage in religious ritual, doesn't that make them stupid for not being atheist like you? I expect several YouTube videos of you trying to convert your cat into the enlightened ways of post-theism.
You did nothing of the sort. You provided a link to a video of an anthropologist describing chimpanzee behaviour and indulging in a little light speculation about why they're doing what they're doing.
Nobody is claiming that animals believe in god(s), rather I'm disputing your bizarre assertion that you know that they don't. There is no way to tell.
None of the voice recognition systems on the market learn my voice distinctly from my wife's or sons, and I don't want their speech triggering things on accident (especially my son's), so I don't use any of them.
I'll be more impressed when I can restrict Amazon Echo or one of these assistants to ignoring any voice that isn't at least rather similar to my own, not merely recognizing the words I'm speaking.
My phone rarely listens to me until I hold down the home button but one guy I know, who has a slow, deep voice, triggers Google to start listening in normal conversation all the time.
I hope it will be integrated soon with the Speech API as well ( https://msdn.microsoft.com/en-us/library/hh361633(v=office.1... ).
I rather like Cortana, but it seems to get a lot of hate - comparisons to MS BOB and what not.
How low would the error rate be for humans that can fully concentrate on listening instead of writing at the same time? Unfortunately, that cannot be tested.
Although the article recognizes that perfection has not been assumed, parity might not even be a capacity.
Conversation is difficult to measure. Take a look at the philosophical viewpoint of Deconstruction. Food for thought.
Don't confuse handwaving with science.
We have this at work (alas), and it does "transcription" of voicemail, which it sends as an email. It's easily 90% wrong, regardless of speaker, unless it's a slightly bad connection, when it's worse.
Wake me up when they can match human recognition of context.
Seriously, great work, but just like facial recognition, this will cut both ways.
If I may ask, what software does your team use at Scribie?