Speech Recognition Is Not Solved
awni.github.io
awni.github.io
I've been thinking about this because my son has Auditory Processing Disorder, APD. He can hear great, even a whisper across the house. The trouble is the words don't always make sense. He can tell me the words he heard and they are correct, but assigning a meaning to them doesn't work like it does for most people.
After playing him a bunch of audio with missing words, the lady who tested him was blown away by how smart he is - explaining that he is quickly filling in almost every missing word based by guessing from all thing possible things it could be, narrowing it down on context and getting it right. I guess that is a normal, all day long task for him.
Since then I've thought about voice recognition differently. The AI to understand the context or fill in the blanks is what will make or break it.
Interesting, all my life I've struggled to follow a conversation in a crowded environment, so much so that I actively avoid background noise with words in, I work with silicon ear plugs in or headphones and music with no lyrics.
Looking at NHS symptoms they describe me as a child, didn't learn to read until I was 8.
I nearly ended up in the remedial track but for a single awesome teacher who spotted that it wasn't that I was incapable but that I was struggling to understand in a noisy classroom, she gave up her lunch time for 6mths and taught me one on one by the end of that year I could read much better.
One distinction may be that the star teacher scales better. What if there were two students with this problem? Should the teacher stay late? Should teachers not take lunch? What if someone needs help in math, etc.
A star teacher may be able to reproduce this effect for a large number of students. Perhaps using technology, recorded lectures, practice apps, etc. Perhaps organizing study groups for kids to teach and learn from each other.
A caring teacher may sacrifice her lunch. A star teacher may not need to.
But one is enough; some of the most famous scientists only got where they were because they had a specific advisor. Neither the teacher nor student would ever have accomplished half as much working alone.
I can tell you the exact moment, problem, response, and teacher - that changed my life forever. It was at that single moment that I understood mathematics. I wasn't suddenly good at it. No, I still pretty much sucked at math. But, I understood it was a language and it was descriptive.
Many years later, that same teacher would be there to help me celebrate defending my dissertation. I didn't actually march, I was already to busy by the time spring rolled around. He's an old man, now. I still stop in and we are still fast friends.
That one comment changed my life. There's a not insignificant chance that it has had a small impact on your life, assuming you've been impacted by modern traffic engineering, or in some pedestrian venues, or had some goods delivered. All that from a single comment that is entirely obvious, in retrospect.
You might be a star programmer, but only if you re-wrote every single program exceptionally, because clearly it's not making a fundamental change to someone's life once that counts, you're not a star until you've done that for everyone ...?
Going above and beyond the expectations of your role in order to make a fundamental change to someone's life is enough to call someone a "star" IMO. It's not like someone asked you to personally pay them a bonus, you can award all exceptional effort s a "star" without problem.
And then buy them a holiday?
An interesting experiment might be to include a speaker's native tongue when trying to recognize their speech in a different language. I bet that speech recognition would be a lot easier on e.g. native Spanish speakers speaking English if you know to ignore, for example, the sound "e" when spoken before "st" or "sp"
As he is getting older he relies less on seeing lips. I'm hopeful at some point he'll outgrow most issues.
Not sure I have Asperger's though, I'm just a programmer who likes his own company.
If you don't suffer in any way shape or form, it doesn't matter how many boxes you tick, if you feel fine there is no compelling reason to go through with diagnosis, except possibly if you have kids.
Autism spectrum disorders are quite heritable, and it's a tremendous advantage to be aware of ones own psychiatric oddities when you raise a child with his or her own. I speak from experience, I wish I known about my ADHD-PI/ADD with a dash of Autism earlier. I am however very happy I got to know about it early enough to help me guide my daughter into a way of thinking that has helped her a lot.
In general, without the knowledge that people on average might have a WILDLY different cognitive landscape than yours, it's very very hard to know which advice to take to heart, and what to completely ignore.
With that said, I'm fairly well aware of what type of questions are on those questionnaires. If the test is legitimate, a result that high over the limit would probably include at least a few 'yes' answers to questions that usually have a negative impact of quality of life.
Adding to that, in my experience people actually on the spectrum are the least likely to self diagnose prematurely, even when the answer is essentially staring them in the face.
Anything that can improve quality of life a little is worth considering, because sometimes it's really hard to judge the magnitude of how it can help.
Mostly, I don't. Earbuds help for some weird reason, as does pressing your ear shut. Looking at their mouth does 80% of the work. And most conversations I have in clubs are pretty mundane and uninspired. “What do you want to drink?” “This DJ is good. Do you also think this DJ is good?” “Do you agree with late justice Earl Warren that Baker v Carr was the most important supreme court decision of his era?” etc... Low entropy, easy to error correct.
Of course, and all humans rely on this as well. No one hears every word perfectly all the time --- it's impossible, because the source person doesn't pronounce every word perfectly all the time. Context clues are a huge part of speech recognition, as well as gestural typing recognition and other forms of machine interpretation of human input. While it's always been a component of NLP, you can clearly see it in action with the Android keyboard these days because after you type two words in a row, the first may be corrected after you enter the second one, based on context provided by the second one.
Google is the best of all the big players at figuring out this context (I once asked it what a Dead Left Shrimp was, apparently I was mishearing the name of basketball player Detlef Schrempf), but it will still search for "Pizza Cake" rather than correcting it.
As an example, I was being directed by Google Maps to a new place, and I asked it "What is the ETA?" It responded, "From Wikipedia, the estimated time of arrival or ETA is the time when a ship, vehicle, aircraft, cargo or emergency service is expected to arrive at a certain place." It was a completely valid answer to the question, but not one that any human would give.
s/Microsoft/Google/g s/Seattle/Mountain View/g
[0] http://alunthomasevans.blogspot.com/2007/10/old-microsoft-jo...
Having a "the" in there like you did should remove ambiguity though. I guess Google Maps defers to Wikipedia when it doesn't know what a term means.
I don’t know about you, but I wanna try this: https://g.co/kgs/4D4er9
The background noise seems to be key. He was doing great in karate and a new teacher started blasting music. He went from doing great to being unable to know he should put his right for forward when told to. That was one of the clues that something wasn't normal.
Interestingly, these disorders all seem to have pretty direct neurological causes: there is something causing a difference in the person's brain between where they sense the information and where they process it, whether it's genetic (i.e. affecting the layout and growth of the brain) or due to trauma such as a head injury.
I'm curious as to whether your son is able to enjoy music, especially music with multiple different instruments playing simultaneously. I have a theory that the same process on a neurological level that makes food pleasurable is what makes smells smell good, music sound good, certain surface (e.g. soft fur) feel good, and certain sights (e.g. a mountain stream) look good. I mean that not from a neurotransmitter perspective, but in the actual neural processing. It seems these configurations are programmed into our brains genetically, just as our propensity toward recognizing faces/speech are, which is fascinating.
Sure it is; it's an essential part of speech recognition for all humans.
Not to say that there aren't open source speech recognition systems, they just aren't completely usable in the way proprietary solutions are. A lot of research goes into open source speech recognition, where they are lacking is in the datasets and user experience.
Hopefully, Mozilla's Common Voice project https://voice.mozilla.org/ will be successful in producing an open dataset that everyone has access to and will spur on innovation.
I don't think this level of computational power can be achieved on a modern CPU, or even a GPU! But GPUs are probably the closest analog to Google's absurdly parallel architecture.
To get a GPU working at maximum performance, you either have to go OpenCL2.0 or CUDA. Compared to OpenCL1.2, OpenCL 2.0 has a better atomics model, dynamic parallelism (kernels that can launch kernels), shared memory, and tons of other features.
NVidia of course supports those features in CUDA, but NVidia's OpenCL support is stuck at 1.2. So in effect, CUDA and OpenCL are in competition with each other.
Anyway, that's the current layout of the hardware that's available to consumers. I think its reasonable to expect a graphics card in a modern machine, even Intel's weak integrated-GPUs have a parallel-computing advantage over a CPU.
So for high-parallelism tasks like audio analysis or image analysis, it only makes sense to target GPUs today.
Edit: I should say that in state of the art results there tend to be multiple components, including multiple neural nets and the tricky "decode graph" that gok and I are talking about. These are trained separately then get stuck together, as opposed to being trained in an end-to-end fashion.
Its "weird" in the ways that matter: there's no commodity hardware in existence that replicates what a TPU does. The only place to get TPUs is through Google's cloud services.
CPUs are basically Von Neumann Architecture. GPUs (NVidia and AMD) are basically SIMD / SIMT systems.
Google's TPU is just something dramatically different, optimized yes for Matrix Multiplication, but its not something you can buy and use offline.
But you will. The entire point is to put this in a phone, so you can distribute a trained neural net in a way that people can actually use without a desktop and $500-$4,000 GPU.
So, in that regard its no more "weird" than other common accelerator/coprocessors for things like compression.
So, in the end, what would show up in a phone doesn't really look anything like a TPU. I would maybe expect a lightweight piece of matrix acceleration hardware, which due to power constraints isn't going to be able to match what a "desktop" level FPGA or GPU is capable of much less a full blown TPU.
As far as I can tell, they put a microphone on your phone and then relay your voice to Google's servers for analysis.
Or Amazon's servers, in the case of the Echo.
I don't see any near-term future where Google's TPUs become widely available for consumers: be it on a phone or desktop. And I'm not aware of any product from the major hardware manufacturers that even attempt to replicate Google's TPU architecture.
NVidia and AMD are sorta going the opposite direction: they're making their GPUs more and more flexible (which will be useful in a wider variety of problems), while Google's TPUs specialize further and further into low-precision matrix multiplications.
Also, I just turned on airplane mode and google assistant recognized my voice.
I'll look into them, thanks!
The issue previously has been a lack of large enough high quality annotated datasets, and open source ASR libraries being a bit behind or not well integrated with cutting edge deep learning. I think that's changing now though. I hope it won't take too long until pre-trained, reasonable size and high accuracy TensorFlow/Kaldi models for many languages are common.
The dataset is only going to grow as time goes on. They just need to get their website in order (a stats page to show where they're currently at and how far they need to go would help a lot) and throw a little advertising behind it. It's just a matter of time
Just yesterday there was a Show HN built with the https://github.com/kaldi-asr/kaldi project, emscripten-ized: https://news.ycombinator.com/item?id=15534531
Snips was mentioned here recently but I haven’t taken a look at it.
It showed promise, but I don't know if any work is being done on it.
However I will say that for my company's use case a properly configured sphinx install produces better results than kaldi. However, getting to a point where you can say that was not an easy task.
Additionally, I actually believe that for most workloads that kaldi is likely better. Not ours though.
I’m under the impression that Google is mostly dogfooding its open source tooling for machine learning in GCP, and actually differentiates based on trained models and compute power.
https://github.com/kaldi-asr/kaldi/blob/master/egs/fisher_sw...
Kaldi hasn't been in first place on that dataset recently, but it was a few years ago.
On other more researchy datasets (eg. for distant speakers or languages other than English), the best system is often based on Kaldi.
https://www.quora.com/Is-anyone-working-on-an-open-source-ve...
Apparently, the business model for voice services in the past has been to snap up as many broad patents as possible to keep competitors at bay. I read an interview with a google engineer a couple years back claiming the same thing. They have to carefully work around a patent minefield with their own services and how these patents are holding back better voice search on mobile technology.
As the article says latency is still a problem and it's a huge problem in current open source solutions, some stuff I was testing was easily 5 seconds. I know that can be improved with configuration, but when dealing with libraries of 10 words or so, that's pretty bad.
I feel like anyone who is seriously interested in this space has been scooped up by all the big companies and the open source solutions have really seemed to linger because of it. It's one of the first areas I've seen where open source alternatives are really behind the proprietary solutions. Kind of bummed me out.
Here's a blog post - http://www.googblogs.com/kaldi-now-offers-tensorflow-integra...
The easiest way to deploy Kaldi is this - https://github.com/alumae/kaldi-gstreamer-server (or a docker image of the that)
(edits: spelling/grammar)
Like, adding a mic button for voice search next to their main search toolbar on Firefox, and then ask for permission to use that data for research.
You can build your voice assistants and run them for free on a Raspberry Pi 3, or Android
1. speech-to-text-wavenet (https://github.com/buriburisuri/speech-to-text-wavenet)
2. kaldi(https://github.com/kaldi-asr/kaldi)
3. Speech recognition module for Python (https://github.com/Uberi/speech_recognition)
4. DeepSpeech(https://github.com/mozilla/DeepSpeech)
5. Natural Language Processing Tasks and References(https://github.com/Kyubyong/nlp_tasks)
FYI.
This sentence is crucial and suggests a way to understand how WER can be apparently lower than human levels, yet ASR is still obviously imperfect:
> When comparing models to humans, it’s important to check the nature of the mistakes and not just look at the WER as a conclusive number. In my own experience, human transcribers tend to make fewer and less drastic semantic errors than speech recognizers.
This suggests that humans make mistakes specifically on semantically unimportant words, while computers make mistakes more uniformly. That is, humans are able to allocate resources to correct word identification for the important words, with less resources going to the less important ones. So maybe the way to improve speech recognition is not to focus on WER, but on WER weighted by word importance, or to train speech recognition systems end-to-end with some end goal language task so that the DNN or whatever learns to recognize the important words for the task.
I've repeatedly seen improvements to traditional measures that make the subjective result worse.
It's incredibly hard to measure and solve (if anyone has good ideas please let me know). I check a lot of sample data manually when we make changes, doing that (with targeting at important cases) is really the only way I think to do things.
I guess a problem would be if people become so used to errors that they send messages without corrections. I have some friends who do this: they send garbled messages that I have to read out loud to understand. But there will always be a subset of people who want to get it right.
Speech separation - Mitsubishi Research has done some pretty impressive stuff on that - http://www.merl.com/demos/deep-clustering. Haven't seen equivalents of that in open source ASR
http://spandh.dcs.shef.ac.uk/chime_challenge/chime2016/resul...
I'm really fascinated by the whole idea of blind source separation and the fact that speech signals are "sparse" in frequency space.
We've had a similar experience looking for hardware / open source beamforming. There's a package called beamformit, but I think it's pretty old.
https://medium.com/snips-ai/benchmarking-microphone-arrays-r...
https://developer.amazon.com/alexa-voice-service/dev-kits/co...
Amazon's 7-mic hardware has its own OEM program.
Similarly I had friends for whom English was a second language who had lived in the US for years and were definitely fluent, but had to enable subtitles for a movie in which the characters had a strong southern accent; in general non-rhotic accents were very troublesome for them having only spoken english with midwesterners.
The article mentions the Scottish accent, and I would call that the hardest accent of native English speakers for those in the US to understand.
I’m currently where speech recognition software was in the 90s.
Personally I've had to "translate" between a South African and an Ulsterman before, both of whom were speaking English but with extremely different accents.
This works well both in speech and sign with our 1 year old too (oh not "poo" but "boots"; {moves open hand forward and backwards away from body at chest height} oh right "lawnmower" ...).
But, other Scottish people certainly don't have trouble with understanding a Scottish accent. So I view that as a certificate that we should be able to build a speech recognizer which can recognize Scottish accents.
As a Scottish person, I'll say there's a huge amount of variation between Scots dialects. As someone who grew up in Fife, it took me well over a year of living in Glasgow to be able to reliably understand people there—and both of them are typically classed as Central Scots.
I also grew up in Fife, although my parents paid good money so I would have an Edinburgh accent. Glasgow was like a foreign country to us...
I grew up in St Andrews, both of my parents having grown up in England, and went through speech therapy as a young child (due to dyspraxia); unsurprisingly, with that, you can imagine my accent is much closer to RP than any broad Fife accent, though most of my speech is definitely Standard Scottish English.
Whether or not gathering a sufficiently large corpus of other dialects will solve the problem would be interesting; also it might be uneconomical to gather a large enough corpus of some dialects, leaving minorities out.
According to Wikipedia it's "a direct continuation and development of the language spoken by the Anglo-Saxon settlers" of the region. https://en.wikipedia.org/wiki/Geordie
They were having dinner at the ships restaurant, and the waiter asked my grandfather something. My grandfather just didn't understand the guy and asked the waiter to repeat several times. After the 4th time, my sister told my grandfather: "Grandpa, he's Swedish". My grandpa paused for a second and then immediately recognized what the waiter had said.
Turns out he had assumed the guy was Danish, and thus failed to interpret the limited sounds he could hear, given the hearing loss and background noise.
As anyone who's had to use langid in practice will testify, it's solved only as long as:
A) you want to identify 1 out of N languages (reality: a text can be in any language, outside your predefined set)
B) you assume the text is in exactly one language (reality: can be 0, can be multiple, as is common with globalized English phrases)
C) you don't need a measure of confidence (most algos give an all-or-nothing confidence score [0])
D) the text isn't too short (twitter), too noisy (repeated sections ala web pages, boilerplate), too regional/dialect, etc.
In other words, not solved at all.
In my experience, the same is true for any other ML task, once you want to use it in practice (as opposed to "write an article about").
The amount of work to get something actually working robustly is still non-trivial. In some respects, things have gotten worse over the past years due to a focus on cranking up the number of model parameters, at the expense of a decent error analysis and model interpretability.
[0] https://twitter.com/RadimRehurek/status/872280794152054784
It shouldn't be hard to improve on by treating the score on a tweet as bayesian evidence to combine with a prior from preceding tweets.
I live with a native French speaker, so my conversations naturally include a lot of French proper names, as well sometimes switching languages mid conversation or even mid sentence.
Lots of recognition engines can handle English and French, but treat them as mutually exclusive. It frustrates me to no end when I know that Siri can recognize a French proper name just fine if I switch it modally into French, but will botch it horribly in English.
Speech is[1] a fundamental component of Language, which is a fundamental component of Intelligence[2]. This is addressed somewhat in the conversation around semantic error rate; that there is more to processing raw audio speech than the calculus of mapping signals to tokens; some understanding of semantics and context is required to differentiate between otherwise indistinguishable surface forms.
I find it doubtful that there's a clean interface that separates the 'intelligent' parts of the brain from the 'language' parts of the brain from the 'speech' parts of the brain. This leakiness (or richness, really) means that you can't neatly solve any one part of this chain to the level of competence that the brain solves it. That means to 'solve' speech recognition, you have to 'solve' language, and thus 'solve' general intelligence. And to 'solve' general intelligence, you have to understand it, in a theoretical sense, which we don't. Indeed, it will likely involve solving all the other modalities of sensation as well. It's definitely the case that you need to have a model for prosody to understand speech. It is entirely possible that vision is a large factor as well, in the form of body language, lip reading, eye contact, and so on.
Speech recognition is quite good for what it is. For many practical applications, especially to do with young, white, newspeaker-accented English speakers who sound the most like the people who develop it, and the data sets used to develop it, it is good enough in the 80/20 sense. But it is nowhere near solved by even the least rigorous definition of the word.
----
[1] according to the philosophy I subscribe to, at any rate
[2] according to the philosophy I subscribe to, at any rate
This is, I think, one of the failings of the Turing test. It's easy enough for us to make new humans; making a machine that acts exactly like a human seems like a silly endeavor. I want a machine that can assist us and reinforce our failings. Which means that we can necessarily differentiate it from another human. I vastly prefer that over a machine that has learned to lie to us.
I know, that's why I'm saying the precise setup of the test is important. That's why, for instance, it's usually presented as a text messaging setup, because the goal is to test intelligence via language and conversation skills. A "face to face" setup wouldn't make much sense, unless we wanted to test that aspect of robotics.
Humans use different sub-languages to speak to friends, bosses, dates, lovers, spouses, teachers, students, and so on.
Each social context has its own vocabulary, its own set of expected conversation starter statements, its own set of likely responses, its own set of problems that may need to be solved - and so on.
A lot of what passes for intelligent interaction among humans is really just this social awareness.
Machines will fail the Turing Test without it. But you can't teach a machine to mimic social awareness by throwing 100,000 hours of speech samples at it.
Nor can you expect Echo/Siri/etc to know the context you're working in with no input from social cues - location, dress, time of day, social relationship, facial expression, etc.
So practical ASR turns into spoken-command-line-plus-guessing.
That turns out to be a pretty poor imitation of even the simplest social relationship.
It's not a useless imitation. Even today's limited voice assistants do a fair job at providing a useful service.
But the idea that you can drown neural networks in sample soup to train them, and build yourself a machine capable of intelligent-seeming conversation is just naive.
That's not how baby humans learn to hear and speak, and it's certainly not going to teach machines to converse at a human level.
I think where computers fall short is in two areas:
1) The rate of errors hasn’t hit the inflection point of being comparable to day-to-day intra-human interaction, and
2) There is no good mechanism for detecting and correcting errors. At least, none that I’ve seen.
That second one is important. When I hear my wife ask me “Please, hand me a tractor”, I realize that I must’ve misheard, and ask her to clarify “what?” With speech recognition, I either have to manually re-read and modify recognized text, or cancel and repeat the entire request. Both take time, and negate some of the efficiencies of using speech recognition.
Something like Rhymezone (searching for similar sounding words) would be a solution. It would still need a human who's deciding which word to choose, but after some time it could learn which words you prefer.
[1] https://www.rhymezone.com/r/rhyme.cgi?Word=tractor&typeofrhy...
https://news.ycombinator.com/item?id=15429287
Wish this article was written a few days earlier.
An analogue of this article exists for most other domains claimed to have been solved.
No, it's not learning new classes of objects from a single image or a few images is very difficult. See
http://www.sciencemag.org/content/350/6266/1332.short
The machine translation is a joke.
Put comments on this page by translating Google into another language and go back to English and see what you get.
I did a little part of you. Just human level for just a simple little prayer.
> But even if you do not make this assumption, identifying the object involves spitting the distance from the performance of the human level.
source: https://news.ycombinator.com/item?id=15429862Not to debunk your claims, just interesting to see how good translation works (even if it's not human-level, I can understand what you say).
"spitting the distance from the performance of the human level"
If this were spoken out loud, it might get mistaken for "splitting the distance" which seems to have a different meaning (half as good as humans).
But the point is subtle differences matter a lot in human communication.
But, I do agree with your point that in a lot of cases even without human-level performance one can get things done.
> I did a little part of you. Just human level for just a simple little prayer.
It's just bad. It sounds like it's based on a n-gram model - about something religious, strangely enough.
For many applications, making a transcription seems like an unnecessary step and source of errors. Skipping transcription when the user doesn't need it (most cases where I use it) would seem like a way to get some gain, but perhaps at the reduction of debuggability.
> For example in voice search the actual web-scale search has to be done after the speech recognition.
That's an area where literal exact transcription is usually required. But even then, Siri/Cortana/Alexa might be better off trying to figure out the meaning of what someone's asking rather than figuring out the exact words spoken in order to return the best results.
Most people are quite bad at formulating good internet searches without a lot of trial and error. Let Google listen to a person talk about the problem they have or issue and then formulate the best results for that instead of forcing us to come up with the exact right phrasing to get appropriate results. It would help tremendously with the synonym and homophone issues that are so annoying now.
Agree. What's really needed is research into (and development supporting) how to combine the expertise from a speech recognition layer with the next layer in a machine learning process. That higher layer contains the domain specific knowledge needed for the problem at hand, and still leverage a speech layer focused on a broad speech data set and speech-specific learning (from Google, Microsoft, the community, etc.)
Today, how richly can information be shared? I see with Google's speech API you can only share a very finite list of domain-specific expected vocabulary.
Why not have speech tools at least output sets of possible translations with associated probabilities? Do any of the top tools allow this?
Then you could at least train your next level models with the knowledge of where ambiguity most exists, and what a couple of options might have been for certain words or phrases...
My kids needed 5-6 years of continous daily talking until I could say they understand a conversation almost completly. Every single word or phrase I directed at them was spoken in a certain context and had a certain role in communicating with them when from the context it wasn't clear what my intentions were. Of course, they had their fun, throwing words away and repeating endlesly some funny word or phrase or terribly spelling them. It is interesting that you as an adult need too learn their own prononciation, at some moment in time I even wrote a small dictionary.
Again, speech recognition will simply stay at recognizing phonemes/words only for a very long time, until we have a true AI Assistant that walks with us, sees with us, and hears with us in the same time. Then it can apply some semantic and other context based related NN.
Siri still doesn't understand the majority of my Minnesota relatives.
http://appleinsider.com/articles/17/10/19/how-to-teach-siri-...
You'd need to create a contact if he doesn't already have one.
"Hey Google, play Mika radio" has a 50/50 chance of starting music by Meiko. The additional "problem" is that I like Meiko, too, but I feel obligated to cancel the Meiko music & re-request Mika so that Google (hopefully) learns to recognize the difference between Mika and Meiko.
Maybe our spoken language will start to transform into distinctly unique sounds so we can verbally interact with computers with relative ease. When I was in the US Army, I was trained to speak in a certain manner to help my communication to be more clear. [1] I don't see a reason humans and computers can't each make reasonable compromises to make verbal communication easier.
People mishear a certain fraction of what they hear and they'll ask you to clarify what you said.
You could have superhuman performance at tasks like Switchboard and still have something embarrassingly bad in the field because the real task (having a conversation) is not trained for in the competitions with canned training sets.
Ditto with Google's Google Now assistant, or whatever the heck it's called these days. I have a Pixel 2 phone (Dragon heard "pixel to phone" -- it doesn't have up-to-date context on proper nouns in the news), but when I tried to create a calendar event using "Create calendar event... meet Bruno for pizza", it heard "MIT pronoun for pizza". It has hundreds of samples of my voice, and it already knew I was creating an event! "Meet" has to be one of the most common first words used in events.
It seems to me like there is pretty low hanging fruit, and that we need more focus on flexibility and resourcefulness rather than acting as though we're moving from 99.5% accuracy to 99.6%.
1. QWERTY transcriptionists working brief, 5-10 minute shifts typing as fast as they can and rotating.
2. Voice-mask reporters using Dragon with a voice mask (more commonly phonetic typos)
3. Stenographers using a steno machine, some of them were not trained to do realtime and lack the ability to edit as they go.
It's all up to the individual, unfortunately. Furthermore, the budget for live captions is sometimes smaller than it should be and so the service purchased is the cheapest, not the best.
I pay attention to captions now. I've watched TSN (Canada) and it seems that past 10 p.m. they put in someone who isn't fully trained or graduated from school and they most definitely shouldn't be providing captions as they are near unreadable.
Good captions are priceless for the access that they offer.
I'm not sure how much work it would take to scale down Apple's voice recognition to run on the device or if it is feasible with their model, but currently it can take 5-10 seconds longer to get an answer from Siri during peak times.
For example, try speaking the words: "OK Google, the thick round box jumped over the hazy bog."
All people all the time doesn't remotely work for people, either. That's why the Air Force, for example, uses the Phonetic Alphabet, and aviation in general uses special jargon that is harder to misinterpret, such as "affirmative" for "yes".
http://militarytimechart.com/military-phonetic-alphabet/
Radio DJs and actors tend to enunciate more clearly than regular people do. You'll notice this more if you speed up the sound - the professionals are still clear at higher speeds where non-professionals become unintelligible.
One example: I often give the same commands to Alexa every day, even at similar times. Yet every few days it just utterly misunderstands those commands, picking words I have never even used before. It doesn’t even offer to choose similar commands I’ve used recently. Worse, the misunderstood command might trigger a paragraph of senseless babble about the misunderstood command, forcing me to shout over Alexa to make it hear the command I really wanted.
So please, please, add some “dumb” options. I want to pick some words, have the system learn them, and just obey commands 90% of the time.
1. Google Cloud Speech API
2. Microsoft Bing Speech API
3. IBM Watson Speech to Text
The ranking was as listed above but we had real challenges working with call-center audio recordings. The quality was less than idea but still very clear. We saw a huge reduction in accuracy compared to in-browser testing. Additionally Australian English is particularly not solved.
Because Google's API isn't currently doing speaker-detection, we looked at using Watson's speaker-detection as a secondary step but found it too complex and error prone. There is definitely room for a startup in this area and it also needs continued investment from the bigger cloud providers.
https://remeeting.com/app/benchmarks
Google actually did pretty badly for us on extended telephone speech. Not sure why.
I've not compared them extensively but for streaming realtime, I found that Watson beat the Google api for a specific use-case. Your mileage may vary!
They also provide a handy Mic / File reader interface for browsers: https://github.com/watson-developer-cloud/speech-javascript-...
Training and test samples must be iid (ie independently sampled from the same distribution) to get good performance. Otherwise there is bias.
This applies not only to speakers as individuals, but to all the other factors mentioned, like context, bg noise, gender, age, accent, education, mike, emotion, etc etc.
Many of the issues described can be traced down to this principle.
This is why, possibly, the core issue is data.
Although models and *PUs matter too, the theoretical performance boundary dependends on data quality only (with a model to match).
Sometimes I wonder, if Apps could add vocabulary to devices dictionary, so certain terms in Gaming, Fashion, or Tech or other domain, which are rarely used in normal conversation would be recognized in speech to text as well as Auto Correct.
The switchboard dataset was "challanging" in its time but pointless from current standards. One doesn't have to look beyond trying to do voice dictation in car to see how horrible current speed recognition systems are compared to humans.
It is like say Web Application is not solved. Surely, it is not. But the remaining part falls into the arts category, while the hard barrier of the technologies itself is no longer there.
Nobody expects to be able to type 'find me the files that changed yesterday' into a computer shell and have it do something sensible. But 'find ./ -type -f -mtime yesterday' sure we can expect that to work. Getting from the first one to the second one requires something called an ontology.
I worked on a team at IBM that was building this sort of capability. The plan was that in conjunction with an excellent tool kit for disassembling a sentence into its component parts (verbs, nouns, adjectives etc) and then matching the elements recognized with an onotology associated with that action. 'find' being the verb would select the search action ontology, 'files' as an object of the verb would select the computer files object ontology, and then 'changed' and 'yesterday' would match 'modification' in that ontology and 'yesterday' would match one of the two sub children of 'modification' (time and change-actor). Then you would go back with what was the compiler equivalent of an abstract syntax tree which the 'command generator' would use to emit the necessary command actions and flags.
I have no idea why they cancelled it. It was an inscrutable place to work at in many ways :-).
Even something simple like "They're unhappy, Ness" could be interpreted as "Their unhappiness" unless you know Ness is a person in the converation.
A modern statistical speech recognition system has no trouble determining that "they're unhappy, ness" is a dramatically less likely word sequence than "their unhappiness".
edit: I read your example backwards, but still, a statistical system can easily incorporate contextual words without actually understanding what they mean. Names from the speaker's contacts in particular are widely used in ASR systems for this reason.
As for the OCR problem, try writing one for Chinese calligraphy and get back to me on if context is important or not.
Google Photos is pretty close.
"photos from december" works fine. "Find photos from decemeber" doesn't.
I can type in "effects photos" and find photos Google applied effects to, which is pretty nifty.
It is probably equivalent to the old game parsers in intelligence, but for a given domain it works pretty well.
They could stand to filter out a lot more words, but it works OK in general. If I type in "photos with a blue sky" it knows to drop "photos with a", photos presumably being redundant on photos.google.com!
That said, Google has done some great stuff in inferring things like units for conversion 10 feet in inches is easily parsed for example as a conversion request.
My guess: because statistical methods are all the rage now, and have been for a while. This is famously illustrated by Fred Jelinek's quote "every time I fire a linguist, the performance of the speech recognizer goes up"[1], and cemented with the success of Google Translate.
The biggest failing of things like Siri and Cortana is they're not conversational, they're really not good at learning from example, they can only respond in hard-wired ways. "What's the weather?" will work. "Do I need a jacket tomorrow?" won't because they don't understand, instead giving the closest answer they can based on available information, not knowing how you prefer to dress or what you think is cold.
Until we can have full feedback neural networks we're not fully capitalizing on this AI stuff. Once we can cut the retraining time down to something marginal, maybe we can have computers "dream" and self-reprocess based on their accumulated learnings instead of waiting for a new kernel from the cloud.
We've tried https://www.btwtalk.com for transcribing our internal calls, and while it does mess up a bit, it's good enough for later reference.
If yes, say "yes" or press 1, otherwise say "no" or press 2 to speak your request again, star to hear the options, or stay on the line for the next available agent ...
Computers cannot compete with our guessing ability, not until they are trained with our life experiences.
My vote is for the auditory nerve as the inflection point. Anything above the cochlear nucleus on the auditory pathway is black voodoo magic and anyone who claims they understand it needs to reevaluate their kool-aid intake.
In Italian, for instance, any device will barely be able to recognize what you're saying, and will almost always do something completely different from what you've asked.