This Uncanny Valley of Voice Recognition
zachholman.com
zachholman.com
Right now, every incremental improvement to voice recognition improves its usefulness. It might appear that we're in an uncanny valley because voice recognition is barely usable right now versus completely unusable in the past, but there is no one that prefers worst voice recognition over better voice recognition.
Wording by wikipedia. So apparently, there might be points on the graph where a slight increase in recognition performance actually freaks out some users.
Incremental improvements are very bad in this regard.
These things are hard to reverse, too (people still speak with a very distinct "I'm speaking on a telephone" cadence today).
Siri versus XBox in his post? Maybe you're using different metrics for what's better or worse.
Ummm....
Wikipedia: "The term was coined by the robotics professor Masahiro Mori as Bukimi no Tani Genshō in 1970. The hypothesis has been linked to Ernst Jentsch's concept of the "uncanny" identified in a 1906 essay "On the Psychology of the Uncanny"."
> They ended up illustrating a crate of Campbell’s® Tomato Soup™ in the corner to make it feel a bit more canny.
Copy editing - it's important!
One reason it's bad is that the sounds we make are mush. It's a miracle if a computer system can correctly retrieve the words from an utterance. Another reason it's bad is that the words we say are nonsense. Our sentences aren't parseable, they don't conform to any actual grammar.
So I see it as another example of people selling something that's supposed to be more convenient than what we already have, but for many reasons, it probably isn't. One day it may be, but it wouldn't be surprising for people to be selling it as more convenient for many years before it actually is.
I'm not criticizing the technology -- it's amazing. It's just clear to me that it isn't ready to be invited into my life. I consider it inevitable that we will eventually lose control of technology, but we can at least try to be judicious.
Back in the day when I made my own CarPC, I naturally made my own interface for it too, for, you know... reasons. The input was a mini numpad, but the output was audio - primarily text to speech. I could drive around and use the interface (mostly) safely, but be exactly sure of my input choices using the one-handed, nailed-down keypad. Worked like a charm.
If in the future we nail down the voice input to be more accurate, fast and flexible, we can avoid the author's concerns and get something as useful as touch screens but based on our ears and mouths.
the usefulness of this didn't dawn on me until I read an interview with Andreg Ng where he talked about the huge amount of voice searches in China. Many adults can't type, so searching with voice is very convenient as opposed to drawing the characters. Many are down right illiterate or young children not old enough to read that much yet.
No, I mean parseable in the way that source code is parsed. I'm not aware of any human language that is parseable by a computer. Humans are able to understand each other because our brains are, loosely speaking, magical.
This isn't solved by magic, but pure statistics. Try and ask your friends who has the telescope in the previous sentence and some will say the man and some will say "I". Without context we can't tell, but we can judge which one is more likely given who had the item in previous sentences of similar structure (the prior). Then usually we also have some context.
Now we can add context: "I got a telescope for my birthday and was eager to use it. The next day I saw a man on the hill with my telescope". Now most people would expect the speaker to be looking through the telescope, but the man might have stolen it and taken it to the hill. Even humans have to guess.
I would prefer a keyboard in any situation except where environment/circumstance prevents it, e.g. when driving or wearing thick gloves to protect me from -20 degree temperatures.
Also, I just dictated that entire paragraph while sitting in a noisy Korean restaurants
The same applies to OCR and other photo recognition techniques like faces or red eye. Tesseract is probably the largest free software OCR project but it still seems to do so much worse than proprietary Adobe and Microsoft products. At least the OCR reader that came with my S4 does a terrible job, though it might be using Tesseract behind the scenes since I think its the one from f-droid.
Digikam does all right red eye correction but it does it with a layered filter rather than any recognition of eyes. It also sometimes can find faces, but not nearly as accurately as Google can.
All these fuzzy logic fields are things that take huge code bases and a lot of R&D to get right and nobody in the free software movement has the organization or just the raw bank to make them happen from what I can see. Red Hat surely is not investing in them (kind of outside their enterprise / server domain) and they are about the only company prominent and powerful enough to do it.
There are (at least) three problems that prevent the widespread availability of FOSS speech recognizers:
1. Data. Large corpora are available via Penn's Linguistic Data Consortium (LDC). Big academic institutions can afford their all-inclusive licenses; small corporations have to settle for their expensive a la carte options, and startups and hobbyists have to go without. Fortunately, there is now more and more freely available data, such as the LibriSpeech database (http://www.voxforge.org/home/forums/message-boards/audio-dis...), which is extracted from the LibriVox website (of public-domain audiobooks).
2. Task-specificity. Speech recognition systems need extensive customization to the intended use case to achieve good performance. Conventional wisdom is that your recognizer can be tuned to work with a wide variety of speakers, or support a large vocabulary, but not both. This customization requires lots of time, data, and expertise.
3. Expertise. Speech recognition development is a PhD-level activity. After 5 years of stable or declining real income, smart students either go work for big bucks at a large multinational, or can't get a visa and go back to their home country.
While there's some hope for #1, #2 and #3 aren't going away anytime soon.
Actually I think it's not out of the question now. The recent advances in recognition accuracy are mostly due to deep neural nets. The research is all published open access, and the cutting-edge tools are mostly open source (Theano, Torch, Caffe). Training neural nets is actually a lot simpler than the old methods of doing speech recognition; I think it's much more accessible to a small team. The only really difficult requirement is lots and lots of clean labeled data for training.
It's just hard to make something work really well for a specific use case; when contributors to an open-source project are all trying to scratch their own itch (make it work for their specific use [language, vocabulary, etc.]), the result may not be universally satisfying.
I don't think we're quite there yet, but DNNs have the potential to replace every piece of the speech pipeline with one single net that gets audio samples on one side and spits out characters on the other. All those acronyms you mentioned (with many, many PhD theses behind them) will be irrelevant, in the same way that tons of previously successful specialized computer vision feature detectors (HoG, SIFT, SURF, etc) are now irrelevant to the state of the art in object recognition.
I have no doubt that these methods and their acronyms will become irrelevant (perhaps they already are), but I guess some of the basic underlying ideas about variability will re-emerge in the training regime of DNNs. Sure, the algorithms (and implementations) for training DNNs are the same, but these ideas are incorporated in the preparation and handling of training data (compare that with augmentation, like creating translated images etc.).
You know when your GPS says "recalculating" in a condescending voice? That's the uncanny valley of text-to-speech.
Siri's UI intends to mimic a person that understands what you are saying. In practice, it gets it wrong in hilarious and frustrating ways, breaking the illusion.
So "uncanny valley" seems ok to me.
E.g. something "50% real" is so far off, we psychologically dismiss it as a cartoon, drawing, whatever. It's not trying to compete with real.
Something "98% real" freaks us out though. Surprised I have not yet seen this link:
http://blog.codinghorror.com/avoiding-the-uncanny-valley-of-...
We are OK with stupid computers, we don't expect anything from them - we order them around is very formal language.
But now there is an attempt to use natural language - and it doesn't work right. So it's actually better not to use natural language, and just stick to the formal language.
That's where the valley part comes in.
Once that fad aspect blows over, then usage plummets and its forgotten. See Kinect, or the nintendo power glove, or qr-codes, or google glass, or the cue cat, or a zillion other examples that are in, or now entering, 8-track-hood.
Uhh, I don't think so: https://en.wikipedia.org/wiki/Uncanny_valley#Etymology
> In 1992, while finishing A Bug’s Life, Pixar had to build a digital valley for Buzz Lightyear to drive his Ford® F-150™ pickup through on the way to the hospital so he get a vasectomy.
So I'm pretty sure the author is being deliberately silly.
I'd be interested to learn, though, if / where alternative terms are in more wide-spread use.
Imagine if they could program things into it that would make your overall life better by slightly altering your behavior? For instance, if you asked "Siri" to remind you every 40 minutes for a ciggarette break, I can imagine her slowly weaning you off, etc.
Also, "eat a dick", really?
> Buzz Lightyear walks into a bar called the "Uncanny Valley" and asks the bartender for a vodka soda. The bartender gives him a vasectomy. Voice recognition is important!
To me this seemed like the main conclusion of the post.
And I agree. Voice recognition seems like an AI-complete problem. I think conversations will always be awkward and frustrating until Siri can construct a mental model of my habits, my particular turns of phrases and accent, what I'm up to right now, what I think is important, who else is in the room, etc. etc. I don't think you can (only) throw deep learning at the problem and expect anything but superficial responses. (Maybe if you had one neural net per user?)
And the article is about how voice recognition is just in its infancy.