Voice will be the next GUI.
Voice will be the next GUI.
If we're talking really sophisticated interfaces like in "Her"[0] I agree. If the interface is working so flawlessly, it's difficult to see why you'd still want to occupy your hands with querying some knowledge base, composing blog posts/letters, etc.
However, I don't think we'll get there with current technology. When I think of the often awkward and cumbersome interactions I have with Siri, I find it really hard to imagine how this will evolve to a 'I don't need to worry about this at all anymore'-level in the next 5, 10, maybe even 20 years.
I suspect the giants of today won't be around anymore when truly voice-controlled interfaces come around.
Bandwidth. You have several degrees of freedom with each hand, with each finger. You have one linear stream with voice.
Latency. You can flick a switch in a few milliseconds, but saying "turn off the lights" or "lights off" takes half a second or more.
Privacy. You can overhear a voice command. You can't overhear (very easily) a buttonpress or touchscreen swipe.
Accuracy. Even with perfect voice transcription, people misspeak easily more than they mistype. And mistyping can be corrected within a few letters, while misspeaking will require interrupting the stream to switch the voice UI into editing mode or something.
Same for turning off the lights: Sure, it's faster if you only consider the flicking of the switch. It's another thing if you also incorporate the time it takes you to get up from your couch/bed/wherever you are without a light switch in arm's reach.
You definitely have a point with privacy though. Also, the volume of all people on a train talking to their smart assistants (though the question remains if this is really so much different from people talking to other people on a train).
EDIT: I misread your point about mistyping vs. misspeaking. Still, the interface I'm talking about does not work in modes. It's able to truly understand you and interpret your commands appropriately.
It still really bugs me how bad UI systems are at error correction.
We have a series of conventions to do this efficiently in spoken English, inflections, quick utterances that call out specific ambiguous syllables, context (the hard one, sure).
Even on a smartphone, if the system guessed a certain word when I swiped it, then I delete the word and enter it again, maybe stop guessing the same word every time?
I think you're right that typing will be more effective in general, but I still see so many areas where the gap could be closed a little more.
Correcting typing on a phone in very hard.
Voice assistants as solutions for people with no hands or eyes, or people who are at some distance from a keyboard and/or monitor - fine. Otherwise, they're CLIs without persistent displays (making even simple multiple choice branching far more difficult: see automated telephone helplines; try to remember what the first option was.)
They seem to be good for setting alarms and sending and reading emails (if you receive very few, very simple emails.) That's something. Otherwise their major use is as assistants for catalog shopping, which is why all of these companies want to own them.
Simple voice interfaces suffer the same problems as command line interfaces while being less flexible and slower to use than even GUIs.
The best voice interfaces have made good progress on most fronts, but discoverable voice are still a big problem. Instead of reading a bunch of buttons, you usually have to guess what features might be implemented. Or you go the route of phone system menues, but everyone hates those
AI will make computers able to understand people perfectly.
Gesturing, talking, writing, any input. The semantic gap will eventually be automatically traversed by understood intent.
From the article: "To my mind, translation is an incredibly subtle art that draws constantly on one’s many years of experience in life, and on one’s creative imagination."
This is true, I just think that we'll arrive at the tools to accomplish this in the future.
Specifically I think(hope) it will be through clever application of GANS and reinforcement learning after a few more applications of moore's law.
Advanced AI would be able to learn about us through replaying years of possible generated experiences.