Speech Is 3x Faster Than Typing for English and Mandarin on Mobile Devices
hci.stanford.edu
hci.stanford.edu
The experiment setup is a little weird here. The participants were given a set of pre-created phrases to type/speech. At least for English, the phrase set [1] contains utterances like
"circumstances are unacceptable"
…which contains rare but in-vocabulary words. That makes it hard for keyboard input (hard for humans to spell long words, hard for touch input to predict unlikely words) but very easy for speech recognition (no other word sounds like "circumstances"). And that test set is so old that it's very likely that the speech recognizer from the experiment (or any state of the art speech recognizer) has already been trained on those sentences.
The utterances being pre-selected is also unfortunate. When users are given a sentence to speak ahead of time, they tend not to hesitate or stutter. They also speak faster than when they're trying to think of something to say on the fly, which is more typical of text input on mobile devices.
All that being said, it's certainly true that you can often input text very quickly with speech recognition, and it's getting better every day. :)
That phrase set was explicitly designed for text entry (typing) experiments. While not optimal, it does allow for more direct comparison to a large body of previous studies using the same phrase set (and similar procedures).
Having said that, keyboard input methods that provide suggestions/corrections at multi-character or word level features probably also benefit from longer words since the recognizer has more signal to work with. Revisiting the assumptions that went into that initial phrase set (character at a time input) in light of modern text input techniques might be a good thing.
I'm sure that set was designed for <del>typing on full-sized, physical keyboards,</del> not touch-screen mobile devices. (Thanks for the correction!)
Also, even though speech is faster than typing on a touch-screen mobile device, it's a lot easier to correct the errors that inevitably happen via typing.
Sometimes it's impossible to verbally correct errors or enter unrecognized words (or names!).
[1] http://www.yorku.ca/mack/chi03b.html [2] http://www.yorku.ca/mack/p25-mackenzie.pdf
Edit: one time when it was faster is when I had a string of 4-5 alarms that needed to be set. Of course I used the touch interface for prompting the voice input.
Am I getting old or is this common?
However, I'm not old enough that I enjoy the tactile experience of handwriting a letter for the tradeoff in speed, though that's because my handwriting is so terrible.
Then I just talk my question as I would have a conversation with a human and I get the answer spoken into the earpiece. Of course, sometimes I need to look up the info on the screen, since it doesn't answer straight.
On the topic, something which reads your subvocalizations is really needed, and could even increase the speed!
The problem I have with speech recognition is it doesn't allow for contemplation and correction (easily). I can stop mid sentence, go back, change something, and finish the sentence when I'm writing. Much harder to do in speech recognition. Particularly with phone stuff, it's better with Dragon, etc., but for composition, writing is still the fastest way for me to get something completed.
Dictation is a bit of a performance. It's not composition. Although it can make you a better thinker, I believe, if you practice it.
Charles Krauthammer is, in my opinion, a pretty good speaker. Not many pauses, rarely any filler like "uh" "um" etc. I was googling around about him, and it turns out he was asked about this. He's a medical doctor, and back in the day, they would dictate their notes over a phone in the hospital, to a recording that would later be transcribed.
After doing this long enough, he became skilled at composing his thoughts without writing them down.
In my case, I learned that I tend to speak in very quick bursts - get a thought out at a rapid clip, followed by a brief pause. Actively trying to measure comes across as far more intelligent I think, and I have time to think through what I'll say next.
Really hearing how many pauses, contractions, etc. you include in your regular speech helps you improve dramatically
At the end you end up with something that takes minimal effort to get into a readable first draft state. And after that it's just editing.
I only do it when I want to send a message in a hands free manner, such as a text message saying that I am stuck in traffic.
Even in places with loud background noise, its a non-issue considering everything nowadays has two mics for noise reduction. I'm very, very surprised at how well Google voice-to-speech works. When it fails its almost always because I have a poor signal from t-mobile or am saying something that's just too difficult for machines to parse correctly.
Also, one thing that is relatively hard for voice recognition to distinguish is varying between two different languages. I am always cooking and words like, dashi, kombu, and gnocchi are hard to parse. There are uses of words from other languages that don't involve saying "translate x in English"
I find, like some commentors above, that for longer-form composition, I often want to skip around. But for a text-length message the time to compose is drastically reduced by dictating it to my wrist rather than removing phone, unlocking phone, typing reply, etc. And it can be done largely hands-free.
The local language here is not english, but I use english voice recognition since it's better, which adds extra weirdness.
I only use it to set alarms and reminders on google now.
No sure how the technology has evolved recently, but that's already shameful enough when a cashier don't understand you, I wouldn't want my phone to do the same in public. My pride requires me to avoid speech recognition nowadays.
Bad news though - some countries like Spain have decided to offer more friendly automated phone service replacing "type 0 to whatever" by using speech recognition. That's a nightmare.
Speech recognition works most of the time, but I often end up with one word that Siri just doesn't understand and then you end up spending more time trying to figure out how the device wants you to pronounce the word than it helps you saving time by using it instead of just typing it out. There are also some funny videos of foreign speakers trying to use Siri when it first came out - that has been exactly my experience with it as well.
Auto-correct is the other annoying one, which can wreak havoc when you're texting in multiple languages. Even if you manage to setup your device properly with all the different languages, it can still cause some problems whenever you mix the languages within a single message. However, the biggest problem I've experienced so far is when you send messages between parties where one of them doesn't have the foreign (latin based) language installed/configured yet. The replies I've received from people on their new phones often ranged from funny to cryptic, where you have no idea what they were trying to tell you. They just typed in the foreign text and hit send without checking only to have auto-correct send you something intangible.
"... and that earned him the name of War Thug ... erase previous word ... War Thug ... erase previous word ... WAR ... THUG ... WAR THUG! WAR THUG! WAR THUG! NOT WARTHOG YOU GODDAMNED PIECE [CENSORED]"
It was the most amusing thirty minutes of the drive.
He ended up not dictating the book on the drive.
It was only later that the thought came to me that he should have gone ahead and cleaned up the text later. I'm not sure how successful that approach would be.
... or, for the sake of argument, the same thing by speech recognition:
Unfortunately speech recognition often produces insulin usable and it might be hard to clean up later can be just you and me I want to tell what you're trying to say.
Maybe my voice sucks, or maybe I should robot it up more, but this is about the level I usually get... it's okay if you're doing a quick search on Google maps and don't mind repeating it 3 times but anything more than that and you may as well start typing.
I would try to use it again, but this time not concentrating as much on the speech, and assuming success.
Disclaimer: This works especially well for things like setting timers and reminders, but I don't really use it for replying to texts.
I can hold a coffee with one hand, see where I'm going with my eyes, and still enter text. It became reliable enough for me to use habitually a couple years ago, and it keeps getting better.
edit: clarification on holding phone
It was super annoying to be around him when he did it; at least it was Russian, if it was English it would be very distracting.
This is amazing to me because most people I encounter don't do that for me, even though I am not a native speaker and by doing that they could ensure that our communication is smoother and more accurate.
If we're looking at the rest of these anecdotal comments as a reflection of the larger English-speaking (and perhaps American) public, the conclusion is that speech-to-text isn't very helpful but in a completely different linguistic and cultural environment (i.e. Chinese) we see that it is indeed quite successful.
One interesting thing to me is that, with sucessful enough voice-to-text transcription, the need for typing (and writing) becomes moot (in that context). I wonder if we'll see that typing as only a skill necessary for specific trades/occupations.
I'm not saying the study doesn't have merit, but "Speech _Is_ 3x Faster than Typing for English and Mandarin Text Entry on Mobile Devices" sounds a bit of a stretch.
On the other hand, speech is a lot worse. A lot of times you can't afford to have 10%, 5% or even 1% error rate since messages are usually short and you cannot infer intended meaning easily. So my WPM accounting for correcting speech with swype is <10.
It should be also be noted that their speech tests were done in a controlled, silent environment. I'd expect the error rate and time to complete a phrase would dramatically increase in a noisy room.
(Well of course on-screen keyboards suck: they're a skeuomorphic ugly hack that has been bolted onto a touchscreen. With a slideout hardware QWERTY keyboard on an Xperia Pro, I was typing slower than on a full-sized kb, but still several times faster than any onscreen input - predictive or not, swipe or not.)
Even for the desktop, speech will be roughly 2X faster than typing, but I have no desire to buy a copy of Dragon because I like/overvalue my keyboard.
[1] - https://hbr.org/2006/06/eager-sellers-and-stony-buyers-under...
If I hit a key on a keyboard, it works correctly every time. If I make a typo, it take a fraction of a second to hit backspace and correct it. Mistakes are cheap to correct.
In comparison, if I talk to my phone, it makes a significant number of mistakes.. making me talk slower/louder/differently to use it. And when it does make a mistake, it's more costly to correct.. I either have to start over, or pick it up and hit backspace. So I disable it and never use it.
The privacy device is called "an individual office", and its pretty important for people in "office environments" to be effective, independently of use of voice recognition.
They've been around for quite a while, but their popularity goes up and down periodically.
https://www.extrahop.com/community/blog/2014/programming-by-...
I cannot find the article at the moment, but there was one discussed on hacker news recently that mentioned a privacy device for telephones before the 1950s that achieved the same effect as cupping your hands over the receiver so that others in the room could not hear what you were saying into it. If programming by voice takes off, I imagine such a thing would be a necessity to keep office environments sane. The same goes for regular text input by voice.
Or maybe we'd all get private offices with decent sound-proofing.
Hmmm … that's not going to happen.
https://github.com/melling/ErgonomicNotes/blob/master/progra...
https://en.wikipedia.org/wiki/Hush-A-Phone_Corp._v._United_S...
Edit: Yep, looks like I completely missed my estimation of talking speed.
TL;DR: audiobooks are almost completely unlike voice recognition, you're comparing apples and oblique angles.
I haven't yet seen an input system which would combine speech with touch in nice way.
Text selection on phones usually sucks, but that's only because you need to disambiguate between clicks and highlighting; if you had a dedicated screen that assumes a user wants to highlight a section it should be a lot more responsive.
I found speech recognition to be useful mostly for brain dumps. I've found I tend to think best when explaining or talking. I used to bring along a voice recorder on long drives, capture what I was thinking, then run it through voice recognition later. It was often a garbled mess, but usefully captured a lot of thinking. Modern recognition systems would do a lot better. Yeah, I should try that again.
Probably you could build some neat solution around this. Like a speakerphone style device with array of microphones to make it easier to identify who is speaking and pick up the words. Or maybe a regular smartphone is enough. The device/app would then just transcript what is spoken and annotate it with names.
Compared to audio recording the benefit would be that going through the written raw stuff is much faster and if you were present, then you can probably recall the stuff even if the transcription is not perfect. Also it might be more socially acceptable to use this solution that to record the meetings.
Slowing down to think about and review what you're communicating is a feature, not a bug, of text.
"Ok Google, text Jim Traffic is bad, I'll be late"
https://trialbysteam.com/2010/03/09/d-j-enright-the-typewrit...
(I think people in the comments there missed and/or misunderstood some of the poem's references, a few of which are scatological.)
When you need to get into serious writing or bulk data entry, maybe it would be a keyboard.
I wonder if a front facing camera on a phone can capture your throat in sufficient detail to decipher what you say even if you speak silently if you hold the phone in your palm as people do when browsing?
Calorie tracking is a proven way to lose weight, but all the tracking apps that I tried feel like they take too much effort. So I came up with an idea a few years ago: you should be able to just say what you had eaten out loud, and then use speech-to-text and NLP to search for each item and count up the calories.
I never got around to building that, but these guys did: https://www.nutritionix.com/app
It works amazingly well. Not only is speech 3x faster than typing, it's also much faster to have a free-form text field that is automatically parsed. And they integrate with Amazon Echo, which I look forward to trying out.
I've been thinking about many other ways to automate calorie tracking. For a while I thought the answer would be an AI that recognizes photos of food, but that doesn't feel important any more. I think speech-to-text takes approximately the same amount of time as opening the camera and snapping a photo. There have a been a few minor errors with Apple's voice dictation, but so far I've seen 100% accuracy from the actual text searches.
So anyway, speech-to-text research. It's all important.
I end up having to correct it with typing anyway.
I personally consider typing English and Mandarin to be very different. There are linguistic, psychological, and cultural issues at play.
Firstly, the dominant Chinese input system is phonetic (pinyin), which perhaps implies some kind of different mental state when typing.
Secondly, it is the case for adult learners like me but also reportedly for many native speakers that precise Chinese characters are easy to forget. People have visual memories of 3-10,000 characters, but perhaps can write confidently as little as half of them from memory. The phonetic input system presents a context-based suggested intent shortlist and the user is requested to select the character they intended.
Sometimes, in more extreme cases, particularly for native speakers with heavy accents or new second language speakers, users may be unsure which character to select or may even input an incorrect but close phoneme, scan for visual recognition of the correct character, fail to find it, then type a different phoneme.
Frequently, typing Chinese is the only major creative interaction that Mandarin speakers have with modern Chinese text, since writing is becoming increasingly rare outside of a school or government-form context.
Also when I try to put a period in (we call them "full stops" in Australia), half the time it inserts a period, and the other half it literally writes "period".
I would like to see some documentation on how to use it better, I had a go at googling for some a while back but couldn't find anything useful. I ended up with the impression that such editing commands weren't implemented.
I still have a go every now and then to see if its improved with updates, but it's basically unusable in its current state.
I see a bunch of comments disputing the result. To those people, do you really think that typing on mobile is as fast or faster than speaking?
It's comparing the two as viable inputs on a smartphone. Speech recognition is far from perfect and doesn't yet allow you to talk like you would to a human. Did you read the article?
And there is the tiny problem that I can't even pronounce some of the things I've typed into Google. Examples include: Japanese manga names, company names, and foreign names from news articles.
http://www.wastedtalent.ca/comic/text-what-i-mean-not-what-i...
But at the end of the day typing speed doesn't matter. It's not whats limiting us (In English, not all languages)
But I think many people equate speak recognition with language parsing which is why people seem to be obsessed with it.
I'm using SwiftKey (the non-online version), and I can type pretty quickly after the keyboard has been "trained". I think a combination of prediction and a better keyboard layout would help input speed quite a lot.
So far this has taken me 1.25 minutes using swipe input on Google b keyboard. Couple of errors and issues (probably because i use my fat thumbs for entry).
[3:07 minutes.]
[1:15 minutes]
Then I guess we'll get used to people speaking to their phones very soon, just as we got used to people talking on mobile phones back in the day.