Google launches voice typing in Google Docs
googledocs.blogspot.com
googledocs.blogspot.com
It's getting PRETTY close, with a basic cheap laptop microphone, my Australian accent, and no training.
But I did say pretty close. Not exactly right. And I'm only speaking a few sentences.
I read out a passage aloud, and apart from not quite getting everything right (but getting some remarkably difficult words/names correct), the thing that struck me is how written word, and the art form of such, is in fact quite different from a verbal stream of spoken words. The little sigils and breaking up of a written text conveys a whole bunch of meaning and subtlety that a word for word stream doesn't quite get.
For me, its just inaccurate enough to still be irritating. Everything's working fine, and then I hit a mis-dictated word, have to do a bit of a double take and ask myself "What the...what on earth does it think i'm trying to say here...oh...i said THAT!" which breaks the flow of anything but the most basic and primitive sentence.
But then that makes me wonder, you know what... how many humans can actually follow everything I'm actually saying? It's not like I can peer into their heads and verify the transcription of voice-to-words inscribed on their brains.
On this topic I have never managed to get my typing above 60 women because this is as about as fast as I can think ahead writing.
https://www.extrahop.com/community/blog/2014/programming-by-...
http://ergoemacs.org/emacs/using_voice_to_code.html
I've got more notes on Github:
https://github.com/melling/ErgonomicNotes/blob/master/progra...
We really seem to be at the point where voice is ready to become part of our daily use once a real API becomes available. People who use Dragon, for example, need to use the Windows version because of the hooks to Python.
The founder, Jim Baker, lost out and unsuccessfully tried to sue Goldman Sachs (who had been advising him).
And yet, even today what happened next to the Bakers seems remarkable. With Goldman Sachs on the job, the corporate takeover of Dragon Systems in an all-stock deal went terribly wrong. Goldman collected millions of dollars in fees — and the Bakers lost everything when Lernout & Hauspie was revealed to be a spectacular fraud. L.& H. had been founded by Jo Lernout and Pol Hauspie, who had once been hailed as stars of the 1990s tech boom. Only later did the Bakers learn that Goldman Sachs itself had at one point considered investing in L.& H. but had walked away after some digging into the company.
- a 2012 NY Times article "Goldman Sachs and the $580 Million Black Hole" http://www.nytimes.com/2012/07/15/business/goldman-sachs-and...
recent discussion here: https://news.ycombinator.com/item?id=10994707#up_11014194
Worth noting that kids who start using this young will have much less problem adapting their speech for automated transcription. Likewise, adults often were drawn to "natural-language capable" Ask Jeeves, but kids just searched with alta vista using grammatically bizarre but effective terms.
As far as I know, I have good hearing, and I often struggle to understand people. It would be interesting to see what a human error rate on transcription is.
I also recall when I worked in television, we'd have someone come in to do captions for the Rockets basketball games...they made a lot of mistakes. Maybe every other sentence had a mistake in it. Partly this was the speed at which they had to type (using a Stenotype style device with chording keyboard), but I suspect they were also mis-hearing things quite often.
In particular, technical terms (whether for computers or basketball) are elusive if you don't speak the language. "WYSIWYG" was a tricky one, in one of my videos.
"delete delete last word delete line select line delete paragraph oh I give up"
I actually verified it was kinda picking them up, but it did this weird thing where I managed to insert a period if i remember correctly, but then it took it out after I kept talking?
We don't have the accent diversity of the brits/north-americans, but we do still have 2 or 3 kinds with some local variation.
Plus I know some Australians have asked where my accent is from (err...Australia?), so I accept the possibility that my own little idiosyncrasies might be throwing it off a bit, even though when I travel overseas and talk, I think I sound like Steve bloody Irwin...
My wife just spoke full-speed Vietnamese into a Google Doc, 100% correct, even the tone marks.
It's just another demonstration that Google and other similar companies are fundamentally on a different trajectory. Let's ride this rocket...
phonemic orthographies (These languages should work better) - https://en.wikipedia.org/wiki/Phonemic_orthography
English is highly non-phonemic the only language I would think is worst is French/Greek with silent letters and the different accents and in Modern Greek /i/ can be written in six different ways: ι, η, υ, ει, οι and υι. My Italian and most Eastern European languages would be easier European languages for voice translation.
This seriously reminds me of my days in college and grad school learning/teaching dead languages. The dead languages (Latin, ancient Hebrew, Classical/Koine Greek and Aramaic) we really don't know how they were pronounced so we just made them phonetic, which tells you that we don't speak them correctly since no language is 100%.
Vietnamese is interesting because there are several dialects that would trip this up but doesn't appear to. https://en.wikipedia.org/wiki/Vietnamese_language
I suspect Vietnamese tones (6 tones in total) give an stronger, more consistent, "orthogonal" signal that helps speech-to-text. For example, a word spoken with an "up" tone will always be spoken with an up tone. Whereas in English, depending on the word's position in a sentence, the speaker's emotion, etc, you might have great across-word tonal variation, even within the same speaker's multiple utterances of the word.
When you listen to Vietnamese it's very "short and choppy" with a lot of tonal variation. I suspect that the recurrent neural nets Google is training for speech recognition purposes have better inference over these separable, tonally-consistent utterances compared with harder to separate and highly variable English speech.
This should enable the very young, the very old, the disabled, etc. to digitize (and therefore share) their words, stories and worlds. Powerful.
In case you were unaware, you can listen to past voice searches here: https://history.google.com/history/audio sometimes my wife and I will go back and listen to my son's searches for good laughs all around ;)
But YouTube auto captions? They seem just as horrible today as they were years ago.
In fact, when I wanted to use voice typing at my desk I would just open the phone, open the document on both the phone and Chrome and watch the text appear on screen. Then edit it from the keyboard.
There are cases where it can still miss by a large margin, but general usability has greatly improved over past 2 years.
I would love to see that this thing become smarter, that it can detect you are saying a list of things then format them automatically. That well be truly AMAZING!
I wonder organizationally if it's structured that way, but I digress...
This particular integration is really impressive after some cursory testing.
If I were CTO at, say, Dragon... I would be having a sleepless night.
https://googleblog.blogspot.co.nz/2015/09/google-docs-classr...
Under the first video, see "With Voice typing, you can record ideas or even compose an entire essay without touching your keyboard."
---
It seems that what's new are the "Voice commands" to Select Text, Format, etc. etc. That's really great! Take a look at the commands here:
I believe so much in this that I made a Siri->Windows bridge and dictate 80% of my text with it. (except that hash rocket) http://myechoapp.com/
https://www.youtube.com/watch?v=x0GXX-SJuQM
John Siracusa used Dragon to write his 20,000 word Mac reviews.
Developers have even used Dragon to assist with programming: http://ergoemacs.org/emacs/using_voice_to_code.html
I think voice accuracy is there but we need more integration with our apps and operating systems. Consistency would help too. Google recognizes "undo" while Dragon recognizes "scratch that"
In the hospital I worked in, there were four full-time staff doing just this. Just in radiology.
When I had this job for a couple of months most of a decade ago, it was obvious what would eventually happen to those jobs.
Although i just assumed that google would be running the speech recognition themselves.
Google: please don't mess this up as a science experiment as you do so many products - build it for real.
This is much much slower than I can type. I know it'd be hard to infer punctuation, but even just a one syllable shorthand command for each common bit of punctuation would be nice (you can disable it by default).
I really want to use this stuff too, because it means I could write stories while exercising at home, which is apparently how Randy Pausch wrote his last novel (he dictated it and had someone else transcribe it... I'd like Google to do that last bit for me).
http://whatsnext.nuance.com/connected-living/thursday-tip-ho...
If you simply want a better free product, you'll have to wait a few more years. In the meantime, the solution will work well for tens of millions of people and Google can learn from them.
Only lacks DNA, fingerprints, retina footprint and non-verbal communication.
Oh wait, it's only a matter of time.
Edit: I didn't downvote you, though, and I've upvoted to counteract whoever did. The point you're raising can clearly contribute to an interesting discussion.
By the way, my impression from talking to someone in the know was that Google has been developing client-side voice recognition that is only a about couple of years behind their cloud voice recognition in quality. (The client-side service still requires the huge amounts of voice data to be used to train the software, but it can run offline on a cell phone processor.) Right now being two years behind in quality is quite noticeable, but that will change soon once quality plateaus near perfect, and client-side voice recognition will make sense again.
Just rumors though.
Powered either by mobile GPU's and more efficient ML algorithms than we currently have, or specialist chips like http://www.movidius.com/
I love Google, they make great products, but at the same time I can't help being afraid of this marvelous monster.
Distributed, self-hosting would be far safer than having all of your most personal information owned by a remote centralized authority.
But they haven't even done that for Hangouts yet, let alone for their databases.
I would rather google had my DNA on file if I had to choose between the two. Why? Because voice signatures can be used to passively authenticate/identify someone without them even realising they've been authenticated/identified. You just need control of a nearby microphone while they happen to be talking. At least with DNA, I'd probably notice someone jabbing a needle in my arm or sticking a cotton swab in my mouth.
I you are alive, you already leave traces everywhere. No need for a needle or cotton swab here.
Of course I am being paranoid, but we all should be :)
It's true that DNA can also be 'non-obviously' collected. But there are two key differences between DNA and voice auth, when looked at from the perspective of a national surveillance agency:
1. 'Non-obvious' DNA collection occurs after the fact, and is much more costly. You have to send someone to the physical location, collect samples, and process them in a lab.
2. Because of (1), DNA could never be used as a 'always on' mass surveillance system. As opposed to voice auth, where you just need enough microphones, covering enough area, streaming data back to a server for processing. Sorta like how Batman finds the Joker in 'The Dark Knight' :)
- a full and complete statement of the facts and circumstances relied upon by the applicant, to justify his belief that an order should be issued, including (i) details as to the particular offense that has been, is being, or is about to be committed, (ii) except as provided in subsection (11), a particular description of the nature and location of the facilities from which or the place where the communication is to be intercepted, (iii) a particular description of the type of communications sought to be intercepted, (iv) the identity of the person, if known, committing the offense and whose communications are to be intercepted;
- a full and complete statement as to whether or not other investigative procedures have been tried and failed or why they reasonably appear to be unlikely to succeed if tried or to be too dangerous;
- there is probable cause for belief that an individual is committing, has committed, or is about to commit a particular offense enumerated in section 2516 of this chapter;
- there is probable cause for belief that particular communications concerning that offense will be obtained through such interception;
- normal investigative procedures have been tried and have failed or reasonably appear to be unlikely to succeed if tried or to be too dangerous;
- there is probable cause for belief that the facilities from which, or the place where, the wire, oral, or electronic communications are to be intercepted are being used, or are about to be used, in connection with the commission of such offense, or are leased to, listed in the name of, or commonly used by such person.
- No order entered under this section may authorize or approve the interception of any wire, oral, or electronic communication for any period longer than is necessary to achieve the objective of the authorization, nor in any event longer than thirty days.
In contrast, FISA warrants are indiscriminate and the criteria for acceptance seems to be 'asking for one'. Of the 35,529 FISA warrant applications made up to 2013, only 12 were rejected.
You seem to be laboring under the assumption that your phone company is more benign than Google, which is an assumption I am not willing to make (for Verizon or ATT among others)
HN likes to talk about the layperson non-techies disregard for technology's impact on our rights, but a grep of this thread turned up no hits on biometrics.
[1]: https://en.wikipedia.org/wiki/Biometrics
[2]: https://www.google.com/?gws_rd=ssl#q=inurl:google.com+privac...
Short of surveillance cameras in our homes, I can't think of anything worse for privacy than government coming in to possession of passive biometrics like voice signatures (e.g. a google voice DB dump collected through a FISA warrant). Just in case people genuinely don't comprehend the danger:
1. They allow a third party to identify you against your will.
2. They allow a third party to identify you without your knowledge that you have been identified/authenticated.
3. You can't revoke a voice sig credential like you can with a password credential. Once the government has it, they can do (1) and (2) at will for the rest of your life.
http://www.scientificamerican.com/article/biometric-security...
Your voice is a biometric. Innovation is fine, but Google needs to address our digital rights.
One reason I can think of is that there was no incentive to build speech recognition because in most organizations that use computers, people talking to their computers is likely to disturb everyone around. But with more and more jobs becoming work from home, this may not be such a problem in future.
Thank you for posting this.
Don't forget to see the entire talk (28 minutes): https://www.youtube.com/watch?v=8SkdfdXWYaI
Voice will need some serious software support if it is going to take off. There is zero chance that users will learn those mnemonics, just as now almost nobody knows how to touch type. The usage will be more in the form of general commands, than specifying every action like we do it now with the keyboard+mouse.
I did not know voice typing was possible before, this has made my day. All I need now is to invest in a decent microphone!
If you generally feel like having a better microphone would help in other cases as well (Skype, vlog, what-have-you), then I'd agree with the above post.
I applaud the effort, but spreadsheets are difficult enough to navigate at the best of times, let alone trying to manipulate with your voice.
Running precise commands? Keyboard wins.
No keyboard? I love the voice search of my roku 4 over trying to search via virtual keyboard and arrow keys, or using my Amazon Echo to find artists/songs
Long text? While I prefer typing, I remember one author I heard (Kevin Anderson, I think) who dictates all his books (first drafts) into a mini-recorder and pays someone to transcribe them. I find that hard to comprehend, but I bet he's not unique.
Spreadsheets? Keyboard wins for data entry, but I bet voice would be convenient for analysis. "Sum column B" "What is the average of column D, excluding 0 values"
Sometimes, things pop-up in mind suddenly and we need to note them down right away.This is definitely the coolest feature. Also, on the other side, this is also helpful to people who face problems with typing, or somehow can't type (temporary injury in hand, or slow typing speed, etc.)
Anyone knows if this is available through API? I would love something like this when writing text on emacs & org mode :-)
Tested it for a quick Todo list and it was about 95% accurate. I can see myself using this
But usually such dictating software does a very bad job regarding Portuguese, regardless of the language variant.
Apparently I offended some people, thanks for the downvotes!
I am not sorry if others cannot take critics from those used to be ignored by them.
Heck, Google even have specific doodles for Brazil and even have an office here in my city.
I'm a native Portuguese speaker and I'm not pissed off as you are.
I have recently been reading The Hobbit with my 6 year old son Benjamin stop he has become engrossed in the book somewhat and then just hearing about the Adventures of Bilbo Baggins and Gandalf as well of course as mine growing are a diary Norrie tilly tilly Bailey and wailing not forgetting the king Under the Mountain himself starring oakenshield Gollum Gollum