Real-time Continuous Transcription with Live Transcribe
ai.googleblog.com
ai.googleblog.com
Communication between a Deaf person and a hearing person is a two-way street, and this tool really only addresses one of those streets.
The tool transcribes the audible speech of the hearing person, allowing the Deaf person to read the transcription.
But if the Deaf person wants to sign a response, they're out of luck. Instead, they need to type their response on the device.
That's OK, but in a "this'll make do in a pinch" kind of way. The ideal is that both the Deaf person and the Hearing person are able to communicate without typing.
There is some pretty cool research around signing-to-text translation -- Matt Huenerfauth https://huenerfauth.ist.rit.edu is doing some really interesting stuff, for example -- but as far as I know it's not ready for prime time.
Could have not said it better myself.
But, it just dawned on me that the idea is to type something and then show it to the other person. Which works, it just wasn't what I was expecting.
While I agree pieces are missing when sign communication is needed, just this piece can still life changing for the hard of hearing (who vastly outnumber the deaf). This is for my mother in-law who stopped going to group lunches because she felt so left out, even with hearing aids. If this app works as advertised, there are going to be tablets mounted on the walls running it. And maybe a phone adapter.
The thing I am most excited about is that most of this work is being done in the open, in the very least much of this is being open sourced by Google, Facebook, and other giants. For all the heat they have been taking lately I do think they deserve to be applauded for this and Of course this is mostly happening in order to sell us cloud services but it cannot be understated how helpful this will be to many around the world.
Another thing I’m excited about is what I can build with these things, and even more what others will create with these models as building blocks in their applications.
As an example, and complete self promotion here, I was able to use some open source models by Google on TensorFlow to build a cross platform App that can read Articles to you using these neural networks. The amazing thing is I built it mostly on nights and weekends, which shows how easy some of this is to work with now, you can check it out here if you like https://articulu.com
I love them doing accessibility stuff, and I can see this being useful for specific people in my life. But when I initially read the announcement I thought about using it to transcribe business meetings.
Obviously the accuracy of this won't be "court reporter"-levels, but for casual note taking it would be "good enough."
[1] https://qz.com/work/1087765/how-to-transcribe-audio-fast-and... [2] https://github.com/mattingalls/Soundflower
I'll add that I'd feel MORE comfortable with just a transcript being saved than I would an audio file (which is what happens now). Deniability.
Auto-transcribing likely falls in a grey zone, because you could argue that it is being recorded before transcription (which is true).
Given that the service is free, you have to know that Google has a plan to monetize that data with someone you don't know yet. When you find out who that is you won't be able to 'undo' the transcriptions you have done in the past.
I also think it is ridiculous that this needs "the cloud". I had pretty good speaker dependent transcription working on an Intel 486 processor, and so find it difficult to believe that given literally 500x the compute power and 2000x the memory on a typical phone you can't do this all locally?
https://www.youtube.com/watch?v=zL6ltnSKf9k
Have the HUD have a dot. When you're wearing the glasses, you aim the red dot on the speaker. The transcription could come just from that speaker.
Just a thought.