An App That Lets You Converse with the Deaf, No Sign Language Necessary
techcrunch.com
techcrunch.com
The main problem with general purpose real-time voice recognition is that current hardware is simply way too underpowered to accomplish the task. For instance, running the Dragon 11 SDK on an Intel Atom Z3770 has it about up to a minute behind transcribing the conversation! So I fear Transcene's approach is using the inferior Google Speech API which plainly put, sucks donkey balls, and is no way comparable to the latest Dragon engine. Apple uses the Dragon engine to implement Siri.
There's also the social burden of needing speakers to install an app on their smartphones and also actually have a smartphone in the first place. Will this be a free "remote mic" app such as Dragon 13 provides or does Transcene expect speakers to pony up the monthly cost as well?
I too think businesses and institutions will not allow this because they need overpriced ADA compliant solutions due to regulations. An example would be Interact-AS which is $800 and is essentially a fancy overlay for the Dragon engine (or Microsoft Speech in the low-end $150 version). Dragon itself only costs a one-time $99 to $199!
I'm also skeptical there's a viable business model in this. The vast majority of the deaf are on fixed incomes and not employed, so what is a relatively expensive $30 a month for app access buying them exactly? It better be a superior remote client to server transcriptioning experience! What's to stop Dragon from enabling multiple "remote mic" apps to work all at once with the mothership PC in their next version, etc.? And if not a client server model, what are the minimum hardware specifications to get "one second" transcriptions? A $599 smartphone is a ridiculous and overpriced luxury for the deaf.
As for Google Glass, it is a non-starter. No one wants to look like an idiot constantly staring off into their peripheral vision to read text instead of looking at whoever is speaking -- which is why Google Glass has been such a massive failure. What is truly needed is spatial aware, augmented reality where the transcriptions are placed over who is speaking via beaming text onto normal glasses or directly onto the retina. This technology already exists in various forms; it is just a matter of a real world implemention into a "killer app". Transcene, are you paying attention?
Nonetheless, this is a very important step forward that no one else is really doing, so I'm in for $250... and holding my breath.
Keep it up guys at transcense
The problem of supporting realtime conversation among multiple people is different enough from voice search that there's scope to differentiate.
To put it another way: real-time subtitles. Imagine having something like this in Google Glass. As a person with profound hearing loss, that blows me away.
This technology could easily be repurposed for subtitling videos.
Your explanation about why you think it's special us useful a good example of a positive outcome of the original comment. The way you initially denigrate the question is not.
As pg put it:
Maybe you think you're making some sort of important point here. Or maybe you realize your comment is inane and you think it's witty. But (perhaps without realizing it) you and the people upvoting you represent one of the worst forces at work in the world. The people who ridicule new things when they first appear in incomplete form are one of the worst drags on innovation. [1]
Looks to me like at least one personal already countered that claim, thus it is no longer undeniably dismissive. I would say your response to his question is even more dismissive of any "dismissal" that may had been interpreted from the OP.
As a real life example, I was leaning towards the original poster's interpretation of the product. The reply it spurred helped me see the product in a different light, and I think it has more merit than I originally did, even if I'm not sure the technologies in use, or even how they are combined, is especially new and noteworthy. As is all to often the case, it's the implementation that matters.
In think my initial opinion was dismissive, the original comment was inquisitive (if a bit critical, but I see nothing wrong with some light criticism), the reply it spurred was illuminating, and my resulting opinion was hopeful. I view that as part of HN's success, not something that needs to be overly policed.
Almost everyone in the United States has a phone. If I could download an app that runs this program along with my cousins, and have my Grandma use her 'iPad' (Nook tablet) to understand, with the assistance of something like Transcense, that would be amazing. By linking several microphones, they may be able to cancel out background noise and only highlight the specific speaker, and that would be a fantastic advance.
I'm wondering what their current state of Transcense's speech recognition is, however. From the video, it did see like there were some errors. I'm sure a deaf user can understand what was meant to be said using context of the conversation, but in a business meeting a misunderstood word can change the whole meaning of the sentence or message. I've used Siri, Dragon Naturally speaking, et al and while they're good, they're not perfect. Dragon in particular supposedly can be taught and learn the user's unique style of speech, so I'm also curious if Transcence will be going the route of machine learning and NLP.
Why can businesses not use this, and how does a VPS replace this?
That said, there does seem to be promise here, as there's a reason we video conference (or talk in person) rather than text-message for everything.
That's pretty damn amazing IMHO.