DeepGram (YC W16) Is Building a Google for Audio
blog.ycombinator.com
blog.ycombinator.com
[0] http://techcrunch.com/2011/06/09/nuance-sues-vlingo-again-ov... [1] http://techcrunch.com/2011/12/20/after-years-of-patent-litig...
EDIT: IANAL and the above is not legal advice but I have some familiarity with patent law
http://www.bizjournals.com/boston/blog/mass-high-tech/2008/0...
List of all acquisitions: https://en.wikipedia.org/wiki/Nuance_Communications#Nuance_a...
In the Boston tech scene, it is pretty well-known what Nuance does and why, and is one of the primary reasons I am very against software patents, despite having several to my name.
Everything else is hopelessly out of date. Today voice recognition happens with DNNs or LSTMs. None of those patents would apply to such methods. They seem to have attempted to patent obvious applications of voice recognition.
http://www.google.com/patents/US6839669
Obviously identifies their old product (and I'm pretty sure I can find prior art for this one) where you could report your own voice and have it, say, open notepad when you say "open notepad". It's about making a computer respond to direct commands, and audio prompts. Seems to me like this would (should ?) never apply to a system that transcribes audio. Any modern interaction system would not be about responding to commands, except in an extremely broad sense.
A page filled with a list of prior art: https://books.google.com.au/books?id=f3IV90zLmaEC&pg=PA209&l...
https://www.google.com/patents/US6785653
Is about interactions with users based on prompts. Essentially a scripted interaction with people
Again this was filed in 2000. There has got to be prior art for this one. I remember using a system that did this before I entered high school, in 1994. And that was a pretty mature program, I bet when it comes to demos and papers you should be able to find systems doing this in the 70s.
Obvious prior art, as the computer is obviously using prompts (e.g. the purchase confirmation) : https://www.youtube.com/watch?v=I7q1cE9_AaQ1 (date: 1989)
https://www.google.com/patents/US7058573
Is about methods to adjust past recognized words on semantic meaning. Seems to be talking about this sort of thing "I was about two" => "I was about to". A neural net recognizer doesn't do this at all.
Also obviously prior art : this IBM engineer explains that this is exactly what he's doing. Judging by that computer this is early 80s: https://www.youtube.com/watch?v=cpdm1O-ob48
http://www.google.ch/patents/US7127393
Is about limiting voice recognition results based on a set of rules. Why would anyone work things this way in 2016 ? It is far easier to just transcribe and respond conversationally to the transcription (or just train a network to respond directly, negating the need for actually having any rules at all. You'd just show the system what to do).
Interesting video: https://www.youtube.com/watch?v=uw8XbmzOG5s
> I've been wanting to make something like this for a while. Really nice product you have! I think this kind of service will really take off in the future. Imagine having an app that constantly records everything and allows you to search it later. Questions like "What did James tell me last November about traveling to Europe?". It would also eliminate hearsay, since you would no longer have to trust one person's word against another — you could simply search the transcript of that moment in the past. In the very long run, I almost wonder if such a tool would make lying obsolete.
I'm pretty sure it requires a whole other level of nostalgia to be immersed and get lost in transcribed texts of the audio of an event.
Not that I'm saying the latter is what the tech will remain like, but at least at first that's what it will probably look like.
[1] https://subterraneanpress.com/magazine/fall_2013/the_truth_o...
I wonder if you were to go further. Suppose you let a sufficiently deep neural network see all simpsons episodes for instance. And it manages to make new simpsons episodes. This will at some point become possible. At that point, does it count as a new work ? Does it fall under simpsons copyright ? Where is the line ?
I'm the other cofounder of DeepGram. I'm glad we're getting everybody's feedback!