Show HN: Siri-as-a-Service Speech API
wit.ai
wit.ai
Would you consider offering a version of your program that I can download and run on my home server? That would be cool...
Finally, since the title is "Siri as a Service", where do you expect the microphone to be in a Home Automation setting? Do you envision people using their cell phones for that?
Thanks.
Many home automation systems will have a built-in microphone (or most probably an array of microphones, which is efficient to cope with background noise). But your smartphone might be useful in case you are in the garden for instance!
One step closer to Jarvis.
Looks like they are using it for their project http://cmusphinx.sourceforge.net/2013/09/processing-speech-r...
Wit leverages several speech engines, including Sphinx.
And Sphinx speech-to-text users can send text to Wit to do Natural Language Understanding (turn text into actionable data).
I know this is considered bad mannered, but what's the NLP behind the scenes look like? I'm curious. :)
Can I run an instance of a wit server myself? And perhaps update its speech models regularly?
I know there are a million APIs for this, but most sound awful. I'd love a service that sounds as good as Siri.
But for the life of me I can't find a link of where you could buy it.
Doing TTS correctly, with intonation etc. is really hard. Its not really my field, but I can imagine that getting intonation right is near impossible with just unannotated text.
Alas, all of their SaaS plans have the same price: "negotiable". I never felt like negotiating, so I have no idea if they're affordable; unfortunately custom pricing usually means it's not the case.
Is there any date entity in the wit directory? How do I parse a date for e.g. 17th Feb or feb 27 or 13/02?
Datetime only has contextual date selection i.e. today or tomorrow.
Google's Web Speech API can also be used to build something similar.
Here is the Web Speech API Demonstration: https://www.google.com/intl/en/chrome/demos/speech.html
If you use Google Web Speech though, you receive text and you still have to do NLP to "understand" the user intent. The other problem is, if Google does not know about specific words (like your company or product name), you have no way to customize the engine (no "Add to dictionary" entry point in the API!).
The idea is exactly like Wit and it was supposed to integrate with all kinds of web services out there, to be as close as Siri, but do a lot more than just opening up new app or visit a weather forcast website. Bascially, your virtual assistant operates like IFTT. At the end, I think speech into home automation is the future gold mine.
I guess I'll have to update my ROS wrappers at https://github.com/LoyVanBeek/wit_ros as well.
Also, to my knowledge they rely on Android's speech recognition and the API accepts only text, not audio streams.
http://www.ask-ziggy.com is also similar