Amazon Lex – Build Conversational Voice and Text Interfaces
aws.amazon.com
aws.amazon.com
Amazon making this is great because it's developer friendly and will surely fall in price and improve in quality over time. They also have to make it work well because it's a critical ordering channel for them going forward.
Seems like IBM has been focused on this for years though? Worried they will try to make this a profit center though... rather than continually dropping the price.
Other companies working on this I should be aware of?
Thoughts?
If you want something kick ass, I highly recommend building it yourself. The bar set by IBM's Watson, Nuance's Nina or IPSoft's Amelia isn't actually very high. For non english languages, anyone who has some knowledge about NLP and deep learning will easily surpass them.
https://github.com/larspars/word-rnn
RNN is great approach if you want to play around with text generation with deep learning. But for a chatbot, deep learning alone is not there yet. We create our own intents based on the domain and predict the intent of each question. We have also made it looks smarter by creating an intent hierarchy where we try to do multiple predictions for a question with a goal to drill the question down in the intent tree. In that way we know that the question is about a bank card and can figure out if you want a new one, block it, increase credit, set limits and so on.
That might be true, but the ability to spin something up without all that overhead is any MVP'ers wet dream, yes?
It's also worth pointing out there is a huge difference between using the bot for a critical piece of your product vs as a supplement to customer support.
Disclosure: I work at Microsoft AI and Research.
None of the above? We had desktop voice to text 20 years ago (that was about on par with google now for me) so I don't see why everything has to go through cloud services.
To properly produce speech, the synthesizer needs to take into account context of around words to determine how to pronounce ambiguous words or how to pace the speech. What about the tone? How about making the speech sounds more engaging? These are the things that need data model to work great. However, it's just difficult to cram that into a desktop app. And why would any company do that when they can put the service online and charge for it (which is totally fair)?
Apple seems far and away the best desktop solution. For other OS's though, it's frustrating that they haven't improved on what computers of the early eighties could do. Now you've made me nostalgic though, I'll never forget the first time I heard my amiga 500 read my words.
Same scheme that started AWS.
Basically, Alexa as a service?
Is there something out there that does this?
For certain things that require hands-free usage, this would be a killer feature.
ex. a workout-app that tells you your next rep and weight, and you can respond with what weight/reps you did to add to your log!
(disclosure: I work for AWS and my team built the Mobile Hub - Lex integration)
There is some case study at the bottom of the Lex[0] page that had a disconnected "Polly" thing that I didn't' quite get, but that makes sense.
And it looks like Lex takes voice as input, although I don't know if you can put your app into a "always listening for trigger word" mode for truly hands-free operation or not. Will have to play with it.
I too, am very interested in a domain specific application of lex+domain knowledge+Poly
> It's pretty bad to have page after page of text about a
> voice synthesis service and not have "Click here for
> demo" above the fold in the first page.
There's a table of demo clips in male/female voices for a few languages at the bottom of the first page.Personally, I was thinking of something cheaper, such as a Raspberry Pi with a stepper motor.
I even put my son's furry dog halloween costume on it :)
Just over 10 years ago when I was a 6th grader I wanted to do this, but I was out of luck. There were barely any products out that had 'conversational interfaces', let alone the commoditisation to enable just about anyone to do it as a hobby.
In past years, some people on HN have wanted to combine them all, but it's never happened.
I think it makes sense to leave them separate. They are all different products, and if they were released weeks apart they would warrant their own story.
No. You'd have two thousand comments about totally disparate systems to filter though.
This is so comical. Alexa is phenomenally bad at conversation, it is so bad that there are almost no successful "apps" built around it despite having an API and app platform.
Alexa is decent at single command/action model nothing more than that.
Every interaction with Alexa is command, response, and that really limits what you can do with it.
'Alexa tell app X to start my car'
Something that supports actual conversation which means not having the user respond immediately . For example its impossible to build a cooking app that walks me through the recipe because 'conversation' is over within X seconds. Of course you could do weird hacks like maintaining conversation on the server side based on sessionid, but that still requires user to constantly prefix their conversations with 'tell app X ..'
These basic UX limitations make it impossible to build any serious application that can do more than 'joke of the day' type of stuff. All the apps in alexa app store are totally useless 'fart noise of the day' type of garbage.
Disclosure: I work at the co that builds this
Lex looks alot like Alexa though the setup flow is a bit different. Also has prompts for each slot needed. That's nice.