Almost every programming language has a canonical API that is extremely simple to set up. In .NET and Google Chrome, it's even first-party, and I only mention those because they happen to be the ones I've used.
And text-to-speech is no different. You can get to "good enough" so easily that I'm somewhat perplexed why more people don't do it.
I just think there is a perception that adding a speech recognition feature to your app is "hard". It's difficult to design a good speech-based user experience (I've found longer phrases work better than single-words, and it's good to try to match on homophones as well), but actually integrating the code is not that tough.
So I would encourage you to ignore Amazon. Seriously, you could make this in a weekend in a Google Chrome browser on an Android device that you plug a nice microphone into.