Continuous speech recognition will have to improve, there are currently opensource solutions with the best being CMU Sphinx [1], it does however only listen to certain keywords thus isn't capable of understanding everything you say, and it needs training. There is a Google voice Recognition API, but that's not continuous so you'd still need to listen to certain keywords and then upload the bit to google and wait for result, which isn't perfect. Recently I've come across a few sites which seem to do just that using the google voice api [2,3], but likewise far from perfect.
When it comes to responding to you there are quite some good opensource projects out there such as Flite [4], MaryTTS [5], Festival [6] and of course e-Speak [7]. They don't sound that natural, and if you're willing to pay you can get much more high quality options [8]. I know you could again use google's text2speech, but don't forget there is going to be a delay you can't get rid of.
When it comes to the whole packet together I've seen Jasper, but he has his own problems, to give an example:
"
-> So Jasper, sup?
-< “Today in news... blah blah”
(meanwhile me) NO!! STOP! STAHP!
"
Google now seems to function pretty quickly and accurately on my Moto G, but once switching apps you get stuck.
Alyt at indiegogo [10] tries to market itself as what I'd love to see...
I'm sure there are more of these projects around, but I haven't seen one yet that is actually capable of learning new things on the fly or even having a little bit of personality (why personality? Because the best GTD-tool is still a whining person reminding you about it, not a monotone computer voice that you'll shut-off anyway). I can imagine though that's going to be tough to build; with a lot of machine learning, virtual neural networks, natural language engines, lisp-alike programming language and whatnot.
When I find free time on my hands I'd love to mess around with all these components and try to come up with a solution of my own, or probably after a re-watch of the movie Her (2013) (which was very interesting just for the idea it proposed!).
[1]: http://cmusphinx.sourceforge.net/
[3]: https://www.google.com/intl/en/chrome/demos/speech.html
[4]: http://www.speech.cs.cmu.edu/flite/
[5]: http://mary.dfki.de/
[6]: http://www.cstr.ed.ac.uk/projects/festival/
[7]: http://espeak.sourceforge.net/
[8]: http://stackoverflow.com/a/4721878/1807383
[9]: http://jasperproject.github.io
[10]: https://www.indiegogo.com/projects/alyt-it-s-like-siri-for-y...
Best opensource text2speech results I've had with this blog post so far: http://ubuntuforums.org/showthread.php?t=751169
PS: My apologies for the very long, unstructured reply but it intrigued me and I hope someone will find this small writeup useful!
I wonder if a generative model would be acceptable at producing speech. It works really great on handwriting (http://www.cs.toronto.edu/~graves/handwriting.cgi?text=This+...).
For the underlying AI, this is a great demo of an AI assistant that can do things like you describe (https://www.youtube.com/watch?v=54HpAmzaIbs) . Good natural language processing, it can learn new facts and even asks questions about things it wants to learn more about.
Of course, that's an extreme example of the argument that has been used on HN a lot lately in favor of data collection, but it still gets the point across. Just because there's an industry push towards something doesn't mean that it's right.
This is not an extreme example. This is (I hope, for your sake) a joke.
Random NSA employee reading this, please don't steal my shit. Thank you