We don’t know how to build conversational software yet
medium.com
medium.com
What interface would you prefer?
https://github.com/diesendruck/ggspeak
Or for the demo video: https://www.youtube.com/watch?v=xmtrGxVvVyg
If you're interested in working on it together, get in touch.
True, and this api will be deprecated for those who want more efficient means of communication in the future and willing to work towards it.
>What interface would you prefer?
EM fluctuations in real time of the neurons in my body to machine code via trillions of nano sensors/transmitters/stimulators. Vocal cords and limbs need not apply… will be way more intuitive than "natural language".
Have you ever tried to do a "complex technical task" by directing an unskilled person over the phone? It's a tremendously frustrating experience.
Natural language alone isn't useful without your interlocutor having enough understanding of the problem domain to work out what you mean and help you. So a natural language interface to the spreadsheet would have to be on the level of "perform linear regression on these sales figures" in order to be more useful than mere speech-to-text.
I think software is inventing a lot of new ways to communicate with each other, and that's a pretty awesome thing; I just think speech is a very inefficient and imprecise way to communicate things, and efficiency and precision are what you really want when dealing with a computer.
Ok it's not necessarily a bad thing (especially in terms of accessibility - plus nobody even knows which setting they should use [1]), but I don't think this is what OP is hinting at. Say you are working in HR "How many staff took a greater than average sick days around holiday weekends?", for that task what does voice interaction provide over a report, which can even be automatically generated and delivered for when you arrive at 9:03am.
[0] https://www.kickstarter.com/projects/403524037/autonomous-de...
I don't think there's anything wrong with having a voice interface, I just never see it being anything more than a secondary convenience.
GUIs offer a lot of great short cuts that are simply faster to visually read and interact with than reading text and typing - or even saying - a response.
I think there are aspects of chat bots that really do offer a new and better user experience, but they're not necessarily the ones being highlighted in demos.
Asynchronous interaction with an interface for potentially long-running tasks is wonderful. I can tell an app to, say, find a list of hotels I might like in Argentina and then flip to another app knowing that I'll be notifieda few seconds later when the response is ready.
The way forward may be a hybrid interface, where bots can respond with messages that are little mini GUIs. There might be photos to scroll through and a few buttons to click. FB showed off a feature like this at F8 today.
Ultimately what would make a bot a really improved experience is if it knows you really well, but that doesn't imply a conversational interface. If you could say "make dinner reservations for tonight" and trust the bot would find the perfect place that would be awesome. That's hard. And it has nothing to do with chat bots per se, just a much better recommendation engine.
Functions?
* level 2 — hard-coded conversation flows
Scripts?
* level 3 — fuzzy/continuous/fluid state.
Threading?
I can't shake the feeling that deep learning is being shoehorned into a bot with a specific purpose (ShoeBot, etc). Talking to a ShoeBot, I would expect it to have predictable responses or state changes. I don't want it to suddenly infer that I'm looking for shoes for my wife because of an opaque model it's formed from an insufficient set of training data. Many applications of deep learning seem to be just a lazy way to have the machines do the work automagically.
If the end goal is a human-like level of conversation, isn't this a desirable outcome? Inference and allusion are two elements that are used all the time during conversations between people -- with positive and negative consequences.
I guess it depends on the ultimate outcome you are seeking from the conversation.
This usually results in improved performance along with "ease of use" in transitioning to new but related applications. A model just happens to be a ShoeBot when trained on specific data, but ostensibly a person or company could make ShoeBot, CarBot, ApartmentBot, etc with the exact same approach, given enough data. This is very different than a workflow of "craft tons of features for domain X, write custom scripts/conversation logic for domain X, etc.".
These choices between feature based approaches and "deep" techniques are tradeoffs reminiscent of "you aren't gonna need it" versus "room to scale", but in the ML algorithms you choose rather than the software stack/implementation.
It depends on things like how much data you have, how much compute you are willing to pay for, how many users you expect, and so on - but neither approach is necessarily wrong.
In general if deep learning approaches don't roundly beat feature engineered or hand crafted approaches, you don't have enough data or are trying to shoehorn (pardon the pun) a solution that doesn't fit. Right tool for the job and all that.
Have a look at the conversational examples Voicebox lists on their website: http://www.voicebox.com/technology/ (disclosure: I worked there for a while.)
The wide-open "talk about anything" software is still a few years out, I think. But, having a "human-like" conversation/interaction with a bot/AI already happens today.