Facebook’s Virtual Assistant M Is Dead
wired.com
wired.com
Imagine if this was a realistic conversation with an assistant:
> "Hey Google, I'd like to order a pizza."
> "Sure, what kind?"
> "Let's see... cheese, pepperoni, sausage... and maybe some green pepers?"
> "Alright. What size?"
> "Hmm, so I need to feed 4 people..."
> "Sounds like a large?"
> "Sure, let's go with a large."
> "Alright. There's a Dominos nearby, I can order that for $8.99."
> "Sounds good."
> "Alright, I've ordered your pizza. Expected delivery in 15 minutes."
No need for the human to understand what the "entry point" is, because you can approach the assistant with pretty much _any_ entry point and it'll give you a useful response. We're still not there yet, unfortunately, and I think it'll be quite a while before we are.
In fact, the pizza order is the No. 1 scenario looked at by the chatbot providers. In fact, it was exactly the scenario my old startup took as a case study and the first application we built with it. It could handle the different toppings, the sizes, and more. You could submit all your requests in one move, it would be parsed and sorted into its little slots.
The problem? There is only a handful of scenarios similar to the pizza. In most cases in the real world, you have to select from an external database, look at proprietary product names, and more. Another staple of chatbot demos, plane ticketing, only works well when limited to North America (in the English word). Good luck asking for a flight to Kinshasa, Kuala Lumpur, or even Wagga Wagga in Australia.
I am not even talking about the switchboards for multiple domains, like in Alexa. These ones only work with "leaky abstraction" (making the user learn magic keywords).
Another problem is really stupid. It's the availability of the datasets. The funny thing is, ye olde style semantic frameworks fare better than the machine learning ones, because there is not enough data for the machine learning chatbot frameworks, and without it, their mighty capabilities are pretty much the proverbial spherical cow in vacuum. But because the semantic paradigm is not kosher/kewl anymore, very few enterprises agree to deploy it.
None of that matters though because the users never liked typing a lot. Back in 1980s - 1990s the adventure type computer games switched from (mostly working) command line interface to point-and-click, and very few users objected.
My take is, the key is a conversational UI with strong visual feedback. For the pizza scenario above, I would draw icons of cheese and numbers, so that the user can be sure it worked.
And when limited to text. I once saw a demo for a voice chat for ticket orders. It took a full minute for it to tell you the multitude of options for flying 'from the Netherlands to New-York next Monday'. A human assistant can reason on what an acceptable flight may be based on a small set of parameters. A chat bot would need to know every preference in detail, like destination airport, time of day, budget, etc.
The same thing happens for McDonald's and other places replacing cashiers with touch-screen terminals, not only I am sure of how the system understood me, I can easily navigate the several options without annoying an employee with dozens of questions, going from that to conversational UI is a step back in the wrong direction.
Just this morning I stopped at McDonald's for breakfast. I gave my order, they entered it, and it showed up correctly on the screen at the drive through. However, the order taker read back something completely different (than what they just entered). Since I saw that the order in the system was correct I ignored that and simply said "sure that's right."
This is why I prefer an interface that bypasses the human order taker. They may say one thing and do another. Even in my case, they may have then thought "well he confirmed what I just said back, and that doesn't match what the screen shows, so I'll fix it to match what the customer agreed to." Then I would have the wrong order. That can't happen if I interact directly with the order system.
Once I went to a McDonald's and ordered some double cheeseburger with no onions, when the order came it was just the buns with no meat.
Especially in a place like McDonald's where they serve several people familiar with the menu, the staff seems annoyed if you don't know what the ingredients of one of their menu items are and often don't handle specific requests well.
It's not just language issues, my ex-girlfriend always asked a lot of questions when ordering at restaurants, sometimes checking with me, you can visibly see the waiter getting tired or annoyed because of this, not coincidentally, there is always some minor screw up with the order because they forgot to write something down. The same happens if you go to a restaurant with 20+ people and the waiter comes in to take everyone's order, the chances of making a mistake goes up.
Personally I'm techy and I still prefer the human aspect. In fact part of the appeal of going to a restaurant is to be waited upon - otherwise I might as well just order a takeaway online. Sure they might occasionally screw up your order but this doesn't happen nearly as often as this thread would suggest. I do eat out a lot and I honestly don't think I've had my order messed up in the last 2 years. I wouldn't say I eat at particularly posh places either though I do actively avoid most fast food establishments (not a snobby thing, I just don't like the taste of McDonalds et al) so maybe the issue of reliability is more subject to the lowest paid positions in the food service?
In any case, even if you did have your way and entered your orders directly into a computer, you'd still have to deal with the fallibility of humans with the chef cooking your food, waiter / delivery driver distributing your food, and anyone else who exists along the chain. In fact I wouldn't be at all surprised if many of the mishaps described in this thread were actually failings of those individuals rather than the order takers whom you assumed had messed the order up.
But if I'm at a good restaurant, a server is actually part of the service. They can give you recommendations on dishes, and wine matches, and help you get exactly what service you want. Also the human experience is just part of the fine dining experience.
It would be really strange if I want to Eleven Madison Park or Bouchon and they just handed me an iPad to order my food from.
That just sounds like a very poor database.
It's not just simple pairings, but includes nuances that would take many different fields to capture: Dish X can be made vegan, but tastes better if you then order with extra seasoning.
With a tablet, you could filter the list with a single tap. I've thought about building such an app a few years ago because I'd love to use it, but I have no idea how you could sell it to restaurants. It seems most places are too conservative and cash-strapped to tie themselves to proprietary tech.
It’s awful — probably an order of magnitude more awful than dealing with a non-English speaking cashier. As with anything digital, it’s a sales funnel designed to trick you into ordering whatever their profit driver is. Have fun not buying a value meal.
But what if you don't know what your options are and how much they cost? You have to ask the person to list off the possible pizza styles, sizes, toppings, and the prices for each one. Then they have to tell you about all the available side dishes and desserts (with prices, again) and how there's a half-off deal if you get THIS side with THAT pizza on a Tuesday, and on and on and on. It'd probably take a good 20 minutes to convey all that over the phone (I hope you have a good memory or are taking notes), and by the time they're done, the poor employee is probably so frustrated that they're ready to strangle you.
Or, I can suck up all that information at a glance on dominos.com, and I won't have to repeat my credit card number over the phone 5 times before they get it right.
Restaurant websites always, without exception, suck. Dominos, while the pizza is garbage, has a wonderful ordering website. But even then the actual menu is awful.
They are always a sales funnel first, menu second. They cannot give accurate ETA, ever. It’s harder to display multiple choices well on a screen vs a sheet of paper.
Unless you are a place with 5 menu items, the paper menu is superior in almost every scenario.
Otherwise, they're just acting as a voice interface for a GUI that I'm perfectly able to navigate on my own, and I'm not a fan of voice interfaces when dealing with technical systems.
I think this is an assumption that might not hold true, outside of the tech sphere.
Us? Of course we want imperative interfaces! E.g. "Order me a pizza using JSON with my pizza_now script with options -g and -b and promo code CASEOFTHEMONDAYS"
People who are not technically inclined? End-state declarative. I see most as fine saying "Order me a pizza with pepperoni" and having it show up at their door.
Ordering a pizza as shown in the above example is very contrived, no one needs that, a GUI is much better to execute this use case. But the power of chatbots will light up if it can answer 'Would this pizza be too spicy?", "can you deliver this after 4pm?". What I mean is when the chatbots can take over more of customer queries which otherwise might be directed to the store via phone call. Or something which requires deep knowledge of the product and when not every corner case can be put on the GUI menu.
"Ok Google, get me 10 Margherita pizzas by this evening" is super convenient.
Wait...
A while back, one of the big features Google advertised was the ability to start a transaction by voice, and then continue it on a device with a screen.
I'll pay but I prefer when someone else orders a pizza.
Correcting that means now you’re arguing with Google/Alexa/Siri, which is an infuriating experience.
[1] https://web.stanford.edu/group/cslipublications/cslipublicat...
For the average consumer who otherwise didn't specify, Dominos is a good choice: it's cheap, reliable, they understand the menu and deals, and it's consistent across time and place.
What if Dominoes doesn't have green peppers? What if they don't like the pizza brand? What if they're asking because they have a coupon? You may end up repeating sections of the conversation multiple times, and in the end the user will just end up so confused they give up.
The only difference is that you would want a confirmation step with the VUI - i.e. "Alright, you want a large cheese, pepperoni, sausage, & pepper pizza from Dominos, which will take 15 minutes to deliver. Place the order?"
I do see the benefit for employers.
> take lamp
You can reduce it to a flow of information. And chat is not always the best way to convey that information.
If we had a secretary (think 1950 when they were common) doing this job mine would choose something nice, yours would come back with your choices. Either way flowers would arrive. Your secretary would probably learn your preferences eventually and soon come with one picture for final approval (which would be right). Mine would remind me a week before our anniversary to start working on the poem.
Both perspectives are correct. There is no reason an AI couldn't handle both, and every reason they should - but of course it will be years (at best) before they do. The first time the AI comes with "I found these flowers, see the pictures [on whatever screen is handy]". I respond with "whatever, just pick something", and the ai decides to not bother asking again. I'm depending on the AI flower ordering algorithm to not make a bad choice - making me happy is harder than making you happy.
Imagine a fashion designer having a chat with a human assistant. The designer might tell them a kind of shirt they want to use for a shoot, and if it’s not obvious what to buy the assistant will come back to them with a list of options (pictures of each on a piece of paper, for ex). The designer can ask questions about each one, ask for different shirts, and simply speak when they decide: points “this one”.
This all seems very doable with current AI assistants. Google for ex already has results from recent voice searches waiting for you on your phone.
Real AI would be fully capable of doing this just as well as a person, but it would take a lot more "intelligence" than currently exists.
But if I ask Facebook/Google/Amazon to order some flowers for me, why _wouldn't_ they monetise it? Gotta pay for those tens of thousands of engineers somehow.
This is something a good personal assistant can do easily.
Though eventually AI will gain those capabilities, one way or another.
Chat bots are on a spectrum between command line and a real human assistant. If I hire an assistant, I probably don't need a manual telling me how to interact with them. A command line is pure mystery. I might be stretching a bit for the metaphor, but the difference is how well they understand human language.
A commandline requires, essentially, learning a new language in order to converse with a computer. A human only requires a small amount of learning in getting the social dynamic relationship right.
Chat bots, virtual assistants, etc. are all pretending to be much closer to the human side of language understanding than they are and they won't be all that useful until they make quite a bit of progress. (and actually I would appreciate a much more commandline-like interface with a voice assistant for now because I'd bet it would be much more likely to give me what I want)
"What would you like to eat today?"
"A bacon cheeseburger with cheddar, with lettuce, ketchup, and onions. Fries on the side."
"Sorry we don't offer fries with our burgers. What could I get you instead?"
"I see french fries at the next table!"
"Oh, you asked for fries, not french fries. Sure, we can do french fries. What would you like for your side?"
"... I want french fries."
...
It's not actually true that any input is accepted. Pretending that it is just creates frustration.
Reminds of those text input adventure games like King's Quest, where you had to figure out a valid verb + noun combination to use.
https://support.microsoft.com/en-us
Screenshot: https://imgur.com/a/Nh0Nn
Anyone who is dealing with this chatbot is probably aware of its main scope -- to provide tech support/answers. Look at the example prompt that is given: e.g.: Reset my Microsoft account password
How try to find that on the Microsoft Support page as it is:
My first instinct is to do a Ctrl-F for "password" or "security"; the first term brings nothing, the second term brings up topics about viruses and ransomware. Even if you realize that you have to go to the "Microsoft account help", you still won't find it. Right now, there's a link for "How to reset your Microsoft account password" but only under the subhed of "Trending topics" -- and I would have never found it if I didn't use Ctrl-F. And past research has found that the vast majority of Internet users do not know how to use Ctrl-F [0] (and I bet it's much worse today, given how most people now just use phone browsers).
Of course the smart thing to do is to use the Search box at the top of the support page -- in fact, I wouldn't even use that because I just use Google for everything, even for finding Stackoverflow answers. But I'm guessing most people don't.
I'd argue that the average user, unlike most anyone who even knows what a command-line interface is, does not care about an interface having "acceptable entry points". And in many such cases, it may be near impossible to design an interface for such users, whose needs and means of expressing them are infinite.
For example, I know that if I can't login to my Microsoft account because my cookies expired and I forgot my password, that I should be searching for "reset microsoft password", or even, "reset microsoft live password". However, I bet there are a lot of people who will state this:
my email doesn't work or i cant see my email
The chatbot helpfully provides a list of options, including "Signing in to Outlook.com", but just in case, options like "Unable to send email". If you search for "my email doesn't work" using the support.microsoft.com search bar, you get a list of Google-like listings pointing to the answers.microsoft.com forums, for topics like "Why doesn't my Hotmail work any more?"
I guess there's no reason why the support searchbar couldn't be tweaked to return the kind of options that the chatbot does, but the chatbot has one key advantage: there's always a "None of the above" option, which takes you to a secondary list of options. This kind of funnel IMHO is way more useful/comfortable than a search interface, in which it's not clear whether you need to rewrite your query, or keep paging through pages of increasingly irrelevant or confusing search results.
Theoretically, the chatbot funnel gets better at funneling people, based on analysis of past users' paths through the decision tree. But even a very dumb bot that doesn't intelligently respond has one more advantage: if you hit "None" a few times in succession, you'll be given a link to "Talk to a person", which signals to the user that, "OK, you win, no need to keep trying search queries, let's get you some human help".
That signal doesn't exist in a traditional webpage or search interface, so the risk is that the user keeps searching the pages in vain until they get so pissed off they just give up. Not everyone has the wherewithal to demand manual help.
[0] https://www.theatlantic.com/technology/archive/2011/08/crazy...
But not that sad -- I didn't use it once after a few easy questions the first day I got access. It would occasionally pop up some suggestions as I was typing messages, but I would always ignore them as they were irrelevant.
I think the biggest issue is that they weren't up front about it. They tried to make it seem like it was an AI doing all the work.
I think if they had straight up said, "this is a human and we're training an AI", it would have been a lot better. It would have allowed them to do things to get stronger feedback, like asking, "was this the right suggestion?" Then when I got irrelevant suggestions, I could give them feedback as to what was wrong and why. But it never asked me for that so I never gave any feedback.
I thought I was just one of millions training the AI and that they would get plenty of signal with all those users. I had no idea it was so limited -- I definitely would have been much more active in giving feedback had I known.
The verge did a better job of explaining what was actually going away: https://www.theverge.com/2018/1/8/16856654/facebook-m-shutdo...
But in that case, most of what I said still applies!
But perhaps the perception that your question was being interpreted in an intelligent human way caused users to think differently and rephrase their questions in a way that made it easier to find the most relevant help/support links? I remember how interesting Ask Jeeves seemed to be -- though to be fair, Google wasn't much of a presence in 1997.
That's about it.
(To be sure, the tech for a lot of this already exists they're just not exposing a text-based version.)
I wouldn't want to be quite as verbose for a text-based version, but oftentimes it really is easier to type more versus less if you're confident the recipient will read and parse the intent of the whole phrase.
All of this is gleaned from fawning articles in the western tech press though, I'm not sure what it's actually like from the average Chinese citizen.
I love chat as an interface to deployments! Hubot is a great framework/bot for hooking into your own environment. Typing deploy prod master in a Slack channel is great. Why is that better than ssh'ing into a jump box and typing cap deploy prod? It's multi-user! Everyone else can see what's going on.
1. When reasonably scoped (i.e., to a specific use case) and iteratively optimized over time, chatbots can meet user expectations quite well.
2. If ostensibly intended by their builders to handle every type of request on the fly via ML pixie dust, chatbots can be miserable failures.
Both things can be equally true.
(Disclaimer: I work for Chatbase, a service for analyzing and optimizing bots. Maybe Facebook should have looked at that. :) )
What purpose would an optimized, limited-scope chatbot for Facebook even look like? Though come to think of it, I can think of a few usecases if Facebook wasn't out to dominate everything about real life. For example:
- When traveling to a new city: "Do I know anyone who lives here or is currently visiting?"
- When wanting to read about or discuss news topics, but only from my current network: "Are any of my friends talking about the election?"
- When bored: "What games are my friends playing?" (I'm thinking back to the time when FB was a games platform for things like Words with Friends)
All of these may be findable through a combination of searches, but I'm not a power user, and I bet most people aren't. I think if I go to the "New York, NY" location page, there's a section that lists friend connections, but a bot that processed a natural language query would be so much smoother.
And what about queries like: "What are my friends doing this weekend?". Searching that exact question brings up nothing of relevance. When I do a search for "weekend", the top results are for things like "Vampire Weekend". I have to scroll down to find a section for "Posts from Friends", and that only contains posts (even from months ago) that contain the literal word, "weekend".
I don't really know how to improve those results, without hurting some other kind of expected functionality. But a chatbot that purports to deal with everyday human questions might be the right interface for everyday quality-of-life questions
Those flaws derive from a wildly optimistic use case for the technology, though. A much cleaner use case would have been a bot intended for Facebook Help (instead of, or to complement, a KB -- assuming people still need that).
More ambitious maybe, but perhaps not impossible, would be a bot that looks for signs of suicidal tendencies in posts or comments and engages the user in therapeutic conversation. (?)
>"Messaging app Kik staked its company’s future on bots and “chatvertising.”
then:
>"Kik pivoted to blockchain technology."
Is there actually any logical pivot from chatbots to blockchain? I am wondering what of your core tech in the former could allow you to pivot to the latter. Or is this simply grasping at funding?
When I heard Kik does blockchain I did a double-take.
Where's that coming from? There's certainly been some important advances recently, but to claim that no progress was made is strange.
Just to give one blatant example, Deep Blue defeated Kasparov in 1997; Chinook had fully solved draughts (checkers) by 2007; TD-Gammon played backgrammon consistently at world champion level by 1995; two computer programs, JACK and WBRIDGE5 won five bridge championships between 2001-2007. All of those are 10 years older than AlphaGo/Zero and each has a very long history going back to the 1950's in the case of draughts AI [1].
You probably haven't hear dabout most of them because they were not advertised by a company with the media clout of Google or Facebook, but they were important successes that solved hard problems. There are many, many more results all over the AI bibliography.
And, just to settle this once and for all- this bibliography starts a lot earlier than 2012.
_______________
[1] All that's in "AI: A Modern Approach", 3d ed.
> the thinking is that now all problems are going to be solvable
...and then failure and the "AI winter" for a generation after the initial promise had been discredited.
What seems missing in a lot of these threads is the idea of "context", and I think that's where there's lots of room for innovation. Current voice-assistants work "okay" for single-sentence queries, but if the device doesn't understand (or if I fail to phrase things in a way that it's expecting), it doesn't ask clarifying questions, and it doesn't use past exchanges to inform future ones (beyond perhaps some voice training data). It also limits the kinds of things it can do by requiring that all of the necessary information be presented in one utterance. It also raises the "mental tax" on doing "real things" because I know I have to say a long phrase just right or start over, and that's sure to raise anyone's anxiety-levels...
They might work "okay" if you're native speaker. As ESL speaker with an accent, it's nowhere near close. It's a total PITA beyond "what time is it".
"It was easy for M’s leaders to win internal support and resources for the project in 2015, when chatbots felt novel and full of possibility."
Chatbots were new in 2015? I think it might be more accurate to say they were new in the early 1990s, but they had a revival of interest around 2015, driven by the possibility that advances in AI and NLP would allow them to do more.
One place where chatbots still have a large opportunity in front of them is in automobiles. The driver is not suppose to hold their cell phone while driving. But they can talk to the phone, and voice-to-text allows them to interact with chatbots. Someone in the industry told me that Toyota has inked a deal with Pandorabots:
Likewise, during and after my time at Celelot [1], I talked to a lot of salespeople, and they told me that was the #1 thing they'd like to see, as an interface for SalesForce. They wanted to be able to meet a client, make a sale, and then drive home, and while they were driving, they could talk to their cell phone and the Celelot service would put all the data into SalesForce for them.
[1] https://www.amazon.com/Destroy-Tech-Startup-Easy-Steps/dp/09...
When our voices are processed in systems like these are the results compared to a corpus of our own speech, a local (geo) population, or language speakers as a whole?
I’ve been curious how slang and people with poor grammar affect results of other users. Will we start seeing “thicc” instead of “thick” and “dat” instead of “that” over time?
One of the reasons I’ve been pondering this (anecdote alert) is that I frequently see iOS dictation spelling bizarrely.
The benefit for text chatbots is you can concurrently interact with multiple bots. So one can be talking to a customer over the phone and use multiple chatbots to find shipping, products, place reservations on stock, verify a credit card and so on.
IMO chatbots are still hugely valuable in Enterprise app space, but also any traditiona multi-tasking environment like customer support, telesales, and environments where is too complex to get a bot to handle everything .... sometimes is better just to let the human brain be (literally) the 'meat in the sandwich' to glue all the chatbot feeds together and create the outcome required. Production line planning is potential example - ask bots to tell you about current environment, stock levels, backlog, shift resources, cashflow, and then human decides which work gets done today. Most good planners i have met can do the planning vastly better, and quicker, than powerful systems with optimisation algorithms and Tb of data.
> That’s because most of the tasks fulfilled by M required people.
> One source familiar with the program estimates M never surpassed 30 percent automation. Last spring, M’s leaders admitted the problems they were trying to solve were more difficult than they’d initially realized.
> But as it became clear that M would always require a sizable workforce of expensive humans, the idea of expanding the service to a broader audience became less viable.
Until nearly everything has a publicly accessible API I don’t see how something like M could ever happen without extensive human interaction. I’ve been giving this a lot of thought lately. I pull information from a few intranet sites routinely for managing travel for my job. It’s almost exclusively online and because of the scattered sites it can be annoying to manage at times. It’s perfect for automation and an AI assistant but without any kind of programmable interface how would an AI assistant ever work with the sites?
I think Alexa and Google Assistant ultimately have the right idea, privacy issues aside. Start with a small enough scope and attract so many users that eventually services are compelled to support the devices. Over time the scope covers almost everything.
I think there are a lot of good use cases that can be automated today.
We asked that it drew pictures of us, and tons of other queries around memes, ect.
Outside of that the only other use case was having it remind me to wish friends a happy birthday when they otherwise didn't list it on facebook.
I wonder how many other products are secretly hand cranked. When you talk to Alexa, and the algorithm can't tell what you are saying, is there a guy in a call centre somewhere typing up what you are saying? Either to service your request or to provide a dataset of transcribed tricky samples?
Alexa is bad enough, the idea that a human really is listening to me in my home is somehow worse.
* https://www.fin.com/ * https://mysecond.com/ * https://www.perssist.com/ * https://getmagic.com/
I have only played around with Magic. Really like the idea of it but never got any real use of it. They executed all tasks I threw at them so slowly which in turn costed me too much money (asking them to warn me each morning if it is going to rain.. well, that cost me 40 minutes the first day, then I canceled that task).
Since then I got a credit card with concierge service, which I pay around 10 USD/month for. Solves all my easy tasks for a much better price.
I use the concierge service over email, and they usually take a few hours to respond. Magic started to work on my tasks within a few minutes, often faster, which already here changes what you can use it for. It is possible to call the concierge over phone when you need it faster, but email has served my need.
Here is a few things I have used them for the last month or so:
* I traveled away two weeks recently, wanted to leave my car at a car service the day before and have them fix it and store it until I was back home. Had the concierge call around and book that for me. Booked change of winter tires this way as well.
* Wanted a get a haircut a certain time on a holiday day. Had them book that as I did not no anyone that was open.
* Called them 30 minutes ahead of a full-booked train and they managed to get my a ticket (still not sure how they did this).
* Tried to buy outsold tickets to a concerts, which they did not manage to do. Wrote back that they were sorry.
* Investigate the ability to book a meeting room within a 500 meters-area.
All this costed me around 10 USD/month, and no premium when you purchase things. I have just started, possible I find a better use for it in the future.
I live in Sweden and use https://www.supremecard.se. Most credit cards buy the same service from the same third party concierge service, so I just got the cheapest credit card as the service is the same.
So Google, Amazon, Facebook, Apple, etc, are all trying to win the AI game, but in the end, to truly have an AI that is flexible and can learn on its own, we would need to go beyond APIs and create a way for machine to understand things, concepts and ideas. This is a massive task that cannot be done by one company alone. It is a new way of writing apps and services.
Google Assistant OTOH is really very good and pretty useful.
I usually I prefer the chat interface rather than a clunky browser UI where companies and govt agencies provide it, and also that in some cases you can add them to your address book like any other contact.
I guess I'm one of those computer technologists who haven't bought into the whole AI taking over the world hype.
My take on it atm is to have a flexible (should work if you mistype a letter) command-based interface which is both voice and text based, where you can perform commands like:
play songs from coldplay
set alarm to friday at 3pm
make list with words a, b and c. Give me a random item from the list
There are some tricky parts though: - Should it be context aware? Notice how I, in the second part of the last command mentioned "the list". I think it should, and maybe even ask "which list?" in case there is more than one.
- How do you define commands in a way that makes it easy to add and compose commands, and reduces or eliminates the ambiguity for the parser?
- Is the kind of parsing you do in voice recognition similar enough to be compatible with text parsing?
If someone know about some tool similar to this, please let me know.
EDIT: Fixed typo.
Now, how would the UI go about specifying which previously defined list you want to refer to? Sometimes you'll want to pick the most recent, sometimes the one most closely matching the definition, sometimes the one matching the "alarm clock" format most, sometimes you'll want to offer the user a choice among all objects similar enough to the description (What about if the object was built iteratively, do you offer intermediary objects as possible choices?), sometimes you'll want to ask a short question to restrict possible choices, if there's seems to be a clear criterion that probably improves understanding fast enough more relative to the time it takes to ask (this can depend on the user/environment/situation to choose between fast/precise answer).
To me the difficulty is that the choice of strategy can be built on the fly depending on context by humans, usually without building an understanding of all possible strategies but instead by just magically guessing a strategies which seems to fit well enough. This means being able to learn strategies based on previous experiences and building an evolving understanding of contexts.
Now this is probably not necessary to build a functioning interactor, but this is a reasonable description of normal human interaction, and the capacity for systems to adapt to contexts without much more outside help than humans do is going to be a good way to rate them.
"Play" gives away the fact you're looking for a song or list of songs, songs which you'll hopefully have in an internal knowledge base.
"all Coldplay songs": just filter by Coldplay. If your library is big enough, you could figure out that it refers to the artist.
"that": we need to filter again by the condition which follows.
Speeding up... "were the most popular of their album": take each song and corresponding album, sort album by property "popularity" and check it's the first one. You'll need to know "popular" refers to the property "popularity".
The third one is pretty similar. The second one could be easy too, but the piece "about love" is complicated on it's own.
IMO this means three things: firstly, words don't map directly to commands / capabilities, which means that having a composable way define capabilities is hard. You'll likely need to define many ways to do each thing, but you could add them one by one, over time. Secondly, the tool should be able to tell that it doesn't know what you're talking about (what is this "about love" thing about?!?). Lastly, it should be interactive, so that it can ask/tell you when it doesn't know ("what list do you mean, A or B").
Your comment about context is on point. We can't expect the tool to understand context it doesn't know about, which is why we cannot expect human level from this. But it doesn't have to be human level, it just needs to be good enough to be useful, and we do have a lot of space to improve.
We spend years learning this stuff, we could slowly teach a few tricks to the computer ;)