What this culminated in is a platform where 80% of request, and pretty much 99% of "commands" are served by rules built with a team of linguists.
What this culminated in is a platform where 80% of request, and pretty much 99% of "commands" are served by rules built with a team of linguists.
It feels very fractal but on the other hand if Alexa has only a specific gamut of responses it's not exactly a limitless state space right?
Very curious about how those rules look like though
the rule were really just strings and we had efficient matching against it. I didn't work on that, I would assume some sort of LHS.
May just be a long list of if/else and/or switch statements or isomorphism.
The problem is the space of possible commands is waaaaay bigger than the space of commands you can manually handle, which means if you just randomly try stuff 95% of the time it won't work. Users learn that very quickly and end up sticking to the few commands they know work.
The one exception is "search" queries - "how tall is Everest" and so on, but that only really works well on Google's platform because they've done all the work for that already.
Contrast that with LLMs which basically at least understand everything you're asking of them. If you give them a simple API to carry out actions they can do really complex commands like "send a WhatsApp to my wife telling her how when I'll get home if I start cycling in 10 minutes". That's impossible without LLMs but pretty trivial with them.
Obviously the downside is they are prone to bullshitting and might do completely the wrong thing.
Isn't this what thumbs-up/down RL is for? To improve the quality of the results.
Had a small Google Assistant thingy for years, and that search stuff works great, until it doesn't, and completely misses the mark. This immediately kills trust and reduces it to a gadget that I will only use for non-critical stuff, always expecting it to break anyway.
> The problem is the space of possible commands is waaaaay bigger than the space of commands you can manually handle, which means if you just randomly try stuff 95% of the time it won't work. Users learn that very quickly and end up sticking to the few commands they know work.
This is not strictly true. Context free grammars can be written to handle (finite) sentences of arbitrary length! if you have a rule like "play me <song>" and then <song> can be "a song that lasts longer than X" or "a song by <artist>" (then you have <artist> be "<some name>" or "some German singer" or whatever....). You can just keep on going.
I don't think even pre-LLM technology allows you to do this.
I can't do something as basic as goto Spotify's search page and filter "only genres I like", neither a smart version of that filter or a manual version of that filter is possible.
“Who has birthdays today?” And I would get a list of famous people with birthdays today. I could also ask if Alice and Bob (two names in the list) had the same birthday and I would get an answer (one time I think I got back some internal query language for it instead… but that’s lost in old bug reports).
Now any interesting question starts its answer with “according to an Alexa answers contributor…”
Kind of says it all.
At the end of the day if you have a complex product but don't have comprehensive test cases, it's just a matter of time until your users notice your product sucks.