Amazon virtually kills efforts to develop Alexa Skills
arstechnica.com
arstechnica.com
Fast forward to 2023 when OpenAI came up with GPT Store, it felt like a deja vu of Alexa skills in some sense. My understanding is that GPTs are also struggling to gain good traction.
There are some fundamental reasons why conversational 3rd party platforms are hard. If there’s any interest I will write about it in the Enterprise AI substack (nextword.substack.com)
But yes, natural language as an interface requires higher activation energy versus GUI which hampers how much value Alexa Skills (and GPT) can provide.
When I look at Amazon’s skills page, it’s a wasteland. Until there’s something worth surfacing, discoverability doesn’t really matter.
That said, I do have two Echos in my house (and one Home Pod) and use them all the time for basic things.
The problem is that the set of supported operations are always MUCH smaller than the set of operations people randomly try. You might develop a skill with 200 commands or whatever and think you've covered everything, but people can up with thousands of possible commands just by guessing.
This means if people just do "I'll try asking this..." then probably 80% of the time it won't work. That's an incredibly frustrating experience. You quickly give up and just stick to the features that you know work, and never try to find any new ones.
But I also disagree that OpenAI has the same problem, because LLMs means you don't need to manually add thousands of possible commands, so any random request that people make is MUCH more likely to work.
I've seen so many Alexa projects at hackathons that were exciting at first glance but didn't go anywhere.
My favourite is event recommendations: lots of people try building an Alexa thing to help you find events, but it turns out listening to a bot read out a list of 20 things happening this weekend is way less useful than browsing a web page with the same information all on a screen at once.
The one main situation where NL interfaces are superior is when you are mobile (like driving) or hands are tied up.
I think this affects GPTs just like it did Alexa. Which means that GPTs aren’t the final UI. The real innovation will be in the right AI UX.
In the parent situation, I don't want a GPT spoken interface to give me the top 20 events: I want it to give me 1-2 events that I am most likely to enjoy.
In the same way that actual conversations take into account tone, facial expression, etc. to jump straight to important information.
I thought that's where Google was going with their "we have all your data because we run all your services", but it seems they Microsoftified before they could get services cooperating for the larger good.
Why? I find conversational interfaces poor for common data retrieval. I can read faster than you can speak. I can type faster than I can speak. I'm staring at a screen 14 hours a day anyway. Just show me the list of 20 events and sort it by what I am most likely to enjoy. Provide links for more information. Show me visual promotional materials. If I need to cross reference it with my calendar it's easier if all the information is visual.
When I got to this point in your comment, I remembered the Seinfeld episode where Kramer was recreating moviefone and tried to speak a trailer, with mouth-sound-effects.
This is fundamentally impossible for a computer though, because even if a computer has perfect historical information about you it can't know some random things that would change your mind in the moment. For example, if you've been to every gig a band has done for years, but at the last one your girlfriend dumped you, a recommendation engine is still going to suggest that band's gigs even though it's unlikely you want to be reminded about them. To most users that immediately looks like a bad recommendation. If the system is only suggesting 1 thing then the whole system looks broken.
The only way around that is to increase the number of recommendations. Hence every system giving out 10+ options - because the people who make it want it to be slightly better than useless.
"I've recognized you removed all future calendar events related to {girlfriend} and your recent text messages concerning her had a negative sentiment. Did you break up?"
Not the world I'd want to live in... but for people less concerned about their data, I can't say it wouldn't be useful!
I mean sure, but just think about it, wouldn't the same happen if you have a friend telling you about the event? Or if you had an attentive concierge trying to organise programs for you? How would you like them to handle it? Not by blindly listing more programs that is for sure.
"Hey you love Blasted Monkeys. Did you know they are having a gig this weekend?"
"Nah, man. We had a bad breakup with Samantha at their last one. And besides it was really her thing and I was just tagging along."
"Oh, that's rough. I didn't know. Blast those monkeys then. How about a day a the beach then? There will be a surf class at ..."
This is the kind of interface a spoken event recommender should have. Is this much harder than just listing events? Yes, it is much harder. The problem is that if you don't go all the way then it falls into a weird uncanny valley. It feels like you are talking with a human, but a very stupid one.
That's the core property of voice interfaces and I see surprisingly little awareness for that: it does not matter if it's a phone menu beep code tree or a GPT or a star trek ship's computer: the low bandwidth linearity of the readout will never go away.
This is what makes voice interfaces so hugely attractive for the "searchy advertisial complex": if you haven't bought enough ads that the almighty relevance algorithm (1) puts you in the top spot you're out. What used to be the first page on the web is the top spot in voice, second place is first loser. No amount of intelligence can ever change that, voice interface implies handing over control in an unprecedented way.
((1) technically, claims that result ranks are not sold aren't exactly lies, when result ranks don't go to the highest bidder. But that does not mean that ad spend isn't a contributing signal in any number of deeper layers so in the end results appear indistinguishable from highest bidder, only that buyers don't get any contractually guaranteed result list exposure for their money)
The core problem is that these systems are just so incorrect in fundamental ways that they're effectively useless.
Imagine a buddy of yours tells you about an event he's pretty sure you'll be interested in. Why does he tell you about this event? Well, he knows your interests, what kind of things you enjoy, when you're free, who you might want to go to the event with, how much money you're willing to spend, how far you're willing to travel, when you like to go out... So when you're on the receiving end of such a suggestion it often feels great! It's like you've struck gold.
Now imagine your average 'AI' powered recommendation engine reading you a list of events. It doesn't feel magical. It doesn't even feel like it knows what the hell you enjoy doing half the time. Forget about knowing about your free time, budgetary restrictions, family restrictions, who you'd be able to go with; None of that stuff is even sort of in the picture. And it's all delivered to you in a voice that sounds like it would be as happy to kill you as give you advice. There's no lively back and forth on the logistics of the event. No feeling of discovery as you two talk it out, honing the plan that brings it from an abstract concept to reality.
It's just dead and lifeless and shitty.
> There are some fundamental reasons why conversational 3rd party platforms are hard.
In my mind the big fundamental problem here is the "3rd party". I'd love to have an "AI assistant" or an "AI buddy" that could watch everything I do and say and write and really get to know me super well... as long as I can be confident that I own and control everything it observes and learns. I sure as hell don't want a 3rd party involved! But alas, I don't see a way we get there that doesn't involve Amazon or Meta or Google or OpenAI sitting between me and my "AI" tools, at least in the short run.
Let the hype-funded unicorns fight to develop (& end up commodifying) the tech and then design/sell selling devices that can support it locally. In that world, the AI assistant that you buy is a discrete piece of hardware, rather than a software treadmill.
Of course, this could mean that you end up on a hardware treadmill, but I think that's probably less bad, granted we can do something about the e-waste.
I wouldn't be comfortable running that as a cloud-service though. Should be open source and run at home on my own machine.
GUIs provide information in 2D, letting eyes skim and bypass information that's not useful.
VUIs provide information in 1D, forcing you to take information in a linear stream and providing weak at best controls for skipping around without losing context or place.
Not coincidentally, this is why I absolutely hate watching videos for programming and news. If it's an article or post or thread, I can quickly evaluate if it has what I want and bypass the fluff. Videos simply suck at being usable, unless it's for something physical like how to work on a particular motor or carve a particular pattern into wood.
Alternatively it could summarise into “there’s a standup comedy gig, a few bands and various classes - what sort of thing are you looking for?” and then discuss with you to find the 2-3 events that are most relevant rather than reel off one big list.
This may require a level of accuracy and intelligence that is unobtainable to work.
It may be fantastic as an aid for low and no sighted people, but so long as I can read, a VUI is strictly inferior.
If I’m ironing and want to know when my first meeting is, chances are a VUI is better.
If I want to see my whole days itinerary, a GUI is probably better.
None of these basic realities are accounted for in current technology. Instead we have these dumb robot voices reading us results from a preprocessed script that it thinks answers our question. No wonder the monkey part of our brain immediately picks up on the fact that this whole facade isn't just a lie, but an excruciating lie. It's excruciating because it's immediately obvious that there's nothing else 'there' to interact with. Even when speaking to another person over the phone, there's a huge amount of nuance you can pick up on. Are they happy? Are they sad? Are they frazzled? Are they in a rush? Are they relaxed? And you automatically calibrate your responses and what you say in the conversation based on all of these perfectly obvious things. Normal humans automatically calibrate what they say, how they sound, what they suggest based on these cues. It works really well!
There's no reason voice stuff has to suck. It has worked pretty great for humans for thousands of years. We're evolutionarily tuned to it. It's just that all the technology we've created around it totally sucks and people are delusional if they think it's anywhere near prime time.
1. Have an API (https://www.eventbrite.com/platform/api)
2. Include JSON-LD Events data on their events pages (https://developers.google.com/search/docs/appearance/structu...)
I think the issue is that having AI just fetch information catered towards any human is not using AI at all. I'm sure the hackathon groups pitching it all started with an idea of building a highly trained AI system whereby the recommendations are meaningful reflections of whatever information it has from you. Unfortunately for them, the most lucrative part of their plan neglected problems with both how to create an AI pipeline that takes many piecemeal inputs, along with millions having missing values represented some way, and renders meaningful outputs and neglected that success would only reap a massive backlash from privacy advocates.
In the end, their plan for a super-intelligent life assistant turns into just fetching event lists from facebook (or elsewhere) without even using the demographic data it has.
You also couldn't really interrupt it, it wasn't a "natural" conversation. I knew what it was going to say and just wanted to reply already, but it only listens when it finishes speaking.
It also talks incredibly slowly, you could 5x the playback speed and it would probably still be understandable.
If all those things were fixed, I'd say it could make an okay product. Not great, but at least okay.
I actually think the core problem is that Amazon just isn't competent at deriving insights from consumer behaviour data.
If I buy a vacuum cleaner on Amazon, based on the Amazon web store recommendations I fully and without a shred of sarcasm expect Alexa to think that I've developed a vacuum collecting habit and recommend a vacuum conference if I ask it for events.
If it looks stupid but it works.... (Amazon has all the data to check that it indeed works)
What percentage of Amazon shoppers make repeated, back-to-back purchases of washing machines or kitchenaid mixers? I'm certain it's vanishingly small.
What percentage of people who just bought a kitchenaid mixer would be interested in baking pans, or whatever? Probably more. But if you buy a kitchenaid mixer, tHe AlGoRiThM just sends you ads for more kitchenaid mixers.
That is the value, not in having conversations with the bot -- google got closer with its assistants, but only because it has a creepily deep profile on its users, so conversations with google had more 'context' than alexa ever could.
I tried to ask alexa for the weather yesterday (I wanted to know if it had rained overnight to decide which shoes to wear to walk the dog), and it first gave me todays weather, then told me it only knew weather for the next 14 days - yesterday was too hard to predict I guess?
But that's the point - simple tasks where the interface is circumstantially superior? Awesome. But if I want to just chat with my computer, that novelty wore out with Eliza & Dr. SBATSO. The conversations with ChatGPT are deeper, but no more meaningful.
So what can alexa do to make my life simpler? Don't read me wikipedia pages in response to -anything-, if you can't summarize it in two sentences, say it's a long answer and you can send it to my phone if I'd prefer. Make the interactions short and sweet. Control the lights, make me coffee, walk my dog, order more coke. I don't need it to have "new skills" - I just want it to be better at the ones I actually want to use.
One of their best devices appears to in the "Temporarily out of stock" category now. The echo wall clock. It pairs with an echo device and provides a visual representation of the count down timer.
This reminds me of Ambient in the design philosophy.
I fondly remember the Ambient Orb ( http://www.ambientdevices.com/about/consumer-devices ). They had an umbrella at one time where the handle would glow if it would rain that day. They had a LCD weather/clock ( https://ambientdevices.myshopify.com/collections/vendors?q=A... ) that I really liked.
With push (notifications), it's interrupting. Alexa isn't too bad about that since it's a ring color / icon on a screen. With pull, it's "I need to fetch this data". With ambient, it's there if you want it when you want it.
I don't need to ask how much longer on the timer with the clock - it's there at a glance.
Unfortunately, Amazon has been making the devices (especially ones with a screen) into an advertising channel to the point I'm looking at replacing the various echo show devices with just echos (and a clock if they ever come back into stock).
On reading Wikipedia... Alexa used to have a knowledge engine somewhere in its code. You could ask it what color a black cat was and get back "a black cat is black." You could ask "what color is a light red flower" and get back "a pink flower is pink." Asking "what color is a blue bird" gets back "a blue bird is blue, brown, and white." That hinted at a deeper knowledge engine. There was also an inventor <-> invention knowledge base. One time I even had it return back part of the query language by asking it if two people (who were born on the same day) were born on the same day. The knowledge base functionality appears to have been delegated to "search Alexa answers".
It has gotten decidedly worse over the years as useful functionality that had no revenue associated with it got removed while "revenue enacting" features (pushing fire tv, product advertisements and such) have been prioritized.
Many of my echo shows are now "face down" because the screen cycle of stuff I do not care about is out of the corner of my eye distractingly fast. Time and weather are better served by my watch now.
It still does timers and reminders acceptably well.
Listening to it? What a waste of my time and focus on something trivially and already solved.
And use this during driving as some other mention? Sorry not a fan to say at least, it definitely impairs everybody who is driving to certain extent, humans simply don't have efficient parallel processing of things that require focus in this way.
The interface with smart devices is good but clunky. While I can say "Alexa open my curtains" and I can go to the slow and bloated Alexa app to open the curtains at a certain hour, I cannot say "Alexa open my curtains tomorrow at 6:45am".
I would love an app that would allow better automation, like doing an action at a certain time and playing an alarm X minutes later. However, skills are voice-only and nearly useless.
I can say "Alexa, set my alarm tomorrow at 7 in the morning". I wish I could say "Alexa, open my curtains at 6:45 in the morning tomorroe" instead of having to go to my cellphone and use the clunky Alexa app.
It just seems like a wasted opportunity for a routine. I'm sure a lot more things would be available if Amazon let developers use more features.
We're not asking to construct a complex schedule via voice command, the app is sufficient for that. Sometimes you just want your curtains opened at 6:45 tomorrow because something is happening and you just thought of it. That's would actually be convenient.
The "funny" thing is Google Home could do this (Turn on/off device after 8 hours) but recently they've removed the option & the assistant refuses. I assume it's for liability purposes, but it feels stupid.
Will the Google Home tell you the device names it can see?
Works for me; however I often say the time before the action, so something like: “Alexa, at 8pm turn on the living room lights”, but “Turn the heating off in 15 minutes” also works for me.
It then creates a timer to activate said device/group at the time requested. (Edit: I also use commands like "Alexa, turn on the lighthouse lamp until 11:30pm", which turns on that lamp and sets a timer to turn it off at 11:30pm)
However all my devices are using the native smart home stuff in Alexa and they are exposed to Alexa via home assistant (So dunno if HA is doing some magic sauce to expose the devices in a certain way that allows such voice commands), So your mileage may (and indeed appears to) vary.
Granted sometimes it mistakes my intentions and I need to repeat myself.
I’ve also been able to create routines via voice, I first discovered that when Alexa suggested to me it could create the automation for me after it perform a request.
Another Edit: When saying "At X do Y" Alexa can only do that action if its within the next 24 hours, guessing thats its timer limit. I also tested saying "Alexa, Every day at 6pm turn on living room lights" and it created a routine for that action, I was then able to disable that routine by saying "alexa delete the living room lights routine", however it just disabled it rather than deleting it like I had requested (and checking the transcript in the app it did pick up i said the word delete and not misheard me).
e.g. As a routine writer, you could include a command within a routine to turn on lights at 6am. However, it would be fixed at 6am whenever the routine runs, with no flexibility since there's no concept of variables or voice input available to routines. A more sophisticated platform would offer more flexible commands, i.e. "turn on lights at {voice_input}".
“Discoverability” here means the ability for the user to learn the systems abilities. Historically, you could read a menu or list of app icons.
A chat box is empty - it tells the user nothing about the systems abilities, just its interface. A smart speaker is even worse - there’s minimal UX hints. Asking a chat box “what can you do” is not likely to be exhaustive, and and will likely require a series of “can you do X” queries.
Ask people how they feel about Alexa's "follow up" suggestions ;)
It's not easy to do discoverability in general. Especially not through an intentionally limited modality. People study this stuff, many businesses and product researchers have spent years workshopping ideas. If it was a quick and easy idea like "just tell the user" then it wouldn't be a challenge in 2024.
The problem with "show messages... suggest followups" is that you can't teach the user about new features that are unrelated to their current interaction because it feels like advertising, and its distracting to their current task at hand.
I think that for this segment to really take off Amazon/Google/Apple all need to start releasing usable stuff to the open source world. A good open source home server voice based back-end that can control open-source devices will start creating apps and drive exploration the space. Eventually it will find the killer apps that will truly kick-start things. The houses lost to people setting up their own back-end is likely very small and the returns could be very big.
It probably won’t be as good, but it’s time to jump ship.
I’m increasingly infuriated with the Alex Echos. Specifically the “while you wait..” or “did you know…”. The only thing I trust the Echos to do is turn on/off devices, set timers, and set reminders (same as what I trust Siri for). I imagine with Atom you can probably hook up an LLM somewhere in the chain which would be awesome, especially if I can use a local model.
Another thing I want to banish from my life is “An item on your subscribe and save is not available”, that’s it, that’s the message. What item? Doesn’t matter. What about if you ask? Alexa has no clue what you are talking about and the app is useless for this specific purpose and in general. In fact the app is horrible to work with, it’s slow and buggy. I want to find the developers of both the app and this “feature” and just ask them why they hate users.
Alexa was neat for a while but I’ve replaced the music playing ability with Sonos which does a way better job and now I really wish Sonos had a HA Voice integration. I can stomach setting up these Atom boxes everywhere instead though.
https://www.home-assistant.io/voice_control/thirteen-usd-voi...
https://shop.m5stack.com/products/atom-echo-smart-speaker-de...
Sounds like they enshitified where they eat.
It shows me triumphant notifications (!) that turn out to just be a request for me to review a product. Yes, that's why I _bought_ this device, so the company I bought it from can use it as a portal to ask me to do free work for them.
It has failed to add any value other than as a device I can use to ensure that our Alexa skills are in fact working or not when I take bug reports out of our queue.
For the amount of counter space this thing swallows I constantly wonder why they couldn't have put a few USB charging ports somewhere on the side of it.
Speaking of USB and/or Bluetooth, why aren't there any "peripherals" for this device? It's just a useless screen connected to a cloud of corporate garbage that no one would ask for and it has zero ability to step outside of this sphere. So, yea, jam it full of "generative" content, how could you possibly make it worse?
Thing is durable, she couldn’t actually crack the casing but broke something inside so it stopped working.
(To be clear we do let our kid watch YouTube supervised after she’s finished lessons, although not very often)
I swapped it back to my old Echo Spot (the round screen one). It's great at being a clock.
Although sadly now discontinued, when mine breaks there doesn't seem to be any comparable replacement from any brand at any price.
They made a bunch of peripherals but they never caught on.
They made media remotes, game pads, a sticky note printer, scales, and other things.
Trapped in a 3 ton metal box for hours seems the perfect time to force feed ads, and yet… nothing
https://old.reddit.com/r/Justrolledintotheshop/comments/1c1g...
There used to be something like 3-4 checkboxes I turned off on my Alexa Show to disable stuff like that, but they gradually will break out stuff into it's own, separate category that is enabled by default and you have to turn off again. I went through my Show again last month and had to disable more than 20(!) new checkboxes to get it to stop showing me things like "New Prime Shows", "Popular Amazon Buys", "Trending Pop Culture News" and other crud like that.
Every few months, a new category enabled by default and a new checkbox to turn off.
This is every Amazon hardware product unfortunately. Kindles are the least bad, but they're not great. They sell really cheap hardware (TVs, Alexas, tablets), but you pay in data and advertisements. It's probably a choice many consumers are happy with.
I bough a FireTV (the entire TV uses it for the US, not just the stick) because it was incredibly cheap. $150 for a 42" 4K television, shipped to my house in two days. That's absurd. It has ads that like to autoplay on the home screen and its OS (I hate TVs have OSs now) runs like trash. I got what I paid for.
I'd pay a serious premium for a product like an Alexa, but with in house processing of voice. Maybe I heard that the Apple version did this, but I heard enough negative things about it to no care enough to look into it.
I've been able to avoid this by just hooking up a computer to a big screen TV with HDMI. That's it. I watch stuff via the computer. Slightly klunkier for some stuff but has worked well enough. Might not be the best interface for kids or folks not comfortable with a full keyboard/trackpad though.
That commoditizes the frontend, which is the part they would actually control. There's no lock-in if you can swap your Amazon Echos for Google Homes and everything still works.
It would drive interest in the space, but probably not money into their pocket.
I don't think there's any inherent advantage for Amazon/Google/Apple to make all the homes smart. It's only to their advantage if they're the ones selling the devices.
For example, I would ask Alexa "can my dog eat <X>". In the beginning it was not clear to me this was a third party skill. Then one day Alexa would respond with something like "nay, Dr. Wolf says ..." (I don't quite remember). I was surprised because who speaks like that? I then realized it was an Alexa Skill made by a third party developer. The developer later attempted to monetized this skill so you could only ask the question 3 times and after that it would bill you.
Alexa has a lot of other issues. Every room in the house has an Alexa Echo Show. But the screen real estate is entirely wasted with useless stuff. Very limited customization.
I'm in process of replacing my home automation with Lutron products. And the Alexa will soon be gone from my home. It seems like everything Amazon does is half-baked, but enough to get out the door. There is no polish at that company.
Please try to donate to a local hacker space, nephew, give away on craiglist, but this just makes me depressed every time. :-(
> It seems like everything Amazon does is half-baked, but enough to get out the door. There is no polish at that company.
Isn't that more and more becoming the norm? Google immediately comes to mind too, and the only exception is probably Apple, but at times even they throw shit on the market that just doesn't feel polished at all. It's just that more often than not they actually try to make it right then, instead of letting it linger...
I literally went back to pen and paper because of this endless closing off and injecting ads garbage.
shared, synced family grocery list is a must in our household now. yes, items must be added 'manually.' but they can simply be checked/unchecked rather than deleting (which also provides a convenient pantry inventory review just before shopping trips).
Amazon, if you really want to make money from Alexa, just charge me a couple bucks a month, it's worth it. It was never going to be a profit center.
Just imagine what a good AI assistant could do with todays LLM tech. Actual things I would have used in the last two days:
- Hey Google, what's the song that goes "da da da da da daaa da"? There are skeletons and elephants in the video?
- Hey Siri, can you move all the windows of project XY to a fresh desktop?
- Move all the shopping tabs to a new window.
- Find the last contract draft for customer ABC.
But for everything else they do suck.
LLMs changed everything.
Second, giving access to APIs... gives you access to APIs, and that's it. Even the most advanced LLMs are dumb as shit when it comes to anything non-trivial in programming. What makes you think that a modern LLM will manage to translate "move tab X to a different window" into a dozen or more API calls with proper structures filled in, and in the correct order?
I don't sense that GPT4 is "dumb as shit". I sense that it's extremely capable and very close to changing everything if, for example, macOS completely integrates GPT4.
Of course it's not trivial if you think for more than a second about it.
LLMs produce output in exactly three ways:
- text
- images
- video
What you think is trivial is to convert that output into an arbitrary function call for any arbitrary OS-level API. And that is _before_ we start thinking about things like "what do we do with incomplete LLM output" and "what to do when LLM hallucinates".
You can literally try and implement this yourself, today, to see how trivial it is. You already have access to tens of thousands OS APIs, so you can try and implement a very small subset of what you're thinking about.
BTW, if your answer is "but function calls", they are not function calls. They are structured JSON responses that you have to manually convert to actual function calls.
You could in theory send ASTs as JSON.
Please show me how you will do that for the tens of thousands of OS APIs and data structures.
Edit: because it's not just "a function call" is it? It's often:
struct1 = set up/get hold of a complex struct 1
struct 2 = set up/get hold of a complex struct 2
struct 3 = set up/get hold of a complex struct 3 using one or both of the previous ones
call some specific function 1
call some specific function 2 using some or all of the structs above
free the structs above, often in a specific order
So the question becomes how can this: { "function_name": "X", parameters: [...] }
be easily converted to all that?Don't forget about failure modes. Where you have to check for the validity of some but not all structs passed around.
Repeat that for any combination of any of the 10k+ APIs and structures
It's the trivial part of all this. You're getting too much into the details of what is available to developers today. Instead, you should focus on what Apple and Google would do internally to make things as easy for an LLM as possible.
Up until now, the hardest part was always understanding what exactly users want. GPT4 is a thousand times better than what Siri and Google Assistant are currently doing.
Again, mapping OS APIs is the easiest part. By far.
You keep skipping the question of how.
> It's the trivial part of all this. You're getting too much into the details of what is available to developers today.
Somehow you have this mystical magical idea of "oh, it's just this small insignificant little thing".
Literally this:
LLMs understand human requests
* magic *
Things happen in the OS
You, yes you already have basically the same access as developers of the OS have. You already have access to tens of thousands of OS APIs and to the LLMs.And yet we haven't seen a single implementation that does what you want.
> Again, mapping OS APIs is the easiest part. By far.
If it is, it would make it trivially easy how to do this trivial and easy task for a small subset of those APIs, wouldn't it? Can you show me how you would do it?
AI assistants only barely had enough data to understand your requests, and now have even less.
For instance, I recently asked her “How many cigarettes does the average jazz band smoke per night.” and instead of the prior usual “I’m sorry, I can’t answer that” she replied a very confident “15 cigarettes”.
I was honestly really surprised and had to check the log in the app to be sure she had heard what I had said correctly. She had indeed.
Another one I have asked her many times is “How many goats are in a goat boat?” (I really just ask endlessly insane crap every time I am in my kitchen, it’s some sort of addiction) and again where she used to reply the “I’m sorry” response, she now replies with the confidence of an LLM “2 goats”
Just an FYI this type of instruction has worked for about a decade, actually.
Very concise and to the point. I might print and hang this!
I've done consulting at F500 companies and was consistently not-impressed by director levels and above, dudes in tech for 30 years and didn't have an understanding of what "production" meant. Outsourcing literally their basic day to day job responsibilities to Accenture and McKinsey to the point where pretty much anyone reading HN could have been in their role.
A lot was "flash" -- looking good, speaking grandly, and sticking to broad approaches that they could assign to a technical senior manager or contributor. And a lot of the time, to be honest, that worked: sometimes you just gotta motivate people with big gestures. But once it became time to actually get outside of the box and generate new ideas they were stuffed shirts.
Meanwhile, AI is the only thing that could possibly make a voice assistant product decent. But it will require a complete change in approach and these "skills" made the old way will be irrelevant.
Going by the rumours the layoff axe fell the hardest on Alexa division. And now we see this news about 3rd party apps. Next Alexa will be put on life-support and then stop it completely 3-5 years from now. It's good that they do have bluetooth and audio cable interface so the hardware won't be bricked.
Either way they need to figure it out soon because the cost will keep accumulating and they can write them off only for so many quarters.
It seems they had sold 500 million Alexa devices until last year [1]. Just the hardware infra to keep servicing them is massive; not to speak of developer cost.
[1] https://finance.yahoo.com/news/amazon-has-sold-more-than-500...
It just starts to feel egregious.
Drives me up the wall that people pay monthly for an otherwise static service just to keep the hardware functioning.
2015-2017 was the golden age at least of Google assistant (or was it still Google Now?). Fast, responsive, and accurate.
Now in 2024 it's grindingly slow and unreliable.
LLMs require either more local power (=way more expensive devices) or more server power, which someone has to pay. Power users aren't paying, as they;ll get better results through their computer or phone to access the LLM of their choice. Casual users aren't paying either, as what they want to do is too low stakes to warrant higher costs.
Amazon will probably never completely kill Alexa, but I don't see them investing much more to advance the product in any significant ways.
I stopped trusting Amazon years ago when they banned my seller account and my business was shutdown overnight. After 15 years, I finally got my account back with no explanation. At this point, I no longer care and have moved on.
"Sad" feels a little strong here
Hell, let most's problems be that a developer incentive program didn't live up to be a lifetime career replacement and I doubt they'll consider it an upset at all.
Diversify. Never rely on a single large monopolist for your only revenue stream.
At this juncture, I have to be sceptical that the reason they now want to kill it is because they effectively want to Sherlock them all by shipping LLM based AI functionality instead, and they want to ensure that nobody else does it on their platform before they can.
For a very long time (probably more than a year) I have been seeing the same link at the bottom of every single Ars Technica article; it is a link to a video titled "SITREP: F-16 replacement search a signal of F-35 fail?"
Are other readers experiencing the same? If not, what could be the reason for targeting me? I am not involved in military tech or aviation, and I usually block tracking cookies so I don't expect the site to know much about who I am or what my interests are. I just find it odd that the footer of a publication like Ars Technica would remain unchanged for such a long time.
Edit: I see now it's a video, meaning it's really just a flashy headline that gets you to click and immediately watch though an ad. They probably make quite a pretty penny off of that.
I think it’s just that footer was designed to promote their video content and Ars doesn’t produce much video so that’s the most recent thing to put there.
But if it's all local? It would be a great product. No internet unless I turn on a physical switch, no ads, no monthly fee. Just do the thing I want you to do. Maybe at the moment it will still need to talk to my GPU. But we're getting there.
But yea this seems like a great vision generally. I think that we’ll progress past a chat LLM to one that autogenerates UIs - several companies have already demoed this.
It wasn't bad; I'm on a 3070 so it was slow, but tolerable (a couple seconds of latency up to 5 or so seconds for the full Whisper+Mistral7B pipeline).
I named mine Jeeves, so at any point I could just "Ask Jeeves" (lol) and it would talk back to me. The TTS spoke slower than I would prefer, but I probably could have fixed it. I also should have prepended the prompt with something to encourage it to be brief. It often started reading out a couple of paragraphs of text.
The only reason I didn’t get rid of it is, that it’s an easy way for my kids to play their music. Otherwise I would have ditched them all
My wife and I are NASCAR fans so I did have a skill where you could just ask "what's the weather at the track" and, based on today's date it would know where the next race was and give back weather for that location. But beyond the things mentioned here, other places have scaled back their free APIs for things like weather and other useful services.
What really killed the interest was the nagging feeling that we had that Alexa was listening more than it should, so we keep it off most of the time so you have to go to it, have it listen, ask your question then turn listening back off.
It's here now, they just haven't announced it. The Alexa devices at my home can answer arbitrary questions. "What is the round shape on top of the world trade center?" "That is the one world observatory" etc.
There's a weird thing where some questions get picked up by the non-LLM. "how much will it rain today?" gets answered by the old weather app with a nonsensical answer like "it will rain today at 4PM"
Skills were always dead in the water the minute they required an AWS account and a Lambda deployment to get going. That's a miserable thing to develop (and an AWS account itself being a looming financial liability) and maintain even if you're familiar with all the inner junctions of Lambda and IAM.
https://github.com/rhasspy/rhasspy
Was discussed last year: https://news.ycombinator.com/item?id=33705938
My only real issue in setting it up and running it has been finding appropriate microphones.
An API call to OpenAI whisper costs around 0.01ct per minute, so really cheap. Feeding the text-from-speech and the possible function calls with their JSON schema into the LLM costs another penny. Everything works.
But now, where's your business model?
On the first part, if a skill doesn't support a precise sentence structure, it becomes difficult to use or simply doesn't support features users expect to use using natural language.
On the second point, not all Alexa devices work with IPv6. I have an Echo Show 15 here that simply refuses to play BBC News and other "Flash Briefing" items when IPv6 is enabled. Really? It's 2024 now. Did the whole factory reset, reboot, and so forth dances.
An LLM startup with enough partner biz dev/sales talent could combine IoT API integrations from enough manufacturers to side-step the boring and limited Apple Siri, Google Hello, and Amazon Alexa oligopoly hegemonies and self-destructive platforms like IFTTT to execute on a compelling, competing replacement making good on all 4 points. However, I don't think giving away computing resources for free is necessarily a scalable business model but it's fair to charge a reasonable micropayment subscription to host and run your code and data.
lol ya, I thought the title was a joke, but apparently their best guess is that there were dozens of developers. This is like a punchline from XKCD.
lol, is this a burn?
> The news has left dozens of Alexa Skills developers wondering if they have a future with Alexa, especially as Amazon preps a generative AI and subscription-based version of Alexa. "Dozens" may sound like a dig at Alexa's ecosystem, but it's an estimation based on a podcast from Skills developers Mark Tucker and Allen Firstenberg, who, in a recent podcast, agreed that "dozens" of third-party devs were contemplating if it's still worthwhile to develop Alexa skills