Amazon to kill off local Alexa processing, all voice requests shipped to cloud
theregister.com
theregister.com
The option to do some on-device processing came on later devices and, as I understand it, wasn't even enabled by default. Furthermore, on-device processing would still send the parsed commands to the cloud.
The headline is vague, but it's misleading a lot of people into thinking that only now Amazon will start sending commands to the cloud. It's actually always been that way. I suspect the number of people who enabled on-device processing was very, very small.
I don't love Amazon, but I love ginned up outrage over tech the author never bothered to understand even less.
And 99% of Echo owners disagree with this. No one cares if a reporter mixed up "amazon is starting to spy on you tomorrow" vs "amazon has been spying on you since the first echo was launched". Only amazon would make an argument like "yeah but we've been doing this for years now and no one made a big deal about it..."
Thanks for saying something anyways. I wish more people would follow your example.
The Sauron's Eye of public sentiment can only pay attention to one thing at a time. If you justify today on the basis that it wasn't paying attention yesterday, you can rationalize anything.
The search engines are crap. There was a story some years ago where Amazon employees from an Eastern European country were actually listening to your Alexa and sent the relevant commands back to the device.
I'm not disappointed that this "non-event" is drawing attention to this comparison. Even if it's farfetched to dream that bringing privacy more to the forefront of the news zeitgeist will result in a shift of the status quo for our industry - heaven knows that if privacy stories don't get mindshare, the status quo could get far worse.
The direction of major operating systems neutering themselves in favour of deep service integration does not fill one with hope
Word of warning: businesses, possibly including your employer, may consider you to be a robot, or a terrorist, or tell you to install Windows, or tell you to use your phone, or tell you that you have no right to privacy.
I recently had to go through a background check for a prospective employer. The background check website wouldn't even load. The support agent told me that it's a Firefox problem and told me that I needed to open the website on Chrome or Edge or on my phone, and that the website is working "perfectly". Alas, it worked fine with Firefox on Linux, as long as the user-agent reported that it's actually Chrome and Windows.
Yeah the website is indeed working "perfectly": perfectly enough to block employment of people who care about privacy.
You gotta accept Microsoft's licenses and EULAs and crap like that. Plus it's still spyware. No thanks
https://developer.microsoft.com/en-us/windows/downloads/virt...
>Due to ongoing technical issues, as of October 23, 2024, downloads are temporarily unavailable.
They are morons.
Maybe someday ReactOS will be usable fully as a Windows VM OS.
And in many ways, I'd be fine with that if they'd be up-front and honest about it.
Alas, Microsoft certainly is not. Cloudflare is not either. Many of these services just sit there and pretend like they're loading without showing any sort of error indication. Much like a tarpit but it's malicious on the business side and with little to no recourse on the real human side.
Cloudflare is not malicious, but is between a rock and a hard place. By caring about privacy you are, indeed, looking more suspicious. In a perfectly anonymous world, reputation-based captchas couldn't work. It's OK if you think they shouldn't exist, but Cloudflare customers and most people like them.
Everyone else is not on some secret plan to destroy the Linux Desktop, they just don't test their websites on linux/firefox (because "nobody uses that"), which makes them unusable, which causes people to drift off linux/firefox.
I have no intention of returning to Windows.
- Things that should be system settings are instead apps (amphetamine, rectangle).
- There's no way to move focus around directionally between windows with a keyboard.
- "Open file" windows give you nowhere to paste a path (tip: cmd + shift + g summons a path prompt).
- Full screen windows now must be managed like they've just become an entire workspace, and the underlying app may or may not support the un-full-screen button.
- The error messages don't give you enough info to actually act on them (apparently "Docker" will damage my computer, and I should uninstall it, but it won't give me a path to the offending file, so I don't know how to install it, also this warning returns if I close it so it's just been hanging around for months.)
My strategy for maintaining sanity is to do as much as possible through zellij (a terminal multiplexer), that way I can use the same muscle memory on Linux as well. As for the rest, I just try to ignore it.
Apple is infuriating in rather different ways.
I feel that it just works.
And am glad.
I installed a copy of Windows 11 the other day for a new machine and it was INFURIATING.
In order to install without internet or an offline account, you MUST know a voodoo command and how to enter it. Used to be their dark pattern was at least on the screen, they’re out of their damn minds now.
Everyone has their breaking point with Microsoft, I hit mine and it’s been nothing but good for me.
As is, I'll be sticking with heavily tweaked Windows to work with my several HDDs full of old software, and avoid the Linux headaches of repos disappearing, deciding between Snap/Appimage/DEB and general incompatibility with office documents, industrial tooling and Adobe software.
I'll only use Linux where I'm paid to at work. Thanks to Linux Torvalds' terrible software distribution model, I've had to do black magic to work around Anon's deprecation of Debian/Raspbian Stretch on which our industrial network gateways run.
What interface do you use for using local whisper?
The popup on my MacBook says specifically:
> When you use Dictation, your device will indicate in Keyboard Settings if your audio and transcripts are processed on your device and not sent to Apple servers. Otherwise, the things you dictate are sent to and processed on the server, but will not be stored unless you opt in to Improve Siri and Dictation.
I have a reasonably recent (<2yo) MBP, which feels like it ought to be "modern enough", and in Keyboard Settings, it says "Dictation sends information like your voice input, contacts, and location to Apple to process your requests." It doesn't say anything about processing happening on my device. Yes, off-line dictation does work for me (with Wi-Fi turned off), but I'm curious under what conditions Keyboard Settings would say something about transcripts not being sent to Apple.
What does this even mean?
Well, that’s impossible though. Of course it is stored as temporarily as that may be.
So the statement already has a logical failure. Possession for processing is still storage.
So the debate is then “how long do we store something until we have to call it “storage””.
And further, they may delete your voice, but the log of the thing you asked for I think is unlikely to go away.
I like Apple as at least they try and place the privacy field. But I think this wording is a bit weasely.
And even if it were, the meaning is obvious that it won't be stored after processing.
There's no logical failure, it's not "weasely". It's a completely straightforward legal statement. The intent is abundantly clear.
You were making best case assumptions, which is a laughable concept when it comes to a fortune 500 company.
The options are different for me in the settings app, and using dictation is for me right now, impossible without agreeing to a modal displaying an agreement to send audio data to Apple.
Are you sure this isn't the local model download prompt? The first time you use it, it does need to download some content to make that work.
[Enable] Dictation Privacy (a policy) [Cancel] ---
I also highly doubt that they will somehow later down the line magically change their minds given the millions (maybe over a billion at this point) of dollars invested in things like private cloud compute (1) and challenging the U.K. government (2) in court over E2E cloud backups.
(1) https://security.apple.com/blog/private-cloud-compute/
(2) https://www.reuters.com/technology/court-hearing-reported-be...
--- Do you want to enable Dictation? When you dictate text, information like your voice input and contact names are sent to Apple to help your Mac recognise what you’re saying.
[Enable] Dictation Privacy (a policy) [Cancel] ---
I think it is entirely fair to put Apple in this category, because they have effectively disabled offline dictation for me except if I agree to sent my voice data to Apple. I have used this offline dictation feature for years by the way
This kind of makes sense, at least to me: local processing will always be limited. The entire premise of the original Echo devices was that all the magic happened in the cloud. It seems like not much has really changed?
The local processing has less latency and works on unstable internet. It's perfect for tasks like 'set an alarm for 8am', even if offline.
The remote processing is good for better accuracy of complex words and queries.
The results are combined in the UI, making the whole thing feel less leggy (although IMO it still feels laggy to have to wait 1-2 seconds after asking a query to get results).
Updating Echo Dot V1 to newer kernel: https://andrerh.gitlab.io/echoroot/
Echo Dot V2 Android tinkering, https://github.com/echohacking/wiki/wiki/Echo-Dot-v2 & https://andygoetz.org/tags/dot/ & https://www.youtube.com/watch?v=H0IEMVDebzE
Unless of course this is hyperbole, and there is in fact every reason to doubt this because it's based on conjecture and anecdote.
Be advised it's not instant.
Why? It sounds like it was really interesting and valuable to observe those patterns.
I disagree about things being less consistent. Let's imagine a 100% LLM world - in this world, you use a bunch of training to try to get the LLM to match your hardcoded responses for common inputs. If you get your training really right, you get 100% accuracy for these inputs. In this world, no one is complaining about consistency! So why not just hardcode that behavior?
LLMs are actually the only reason I'd consider processing voice in the cloud to be a good idea. Alas, knowing how the aforementioned companies designed their assistants in the past, I'm certain they'll find a way to degrade the experience and strip most of the benefits of having LLMs in the loop. After all, as past experience shows, you can't have an assistant letting you operate commercial products and services without speaking the brand names out loud. That's unthinkable.
But now, in the age of LLMs, there are no "simple rules". An example. Say you live in San Francisco:
"Alexa, I'm thinking of going to New York, how many flights are there each day?"
This is a hard question, and one that will go to the cloud.
"Alexa, what's the weather"
This seems like an easy question. But with local only processing, you'd get the weather in San Francisco. But with the LLM, it will probably give you the weather in New York, which is most likely what you wanted if you asked these things just a few seconds apart.
[0]https://www.macworld.com/article/678307/how-to-use-siri-offl...
Echo Plus is a Zigbee hub with US-origin firmware.
Also ignoring the dark patterns in which big tech will make it annoying/obfuscated/unknown that better privacy options are available.
I think it's likely that they looked at the numbers and realized they were spending a lot of money putting NPUs on devices and maintaining separate voice parsing models for a very small minority of users.
Is Alexa hearing a gunshot a request for assistance? It’s not a voice command, okay, but where does “oh that’s not our business” really end in vast data collection platforms such as this? Does Alexa have any duty to report voice requests about self-harm?
In the age of LLMs, even a simple request like "set a timer" is sent to the cloud so that it can be processed in the context of what you've said previously, what devices you own, what time of day it is, etc. etc.
And FWIW, you will get a better voice to text in the cloud because that model will know about your device names and other details. For example, if you say, "turn on the kitchen light", the cloud knows you have a light called "kitchen", so if you slur a bit it can still figure it out.
It cannot. Keep in mind that the Alexa devices are built to be as cheap as possible, so they have minimum amounts of RAM and CPU. The tiniest of models can barely fit on the device.
> The Echo speaker is likely the hub for the IoT device controlling the kitchen light.
Generally the devices that are controlled by Alexa are done via Wifi from the provider of the device using their own APIs. Very few "Works with Alexa" devices can be controlled locally. But yes, some of them can. However, the Alexa device doesn't know it is called "kitchen".
> None of the context you speak of requires "cloud".
I just gave you a simple example. Here is a better one that I used down below:
Say you live in San Francisco:
"Alexa, I'm thinking of going to New York, how many flights are there each day?"
This is a hard question, and one that will go to the cloud.
"Alexa, what's the weather"
This seems like an easy question. But with local only processing, you'd get the weather in San Francisco. But with the LLM, it will probably give you the weather in New York, which is most likely what you wanted if you asked these things just a few seconds apart.
Every point you made is a choice that was engineered into the system to impose dependence.
A far cry from your claim of "it cannot be done". And the $40 retail sticker for that board doesn't seem to match up with your claim about pricing.
Even low cost modern SoCs have NPUs in the double-digit TOPs, and memory densities being what they are, there's very little excuse not to run a special purpose language processing model on device. A 128gbit (8GB) memory module goes for as little as $0.04 ea on digikey: https://www.digikey.com/en/products/filter/memory/774?s=N4Ig...
Lots of discussions: https://news.ycombinator.com/item?id=43365424
Even Apple did open their AI system to OpenAI (and potentially other vendors in the future).
As long as there is explicit consent to do so, it is fine. Nobody forces you to buy an Alexa product.
So ignorant, especially with music. Whoops, no Internet access on the plane? No music for you.