Amazon Echo Flex: Microphone mute button appears to be real and functional
electronupdate.blogspot.com
electronupdate.blogspot.com
> But, basically if one presses the switch it causes the flip-flops to change state and to then either add or remove power from the microphones.
> The 'mute' button appears to be very real and functional. When the button glows 'red' the power is removed from the microphones.
Or would it be better to reduce the need for comments like this one by putting the conclusion much earlier in the piece? I understand not putting the conclusion in the title, but perhaps in the first paragraph, in the spirit of BLUF [0]?
It's a similar concept to the 'Topic Sentence' which would be used to sum up the ideas of the following paragraph.
The only reason I could see for actively choosing not to use BLUF - would be to force users to read more of your article for the sake of it, which seems (to me) like a 'cheap' tactic.
I personally think for informative reporting, it's maximally respectful of the reader's time and attention to give a quick summary up front. Then they can decide if they want to continue for the full details and story-telling experience. Many people won't, but those people would probably skip the story in the first place and click through to the comments for the punch line.
Thanks for bringing up this discussion!
Japanese people can learn to use BLUF, but it's very much not the culturally expected approach.
How does that fit with the culturally expected approach in Japan? If you're sharing a revelation, is it still normal to build up rather than start with the answer?
YOU are not trying to make a point. If a point is there, your role is only to help to reveal it.
In other words, it should be in balance and objective.
On a tangent, I find it quite refreshing in film when they start with the punchline. You know what happens in the end before you even start; if it's a good film, it's still worth watching, even knowing the end.
Why?
There is a reason academia uses abstracts, and it's not because the attention spans of the readers are unusually short.
Especially if it's an article and not a video when I could easily skim to the end to read the conclusion if I don't want to read the story. I feel disrespectful towards the author if I want the conclusion up front.
But point taken and I've changed it now. Thanks!
It sounds like you could press mute, walk away and it could reset unmuting itself in the process without you realizing. You would need to constantly monitor that LED for assurance it is still muted. It’s not clear to me from the article how the SOC is connected to the flipflop reset. i.e. can the flipflop reset be initiated via software?
A mechanical switch that can’t so easily and unexpectedly lose its state would be preferable. There are probably good reasons they didn’t do that (aesthetics, cost, reliability?)
To be clear, I am not saying we should trust Amazon & Co, I'm saying that if that's the level of trust we have with home assistance devices, why even bother with the stuff ?
What we need a completely independent organisation like EFF to audit and certify basic guarantees around privacy/safety. The companies should also prefer this route as it would give take some heat off in EU/Congress without creating bureaucratic bottlenecks.
One problem I see that such an organisation can act as front for the cartel to prevent entry of new startups/players.
I'm pretty sure we already have laws to ensure honest advertising. The problem is that it's hard to tell the difference between malice and incompetence.
I assume it's like reading the side effects of medication and think they are other people's problems somehow. How do people trade core privacy for setting the alarm or getting the weather report by voice command. I don't respect them. Same with off the shelf IP cameras in one's bedroom. Hello?
Air gapped speech recognition and FOSS, or this should never, ever become a thing.
Of course they both have microphones hooked to data centers, but the Alexas/Echos are almost worthless without the microphone enabled. Many people I know have turned off listening for trigger words on their phone.
That being said, I have a HomePod that I have to manually press the touch button on top for it to listen. I use it for music, an occasional family conference call and occasional for generic assistant stuff like weather.
I tried a Google Home but couldn't get it configured in that mode -- listen on physical touch only (that was 3-4 years ago). And doubly couldn't get it setup without deep integration into a google account. The HomePod complains occasionally that it can't answer because it hasn't been personalized to an account, but is otherwise totally fine operating without access to my every inner thought (web/location history, etc.).
I'm just as concerned about every modern car essentially being a microphone hooked to a datacenter on wheels. And new TVs. And a few refrigerators and kids toys. And light switches. At least phones (well iPhones -- haven't used the last few Android versions) are easy to setup without listening enabled. And on iPhones it's opt-in during the setup process (the Siri checkbox is default on, but you are prompted to decide with no other confusing decisions on the screen at the same time).
Yes.
What's the difference? You trust it not to listen without the button press, I trust it to not listen without the wake word. To think that the distinction is meaningful is wishful thinking.
The only way to be sure these devices aren't listening to you is to not have one at all. Or, as the person in this article did, take it apart and de-cap the chips to make sure.
In the case of smart phones, I think it's a mistake to assume that switching some of these functions off in the settings will actually have the desired effect: preventing the phone from recording audio without making the owner aware of this recording.
Back on topic, it's not a phones intended purpose to snitch on you, they usually do not listen into the room, without extremely malicious manipulation, and more so you simply can not realistically avoid having a phone in your reach, today. If I were talking business I would probably leave the phone outside too, for that matter, and I don't trust cheap communication hardware, either.
In real life drawing analogies by principle doesn't make for a good argument, most of the time, as life is adaptation and compromise, not a path through a logical circuit.
Do you honestly not see an obvious difference between a phone and an Amazon wired device intended to record and remotely analyze what's happening in the room?
I sure don't, from a security viewpoint. They are both cloud-connected microphones. Google is just as likely as Amazon to listen in surreptitiously. And my phone is 1000x more vulnerable to someone _other than_ Google listening in due to installable apps.
Do you honestly not see that there's no fundamental difference between the two? A CPU, ram, storage, wireless modems, microphones and a battery all running under software we don't control.
As far as intent, it's not possible to be sure about intent, you have to judge based on the information we have which is the features and capabilities of the device.
I can't really get away without having a cell phone. The slope has been slippery on this one; when I purchased my first smart phone this wasn't a big concern (but maybe it should have been). It's been several years and it's a bigger sacrifice, in my opinion, to give up a smart phone. Even if I wanted to, my employer might buy one for me, for instance.
I do think the expectation today is that everyone have a smart phone.
But I have Amazon smart speakers in every room (on their own network that can see the network my smart lights are on.) They're great. I don't worry about Amazon listening in.
It's routine that I will just leave an item and forget where it is several times a day that I take it almost for granted up until it causes me to be late going out the door. Now I can yell "Find my phone" (or wallet, keys, etc) and my phone is there. Yes I could go to my computer, go to the find my phone page, but there's just more friction and these issues happen to me all the time. The more things I link up to voice control, the more options I have in regards to accomplish a task, and sometimes voice is the most efficient means of going about it.
I can get entertainment like a podcast or music without a single screen being involved, which given that I inherently have low impulse control, is a nice separation of things. Or I can again, find my stuff without looking at any screens. Voice input can act as a lesser evil to screens.
Saying to just use FOSS and air gaps, because of how immature such solutions are, is effectively saying you shouldn't use these systems at all. The benefits of using the systems that exist today makes it so that the average person, like yourself, judge me less because I appear more punctual and less forgetful and more productive and better rested and so on. Being human, you are likely to disrespect me for being late and disorganized no matter what my excuse is, whereas you'll only disrespect me for using voice assistants if you know I'm using them, so it's a winning tradeoff. Just like how some will judge me for taking medication but only if they know I'm using it.
I’ve dealt with many people from Amazon in my professional life and based on their handling of customer data I believe Amazon take seriously the trust customers place with them. Amazon believe the Alexa and Echo platform will grow their retail and service relationships with customers.
I don’t think they would throw that away to fool us all into believing a fake mute button - for what?
On the other hand, it's important to verify claims, and Amazon should be very happy about this verification. It proves that their claim could be trusted. Some other conclusion could have just as easily been true.
The fact that they used a hardware switch means they did it the right way.
By contrast my Logitech camera has a light to show when the camera is running - but the light can be disabled via their software which makes me uncomfortable with fully trusting the light.
That they went ahead and implemented in hardware is encouraging. Someone over there took the feature seriously enough to take it this far. I am sure it added cost over a software-only solution.
No one I know would trust a software mute.
#doppler4lyfe
Still proud of this product!
Not that I think you made the wrong choice, I think it's a good one, but I'm just curious how many consumers you thought would actually understand it, or how that information would be distributed
As an example, we were very clear and upfront in the original product detail page that users' voices went to the cloud, as we were not trying to hide that. Even so, some people tried to play a bit of "gotcha" and fear-mongering about the fact, as though we were not being transparent.
Particularly with software auto-updates, a software mute is simply untrustworthy.
True, but what is the incentive for Amazon to behave given their dominant market position? Customers might not have viable alternatives.
You really believe this? Still?
Amazon is so big that the loss of outraged users is a rounding error to them. Ford used to let cars off the line missing bolts. They're still here because their pockets are deep enough to wait until people forget.
Of course, technically, not even this hardwired circuit is complete proof that there is no trickery going on. If Amazon really wanted, they could simply hide a second mic at some other place inside the device.
I think the only solution that would really solve this was public auditing of the whole device.
It's not a problem without prior art though, and other industries have had to find solutions such as certifications, regulatory bodies etc. Think of medicine, food, vehicle saftey standards and the like. I'm not an electrical engineer but I can trust my fridge to draw what it says it will, because I know it's been through the regulatory processes to get to market and be stickered with the 4 star energy rating.
I think at this point we might need a privacy and ethics regulator for technology, with a certification program.
The real problem is that we don't have something analogous for microphones. The best we could do is some kind of humming device that you could turn on near the sensor that would drown out any noises it might otherwise pick up (and also be unobtrusive).
It is definitely not low cost.
In the hardware business, we count pennies. We count fractions of pennies with the costs of various components.
A shutter means at least one extra mechanical part, and a more complicated housing and increased assembly time at the factory.
You might look at a piece of plastic and think it is only a few pennies of plastic. But by the time you get it integrated into the product, and shipped out to consumers, it can add up to a dollar of cost. That really eats into the profits of low-margin hardware.
Only if enough consumers care about hardware shutters does it make business sense to include it. If most don't, and it is not a part of their buying decision making, then you're just throwing away money. Especially if it leads to increased post-sales tech support costs: ("My camera doesn't work! Is the shutter closed? Oh it is!").
Note that this is not an argument not to have hardware shutters. They are a good thing. But they come at a significant (to the final price) cost.
In any case, I don't think that's relevant to my point: I specifically said "important for laptops to permit a physical slider to block the integrated camera". That phrasing was deliberate: you don't want to trust the OEM's slider in the first place, just be able to add you own.
So, am I addressing a strawman? What OEM impedes you from adding a slider? Well, actually, Apple says it's not recommended to use a webcam cover because of how they designed it. That is poor design and security practice.
Isn't it also the most obvious thing?
To be confident something is true, you need to confirm it. How else would the universe and human perception work?
Still, this is not how humans usually interact. I trust there is no poison in my food without analysing it, I trust people to no shot me as I walk in the street, I trust the shopkeeper to accept that I have paid if the card reader says so… As ethno points out [1], some social work is needed here for our collective trust no to be misguided. Crucially, if the social work to buy trust does no exist, it entirely falls on you.
I consider myself that trusting the Echo (or other home assistance devices) is too much work for the meager benefit it provides.
In my opinion, it's only in the case of electronic and software products where we are supposed to take so much functioning of the device on faith. Perhaps this is a bad idea and we need an organization to vet these products as well.
Jeff (and others) have spoken publicly about this in the past, and this teardown is correct. The hardware is designed such that, if the privacy LED is on, the microphone is unpowered. In order to compromise that, one would need to physically alter/damage the hardware.
Yes the light will turn off, but who constantly monitors the lights on their devices?
Because yes, this can be used to justify close to anything that is in the interest of the company even if it's to the customer's disadvantage. Those features usually bring money to a company so removing them would be a loss. On top of that you can always attach that generic claim that any change will cause the system to implode and will also make your cat feral.
Well, some of that may be true but it doesn't mean we should use it as a justification to keep doing things like that. Products can be built with the consumer's best interest in mind at little additional cost if any. The problem is the consumer's best interest usually brings in less money.
I love how people on HN can rationalize every imaginable bad thing, and figure out reasons why the good thing is really, truly impossible. Qwerty is superior to any alternative, healthcare in the US is completely rational, we can't build cheap tunnels but at least we build the best tunnels, English spelling is a feature not a bug, as are Imperial units by the way, and college education is optimally priced.
That said, I think companies also have to consider the saleability of privacy features. How many Echos would you, personally, buy if they had a mechanical switch? How many people do you think would make the same decision? How does this stack up against the change in failure rates from a physical switch?
Again, you're completely correct. It's cheaper to design in privacy features up front. I can see some wrinkles that make the question more subtle, though.
Looking at this thread, there is no amount of money nor action they could take that would convince people here to buy an echo, so honestly why bother?
You could use this logic to shun literally any improvement to the device not done in software.
Amazon's Software Engineer level progression is:
SDE1->SDE2->SDE3->PE->SPE (Senior Principal, Director Level)->DE (Distinguished Engineer, VP Level).
He even has the jeff@ email address.
Which we already knew.
The problem is that many people don't trust the good faith of the company you work for, and by extension they won't trust you either.
Suppose for example that Amazon wants to know how many people talk about some topic or brand, and wants to sell that information, but they think that the mute switch will get in the way of collecting the data.
All it would take is a relatively small sample - maybe 1/1000 Alexas randomly built in such a way that the mic can't be muted - and they would be able to sell these aggregated statistics with reasonable confidence intervals.
Not that I think this is happening, but it is a possible reason I might not trust a single teardown, and it doesn't require worrying about a targeted attack against me personally.
To compromise it, one just needs control of the SOC (which Amazon has) to turn the light off and the mics back on.
A physical switch would have been a better choice, if it was actually for security instead of security theatre.
A compromise would be to play a message after turning the device on, requiring user to take action and unmute it or leave muted. This shouldn't be a problem in most scenarios as Echo is meant to be powered once and run for a long time, so it would be just a minor inconvenience.
It does this already
Doesn’t even need to be at Amazon’s factory in China of where ever. But that is largely feasible too.
The difficulty is more about data exfiltration. Has to be appended to some real ‘normal’ request the next time a legitimate use of the mic is made, then piggybacks uploads on that connection.
Recording all times when a voice is heard would also massively increase the payload sent to the server. If you were just passively network monitoring it could stick out. Which is not something nation states want.
Forget these smart home mics, if I was a nation state I’d just stick with the mobile phones and laptops every person uses.
People are already paranoid about their Alexas inherently (and far more likely to potentially care to look) making them a poor target for hacking. It’s far more useful to turn the thing they carry around with them 24/7 into an active microphone with GPS, and plenty of more exfiltration possibilities with any mobile and desktop OS network traffic.
First, this generally means that when depowered, the output signal line is connected to ground by the ESD protection path.
Second, the actual mechanism which picks up the audio is generally variably capacitive-- a few picofarads-- in MEMS mics, and requires a pretty high bias (~15-200V) voltage to induce a usable signal for the in-package amplifier to be able to pick up the audio.
To not only have the in-package amplifier turned off, but also no significant bias voltage across the microphone electrodes, and somehow pick this up on the other side of the trace would be somewhat miraculous.
Otherwise, all this hardware protection is moot (mute? ;), because you can listen for a few seconds, and then turn the mute light back on.
Or you can even toggle them back and forth really quick, and turn the light half on and keep the microphone's power pin capacitor charged.
Is there any reason to believe that there's no processing or recording going on?
I have wireless headphones that work like this and it's a terrible design. If they go out of range temporarily, they unmute. It makes the 'hardware' mute functionality absolutely useless since it cannot be relied upon to stay muted.
This is a long solved problem, a lever switch. Laptops (usually enterprise laptops) sometimes have them to turn on/off wireless communication.
https://ux.stackexchange.com/questions/34956/who-needs-an-ex...
> The SOC (system on chip) controller seems to be connected to the flip flop resets
> The 'mute' button appears to be very real and functional. When the button glows 'red' the power is removed from the microphones.
Does the microphone take any time to boot up? How far could a malicious SoC get by toggling "reset" on and off very fast? If you toggle the microphone on and off at 96kHz with e.g. a 1% duty cycle, then you'd be able to sample the level from the microphone at 96kHz, and the LED would still be glowing at 99% of its usual current (which would be visually indistinguishable from 100%). This would allow the SoC to record audio at full quality, and still leave the LED glowing.
..and it says 10 ms max turn on time, so 100 Hz.
Of course, a typical or minimum turn on time could be better, but I imagine it it were better, it would be advertised as well.
Also nice to see that the led is a real one, even if the micro does have the ability to override the mute switch... I guess the 80's scifi movies with red glowing eyes to indicate the 'evil' subroutene was activated got it pretty close!
His videos are definitely not scripted (and barely edited), which can be off putting to some (hearing “umm”s among other things), and he doesn’t go into the detail that Ken Shirriff (@kens) does, but they’re interesting nonetheless.
Part 1 is not "ICFJ", it's "CFJ". The line is the pin 1 marking. This one is the harder of the two; however, if you've been doing this a while, you'd recognize this SOT23-5 as coming from TI, whose online marking code lookup returns "SN74LVC1G14DCK" in just a few seconds. (Incidentally, that means the actual marking is "CF" with a lot/date code marker of "J"; the random-looking lines above and below the text are also a lot or date mark.) No acid needed!
Part 2 isn't marked "COOR", it's "C00" with lot/date code "R". That's decodeable by eye as a TI LVC2G00 dual NAND gate.
The third part, the one marked "JW", is probably an ON Semiconductor dual transistor, but those guys are always harder to track down. The fourth one, the SOT-323 (?) marked "39R"... that one will be hardest of all.
If the transistor is a PMOS like the author says, when the output of the NAND goes low wouldn't both the LED and the transistor turn on?
The way the LED is drawn it should turn on when the NAND output is sinking current, right? Maybe the LED is drawn upside down? Or am I misreading/misunderstanding something really simple?
Yup, and even if the same pin drives both the LED and the microphone power switch, there's various attacks (e.g. high frequency PWM 75% duty cycle-- enough to light the LED to nearly full brightness, but also probably enough to keep the mic bypass capacitor nearly fully charged).
Since the SOC can only turn off mute, and not turn it back on, any attempt to eavesdrop would leave evidence (a mute LED left off).
I'd expect the light to be green when I'm safe from spying.
The red light should flash or pulsate when recording, in fact. Otherwise, people without good color vision might miss the cue.
Otherwise the LED could fail and you wouldn't be able to tell if the mic was on.
As is, you assume the mic is on if the LED fails, which is preferable from a privacy point of view.
This is especially jarring as the rest of Alexa LEDs are all blue.
From an end user mental model perspective, I think Echos are closer to that than to recording studio equipment. “Normally, to use this device, I speak to it and it warns me when that function is disabled.”
Even the most basic form of digital audio stored as PCM (redbook, etc) is actually nothing but a file where every data-point in the file is an integer representing the 'magnitude' of a disturbance during that short milli/microsecond of time. Any vibrating 'disturbance' is the very definition of "sound" itself too.
What I'm getting at is: how do you capture those signals without the use of obvious silicon devices (so no MEMS or opamps or whatever)? I'd imagine you have to hide it in software somehow. Then how do you acquire that signal? A simple analog input pin to the ADC of a cheap microcontroller is not good enough.
To pick up a better signal you need either a large flat or concave object (like even a computer case), or else something small and flimsy. You may know that even shining a laser pointer at the window of a building can pick up sound from a mile away simply by using a telescope to capture video of it and then let video processing detect the miniscule jostle in position due to the vibrating window. That trick has been used for decades, and works well.
But yes recording stuff secretly from an electrical device means you need either computer microcode or software secretly onboard that is recording signal variations that emerge from these physical vibrations, but there is almost no way, even after full disassembly, that even an electronics expert can necessarily detect this setup.
How do you arrive at the conclusion that microcode or software cannot be detected even by experts?
I was just saying it's a practical impossibility for even an electronics expert and even after disassembling something, to detect if a device had been recording or not, just by examining the device physically.
Once you start reverse engineering the code itself you're likely to find any hidden surveillance, but even then not guaranteed, because sufficiently obfuscated code can camouflage what it's true purpose was.
I think we should appreciate that while this isn't the best or most secure design (I'm not that happy with the host controlling the reset), it at least has an LED that shows the microphone status, which can't be controlled by software.
I think this is the very bare minimum we should require from all our digital devices.
Could the feedback from the transistor be used to debounce via the data line as well?
Measuring the current of the microphone would only tell us that at a specific moment, the microphone is off and the LED is red. But that might be a software switch, in which case nothing would prevent Amazon to turn the microphone from time to time while keeping the LED red.
Or, if possible, dont connect to a network at all.
sed -i 's/impossible/inconvenient/g'
The power of convenience is probably also the reason to place a megacorp microphone in your home in the first place.If I want to do what generally amounts to a google search it's generally easier to pull out my phone than walk over to the device, flip a switch, and speak.
Old Logitech Pro 9000 is the exact opposite. You can control the LED independently of the camera state.
It could have been wired to indicate recording for a minimum of a second after each image capture. But it wasn't.
TLDR, it appears that mute function can turn itself on by itself.
However, it is clickbait, by definition, but it’s not clickbait. If it was, I’d expect: “You’ll never guess what your Echo Flex does when you do this one thing!”
They don't owe us some nice title, or some convenient summary so we don't read their post...
"Otherwise please use the original title, unless it is misleading or linkbait; don't editorialize."
Use the original titles.
This has come up before, and titles have been rewritten accordingly.
Me? Title like that, I just skip to the comments, there's always someone at the top to sort me out.
If it was not a real mute, it would not be a question in the title but a statement such as "Amazon Echo Flex Teardown: The Microphone Mute is Fake".
Looks like when you press the mute button, you both light the led up and pull-up the transistor gate (It looks like PNP so, it turns off when you apply power to the gate). This effectively cuts off power from the mic itself.
So, the LED is not controlled by firmware but, by the behavior of the circuit.
AFAIK, OTA updates cannot re-wire PCBs or do we have nanobots now?
[0]: https://1.bp.blogspot.com/-vVhoXpEtTXo/X_0iLmz9gcI/AAAAAAAAB...
Not as good as a physical switch I guess, but better than some implementations of laptop webcam activity lights at least...
That's literally what you're looking at.
The schmitt-trigger and the DFFs debounce the switch in hardware, and then the FET that controls the power input to the microphone also controls the power input to the LED. What do you want, a relay?
Firmware can only reset the DFFs (which does turn on the microphone), but it can't change the physically coupled LED / mic-power behavior.
"Even better, there is no microphone", is about the only "even better" scenario than the signal lines being broken. A powered down active component is only "almost as good" at best.
I'd have a decent chance at recovering the audio in the former case, but nil in the latter..
For instance, if it's gated with a FET, and then hooked up to a SOC ADC pin that's multiplexed with other functions, one could use those other functions to provide biasing opening the FET.
If it's gated with something better-- e.g. an actual relay-- the audio waveform may still be recoverable: it's very, very large in magnitude compared to how it is with the device unpowered, and will be consuming variable amounts of power with the audio waveform. You might be able to recover the signal by measuring power nets AC coupled with lots of gain.
That MEMS microphone is producing analog output, is it not? I did a cursory search for the printed part numbers but came up empty handed.
Sure, cutting both would be even better.
But-- as I posted earlier-- recovering a signal from a depowered MEMS device would be somewhat miraculous. https://news.ycombinator.com/item?id=25743822
Surely the most valuable data is just after the mute is turned on?
It's gonna be pretty damn fast. Standard ceramic decoupling cap-- very likely 0.1uF-- right next to it. 1/2 C V^2 at 3.3V is 0.5 microjoules. If the entire energy in the capacitor can go to effectively power the microphone (it can't), and the typical power consumption of a MEMS microphone is about 1 milliwatt, it might keep working for about 500 microseconds.