Living with and building for the Amazon Echo
medium.com
medium.com
I build my own version of this, which I can customize as I want. I can get the weather (my central home automation server pulls forecasts from wunderground.com, current temp from my own sensors, the voice control unit then pulls data from there), open/close curtains and turn lights on/off, and it says the time. I have a small list of other 'dialogues' (as they're called in my system) I'm going to add when I have some time, but I'm still figuring out what functionality is worthwhile.
My software is based on pocketsphinx on raspi, so it's easy to put one in each room (I have only one now myself). I'm using a mxl ac404 teleconf mic, which works ok; you do have to speak up for it to pick up the commands. I'd love to get an echo to see how much better it works. I have a primitive 'tts' system that plays back prerecorded messages, and falls back to festival for unknown words. I paid a voice actor a bit through fiverr to get the 100 or so words I need. Sounds better than full synthetic tts systems, although I need to work on improving the timing between words.
This is only a few days worth of work, too. It's not hard to make. I'm not claiming my systems works as well as an echo, but I prefer the control mine gives me.
I worked on speech parsing software at Audible, where much of the data was originally gathered to make Alexa and the Echo possible.
They are not telling the truth when they say that they only send clips of the actual commands.
EDIT:
This is not to say that they constantly stream audio data, either. But they send much more than just the voice extractions of the commands themselves. They have to in order to build a profile of the users' voices, habits, etc., which aid in the quick processing of incoming speech data.
The service is not selective enough to only pick up on voice information proceeded by a correct utterance of "Alexa." Amazon is "customer-obsessed" - and one of the product execs asked on a phone call:
> I am an Alexa user, and call out her name but stutter slightly. I still intended to say "Alexa," so why shouldn't she respond to me? That's a bad user experience. Customers want to be understood, not ignored.
This is paraphrased, but the consequences of this question were enormous. It basically ensured that Alexa users would not ever have privacy again.
So, I'll try another pass:
Dear AndrewUnmuted: would you kindly elaborate on what else, other than the audio that follows the wake word, is transmitted to Amazon that ensures "[Echo] users would not ever have privacy again."
I think we can all stipulate to the fact that there is an audio transmission beginning after the wake word and ending when we see the spinning light pattern that indicates Echo is waiting for a server-side response.
You can confirm this yourself: if you visit this page at Google, you can hear every Google Now voice query you've made, and verify it includes audio before & including the "OK Google" trigger phrase:
If you've ever worked for Amazon, you'd know how dangerous it would be to go into the level of specificity that you are demanding. I refuse to satisfy some loud-mouthed internet user just because of the "seriousness" of my claim.
It would be extremely easy for you to test my claims yourself.
I, as well, do not know you, so I'm unable to accept a drive-by accusation that having an Echo means I'll "not ever have privacy again", a bold claim indeed. Lots of people shout out things that seem plausible, even flash a bit of confidence-inspiring bravado, and have no intention of actually helping clear the air on a topic like this.
I stand corrected on one thing; I do see that Echo transmits a fraction of a second of audio before the wake word. I also know, from when I was doing my own debugging/hacking a while back, that there are requests made in cleartext to download packaged code during firmware updates, and based on the limited console hacking I did way back then (though I don't know enough about signal processing to have gotten more than that from the test pads on the bottom), Echo appears to have two bootable partitions, and swaps between them upon a successful upgrade - pretty standard stuff, it didn't appear that one partition was hiding anything, just that it had older packages.
My network analysis at the time was inconclusive in some respects due to some traffic being encoded/encrypted in a way I didn't immediately recognize -- frankly, my efforts at that time were focused on what was coming back to the Echo, and I will certainly investigate more of what's going out.
Your commentary thus far does not suggest anything privacy destroying (they know who I am and what I've bought? So?) but I will be happy to share what I find if I find anything. I know of many people smarter than I who've intercepted and analyzed the traffic and didn't find anything that suggests users "would not ever have privacy again", so I don't have high hopes of finding enlightenment without the assistance of someone who professes to have the answers.
What I can say, though, which may give good insight into your question is this: The "Hushpuppy" project, which became known as the WhisperSync and Immersion Reading technologies, were not self-trained or self-reinforced. QA engineers in India were contracted to ensure synchronization between the book and the audiobook. These technologies, though, were the building blocks upon which Alexa was developed. The data collected from these operations became essential knowledge for the engineering team that developed Alexa.
EDIT: As an aside, Amazon bases all of its services around that Amazon account that almost everyone in the civilized world now has. Alexa is not the first, and certainly not the last Amazon service that learns and improves based upon the user's Amazon account.
In this context, do I speculate in the right direction when I think about isolating voice from background noise? Any tips to share that could help my diy tinkering?
To be fair your phone has the same capabilities and it even follows you around from room to room!
Who wants to host their content in the cloud on servers they don't control?
Who wants to buy a netbook with nothing but Google Apps on them, collecting data to show even more targeted Ads?
Who wants to install Chrome (a closed-source browser) when Firefox and others provide much more transparency and privacy-controls?
Who wants a phone that tracks your location and reports it to World's most powerful AI company at every instant?
Who wants an on-demand-taxi app that tracks your movement inch-perfect, stores your payment and personal information, and keeps the record of your usage till, probably, end of time instead of... well... hailing a cab right off the corner?
Who wants to trust entities that hold your money in their coffers and in return show "digits" on-screen as a proof of presence of "actual" money?
Who wants to fly in giant metal-tubes, over which one has has no control whatsoever-- from who's flying, to what food that would be served, to what kind of hygiene is maintained, to the safety procedures followed, to the transparency into ways in which your baggage is handled, or your booking information, for that matter?
Who wants to drive a relatively smaller metal-tube on roads with other drunkards and addicts? Why are tax-dollars wasted on initiatives that present such grave dangers-- where someone else's mistake ends up costing someone else entirely?
Sure, there's always a real and present danger. You gain some, you loose some.
--
For me, it was about time someone made ubiquitous computing main-stream. It can only be good for the tech-ecosystem. Google seems in a prime-position to better Amazon's offering, if it hasn't done already through its home-products division, Nest.
That said, regarding the RasPi project, I did something similar for home automation before Echo/Hue/SmartThings/etc were around, and I used the Acoustic Magic Voicetracker I array microphone, which still works well for an updated use of that (now it sits on my desk for dictation and VoIP calls). Perhaps that microphone would be useful to cover an entire large room for you? The manufacturer claims 30 feet of usable distance between the array and the speaker, and in practice, I found that to be pretty accurate.
EDIT: it should be pointed out that the microphone (https://www.acousticmagic.com/products/voice-tracker-i-detai...) is more expensive than an actual Amazon Echo. The price hasn't changed in the 12 years that I've owned it, either. But I did test the audio samples my Echo sends with audio recorded via the array microphone, and the standalone microphone was far superior at all distances in terms of quality, would would be important if you're doing speech-to-text processing with PocketSphinx.
How well does the Acousticmagic work in such cases? Can you speak away from it, and will it still pick it up? Furthermore, how well does it work wrt background sound filtering? Does it hear you when the tv is on, or when children are playing or similar cases? Those are the main cases in which my setup requires speaking up.
You can use an offline speech recognition engine in that.
I'll try training Julius, though, it sounds like it may be the best solution to the problem.
For jasper (pocketsphinx) you have to manually program the action for all of these. So it's a lot more setup. I still like it and use it all the time though.
I've got a few other ideas: control my roku, add milk to the grocery list, read emails, etc.
Nothing life-changing, but fun stuff that makes small parts of my day easier.
Jasper also writes audio to disk, then runs command line tools on those files. I haven't tested if this is a significant source of latency.
Do you really think it would be easy to hide a persistent audio datastream to a central server?
I think that what you mean is that the monitoring is done in the hardware or firmware using closed-source code that can and will be regularly updated remotely and hopefully securely. And that Amazon told us that it would wait until it thought it heard "Alexa" or "Echo" or anything that sounds sort of like it, or whatever they decide to change the software on your particular device to listen for in the future.
It would probably also be fodder for frontpage HN, in case anyone needed some attention out there.
Amazon has told us that this product we paid to have in our homes won't spy on us, and has (to my knowledge) given me or anyone else ZERO indication that they'd suddenly decide: "Privacy? Fuck that! Let's see if someone is saying something salacious in that bedroom in Watertown, NY; that customer seems to be buying a lot of lube." Or, less sarcastically, violate their paying customers' expectation of privacy to suit their own ends, whatever those may be.
Google, however, has "snuck in" code to actively listen to the microphone in their browser, which we don't pay for. I won't use the old "if you're not the buyer, you're the product" routine here, but I will say that I trust the privacy protections of a free browser with portions of black-box, closed-source code a hell of a lot less than I trust the same protections of a paid-for product with portions of black-box, closed-source code.
What is some advice you can give for me to build my own "Jarvis" like Zuck is doing but be reasonably sure that the components aren't going to "phone home" somewhere or be rooted?
Locking down my network?
As a for instance, it would be nice if keyword enablement could mute the stereo. It'd be nice to Chromecast audio to it. It'd be nice if I could use it to play things on a fire stick.
If it really wanted to be a device of the future, it would link to other Echoes in the house, allowed for intercom, and localized audio tracking.
Some of these things will either come or never see the light of day because Amazon hates interacting with other companies (see: Android).
I'm waiting for an open source alternative...
Essentially, I want a better mic so I can run this:
I think in a couple of years this stuff will become ubiquitous
Baidu open sourced their warp CTC a few weeks ago, give it another few months before someone will release a trained English network for it
Not cheap and a little ugly some would say...
I think it is cool that it's programmable, but I'm not that impressed thus far.
(I should add, the reason I have one is that it was a gift from AMZN after attending an event of theirs last year, I didn't buy it)
audio jack: I like that it doesn't have an audio jack. I already have a Fire TV connected to my tv setup, it should play the audio there.
Plus AFAIK CEC can't turn back off the TV, so after the first use it's just on forever until you manually turn it off.
1) you don't believe anyone is listening, and even if they are, you don't care
2) you don't believe anyone is listening, but if they are, the consequences are worth avoiding
If you find yourself in the #2 mode, don't buy an Amazon Echo.
OTOH, having a platform that accepts questions (even in written form, i.e. an NLP google) on which we could build on would be great.