The $10 Echo
sammachin.com
sammachin.com
I should clarify that there is a bit of artistic license on the title calling this a $10 Echo, there are 2 caveats;
Its only about $10 the CHIP is a "$9 Computer" plus tax & shipping etc. and this hack then requires a speaker, microphone and a button, however you may well have these lying around of have another device like an old speakerphone or something that you could gut and put the CHIP into as a new brain.
The other point is that its not really a full Echo, the Echo is a hardware appliance that implements Far Field Voice Technology, Wake Word Detection, A High Quality Bluetooth Speaker, Support for various media service (Prime Music, Audible, Pandora etc) and finally the Alexa Voice Assistant. This hack is really just access to the Alexa Voice Assistant without the wake word or far field tech.
So its about 90% of the functionality for a much lower price. Personally I have an Echo in my dining room which is great when walking around the house but I built this to use on my desk when I want to trigger home automation features without shouting downstairs.
Very specifically, you can speak in a conversational tone with Echo, vs barking commands at Siri.
I don't have experience with Cortana, but the Voice Recognition capabilities in the XboxOne is garbage.
High end smartphones have a similar high quality mic. It's no problem at all talking or using e.g. GoogleNow.
(obviously being silly and oversimplifying but same idea)
If you think of it as a device that performs specific actions (i.e., plays audio, gives a weather forecast, sets a timer, etc.) via an voice interface, rather than a general intelligence that can answer random questions, you're going to have a better time.
Many of the commands it "fails" at is when it tries to "guess" what I might have meant to say because it can't find something in its database. For example, if I give it a wrong name for a song (say "Alexa play She's Leaving by The Beatles" it might play some song called "She's Leaving" by some other band because it can't find that song under The Beatles.
It would also be interesting if you could create sounds that sound like commands to the algorithms, but are not easily recognizable for the user.
e.g. you could do it by buffering the request, checking and submitting once you're confident about authenticity, but then you'd be delaying submission and response, depending on how much you'd need to buffer to gain sufficient confidence that the individual is authorised.
It's not "secure" in any sense, but it cuts down on false positives and perhaps some very juvenile hacking attempts. Any professional hacker will generate a close-enough voiceprint that you can't use voice recognition inherently (purely passively) as authentication any time soon.
You've probably not used the Kinect controller? It has a very sophisticated distant speech system.
Also obligatory, why the downvotes?
[1] http://research.microsoft.com/apps/pubs/default.aspx?id=1022... [2] http://research.microsoft.com/apps/pubs/default.aspx?id=2012... [3] https://www.google.com/search?q=tashev+video+lectures&oq=tas...
There have been advances in deep learning based speech recognition to reject noise - http://usa.baidu.com/deep-speech-lessons-from-deep-learning/ for instance, but clean audio source helps massively, especially if you are trying to talk from any distance instead of being right next to the mic.
I haven't seen open source libraries for state of the art RNN/LSTM based speech recognition, anyone know if they exist?
MEMS mics have been taking over electret mics recently, they are way smaller and have better SNR (at equivalent size).
I'm building a 4 mic array at the moment, tetrahedral/ambisonics mic for VR use, the mics themselves are very cheap. Mouser and Digikey stock them - http://eu.mouser.com/Sensors/Audio-Sensors/MEMS-Microphones/...
Haven't looked into the software/beamforming part yet - does anyone know if there are open source libraries for that?
Some MEMS mics can also do ultrasound range, opening up pretty interesting possibilities for cheapo LIDAR type echolocation with a mic array - but that's for the next project :)
PS: interested in your project, anywhere I can follow it?
The Echo's ability to do voice to text is like nothing that's ever come before it based on my experience building, and speaking with a consultant in the field that worked on the Echo team.
That's the problem with progress and technology. It will happen regardless of how comfortable you are with it. In fact, if it helps those with power to maintain or extend their power it will happen. If it tends to limit their power it will be squashed and marginalized.
The trend you cite will happen because of ignorance, not progress.
Citizenfour was eyeopening on this point.
And the real voice recognition should be turned off 99% of the time. It should just passively listen for a keyword before using any cloud service.
"You already have the flu, why worry about getting gastroenteritis?"
>And the real voice recognition should be turned off 99% of the time. It should just passively listen for a keyword before using any cloud service.
On tools like OP's link, or Jasper, etc. you have the source and you can verify that it's actually not sending out data permanently. With the Echo, you're just hoping Amazon didn't do exactly that. You could watch the network activity, but that would just tell you that some data is sent back, sometimes.
and no different than a telescreen (1984). No one will leave home without one.
Note that high-quality and cheap are not contradictory requirements. iPhones and other smartphones have extremely high-quality microphones (and they are tiny) and must certainly cost much less than $5 as a part, but they seem impossible to find. Every off-the-shelf mic under $5 that I've tried has been garbage in terms of sound quality, and even much more expensive ones have not been great.
Good mics certainly exist. Here's a compact analog microphone I own that has fantastic sound quality:
http://www.amazon.com/Sony-Electret-Condenser-Microphone-ECM...
But it's ridiculous that this microphone costs 6x as much as the CHIP and 10x as the RPi Zero. You can't build a product around the CHIP or the RPi if an essential part is that expensive.
http://www.amazon.com/Cyber-Acoustics-100Hz-Microphone-ACM-1...
I agree it'd be useful if there was a brief explanation at the beginning of the article.
http://sammachin.com/hacks-and-projects/alexa-in-the-browser...
However, I would love to have an Echo like device which does the voice recognition/algo crunching/etc on premise, i.e. on a server/device I control.
Is there anything out which does this nicely?
I wish more people were like you and voiced their concerns about this ongoing problem around centralization of services. However...see http://www.cultofmac.com/264381/hardware-siri-runs-puts-new-...:
> Over on Reddit, a post suggests that there are 3 main ‘instances’ (or server farms) for Siri in the United States, and at least one for every other country.
> This information allegedly comes — albeit second-hand — from Apple’s lead cloud architect, who says that every instance of Siri runs on 32 powerful HP servers with a total of 1024 cores and 32 terrabytes of RAM apiece. That certainly makes the new Mac Pro look long in the tooth.
> Specifically, each instance of Siri is made up of 4 HP c7k enclosures made up of 8 HP server blades each, with memory upgrades to 1TB of RAM.
> According to the post, if one server dies, it’s simply removed and another one slapped in, with no downtime.
Granted, there are a lot of people yacking at Siri at a given time, but still. You probably need a fair amount of horsepower to pull this off.
https://mesosphere.com/blog/2015/04/23/apple-details-j-a-r-v...
They say the clusters have 1k worker nodes
Amazon is starting a $100M fund for anyone building voice related services. This is very interesting. Has any company been funded via this fund?
That's just a guess but it's similar to what Google does with AOSP and Google Android.
You could get a dedicated Android phone/tablet, and leave it always plugged into a charger somewhere. I just tested it on my Phone (LG G3 Android 5.0.1), and I'm able to get my phone to respond to "Ok, Google" questions, even with the screen off.
Though the microphone probably won't be nearly as good as the Echo, and from the setup it seems like you have to train "Ok, Google" to respond to your voice, so I don't know if it will work with multiple users.
Too much to hope for I suspect.
I have an Echo and the mic array is fantastic. And the idea of having one in every room for localized interaction is great. But the idea of such a good mic in every room is scary if you consider bad and/or state actors.
OK Google works, but the mic really isn't as good as the array on Echo. I've been looking into external mics for Android - will test with simple omni mics first but ideally an external array mic would be the best choice.
I have to be honest, I got an echo for Christmas (thanks Mom!), and while it's really cool, it has only served to highlight how much amazon prime music sucks. In fact, it has only served to highlight to me how disconnected cloud music services are.
Google music has just about everything (now that it is connected to youtube), but amazon doesn't want to let its speakers play google music's music service because they want to sell you their own.
The future is a strange place right now.
The June2016 date is for orders placed on their store today which I think is the earliest they currently expect to have filled all the current pre-orders and later stage kickstarter backers. IIRC they're currently producing 5000 units a month.
They provide Alex-as-a-service: https://developer.amazon.com/public/solutions/alexa/alexa-vo... ("Build for Free. Using AVS to power speech experiences on your devices is completely free.") (Note: I'm guessing that this will be lower accuracy than the echo, since the echo's acoustic models are probably tuned for their hardware.)
Echo also is a pretty nice wireless speaker.
Intelligence is hard to mimic, but physical form factor is not. All google or apple needs to do is come out with with a similar speaker/mic combo and it will totally blow echo out of the water.
Just a guess though. I'm surprised that wasn't included in that rather pricey wifi router they launched recently. It has a decent sized speaker in it but perhaps they decided not to enable it just yet. My guess is that they wanted to start selling the wifi routers but didn't think the voice stuff was up to snuff for a consumer application/selling point just yet.