> Please, please, please be a completely open, extensible platform...
That's one. The second one is, please make it self-hosted. No cloud bullshit.
I know I'll probably never live to see the second one coming true.
> Please, please, please be a completely open, extensible platform...
That's one. The second one is, please make it self-hosted. No cloud bullshit.
I know I'll probably never live to see the second one coming true.
You could build this on a pi with a mic, speakers, some foss stt and tts engines and some basic training data. But it'll suck.
I'm not buying you couldn't make a decent, self-contained, off-line speech recognition system. Sure, it may not be as good as Echo or Google Now (though the latter does suck hardly at times, it's nowhere near reliable to use, and it doesn't understand shit over a quite good and expensive Bluetooth headset). But it would be hackable, customizable. You could make it do some actual work for you.
Oh, and it wouldn't lag so terribly as Google Now does. Realtime applications and data over mobile networks don't mix.
That's a key limitation, though.
But we're getting close to the point where you can do some of this. For example - http://arxiv.org/pdf/1603.03185.pdf - LSTM speech recognition running on a Nexus 5.
The more serious problem with this is that it's going to be expensive -- and somewhat wasteful. There's a lot of pressure to keep consumer devices as cheap as possible, and the cloud is an awesome way to do that. Having shared cloud-based infrastructure for the speech recognition as opposed to putting it into every device (even though it's only used for ~5 minutes every day) is probably a lot cheaper. Consider the hardware in an Amazon Echo:
https://www.ifixit.com/Teardown/Amazon+Echo+Teardown/33953
256MB DRAM and a TI DSP: http://www.ti.com/product/dm3725 with a single Cortex-A8 core (about $23 + a smidgeon for the dram)
vs. a Nexus 5 (2GB DRAM, 4 core 2.2Ghz Krait 400) -- the N5 has roughly 8x the DRAM and compute of the CPU in the Echo.
Would you pay an extra $150 for a LocalEcho that still had to send most of your queries to a search engine for resolution, or to a cloud music service for music? (You & I might, but most consumers wouldn't.)
> That's a key limitation, though.
Why would it be? Sophisticated exchange of theorems and not essential for this scenario, is it?
I agree. It's not a problem of technology, it's a problem of incentive. There's no money in developing self-contained, off-line speech recognition system, unfortunately.
Nonsense. Self-hosting is highly valued in the enterprise sector. But we're not talking about the sort of products that could be sold to consumers for a few hundred dollars here.
A Pi, though, couldn't do well at all, just like you said. If I wanted to build a system like this for myself, I would target an HTPC form factor.
edit: Another possibility, which was explored elsewhere in this thread, would be to keep the listening device "thin", but have the ability to offload the processing to a machine in my LAN instead of one the "cloud".
Just the other day I was looking at CMU's Sphinx project for speech recognition. It seems quite capable, even of building something like this Google thing, but I haven't tried to actually use it.
Large-vocabulary recognition probably needs something better than a Raspberry Pi... so, just use a more powerful CPU.
Yes, Google has an incomprehensibly enormous database of proprietary knowledge and information. Good for them! If we want to build a home assistant that doesn't depend on Google, we'll have to make tradeoffs. That doesn't mean it has to suck.
Is it PocketSphinx?
I was mostly interested in automated transcription, didn't look much at the live recognition stuff.
Is there something I'm missing?
But, on the other side, if it's not open and you can't use any device with it... I'm going to be really upset on a personal level.
The reasons consumer IoT isn't huge yet are: 1) Disparate connection types (e.g., I could buy Z-wave, Wifi, BLE, etc and they all onboard differently) 2) I can't choose which device I want to use with which platform because of politics.
Some of these devices (thermostats or security systems for instance) aren't impulse buys. If I have a Honeywell thermostat, and Home doesn't support it, I either buy a new thermostat or don't buy Home.
That's a crummy choice for a consumer.
I rather suspect that the knowledge graph it uses is a rather hefty dataset. Probably not suitable for a home installation. And how would you keep it up-to-date without the cloud? Would you have it scrape websites and consume feeds itself?
The more important aspect of it is fixing the problems with said knowledge graph. For instance, Google doesn't have the data on the public transportation in my city. I could easily write a scrapper that would fetch me the bus/tram timetables - but there's no way to integrate that source of data with Google Now. It's one example, but in practice Google's knowledge graph is pretty much useless for me. At best, it can answer me some trivia questions sometimes.
Let me introduce you to PuSH: https://en.wikipedia.org/wiki/PubSubHubbub