Offline voice AI within 512 KB of RAM [video]
youtube.com
youtube.com
In particular, voice AI systems are by their nature not very discoverable, which makes it hard to assess them quickly. What the video shows is a system that can recognize some input sentences and speak some output sentences. It does not show the size and flexibility of the command set, and it only gives a vague impression of the speech output quality and recognition accuracy.
Technically, you could reproduce this demo with the system software that shipped on a Mac Quadra 840AV in 1993.
We are in process of open-sourcing a statistically-significant benchmark for this tech. But this will happen in 2019.
[1] https://ai.googleblog.com/2018/05/duplex-ai-system-for-natur...
It's not as vaporware-y as blockchain, but I still see some potential for the bubble to pop. At some point people will hopefully realize that but everything that can be done with a neural network needs to be done with one (and tons of data).
The neural network stuff is anything but vaporware, it's delivered incredible results, but people keep coming up with silly dismissals along the lines of it not being "real AI".
That is not what I meant at all. For me, it is the nature of the field and has been that way since the days I first learned it (the 1990's). Once an reproducible algorithm or methodology is discovered to solve an AI problem, it generally ceases to be an AI problem.
Imagine if people treated programming languages like that. People would get excited about the idea of communicating with a computer, and then when you finally build Python and show it to them, they say "but that's just parsing, what about a real language?" The bar just keeps rising whenever you get close to it. That's the sort of thing I meant by "silly dismissal".
Inferencing is also highly amicable to hardware optimisation, which we're starting to see in the latest flagship mobile SoCs. I expect to see low-cost microcontrollers with inference accelerators within the next couple of years.
We had to come up with a bunch of ideas on how to fit our stack into the on-chip RAM (512 KB) and leave enough for OS and the actual application.
The demo is done with some noise and there is some reverberation as well. The speakers are also somewhat accented. That being said we would like to open-source a benchmark for this (similar to other products we have). The comment on accuracy is a bit tricky as it would depend on parameters you mentioned and specific task. I will provide more information when we open source the benchmark in Q1 2019.
You can find some information about the wake-word engine here https://medium.com/@alirezakenarsarianhari/yet-another-wake-...
I'll be interested when 'reading' thoughts without phoning home becomes a reality.
Keyword spotting is one of the modules we run in this demo. That's how we detect "Hey Barista". We also run an engine we can "Speech-to-Intent" that infers user request from follow up command.
One thing I wanted to mention is that there are two challenges when running DNNs on embedded platforms (1) limited compute power (CPU) (2) limited memory (RAM). RPi zero is definitely bound by (1) but not (2) as you get 100 MBs of RAM on it.
Z80 I believe was something like 4 cycles per instruction and only a megahertz, so we would be talking at least 16x slower than a M4. You would need something between a 486 and Pentium to get close to the M4, and then evenn further to get to the M7. If I remember correctly you couldn't even decode MP3s in realtime until the faster 90+MHz 486.
It would decode mp3, up to 256Kbps, over NFS, 95 percent CPU utilization.
I was quite surprised.
There was speech synethsis on the apple iie in like 64k of RAM and a 1mhz CPU.
Speech synthesis is far far far easier than parsing a human voice, especially when it doesn’t need to sound realistic (as was the case back then).
In the old papers about problem, there was vocabulary size of 64K words, because nothing was working for bigger vocabularies.
It is done using SVD and improves speed without sacrificing memory access patterns.
I totally understand the need to support makers community. We do have GitHub repositories for engines demoed here which allows you to use these technologies to some extent (not the full set of capabilities). I am working with our partners (both Soc and distribution) to come up with a maker-specific product for evaluation and personal use. It most probably will be a HW/SW product (i.e. a board that comes with our software). The product should allow you to use the full set of features on that specific board. I am expecting this to happen in 2019 and I will disclose the information as I am figuring things out.
https://www.arrow.com/en/reference-designs/imx6slevk-imx-6so...
It was an ARM Cortex-a9 with NEON extension instead of ARM Cortex-M7. It is basically a different family of i MX processors.