If I'm being perfectly honest I'm surprised we got it this far already. If I wanted to be really critical:
- Far-field speech is actually kind of hard. There are at least dozens of "knobs" we can tweak between the various component libraries, etc to improve speech quality and reliability for more users in more environments. We've tested as much as we can considering there's only two of us but we need more testing from more speakers in more environments.
- On the wire/protocol stuff. We're doing pretty rudimentary "open new connection, stream voice, POST somewhere". This adds extra latency and CPU usage because of repeated TLS handshakes, etc. We have plans to use Websockets and what-not to cut down on this.
- We don't really support audio playback yet. For a real "Amazon Echo" type experience you need to be able to ask it random things like "Hey what's the weather outside?" and it needs to "tell" you.
- Ecosystem support. Using the example above, something like Home Assistant or similar needs to know where you are, get the weather, do text to speech, etc for Willow to be able to play it back.
- Other integrations. Alexa has "skills" and stuff and we need to be able to talk to more things.
- UI/UX work. We support the touch display but we did just enough to show colors, print status, add a button, and make a touch cursor that follows your finger around. We also only give audio feedback with a kind-of annoying tone that beeps once for success and twice for failure.
- Speaking of failure, we don't do a great job of telling you what went wrong and where.
- Configuration and flashing. It's very static and has multiple steps. There are all kinds of things that need to get done to make Willow easy enough for less-technical users to deploy and actually use daily without any hassle.
- Local command recognition. It's very early but as noted in the README, wiki, etc the ESP BOX itself can recognize up to 400 commands directly on the device. In testing it works surprisingly well but we have a lot of work to do to make it actually practical for most people.
- Open sourcing our inference server. We plan to do this next week!