Show HN: Local voice assistant using Ollama, transformers and Coqui TTS toolkit
github.com
github.com
These are easy to make and fun to play with and it's awesome to have everything local. But it will take more to build something truly useable. A truly natural conversational AI needs to understand the nuances of conversation, most importantly when to speak and when to wait. It also needs to know subtleties of the user's voice that no speech recognizer can output, and it needs control over the output voice more precise than any TTS provides. Audio-to-audio models in the style of GPT-4o are clearly the way forward. (And someday soon, video-to-video models for video calling with a virtual avatar. And the step after that is robotics for physical avatars).
There aren't any open source audio-to-audio models yet but there are some promising approaches. https://ultravox.ai has the input half at least. https://tincans.ai/slm has a cool approach too.
I think that's not true. See this for example: https://huggingface.co/facebook/seamless-m4t-v2-large It's not general purpose like GPT4o but translation still seems pretty useful
Docker is a great option if you want lots of people to try out your project, but not many apps in this space come with a dockerfile
We are also working on a complete open source stack for ASR+TTS+LLM and will be releasing it shortly.
For others who also hadn't heard of it, here's an overview: https://github.com/rhasspy/rhasspy3/blob/master/docs/wyoming...
Too bad that the project is in limbo after Coqui (the company) folded. The license limits the use of the weights to non-commercial usage unless you buy a commercial license, and there's nobody left to sell you one now.
[1]https://open.substack.com/pub/jdsemrau/p/teaching-your-agent...