93 karma · joined December 6, 2023
Here are the following models I found work well:
- Qwen ASR and TTS are really good. Qwen ASR is faster than OpenAI Whisper on Apple Silicon from my tests. And the TTS model has voice cloning support so you can give it any voice you want. Qwen ASR is my default.
- Chatterbox Turbo also does voice cloning TTS and is more efficient to run than Qwen TTS. Chatterbox Turbo is my default.
- Kitten TTS is good as a small model, better than Kokoro
- Soprano TTS is surprisingly really good for a small model, but it has glitches that prevent it from being my default
But overall the mlx-audio library makes it really easy to try different models and see which ones I like.
Humans write a bit messier — commas, short sentences, abrupt turns.
The insights about VAD and streaming pipelines in this thread are exactly what I'm looking at for v2. Moving to a WebSocket streaming pipeline with proper voice activity detection would close the latency gap significantly, even with local models.
The remote control feature is cool but the real unlock for me was voice. Typing on a phone is a terrible interface for coding conversations. Speaking is surprisingly natural for things like "check the test output" or "what did that agent do while I was away."
The tmux crowd in this thread is right that SSH + tmux gets you 90% of the way there. But adding voice on top changes the interaction model. You stop treating it like a terminal and start treating it like a collaborator.
Here is a demo of it controlling my smart lights: https://www.youtube.com/watch?v=HFmp9HFv50s
Focused Youtube: https://chromewebstore.google.com/detail/nfghbmabdoakhobmimn... Removes all recommendations and just keeps a search bar. No shorts rabbit holes or algorithm-based media consumption
StayFocusd: https://chromewebstore.google.com/detail/laankejkbhbdhmipfmg... I like using the nuclear option. Blocks a bunch of sites I have that are in a list, such that I cannot open them at all.
This is why Anthropic naming system of haiku sonnet and opus to represent size is really nice. It prevents this confusion.