408 karma · joined March 22, 2013
https://en.bouffalolab.com/product/?type=detail&id=16
voice processing is in hardware unfortunately, but it exposes some things like DOA
Orphanes did struggle but most families were not just two person, families were big and supported by community.
Once knowledge is distilled you can build on top of it easily by merging concepts for example.
So no secret here.
Voice sounds robotic and plain. Most likely a lot of audiobooks in training data and less conversational speech. And dropping diffusion was not a great idea, voice is not crystal clear anymore, it is more like a telephony recording.
There are much more clear sounding systems around. You can listen for StyleTTS2 to compare.
https://github.com/openai/whisper/blob/main/language-breakdo...
https://www.smithsonianmag.com/smart-news/ancient-welsh-gold...
So we have numbers on PTB original perplexity 8.79 quantized 9.68, already 10% worse. And PPL reported per token I suppose? Because word PPL for PTB must be around 20, not less than 10.
Any numbers on more complex tasks then? like QA?
For example if you look for singing voice, they might suggest you an adapted model that is good specifically for singing.
The testing process is also not very straight, you need to understand what to test and how to test properly. For example, some of their voices might be better for questions, some for news.
You'd better talk to them.
https://alphacephei.com/vosk/lm
You can restrict the vocabulary the way you like, for example, here is the chess app built with Vosk
You might not understand but there is a huge amount of work behind this simple demo.