It looks like the database will be open sourced later this year: https://voice.mozilla.org/faq
I'm wondering if the format will be easily translatable to the kinds of models that software like CMUSphinx and Julius use
I'm wondering if the format will be easily translatable to the kinds of models that software like CMUSphinx and Julius use
If you poke around github and the Kaldi lists a bit more you can see that they are experimenting with and probably planning to use Kaldi.
I wonder what they plan to do for provisioning. It is one thing to collect data and train models, but quite another to make the service available over the web in an unlimited capacity. And we are not yet to the point where you can reasonably expect to run a high quality open-vocabulary STT system in your browser. The search network is typically in the GBs range.