> The world's most advanced Speech-to-Text, customized for your application.
> The world's most advanced Speech-to-Text, customized for your application.
If you do end up trying the API now and have any feedback definitely let me know. Our goal is to work towards making the API really easy and simple for developers, but we know there’s definitely things we have to improve on.
We’ve been working hard towards launching publicly in 1-2 months but everyone on HN knows you can’t always control this stuff.
Dylan
We're surprised to see ourselves on HN tonight -- but if you do try out the API I would love to see what you think! Thanks for your interest.
DeepSpeech is an awesome open source project, and we absolutely support open source speech-to-text. We're going to be open sourcing parts of our stack in the future as well.
We're a lot more accurate on real-world data (like phone calls, podcasts, accents, noise, etc.) than the current DeepSpeech model. We're actually less accurate on LibriSpeech Clean than what DeepSpeech reports, but we've found Libri Clean isn't very representative to real-world data and as a result isn't a great benchmark.
We're planning to put out more thorough benchmarks comparing our API to other services (including DeepSpeech, Kaldi, and CMU Sphinx) in about 1-2 weeks. If you want to email me at dylan[at]assemblyai.com I can send you the benchmarks once we have them.
Another thing that's hard to do is host RNN based models like DeepSpeech in production at scale. It gets expensive fast, since you need to do inference on GPUs to keep your latency time down. We spend a lot of time optimizing our infrastructure to keep costs down, so there's a good chance we can host RNN based models cheaper than if you were to do it on your own since we specialize in this.