We had a big launch planned for 4-6 weeks from now, and have been working towards getting things ready for that. As a result, we're missing a lot of things we know we need like benchmarks comparing ourselves to other services. Please bear with us!
If you do end up trying the API, we'd love your feedback. We're trying to build a really simple speech-to-text API for developers, that you can get up and running with in just a few minutes, and that doesn't require you to implement <insert big tech co here> into more of your stack.
There's a lot more we offer too, like:
- customization via transfer learning for higher accuracy (language models now, acoustic models soon)
- supporting lossy audio like low bitrate mp3 files from phone calls
- transcribe any audio file without having to specify it's metadata like sampling rate or encoding
We're also constantly improving our models for higher accuracy. Every few weeks we ship accuracy improvements based on improvements to our DNN architectures, better data, better data augmentation, etc.
In terms of our STT stack, we're using CTC based models combined with RNN-LMs -- all built in TensorFlow and PyTorch -- for decoding. Happy to provide more info around our stack if you're interested!
For any questions off HN -- you can reach me at dylan at assemblyai dot com
Thank you!!