Are there any good guides out there for setting up Sphinx like Google/Amazon's API? I'd just rather not schlep audio data off to an unknown party when I have plenty of excess compute on my KVM cluster, accuracy isn't particularly critical.
Privacy aside, the biggest problem with self hosting these kinds of models in our opinion is the compute required. The really accurate models are so large, they require GPUs for inference. You could run the models on fast CPUs if you don't care too much about latency, but the throughput would be pretty low. Either way, GPUs and fast CPUs get expensive fast, so our hope is that by us specializing in hosting these models, we can offer you a price point that would be cheaper than if you were to try to host it yourself.