There are some good open source STT solutions out there like Mozilla's DeepSpeech (https://github.com/mozilla/DeepSpeech) that wouldn't be too hard to put behind an API. We also plan to open source more of our stack in the future so you can run it yourself.
Privacy aside, the biggest problem with self hosting these kinds of models in our opinion is the compute required. The really accurate models are so large, they require GPUs for inference. You could run the models on fast CPUs if you don't care too much about latency, but the throughput would be pretty low. Either way, GPUs and fast CPUs get expensive fast, so our hope is that by us specializing in hosting these models, we can offer you a price point that would be cheaper than if you were to try to host it yourself.