So... instead of disagreeing out of hand, how did you overcome the fact that you had to pay for bandwidth between your terabytes of data and the computation centers?
Terabytes are cheap compared to my time. I've also never required that much training data, so perhaps your use case does not lend itself to these services. I've built automated ML model training systems myself with TreeNet and Mahout, as well as used API-based ML systems, and I'm comfortable saying there is a strong case to be made for services which save much unnecessary effort.
Storing terabytes is cheap. Terabytes of bandwidth is very very expensive.
And that assumes you're willing to host with those providers, many of whom don't provide computing power. Also, the machine learning service you are using has to pay again for that bandwidth too and that cost gets passed along to you.
Transferring 1TB at a sustained rate of 1MB/s will take about 11.5 days.
One easy way to get around bandwidth costs is to just have the clients of your service communicate directly (via Javascript or other client-side code) with your "machine learning as a service" (MLaaS? :D) vendor. Of course this model only works for certain businesses and allows learning on certain things, but you can go a pretty long way with this.