Ask HN: How can I set up an ML model as a scalable API?
I've been searching around and haven't found a clear standard/best way to do this.
Here are some of the options I've considered:
- Algorithmia (came across this yesterday, unsure how good it is and have some questions about the licensing)
- Something fancy with Kubernetes
- Write a load balancer and manually spin up new instances when needed.
Right now I'm leaning towards Algorithmia as it seems to be cost-effective and basically designed to do what I want. But I'm unsure how it handles long model loading times, or if the major cloud providers have similar services.
I'm quite new to this kind of architecture and would appreciate some thoughts on the best way to accomplish this!