Show HN: Cortex – Open-source alternative to SageMaker for model serving
github.com
github.com
* Is this really for more intensive model inference applications that need a cluster? It feels like for a lot of my models, a cluster is overkill.
* A lot of the ML deployment (Cortex, SageMaker, etc) don't see to rely on first pushing changes to version control, then deploying from there. Is there any reason for this? I can't come up for a reason why this shouldn't be the default. For example, this is how Heroku works for web apps (and this is a web app at the end of the day).
As for your second question, we definitely want to integrate tightly with version control systems. Since right now we are 100% open source and don't offer a manged service, we don't have a place to run the webook listeners. That said, most of our users version control their code/configuration (we do that with our examples as well: https://github.com/cortexlabs/cortex/examples), and it should be straightforward to integrate Cortex into an existing CI/CD workflow; the Cortex CLI just needs to be installed, and then running `cortex deploy` with the updated code/configuration will trigger a rolling update.
If you're referring to version control for the actual model files, Cortex is un-opinionated as to where those hosted, so long as they can be accessed by your Predictor (what we call the Python file that initializes your model and serves predictions). If you're interested in implementing version control with your models, I'd recommend checking out DVC.
- Could you share your experiences?
- why would one choose this over docker for instance?
We also have a pretty active Gitter channel: https://gitter.im/cortexlabs/cortex
As for your second question, Cortex uses Docker to containerize models. The rest of Cortex's features (deploying models as microservices, orchestrating an inference cluster, autoscaling, prediction monitoring, etc.) are outside Docker's scope.
Does this approach preclude the need for queuing (a la RabbitMQ) and/or a load balancer?
How do you handle API authentication? Is there a module that interfaces with AWS API gateway? or external API authentication?
To keep the AWS bill as low as possible, Cortex supports inference on spot instances, which are unused instances that AWS sells at a steep (as in 90%) discount. The drawback is that AWS can reclaim the instance when needed, but with ML inference failover isn't as big of a deal, since you typically don't need to preserve state.
If you use spot instances, choose the cheapest instance type possible, and keep your autoscalers minimum replicas to 1 (meaning it won't keep many replicas idling), you should be able to deploy the model pretty cheaply. Significantly cheaper than with SageMaker, at the very least.
There's some more info here: https://www.cortex.dev/cluster-management/spot-instances
1. Size limits. Lambda limits deployment packages to 250 mb uncompressed, and puts an upper bound on memory of 3,008 mb. That's not nearly big enough for a lot of models, particularly bigger deep learning models.
2. As you mentioned, GPU inference is supported on Lambda, and for many models, GPUs are necessary for serving with acceptable latency.
3. Lambda instances can only serve one request at a time. With how slow ML inference can be—especially if you need to call another API or preform some IO request—it's easy to lock up Lambda instances for full seconds just to serve one prediction.
The TL;DR is that while Lambda works for some use-cases, it in general lacks the flexibility and customizability needed for most inference use-cases.
With Cortex, we wanted to build something so that developers can take a trained model—regardless of if it's trained by their DS team or if it is a pre-trained model—and deploy it as a production API without needing to understand k8s. Because Cortex manages the k8s cluster, we can do the legwork for features like spot instances, request-based cluster autoscaling, GPU support, etc, and expose them as simple yaml configuration.