- CLI/UI.
- Main backend for storing projects, user accounts, managing pull requests, forks, runners, deployers, loadbalancers.
- Data backend (dotmesh).
- Auto provisioning of VMs with jupyterlabs running and data synced to GCP, AWS
- Runners that configure environment, install dependencies open up tunnels so users can access them and start working.
- Optimized machine imagine builds so the startup takes ~1min (some of the docker images like jupyter lab are very big)
- Model packaging into docker images.
- Model metrics capturing (a proxy that runs as a sidecar and intercepts requests) and then attaching relevant classes for your models.
- Kubernetes operator to deploy the actual models. User didn't have to worry about creating deployment manifests, services or ingresses (they wouldn't even care about docker images). They would just say which model to deploy and they would get a URL. Models could be deployed in a k8s cluster built from nodes with spot instances so would run pretty cheap :)
- Last component that I worked on was probably one of the most fun - an inference router that could allow canary deployments for models and also shadow deployments where traffic is sent to many models at once but responses are taken only from primary. We got a really nice UI for this as well where you could drag sliders around to configure % of traffic and so on. Unfortunately never managed to write docs for this.
- Terraform to wrap everything and deploy to GCP/AWS.
Our team was always quite small so we were stretched thin. In the end sales were going well as well, probably 6 more months and we would have broken even and then profitable :)
- model management solutions
- enormous quantities of data
Well, I have docker containers for my data, and docker containers for the underlying TensorFlow versions, and docker containers for model source code that went into large-scale training. All of that is coordinated using a git repo which has some shell scripts to execute the correct model with matching data on a compatible TF image. So if someone says "I need you to re-run last week's model", I check out the script from that time and run it on an empty server.
I honestly don't get what MLOps is or why I would need it.
Generally things like MLOps are targeted at providing systems for people that either don't know how, or don't have the time, to set something up themselves, especially when the data science team is more than just one or two people, and when they're using a range of tools that need to interoperate in various ways.
Something like Kubeflow, for example, does a lot more than what you describe.
And no, we use TensorFlow and Chainer, but both are python frameworks. Plus Numpy, of course.
I like that KubeFlow has an introduction video, but I find it quite odd, too. They talk a lot about how they will make things simpler, but then I learn that I'll need to run Kubernetes on my laptop, my servers, and potentially the cloud.
Plus they use irritating marketing phrases like "you can just focus on your model" or "let KubeFlow handle the abstraction of running on X". Most deep learning models nowadays are memory-limited, so improving the model may well mean optimizing GPU memory usage. And after trying to port a working model from Ubuntu + 1080 TI to Google Cloud + V100 and/or Google Cloud TPU, I distrust anyone who would treat those significant hardware differences as "just an abstraction".
So at the very least, I'll need to enforce consistent GPUs and Operating Systems among all servers, just to make things run OK everywhere.
What I think might work is auto ml combined with models ops, or rather auto model ops.
Or even better: auto data managmenet -> auto pre-processes -> auto ml -> auto ops.
If they also offer a private Docker repo, fast S3-compatible storage and some pre-built images with Jupyter preinstalled, that might be the entire ML pipeline that I need.
http://papers.nips.cc/paper/5656-hidden-technical-debt-in-ma...
I use MLFlow at my company for solving two specific problems: tracking experiments and model performance.
But if you're small enough, you probably don't need MLOps, Devops, etc...
Having your model in docker generated from git is a good first step, but does not solve the most important issues when you are at scale: reproducible training, tracking of experiments including data, ML-specific observability for your models in prod, etc. See e.g. https://martinfowler.com/articles/cd4ml.html.
More concretely, since the above easily sounds like a buzzword soup:
1. Experiment-wise, you don't want just want to track your ML model definition, but also track the data and everything else used to build it, not just use it. A typical thing I have seen at every company I worked at: we have this model but we don't know how to reproduce it because the lost the data, or there was some magic numbers in training that may be on an internal wiki if you're lucky.
2. For some important use cases, you want to iteratively work on improving the model in production. That often means work on the data side, feature engineering, tracking skew prod vs training, etc. the model is not often changed. In almost every case, a useful model is a model that sees 100s if not more iterations in production. You need 1. to do 2.
3. Some of those tools are useful to enforce invariant or detect data issues, which is again very common when you run models for a long time. See e.g. tfdev.
But to go back to your point: MLOps is a 2nd order kind of thing. The first order is of course that most ML-related projects are useless, poorly conceived, or even lack any kind of quantitative analysis on the business and/or product. Companies should invest there before MLops IMO.