Like anyone working in that space who wants to keep their sanity, we're building our machine learning platform[0]. We shipped
many projects and it can be taxing to work on different stacks, especially given the fact that we built for enterprise and they always want
complete solutions. The model is but a foot in the door, and you must do custom front-end/back-end/model/pipe/data acquisition.
We decided to build tooling around that workflow. We shipped and were paid with that workflow, so we wanted to make it efficient and effective. Other solutions didn't fit our needs. Straight out of https://xkcd.com/927/, except it needs to address our use cases. We backtest with our past projects and use it for our current ones to scale our consulting capabilities.
>- standardizing end to end tooling for special resources, eg queueing and batching to keep utilization high for production GPU systems, high RAM use cases like approximate nearest neighbor indexes, and just run of the mill stuff like how to take a trained model and deploy it behind a microservice in a way that bakes in logging, tracing, alerting, and more.
We do schedule notebooks[1]. We also publish AppBooks[2], which are automatically parametrized notebooks: it automatically generates a form so anyone can set variables the notebook author's chose to expose and run a notebook without changing the code. Extremely useful when you want to have a domain expert tweak a domain specific variable, without them having to know what a notebook is. In some projects, there's someone with deep, deep expertise in a field for whom a variable is really important, but that variable gets dismissed by the ML practitioner because they didn't see a correlation or an impact on AUC or something, so the domain expert has an input, whether on relevant variables, or the real world metrics we're working for. We also added instrumentation for the basic CPU/RAM/GPU, data, servers running, etc. Again, so that our teammates don't bother with this. We use different Docker images for notebook servers with some that are 30GB so members don't bother with dependencies, GPU/tensorflow/cuda and version conflicts.
We automatically track metrics and parameters without the notebook's author writing boilerplate code, because they forget. The models are saved, and the ML practitioner can click on a button to deploy a model[3]. We stressed a lot on self service because its absense put a lot of stress on us: an ML practitioner wants to deploy a model, asks someone in the team who's probably busy. So we said: anyone should click and deploy. This is also useful because a developer downstream will only need to send in HTTP requests to interact with the model. We used to have application developers who also needed to know more than they should have on the internals/dependencies of the models. Not anymore.
We also added near real-time collaboration/editing so many people can work on the same notebook, especially useful when a team member is struggling to implement something, and others chime in to help debug/refactor, and review[4]. Everyone sees everyone's cursor for better awareness of what's being done. A use case is an ML practitioner playing with a paper who's struggling on the algorithms or data structures part of the paper. They can sollicit another team member who'll chime in and help. One of the features that's useful is the multiple checkpoints[5], which allows one to revert to an arbitrary checkpoint, not just the last one. This, again, is porcelain because an ML practitioner playing with git is a context switch, and they don't really like it.
So, we've done a few things to make our life easier. The workflow isn't perfect, and the tools aren't perfect, but we're removing friction.
We add applications for the usual timeseries forecasting, sentiment analysis, anomaly detection, churn, to leverage projects we did in the past.
[0]: https://iko.ai
[1]: https://iko.ai/docs/notebook/#long-running-notebooks
[2]: https://iko.ai/docs/appbook/
[3]: https://iko.ai/docs/appbook/#deploying-a-model
[4]: https://iko.ai/docs/notebook/#collaboration
[5]: https://iko.ai/docs/notebook/#multiple-checkpoints