Run tracking liberates ML teams
dotscience.com
dotscience.com
Regardless of my uncertainty around trying another 3rd party platform, the premise is spot on.
Please reach out to me on luke@dotscience.com if you'd be up for finding a way to work on that together.
At my work, we use Azure ML Studio. I think the solution for deployment is to run some scripts to save the model information to git and automatically deploy from there. It will take a little bit of effort to set up, but I think it should work.
That makes sense about saving the model info to git, but how do you track the provenance of the data? Do you see that as being important? (e.g. to track back from a model to which data it was trained on, when later updating or fixing issues on the models?)
Feel free to get in touch directly - luke@dotscience.com if you'd be willing to try our stuff.
There's a bit of a double standard.
Also, we are deliberately not using Databricks for this to avoid vendor lockin for something that will almost certainly be open-source soon.
However I don't quite understand your point that "data versioning is not sufficient" because "I cannot roll back my datasets to a previous version". Surely data versioning would _solve_ the ability to roll your datasets back to a previous version? Or, are you not convinced of the need for reproducibility in data science? My rationale for this is the following: if you are building ML models that are going to make important decisions in production, it's imperative for debuggability that you're able to re-run that model training run later. If a model makes a bad decision in production, you need to know what dataset it was trained on, which means needing to be able to retrieve that dataset. That's because you can't isolate and fix the problem without being able to re-run it.
Yes, data changes and marches forwards, and that's why you should retrain models. But you also need to be able to go backwards to do robust ML. My 2 cents.
I also wrote a DevOps for ML Manifesto: https://dotscience.com/manifesto/
To summarize:
1. All models must be reproducible by someone else 6 months later.
2. All models must be accountable, that means you must be able to justify the basis on which they made their decision, in particular you need to know which data version was used and where it came from (provenance).
3. Model development must be collaborative. That means I need to be able fork a copy of your project and maintain all the metadata about which runs you did, with their respective provenance history.
4. Models must have a continuous lifecycle. You're not done when you ship - because models are about finding patterns in data, and the world is constantly changing, you need statistical monitoring and retraining to compensate for model drift.
Do you disagree? Did I miss anything? Let me know your thoughts!
It seems to me that people working on really large-scale problems always have these external dependencies that they just can't control and subsequently data versioning only goes so far when you cannot make a full-copy of the data for each version.
I agree with you in general that we should strive for reproducibility.
Lots more detail here: https://dotscience.com/product/ and a super long deep dive demo wih lots of examples (I would have made it shorter if I'd had more time ;))
There is an excellent open source project that nails this called sacred. It's not perfect, but it works, and as far as I can tell it has won the popularity contest.
Please join me in using and contributing back to sacred!
Probably google will pull a tensorflow soon and sacred will go the way of theano but until then...
What are you using at the moment?