Also, we are deliberately not using Databricks for this to avoid vendor lockin for something that will almost certainly be open-source soon.
Also, we are deliberately not using Databricks for this to avoid vendor lockin for something that will almost certainly be open-source soon.
However I don't quite understand your point that "data versioning is not sufficient" because "I cannot roll back my datasets to a previous version". Surely data versioning would _solve_ the ability to roll your datasets back to a previous version? Or, are you not convinced of the need for reproducibility in data science? My rationale for this is the following: if you are building ML models that are going to make important decisions in production, it's imperative for debuggability that you're able to re-run that model training run later. If a model makes a bad decision in production, you need to know what dataset it was trained on, which means needing to be able to retrieve that dataset. That's because you can't isolate and fix the problem without being able to re-run it.
Yes, data changes and marches forwards, and that's why you should retrain models. But you also need to be able to go backwards to do robust ML. My 2 cents.
I also wrote a DevOps for ML Manifesto: https://dotscience.com/manifesto/
To summarize:
1. All models must be reproducible by someone else 6 months later.
2. All models must be accountable, that means you must be able to justify the basis on which they made their decision, in particular you need to know which data version was used and where it came from (provenance).
3. Model development must be collaborative. That means I need to be able fork a copy of your project and maintain all the metadata about which runs you did, with their respective provenance history.
4. Models must have a continuous lifecycle. You're not done when you ship - because models are about finding patterns in data, and the world is constantly changing, you need statistical monitoring and retraining to compensate for model drift.
Do you disagree? Did I miss anything? Let me know your thoughts!
It seems to me that people working on really large-scale problems always have these external dependencies that they just can't control and subsequently data versioning only goes so far when you cannot make a full-copy of the data for each version.
I agree with you in general that we should strive for reproducibility.