Why we do machine learning engineering with YAML, not notebooks
towardsdatascience.com
towardsdatascience.com
You can also set it up as a 'filter', so it automatically runs before any git operations, whether it's add, commit, diff or an interactive rebase.
You can also use ReviewNB [1] that is literally built for Jupyter notebook diffs. You can see notebook visual diffs for any commit or pull request on GitHub. For pull requests, you can also write comments on a notebook cell (emulating the typical code review experience for Jupyter notebooks)
Disclaimer: I built ReviewNB for Jupyter notebook code reviews on GitHub.
(Edit: also, clearly, I don't read the news -- is YAML being "superseded"? By what now??)
I don't think YAML/JSON are being superseded, but I'd really love for something like Cue (https://cuelang.org/) to become the standard for storing configuration.
i hope by something that isn't the kitchen sink... i actually prefer XML and i hate XML.
started looking at https://dhall-lang.org/# which compiles to json/yaml, is seriously strongly typed and explicitly not Turing-complete both as a design goal and current reality.
[0] https://netflixtechblog.com/notebook-innovation-591ee3221233
We don't deploy or execute ML models in production as notebooks. We have many other solutions for that use case. In particular, check out https://metaflow.org
Have you never seen a script in production?
Papermill (https://papermill.readthedocs.io/en/latest/) made it extremely easy, to the point where I question the real value of moving away from this model. But software engineers practically hiss when you mention Jupyter because it's too different from the rest of their tooling.
I really dislike this practice. Code that runs in production should be code reviewed and there should be some monitoring in place to make sure the job is working correctly
You can bet that prod services from companies you heard of are running on something more analogous to versioned docker images. Not a yaml file which says, 'Go run whatever predict.py is in the current folder.'
The moment one of your dependencies breaks your code, or snookers your performance, there will be a lot of head scratching going on.
Reasons 2 and 3 for avoiding jupyter are more justified but easy enough to work around between jupytext and papermill.
The topic of production notebooks often shows up. I’ve seen tools like papermill and dagster being used for notebook prod, just like in Netflix.
I concluded that using notebooks for prod is always a tech decision, often influence by a tradeoff: risk for premature optimization (writing scripts early on in the project that may only be used once) and underengineering (using non-maintainable and clunky code to support mission-critical workloads): https://ljvmiranda921.github.io/notebook/2020/03/16/jupyter-...