Netflix's Metaflow: Reproducible machine learning pipelines
cortex.dev
cortex.dev
Also I'm happy to answer any questions (I lead the Metaflow team at Netflix).
There's a bit in the Metaflow docs that talks about choosing resources, like RAM: "as a good measure, don't request more resources than what your workflow actually needs. On the other hand, never optimize resources prematurely."
The problem is that for memory, too little means out-of-memory crashes, so the tendency I've seen is to over-provision memory, which ends up getting very expensive at scale.
This choice between "my process crashes" and "I am incentivized to make my process organizationally expensive" isn't ideal. Do you have any ways you deal with this at Netflix, or have you seen ways other Metaflow users deal with it?
I have some ideas on how this could be made better (some combination of being able to catch OOM situations deterministically, memory profiling, and sizing RAM by input size for repeating batch jobs), based in part on some tooling I've been working on for memory profiling: https://pythonspeed.com/fil, so would love to talk about it if you're interested.
While it is true that auto-sizing resources is hard and the easiest approach is to oversize @resources, the situation isn't as bad as it sounds:
1) In Metaflow, @resource requests are specific to a function/step, so you end up using resources only for a short while typically. It would be expensive to keep big boxes idling 24/7 but that's not necessary.
2) You can use spot instances to lower costs, sometimes dramatically.
3) It is pretty easy to see the actual resource consumption on any monitoring system, e.g. CloudWatch, so you can adjust manually if needed.
4) A core value proposition of Metaflow is to make both prototyping and production easy. While optimizing resource consumption may be important for large-scale production workloads, it is rarely the first concern when prototyping.
In practice at Netflix, we start with overprovisioning and then focus on optimizing only the workflows that mature to serious production and end up being too expensive if left unoptimized. It turns out that this is a small % of all workflows.
The better your estimator function, of course, the tighter constraints you can use.
Metaflow is rather unopinionated about those types of questions, since they are subject to active research and experimentation. Metaflow aims to make it easy to conduct the research and experiments but it is up to the data scientist to choose the right modeling approach, features etc.
In some cases, individual teams have built a thin layer of tooling on top of Metaflow to support specific problems they care about. I could imagine such a layer for specific instances of transfer learning, for instance.
In general, we are actively thinking if/how Metaflow could support feature sharing in general. It is a tough nut to crack.
Or you're forever tied to the initial hardware+compiler version?
Depending on the libraries you use, the exact results may or may not be reproducible on other architectures. If cross-platform reproducibility is important to you, you should choose your libraries accordingly. Metaflow provides the tools for choosing the level of reproducibility that your application requires.
Also, what tools does metaflow offers to control the level of reproducibility?
Netflix’s recommender system is hands down the worst I have ever seen. Every single thing I watch, it suggests The Queen’s Gambit and two other random Netflix productions.
Even if I watch the first of a trilogy (LotR, for example). How can they be so terrible at this?
The categories in the main browsing view are also hysterically arbitrary. It kind of looks like a topic model with back-constructed titles for the topics.
Finally, they replaced the ratings with “% matching”. I guess so they can recommend their subpar productions even if they get low ratings.
That or a ceding of recommendations to marketing. That feels too cynical, but more accurate.
Netflix doesn't show IMDB scores, so I have to check it always independently...it sucks.
Much later, when they switched to simple thumbs up/down, the recommender system was entirely useless to me. (Not merely because of the dumbed-down rating system; the recommendations were genuinely bad.)
For the time in between, I'm not sure if the degradation was gradual, sporadic, or not degraded at all.
I still think recommender engines should always enable some form of user tuning. If it doesn’t, then the recommender is a tool for services to control your behavior rather than the other way around.
Although even before I started listening to Spotify while working etc, it seemed to have run out of things to recommend. My weekly discover playlist would be half things I’d already liked. So...who knows. But I miss the discovery functionality quite a bit.
Nothing really beats the recommendations of a human curator with exquisite taste, and just listening to new music nonstop and plucking out the gems as you go.
It's also easy to theorize about "big media" controlling what I listen to, so I still feel the need to do my own exploring, even when I'm getting recommended good fresh stuff.
Not a streaming service but Google's Discover news is also very good (probably the best recommendations I have come across).
Similarly, if I play Blacklist in the background as basically noise, I don’t want to see a bunch of related shows. I guess I could give it a thumbs down but I only do that for actually terrible movies.
Also Spotify and Apple music seem to have okay recommendations
That's always what ratings were. People didn't understand that (as you can see), so they changed it to make it more transparent.
https://www.businessinsider.com/why-netflix-replaced-its-5-s...
>Netflix’s star ratings were personalized, and had been from the start. That means when you saw a movie on Netflix rated 4 stars, that didn’t mean the average of all ratings was 4 stars. Instead, it meant that Netflix thought you’d rate the movie 4 stars, based on your habits (and other people's ratings). But many people didn’t get that.
What happened is that the notion changed from "predict a scalar rating" to "predict a binary satisfaction."
As parent poster noted, the effect of this is to push "3 star" and "4 star" acceptable shows to the user, instead of "5 star" great (in the user's view) shows.
Also, in Netflix's defense, users are horribly inconsistent in their expressed ratings (they'll rate a movie based on their personal mood at the time, and they'll binge shows they claim aren't 5 stars while ignoring their 5 star movies)
I remember they stated the reviews would still be available in some form for export.
It's no longer the Netflix of old imo.
This is exactly why they switched from 5 stars to up/down/blank.
Netflix doesn't make money by showing people stuff they don't like.
But the recommendations seem to work perfectly for me - I wonder whether that is the case for the silent majority?
The match % is usually spot on for me, and I've never seen it recommend any titles I've given a thumbs-down to.
In my limited experience, they are all dismal and best ignored entirely.
When it was easier to see something plausibly like a real user review, that helped I guess.
At least with Amazon Prime there's a high chance they'll at least find the title one searches for and if I am really motivated I can pay a few bucks to watch it.
Netflix just draws a blank.
That's not to say I haven't watched some entertaining things on Netflix, but they seem much better suited to TV series than movies, and I almost feel like I found decent things to watch inspite of their recommendation system, not because of it.
An ML tool from Netflix? It feels like the last thing I'm likely to use.
Until you log into prime video.
Can’t manage to give me a “continue watching last thing button”. That’s literally the most likely thing I want to watch. Also routinely suggest starting with S02 even though I’ve not watched S01.
Never mind machine learning some common sense would be greatly appreciated
My guess is that whatever rankings they currently produce maximize some internal revenue target and that target' user base is not me or you.
If I watched S01E04 yesterday I want to watch S01E05 today. The interface should be suggesting that PLUS whatever else the ML comes up with in addition, not instead of.
Prime is full of minor irritations like that make me wonder whether Amazon engineers dogfood enough
I always assumed those three were human selections, not part of the recommendations.
Like the paid ads on top of your search results.
And while saying that, I appreciate their high level technical staff. These decisions are made by bean counters.
What I really want is a single solution, or a set of pluggable, integrated components that offer:
* training data and model storage (on top of a blob store like S3, minio, ...)
* interactive dev environments (Notebooks, dev containers, ...)
* training (with history, comparisons, parameters, ...) with experiments for parameter tuning
* serving/deploying for production
* a permission system so researchers and developers can only access what they are supposed to
* software heritage, probably via Docker images or Nix packages, combined with source code references
* (cherry on top: some kind of integrated labeling system and UI)
Right now you have to cobble this together from different tools that are all pretty suboptimal.
The big players can set up sophisticated systems, but I'm curious to hear how other startups are currently solving this.
disclaimer I am one of the authors of an open-source solution (https://github.com/polyaxon/polyaxon) that specializes in the experimentation and automation phase of the data-science lifecycle.
Our tool provides exactly the kind of abstraction you mentioned:
* Training, data operations, and interactive workspaces (https://polyaxon.com/docs/experimentation/)
* A scalable history and comparison table (https://polyaxon.com/docs/management/runs-dashboard/comparis...)
* Currently pipelines and concurrency management is on the commercial version (https://polyaxon.com/docs/automation/) but several companies use Polyaxon with other tools like Kubeflow (https://medium.com/mercari-engineering/continuous-delivery-a...) or it can be used with MetaFlow for the pipelines part.
I would really like to hear your thoughts and feedback.
Isn't that what the parent you are replying to is talking about with "Right now you have to cobble this together from different tools that are all pretty suboptimal."
If a company is already using a pipelining tool, a visualization tool, or a data management tool, Polyaxon will work and integrate with those tools seamlessly.
That being said, and I fully understand where the OP is coming from, there are several companies not interested in managing several solutions and all the complexity that comes with the infrastructure, deployment, maintenance, upgrades, user facing clients, authn/authz, permissions... Polyaxon provides the right abstractions for covering the experimentation and the automation phase.
It seems that the "all-in" platforms are too "rigid", and all of the point solutions for the things you mentioned aren't proven enough.
That is magical thinking. I prefer best of breed solutions that integrate nicely with other best of breed solutions every day. That way if a tool doesn't suit you tomorrow, you can relatively easily swap it out for something better
https://github.com/Netflix/metaflow-tools/tree/master/aws/cl...
We're also strictly based on Git and other Open Source formats and tools so connecting with other tools you use like Colab for IDEs or Jenkins/Kubeflow for training is super straightforward (we have examples for some)
Open-source:
* https://github.com/logicalclocks/hopsworks
Managed platform on AWS/Azure (with elastic compute/storage, integration with managed K8s, LDAP/AD):
But unfortunately, as you say, most of the pluggable tools are not very good and/or not mature enough.
Here's our attempt at model storage, experiment tracking, and software heritage: https://replicate.ai/
For interactive dev environments, Colab, Deepnote, and Streamlit are all great.
For deploying to production, Cortex mentioned in the post is great.
All are a work in progress, but I think we'll soon have a really powerful ecosystem of tools.
You are very likely going to be using a handful of tools that cover to full gamut of needs. This is heavily discussed in the blog post about an MLOps Canonical Stack and many of the tools being suggested below are included.
https://towardsdatascience.com/rise-of-the-canonical-stack-i...
As a new project we are still figuring out some of major topics you described.
In short, we built a data science pipeline tool that should fit well with existing workflows in machine learning and data science. We chose to embrace and integrate open source projects to create a simple and seamless experience with best in breed solutions for various tasks.
We are particularly happy with our deep integration of JupyterLab building on the excellent Jupyter Enterpise Gateway project from IBM (Codait) for connecting kernels directly to your pipelines. For scheduling we build on top of Celery combined with containerization primitives. For stable and well defined dependency management we built a small environment abstraction on top of Docker. It works really well in our experience!
Feel free to check out the project on https://github.com/orchest/orchest
Self hosting should be as easy as running about two lines of code.
This could be its own “Ask HN” thread.
* data stores: automatic download/upload to/from AWS S3, Azure Blob, Google Cloud Storage, OpenStack Swift or stores that implements S3-like interface
* interactive environments: we do have notebook hosting with automatic orchestration
* training: history, comparisons, parameters, hyperparameter tuning with Optuna, Hyperopt or custom optimizer (https://github.com/valohai/optimo); additionally visualizations about training progress and hardware resource monitoring
* serving for production: our deployments allow you to build, push, manage and monitor HTTP/S based services on Kubernetes clusters (https://docs.valohai.com/core-concepts/deployments/) but you can just as easily download your model and deploy it yourself as your use-case requires
* a permission system: we have organization management with teams and such, but your mileage may vary depending how fine grained control you need
* software heritage: all runs are containerized and how they was ran is recorded so everything is reproducible if the base image and data exist at the original source, we also keep track of data heritage (what files X were used to produce these files Y https://valohai.com/patch-notes/2019-09-03/)
* labeling system/UI: full web UI, command line client and a REST API (https://docs.valohai.com/valohai-api/) but no labeling tools though
Essentially your whole machine learning pipeline under one roof; from data preprocessing and training to deployment and monitoring. Also, we are technically agnostic, you can just as easily run Python/Julia/C++ or Unity engine to generate synthetic datasets (https://www.youtube.com/watch?v=QxMuWuk_W10)
not self-service or free though; our technical support team handles all the setup and maintenance
let me know if you have questions about Valohai or MLOps (https://valohai.com/mlops/) in general, I've seen quite a lot of projects and pipelines as I work at Valohai as an ML engineer helping our customers to setup end-to-end ML pipelines
Kedro has the simplest, leanest, functional-programming inspired pipeline definition and also spits out AirFlow and other formats readily + comes in with an integrated visualisation framework which is stunning & effective.
I've played with cortex before, and it is easy to use, but I am still questionable if automating kubernetes deployments through an easy code interface, without much kubernetes know-how, is safe.
In my experience, even when you have a tool automating a lot of kubernetes for you, you will still run into trouble that will be best handled if you are familiar with kubernetes. I'm not sure what debugging utilities cortex has, but I think the ultimate solution to this problem will be a tool that truly allows users to not think about the fact their deployments are running on kubernetes at all.
I'm also interested in the similarities of Cortex and Seldon-core. Of course, seldon-core does not automate infra provisioning, but based on my previous point, I think many teams are better off being more hands on with this infra.
Lastly, there is a third tool missing from the mix - monitoring. I think cortex offers some tools in this area, but I wish they would make a part two showing how the monitoring functionality they offer can integrate into a retraining pipeline within metaflow. This post shows you how to get started, but it doesn't show you how to maintain applications long term.
https://docs.metaflow.org/introduction/what-is-metaflow#shou...
I was curious to see what advantage Metaflow offered over TFX.
For example, Metaflow doesnt support kubernetes today - https://github.com/Netflix/metaflow/issues/16
so ultimately the scale up story in most of these management tools is iffy.
I previously asked about kubeflow here - https://news.ycombinator.com/item?id=24808090 . Seems people think its pretty "horrendous". It seems most of these tools assume a very specialised devops team who will work around the ml tool...rather than the ml tool making this easy.
Most people already are running k8s of some kind. I beginning to see k8s increasingly as an invariant. Plus not all of us are on AWS.
Secondly, AWS Batch is only applicable for metaflow for ml training.
Based on this thread, the comparison should include
* metaflow (model training on AWS Batch) * polyaxon (model training on kubernetes) * pachyderm (experimentation) * hopsworks (model training/serving/ and more, mostly on kubernetes) * cortex (model serving on kubernetes) * seldon-core (model serving and monitoring on kubernetes)
and likely more that I missed.
I can see why it would be so hard to put together this comparison.
Even with all these tools, there is still a lot of manual work for data scientists or DevOps engineers the data scientists pass their models off to.
It also seems there is yet to be a fully open source DevOps stack. Most companies still build custom software to glue together manual processes (like integrations between different tools for training, deploying, monitoring, etc). This could be one factor why more comparisons of these tools and stack discussions have not been more popular - they can't share them yet.