HNHacker News
TopNewBestAskShowJobs

seeravikiran

10 karma · joined December 3, 2019

submissionscomments
seeravikiran··on Machine Learning at CNN
What were some of the pain points you face(d) - looking back at your Metaflow adoption? Disclaimer: I work in Netflix ML Platform that helped open-source Metaflow originally.
seeravikiran··on Metaflow, Netflix's Python framework for data science, is now open source
Happy to help either through our gitter chat or help@metaflow.org.
seeravikiran··on Metaflow, Netflix's Python framework for data science, is now open source
Thanks for reporting it. We ll fix it. Sorry for the inconvenience.
seeravikiran··on Metaflow, Netflix's Python framework for data science, is now open source
Thanks. Let us know how you like the prototyping -> scaling out & up journey.
seeravikiran··on Metaflow, Netflix's Python framework for data science, is now open source
I wouldn't exactly say that. Jupyter notebooks don't have an easy way to represent an arbitrary DAG. The flow is more linear and narrative like. That said, we do expect metaflow (with client API) to play very well with notebooks to support a narrative from a business use-case pov; which might be the end-goal of most ML workloads (hopefully). I would like to think of metaflow, as your workflow construct - hopefully making your life simpler with ML workloads when involving interactions with existing pieces of infrastructure (infra pieces - storage, compute, notebooks or other UI, http service hosting etc.; concepts - collaboration, versioning, archiving, dependency management)
seeravikiran··on Metaflow, Netflix's Python framework for data science, is now open source
Thanks for sharing the context. Hopefully we can have a (fast) follow up with Kube integration depending on demand.
seeravikiran··on Metaflow, Netflix's Python framework for data science, is now open source
Yes - you should be able to use dask the way you say.

Your first part of the understanding matches my expectation too. Dask single box parallelism achieved by multi processing - akin to parallel map. And distributed compute is achieved by shipping the work to remote substrates.

For your second comment - we leverage pickle mostly to keep it easy and simple for basic types. For small dataframes we just pickle for simplicity. For larger dataframes we rely on users to directly store the data (probably encoded as parquet) and just pickle the path instead of the whole dataframe.

seeravikiran··on Metaflow, Netflix's Python framework for data science, is now open source
I guess (?) - minus the input spec being not YAML but more language native (pythonic for e.g.)
seeravikiran··on Metaflow, Netflix's Python framework for data science, is now open source
Yes - that’s our thinking too. Compilers finding your typos for variable names seems helpful for user productivity.
seeravikiran··on Metaflow, Netflix's Python framework for data science, is now open source
Thanks for pinging on this.

re: Kubeflow - imho it is quite coupled to Kubernetes. We don’t intend to be tied to a specific compute substrate even though the first launch is with AWS. We do follow a plugin architecture - so I’m hoping Kube happens sometime.

re: Flyte - I’m less informed on this but happy to educate myself and get back.

seeravikiran··on Metaflow, Netflix's Python framework for data science, is now open source
I would also add - dependency management (certain degree of reproducibility) as a first class feature leveraging conda.
seeravikiran··on Metaflow, Netflix's Python framework for data science, is now open source
Our hope with metaflow is to make the transition to production schedulers like Airflow (and perhaps similar technologies) seamless once you write the DAG via the FlowSpec. The user doesn’t have to care about the conversion to YAML etc. So I would say metaflow works in tandem with existing schedulers.
seeravikiran··on Metaflow, Netflix's Python framework for data science, is now open source
With many objects under the same S3 bucket - say for a flow or a run (with many tasks).
seeravikiran··on Metaflow, Netflix's Python framework for data science, is now open source
1. Metaflow should best help when there is an element of collaboration - so small to medium team of data scientists. Collaborating with your self is also another scenario when Metaflow can be useful since it takes care of versioning and archiving various artifacts.

2. Keeping the language pythonic, without any additional need to learn a DSL has definitely been key to Metaflow's adoption internally. That said, this is something we are open to hearing back, esp. with this OSS launch.

3. Yes - definitely think so. Personally my favorite is the local prototyping experience part; when everything can fit in memory and is blazing fast. There is an also an open issue for fast-data access, which you can upvote if interested in seeing it open-sourced.

4. We don't think there is an exact equivalent as well. :)