Kedro – Creating reproducible, maintainable and modular data science code
github.com
github.com
Pros:
* Forces data scientists to produce an end product that is not poorly organized Jupyter notebooks.
* Data Catalog is good for well structured systems
* Pipeline visualization stack is great (Kedro viz)
* Config options are pretty good
* Seems stable. Dev team is pretty good on this and avoiding breaking changes.
Cons:
* Data catalog is kind of bad for any non structured setup with flat file data with manual file movement (which is bad to begin with but sometimes that’s life)
* Productivity of making brand new data science code seems to drop when data scientists leave notebooks and
* Most of the time I get brought into a client context because the client doesn’t know anything about data scientist. A lot of data scientists, from both parties, come from academic backgrounds and aren’t great at code. The nice thing about notebooks is that they run. Kedro requires you to create pipeline and node objects to wrap around your code before it runs. It requires some familiarity with Kedro to understand, run, or modify. This makes it seem like a bad idea to dump on a novice client. If the data scientist on their side inheriting it doesn’t really get it, or leaves, there’s unlikely to be enough internal knowledge to maintain it. I try to avoid pushing any new tech stacks on my clients where I can for this reason.
So… I like it but don’t love it for consulting work, which is ironic.
https://medium.com/hacking-talent/production-code-for-data-s...
https://medium.com/google-cloud/migrate-kedro-pipeline-on-ve...
The kedro project is still running nearly 3 years after it started, the notebooks are long forgotten.
Notebooks are fine for single contributors if that is what they are comfortable with. If that is what they are comfortable with its probably because they have not experienced engagements like you are bringing to them that require them to collaborate as a larger team. If you plan to effectively run projects that last longer than a few weeks with more than one data scientist, I'd really challenge them to lean into kedro. The long term productivity of the project will greatly benefit from it. You will be showing the team a more sustainable way of creating pipelines that will lead to continued success after you leave. This leaves a better name for you than a quick cash grab that gets long forgotten in my opinion.
A few years ago, I started working as a data scientist at a big financial firm and reviewed all workflow orchestrator available tools (including Kedro). I didn't like that all of them forced me to re-write my Jupyter code into their frameworks (they're supposed to make me more productive, not less).
True, notebooks have their issues but they can be fixed (I don't buy that "Jupyter is only for prototyping argument"). So, long story short, I started a project with a friend that makes us more productive by fixing the problems that notebooks' problems. https://github.com/ploomber/ploomber
https://github.com/orchest/orchest
What do you think of Orchest? Given your unique perspective we'd love to hear what you think of it.
We've been having success with agencies, especially when the hand-off needs to be something the clients can easily run with.
I'd be very curious to hear if anyone has used it and what impressions are on it's value. I've only heard some businessy talking points about it, and am hopefully understandably skeptical. If someone not from McKinsey had used it, I'd love to hear impressions.
As far as impressions of it as a tool go it’s really just an opinionated way to build a pipeline and structure a project. Which is pretty useful when you have many. If you’re doing something that you know will be one off it is a bit overkill. Like using airflow when a makefile will do. There are some pretty nice plug-ins though if you want to try and have an opinionated ecosystem.
Donation is the normal term for proejcts joining the foundation - in doing so you establish a steering committee where members are entitled to voting rights. To graduate as an incubation project, 5 organisations need to join your board and the project is thus the priorities of the project will no longer be driven by just one organisation.
In the short term there is still a full time internal team staffed and maintaining the project, but excitingly we now have a mechanism for new collaborators to properly come onboard.
Unlike similar projects kedro is just a python framework. It lets me build, deploy, and orchestrate however works best for my team.
E.g. if your default might be a series of Python files connected by a master script and some config files, Kedro organizes them into an explicit Pipeline object. Think loading, cleaning, feature gen, model training, mode predictions, etc.
I mean I guess it's "opinionated" but having worked in this field for some time, ML code and pipelines are often laughably immature, stuck with spit, glue and sellotape. Kedro forces people into a bit of a straight jacket for sure, but man does it make things more readable.
Plus points, it's very extensible so if there's a design element you're missing or one of its opinions you just can't abide by, then you can extend it. I think it deserves more exposure.
Data science is many researchers first foray into programming in any language and, speaking from my own path, was a relatively self-directed discipline in the early days.
I often longed for a 'standard approach' but many of the projects that suggested them were packed full of boilerplate that made code even harder to read and write, especially if your background wasn't heavily into OOP.
In my opinion, the main advantages of Kedro are:
* standardize the development of ML pipelines
* easy to extend with plugins
* easy to share pipelines components between different projects
* data catalog (also easy to create custom datasets)
* kedro-viz
Today, we have a cookiecutter template with a custom Kedro pipeline that is used by new projects. It has the most common steps of ML pipelines developed here in Wildlife. Therefore, speeding up the development of new models. Moreover, it has the Kedro MLFlow plugin for automatic experiment tracking. Finally, we also have a catalog of pipeline steps that can be easily added to a Kedro project.
https://docs.github.com/en/get-started/exploring-projects-on...
This looks really interesting!
The 'finished article' in many cases should be deployed in production in one of those tools which provide specialised bells and whistles like scheduling, monitoring and observability.
Regarding MLFlow, there is also a slight overlap in terms of experimentation, but not things like model serving. Kedro has a mechanism to track experiments, but it's more designed to give users with zero infrastructure something for free out of the box. It's been built in a way that it can be repurposed for more dedicated experiment tracking tools - the folks at neptune.ai built their own plug-in for this purpose: https://docs.neptune.ai/integrations-and-supported-tools/aut...
Kedro and Metaflow make it easier to develop robust ML projects where orchestration plays an important role but it is not everything. They are two separate projects, so the way how they approach the problem differs greatly in details.