Pachyderm Raises $10M to Bring Data Provenance to the Enterprise
pachyderm.io
pachyderm.io
Pachyderm's pipelines are also smart enough to know what data has changed and what hasn't and only process the incremental data "diffs" as needed. If your pipeline is just one giant reduce or training job that's can't be broken up at all, then this isn't valuable, but most workloads include lots of Map steps where only processing diffs can be incredibly powerful
FYI, it's ridiculously easy to get going playing with Pachyderm if you just want to check it out. You can run it on Minikube.
Thanks for the tip. I just started down the k8s path from bare metal cluster and will try this.
Why would I use Pachyderm?
Each pipeline can have separate resource requirements (e.g. GPUs, lots of memory, etc) and gets scheduled by Kubernetes.
Finally, Pachyderm is versioning all of the intermediate steps in your data pipeline so if a downstream step fails, you don't have to restart from scratch, you can pick up right where it left off.
Most people don't need "industrial strength" till you hit a certain scale. They tend to optimize for ease of use/simplicity. It's one thing if you don't have to change their workflow, it's another if you have to not only have people change their workflow but also teach something new.
Convenience matters more. There's a pain threshold of "new tool" vs "this costs me x amount of time".
How are you guys overcoming this? Even though we're also in the infra space, I've never seen you guys out in the wild. Where would I bump in to you and what is the scale I would want to add the complexity of k8s + pachyderm + whatever other deps you guys have over just using an S3 bucket?
Another question: Why hasn't AWS just packaged this up and offered it as an extension of their k8s service? How are you guys going to overcome that?
Usually I see "hybrid cloud" or "on prem" as the response. If it is on prem, are you guys relying on the presence of k8s at customers? Do you guys use something like gravitational?
Not what I was describing. We’ve been using the setup for a while with multiple projects and for the data science end of a clinical trial.
I like building new setups incrementally out of tools I already know, so the question is obviously why is this new tool a radical improvement
Point taken about restarts, but I usually want to see why a thing fails before retrying.
It seems like Pachyderm mainly offers better granularity, especially for the data management.
For me, there's various properties I like about pachyderm (though they're missing some key workflows I'd like). Looking at some output data, I can see exactly what versions of everything was run against what data, and the data is all version controlled. This should also allow experimentation by individual people and easier approval of updates. I work with multiple teams with many components, where each one may impact others - being able to look at some final result that's weird and go back to figure out what's going on is hard. Tools like pachyderm help with that.
But if you need an always-on, always-ready pipeline that can scale horizontally on an as-needed basis, and perhaps mixes many different modules/languages, than Pachyderm might be a good fit.