ML Experiments Management with Git
github.com
github.com
It is a much simpler and much more magical piece of software that truly expanded how I think about writing, exploring, and experimenting with code. Even if you never use it, you probably would really enjoy reading the blog posts the author wrote about the design of the tool https://amakelov.github.io/blog/pl/
https://www.dolthub.com/blog/2021-06-04-flyte-dolt-plugin/
https://cloud.google.com/vertex-ai/docs/pipelines/configure-...
https://docs.aws.amazon.com/sagemaker/latest/dg/pipelines-ca... https://docs.aws.amazon.com/sagemaker/latest/dg/querying-lin...
I 100% agree that managing large datasets by moving them around is not practical, and definitely not in LFS/DVC-style. There should be a level of indirection if reproducibility is needed (pointers are versioned to files, not the data directly, data should be staying in the cloud).
Here, I would love to one more time mention some other cool features that DVC has. E.g. `dvc exp` set of commands where it is creating custom git refs to snapshot experiments, of DVCLive logger that helps capturing metrics, plots, etc. And also VS Code extension [1] that provides quite cool experience for experiments workflow inside VS Code.
Point here is that for DVC the ability to capture some large files and directories (that do not fit into Git) was always a low level mechanism to support higher level scenarios (e.g. you need to save a model somewhere as an output of an experiment).
[1] https://marketplace.visualstudio.com/items?itemName=Iterativ...
I am not sure I understand that correctly. Are you saying that LFS/DVC manage the data suboptimally because they do not use some kind of pointer?
I only have some experience with DataLad[0], not with DVC or LFS. DataLad is built on git-annex, which does a pointer indirection through symlinks or pointer files in git. You basically manage the directory structure in git and can "get" and "drop" specific files as you need them. git-annex keeps track of where (e.g. on what (remote) system, which could be anything from a http server over s3 to a nextcloud via webdav and more) the data is and how it can be fetched. I always thought DVC did something similar.
Instead of trying to store data and ML models in one place (like S3) and code, models & documentation in another place (like GitHub), we are scaling Git so you can version everything in a single system. You can just use git and you don't need to learn a new tool or set of commands.
This way, you can start with a simple experiment tracking approach of folders inside the same branch and then evolve gradually to multiple branches with long running experiments.
We're about to release our Github integration, so data & ML teams can take advantage of this inside their existing Github repos. If anyone wants a tour or wants to chat, my email's in my HN profile.
If you're curious about our tech:
- Here's an example 3.3 TB Git repo: https://xethub.com/XetHub/RedPajama-Data-1T
- We wrote a paper on our solution to scale git to 100 terabytes: https://about.xethub.com/blog/git-is-for-data-published-in-c...
- We created a Rust library to mount large repos to machines with limited storage space: https://news.ycombinator.com/item?id=37573679
A relational database like Dolt seems like a better fit. You want to be able to query by experiment name, date, test results, and other metadata.
Let me know if I'm doing it wrong! What's your use case?
Out of the box you can't. We're taking a different approach at work (XetHub). GitHub sees pointer files but we embed rendered views and diffs (supporting more file types incrementally) using a Github app from our service.
https://www.pachyderm.com/blog/data-versioning-comparing-dvc... https://www.dolthub.com/blog/2022-04-27-data-version-control...
As far as I know, DVC is better than Pachyderm for small datasets, but Pachyderm scales way better
I think it's a shame it took Airflow and similar several more years to realize "each step is a docker container" is the right way to build a dag. It's not clear to me why Pachyderm was left behind while Prefect and Dagster became serious contenders, and Airflow/Astronomer started recommending everyone use it just like Pachyderm (container per step).
Every time I see DVC mentioned I always feel like the idea was so close (and perhaps right in intuition to use git for everything) but the execution had just enough friction that I looked elsewhere. Small DX improvements really do cascade pretty far.
Also I always thought the idea of using Git branches to track experiments was a bad idea. I would never want to only have one experiment "active" at a time. Even if I'm only running one process at a time, I still want to be able to look at outputs and such all side-by-side. Maybe there's some magic tooling they created that makes it workable.
But for a data project it would be a big pain to have separate worktrees just to work around what IMO is a usage anti-pattern to begin with!
[1] https://iterative.ai/blog/experiment-refs
[2] https://marketplace.visualstudio.com/items?itemName=Iterativ...
I'll have to take a look at this. Most/all of my projects use small or medium scale data, and I consider DVC indispensable for tracking data therein. I wouldn't mind having a good system for tracking experiment results, although admittedly I find that a spreadsheet or text file does a pretty good job for what I need to do.
It's neat for research as it stores the data on scientific data repositories like Zenodo and you get DOIs.
HF seems like "GitHub for models and datasets", it has a cool brand and everyone in ML uses it in some capacity. But when it comes to _company_ needs like private datasets/models, experiment tracking, CI integrations, etc. it seems WandB is a superset of HF.
HF has an enterprise offering, but it seems to be de-prioritized, and I think you'd still need WandB or MLFlow for experiment tracking?