I see far less reasons to version data, in fact, I find reasons against versioning data and storing them in diffs.
I see far less reasons to version data, in fact, I find reasons against versioning data and storing them in diffs.
The wheels of data versioning just get reinvented over and over and over again, with all sorts of slightly different tools. Most of the job of "boring CRUD app development" is data version management and some of the "joy" is how every database you ever encounter is often its own little snowflake with respect to how it versions its data.
There have been times I've pined for being able to just store it all in git and reduce things to a single paradigm. That said, I'd never actually want to teach business analysts or accountants how to use git (and would probably spend nearly as much time building custom CRUD apps against git as against any other sort of database). There are times though where I have thought for backend work "if I could just checkout the database at the right git tag instead needing to write this five table join SQL statement with these eighteen differently named timestamp fields that need to be sorted in four different ways…".
Reasons to version data are plenty and most of the data versioning in the world is ad hoc and/or operationally incompatible/inconsistent across systems. (Ever had to ETL SharePoint lists and its CVC-based versioning with a timestamp based data table? Such "fun".) I don't think git is necessarily the savior here, though there remains some appeal in "I can use the same systems I use for code" two birds with one stone. Relatedly, content-addressed storage and/or merkle trees are a growing tool for Enterprise and do look a lot like a git repository and sometimes you also have the feeling like if you are already using git why build your own merkle tree store when git gives you a swiss army knife tool kit on top of that merkle tree store.
This is versioning
Well unless fraud is the goal.
People in ML ops use git because they aren't very sophisticated with programming professionally and they have git available to them and they haven't run into the consequences of using it to store large binary blobs, namely that it becomes impossible to live with eventually and wastes a huge amount of time and space.
ML didn't invent the need for large artifacts that can't be versioned in source control but must be versioned with it, but they don't know that because they are new to professional programming and aren't familiar with how it's done.
Mlops people are very aware of tools that are more suited for the job... even too aware in fact. The entire field is full of tools, databases, etc to the point where it's hard make sense of it. So your comment is a bit weird to me
It's not perfect, and still feels like a bit of a hack compared to something like p4 for the context I uses LFS in (game dev), but it works, and doesn't require expensive custom licenses when teams grow beyond an arbitrary number like 3 or 5.
As an example, a Unity game repo reduced in size by 41% using our block-level deduplication vs Git LFS. Raw repo was 48.9GB, Git LFS was 48.2GB, and with XetHub was 28.7GB.
Why do you think using a Git-based solution is a hack compared to p4? What part of the p4 workflow feels more natural to you?
Some of these have workarounds and hacks for more experienced users. I'm not about to run around teaching people the intricacies of arcane git incantations, while p4 functions, by default, how you'd want to. The programming side is better on git though, yeah.
We're working on perforce-style locking on XetHub, and I believe git already supports things like only cloning the latest version of files. Cloning the full repo without "smudging" (pulling in binary file contents) is already possible, and cloning while smudging a subset is on our roadmap. We're definitely on a path to making git UX for dealing with large binary files as easy as perforce, and there are lots of advantages to keeping a git-based workflow for teams that already work with git.