Oxen.ai: Fast Unstructured Data Version Control
github.com
github.com
edit: Looks like e.g. the Unreal engine has support for Perforce and SVN (https://docs.unrealengine.com/5.0/en-US/collaboration-and-ve...), where they make use of locking a file when editing it to avoid merge issues with binary files. I can imagine that frequently goes wrong.
There are two states, in woodworking for example, the raw material and the final product. The mental model of "cutting longer" or rolling back does not exist.
The idea of a filesystem level understanding of "undo' just has not been communicated or explored.
While not Oxen related - TimeMachine seems to be the closest conceptual understanding. Snapshotting the drive, deltas, etc are generally not considered in my experience. I would be curious how larger or mature groups manage versioning.
Curious, do you know the solutions here? I've looked into integrating Git LFS-like things into Git and it felt like Git doesn't give you many tools. Smudging, as hacky as it is, feels one of a very small pool of options you have. I want to say Git Annex does something slightly different.
Regardless Smudging felt fine due to how difficult Git made it. Are there better ways in your mind?
In later versions, LFS doesn't play nice with misconfigured internal networks based on Microsoft's Azure / Active Directory. The LFS project dropped support for NTLM (for good reasons) but it's difficult to convince IT to switch to Kerberos. Our team have been trying to do that for several weeks now. So the result is that if we want to use LFS we are stuck with old versions of LFS which still supported NTLM, which necessitate using old versions of GIT which supported this old version of LFS, else we run into all sorts of weird error messages. Also the same combination of old GIT + old LFS works perfectly on some machines, and fails with some weird error messages on others. Possibly due to another IT misconfiguration or something.
I've been wanting for a while now to check if I can fork LFS and fix this, unfortunately my knowledge of golang is very sketchy. Hopefully we'll manage to convince IT to drop NTLM support from the on-premise Azure.
What year is this? ;)
On a more serious note - ntlm, really? And with azure ad, not a regular (old) on-premise domain?
I know it can be hard to drag windows setups into the future (hello WINS!) - But why keep ntlm auth in a post windows 7 world? Legacy systems without ldap/radius support? Genuinely curious (and a little horrified).
Apparently, it's related to the few win 7 machines that are still present somewhere in the organization for various reasons. All of them are airgapped (I hope) so there's no real reason to keep NTLM, but nobody has yet made the decision.
Working on getting to the guy in charge to authorize it. Sigh.
If one file doesn't resolve, for any reason, git-lfs will block the entire pull. No way to say "get whatever you could, and the fact that aws is not serving that file is not a blocker"
Git clone does not pull the entire lfs history of an image. Lacking the entire history, lfs will fail consistently on future actions.
A few starting points from personal experience would be
- Ensuring git-aware code editors don't try and mangle blob histories when they encounter them - Modern, high-performance blob downloads/version-resolution - Human-readable/debuggable meta-files, where it's easy to recover from bad states (instead of an opaque empty file with a GUID...?)
You'd also need for the blob encoder to be exceptionally stable - if minor changes to the input lead to significant changes to output, you're going to have a lot of trouble.
Something like Google's Courgette (binary patch generation) could help to simplify things, but an old thread mentioned a patent suit.
It's an early prototype to gather feedback around the ergonomics of the interface and the manifest format, and I would love to learn if people might find it useful enough for the use cases where LFS falls short.
Love to have you try XetHub and give us your feedback!
From your readme it seems like the oxen repo and software project repo are not as closely coupled as in dvc? It seemed like in the current state of oxen, you could do something similar with make files and oxen tracking?
Oxen seems really good for longer lived data and computational science projects, where dvc seems more oriented just at analysis projects. I have a project that I want to try it out on :)
Any other features you would find useful or a dealbreaker?
https://github.com/Oxen-AI/oxen-release/blob/main/Performanc...
oxen push origin main # ~308.98 secs
Where does that push to? Does this benchmark really just measure how well-provisioned various different VC-funded websites currently are?I think a proper benchmark here would be install the server parts of Oxen, Git-LFS, etc on the same machine, and then time how long it takes to commit and push the same dataset from some other machine.
Although of course given that we live in an age where people expect to upload their immense datasets to the cloud for some reason, a "proper" benchmark might not be a relevant one. I'm not sure what a really good benchmark of that would be.
Fundamentally even adding and committing data locally is slower, even before the push. But I agree the remote matters too.
I'd nowhere near the same performance with oxen. The analysis is very biased to help Oxen. I wish people had more integrity before trying so hard to push a half-baked product into the market.
@oxen_ao -> @oxen_ai
We did some benchmarking here: https://github.com/Oxen-AI/oxen-release/blob/main/Performanc...
~TLDR~ 200k+ images from the CelebA dataset take ~6 minutes to add, commit, push into Oxen. Same dataset takes ~3 hours for DVC.
Would be interesting to see rclone.org (and maybe s4cmd) (should be faster access to s3?) - and also mercurial and bitkeeper.org - both should behave a little better than git on large files?
How does this compare to lakeFS?