Storing large binary files in Git repositories (2015)
blog.deveo.com
blog.deveo.com
Also, it's worth noting that git-lfs stores 2 copies of each large file on local disk. git-annex stores 1 copy. Which is a pretty big plus the article left out.
Also, that ks for git-annex! I use it for hundreds of gigabytes of files, and find the one-copy-per-file support essential for how I use it.
Sorry for possible off-topic, but I wonder if anyone knows a way to convert/upgrade older (created an year ago) git-annex repo to a newer v6 one?
I haven't tried this yet myself.
This is what I'm doing after a disastrous excursion into git-fat.
But really, if git's maintainers don't want to store binaries, why shoehorn it in? Use the right tool for the job.
On the plus side, it's stupidly simple. On the down side, it's stupidly simple. The readme explains how it works.
How? Well, the framework stores all assets inside /assets, which normally is a git submodule. Jenkins works pretty fine with this construction, and for developers there are a couple self written PHP shellscripts that execute a shallow clone outside the repository and then do a bit of rsync magic to sync it all.
this sounds like a problem for teams blindly following orders from clueless managers.
> if you can't get a textual diff, or if a textual diff is meaningless, what benefit do you even get having the file on git?
A nice property of checking out whole repository using only single program. All assets are downloaded at once and always proper version of code comes with proper version of binary assets.Using multiple different systems to get myself "up to date" is a massive pain - I should be able to run one command and then have a wholly buildable set of files.
yeah, you don't. :D
[EDIT:] Perhaps we were both unclear! By "code", I mean stuff that can be read and edited by normal humans, which by some automatic mechanism can be used to control computers. E.g., if you had a log of photoshop actions that could be replayed to produce a given image, that would be code, and it would be suitable for storage in git. In general, the image itself ought to be elsewhere.
To give another example - diff is a pretty common action to see changes between two specific versions of the code. Now if I'd treat both code and non-code assets as equivalent (because they are human produced, etc.), then I'd expect the tool which manages them to also show me diff between versions of binary files. But simple byte-by-byte diff* would be pointless, exactly because non-text files need to be decoded before they can be used. So if I'd treat them as equivalent, I'd expect graphical binary assets to show me the picture with differences between two versions; I'd expect video binary file to show me which time points are different in the recording, etc. etc.
* Well, yes, not exactly byte-by-byte, since it needs to be encoding-aware. But still, much simpler than parsing any random binary file format to show reasonable diff.