Announcing Git Large File Storage
github.com
github.com
Compare "Git Large File Storage"'s file spec:
version https://git-lfs.github.com/spec/v1
oid sha256:4d7a214614ab2935c943f9e0ff69d22eadbb8f32b1258daaa5e2ca24d17e2393
size 12345
And bigstore's: bigstore
sha256
96e31e44688cee1b0a56922aff173f7fd900440f
Bigstore has the added benefit of keeping track of file upload / download history _entirely in Git_, using Git notes (an otherwise not-so-useful feature). Additionally, Bigstore is also _not_ tied to any specific service. There are built-in hooks to Amazon S3, Google Cloud Storage, and Rackspace.Congrats to GitHub, but this leaves a sour taste in my mouth. FWIW, contributions are still welcome! And I hope there is still a future for bigstore.
I haven't settled on one for my own use, but I'll compare features of bigstore and git-media before I do. Thanks for making your project available!
git-bigstore is awesome too though!
Git LFS isn't tied to any specific service either. You can install our reference server somewhere, and start using it with your GitHub (or any host really) repositories without having to sit in our wait list or pay us a dime. Though our reference server isn't really production ready, so I wouldn't advise that for real work just yet :)
I'm interested to see how this scales. My feeling when I looked at it was that it was not sufficiently scalable without improving the smudge/clean filter interface. I mentioned this to the git devs at the time and even tried to develop a patch, but AFAICS, nothing yet.
It would provide an easy way for people to host their git-annex repos entirely on GitHub.
lfs-get SHA256 > file
lfs-store SHA256 < file
lfs-remove SHA256 (optional)
lfs-check SHA256 # exit 0 or 1, or some special code if github is not available
Presumably the right way would be to use their http api, but these 4 commands seem generally useful to have anyway.As for git-lfs relative to git-fat: (1) the Go implementation is probably sensible because Python startup time is very slow, (2) git-lfs needs server-side support so administration and security is more complicated, (3) git-lfs appears to be quite opinionated about when files are transferred and inflated in the working tree. The last point may severely limit ability to work offline/on slow networks and may cause interactive response time to be unacceptable. Some details of the implementation are different and I'd be curious to see performance comparisons among all of our tools.
Re the python startup time, this is particularly important for smudge/clean filters because git execs the command once per file that's being checked out (for example). I suppose even go/haskell would be a little too slow starting when checking out something like the 100k file repos some git-annex users have. ;)
The one major obvious drawback for it being fixed in git is that it's not backwards compatible with old clients though. Doing it in go is probably a good improvement over the existing solutions of git-media and git-fat but I don't think is the final one.
Funny enough, although I had thought this since I started working with git-fat, I only recently admitted it[1]. Perhaps if I had admitted it when I first started work on it then there's a chance they would have seen it! :-P
[1]https://github.com/cyaninc/git-fat/issues/41#issuecomment-88...
I also looked at git-annex, and I could see using it if it were just me on the project (or as a way of keeping fewer files on my laptop drive), but I was reluctant to add any more complexity to the source control process, since explaining how to use git-annex to the entire team was too big of a barrier.
I have been enjoying the simplicity of git-fat, but running git diff and especially git grep makes me think I should switch to something else.
As anyone who's worked on project with large binary files(the docs assume PSDs) you need to be able to lock unmergeable binary assets. Otherwise you get two people touching the same file and someone has to destroy their changes. That never makes anyone happy.
It's also unseen how good the disk performance is. These two areas are the reason why Perforce is still my go-to solution for large binary files.
http://mercurial.selenic.com/wiki/LargefilesExtension
http://mercurial.selenic.com/wiki/LockExtension
Like other people have noticed, you can have the good parts of a DVCS and the good parts of a CVCS. It doesn't have to be either-or:
https://blogs.janestreet.com/centralizing-distributed-versio...
http://bitquabit.com/post/unorthodocs-abandon-your-dvcs-and-...
Switching off largefiles requires rebuilding the repository which rebuilds the entire repository. Orchestrating the migration to a new repository for engineering department is also painful which is why we're stuck for the near future (for example ongoing support for a version that's built from largefiles repo with ongoing feature work in non-largefiles repo).
The tooling for mercurial tends to lack behind git's, likely due to git's enormous popularity - so I personally would recommend avoiding largefiles extension.
Some of the issues are likely due to our project which is >93k commits at around 1GB for the repo size - I think we have a commit in the history that has 40kb commit message - I believe the developer mixed up some streams/pipes for some reason or other while committing - I'm not sure why we got stuck with it in our history though I suspect this could be the cause of some issues. I can list a few of the issues I still see regularly, but since about a year or so back I've abandoned using the plugin in favor of doing everything manually on the command line, except for viewing history and resolving merge conflicts.
My setup: - Eclipse 4.4.2 (configured to run with 2GB memory) - though most all of these issues I've seen since 3.something - MercurialEclipse: http://mercurialeclipse.eclipselabs.org.codespot.com/hg.wiki... - OS X (10.10 and 10.9), mercurial installed through homebrew.
1. During some actions (I believe to be cloning/sharing projects, maybe elsewhere) the the config for mercurial in eclipse pops up thinking it can't find the hg binary, giving an error about bad location. Nothing has moved or changed on the system, simply focusing in the location field of the hg location and back out causes it to re-validate. I've seen this one recently, maybe a few days ago.
2. Some files can't retrieve history or show annotations. I suspect because the files have been through numerous edits over the years and the plugin runs out of memory or hits a timeout when trying to load everything.
3. Occasionally refreshing status throws error and pops up dialog informing me. When this happens the plugin becomes unusable until eclipse is restarted, but popups continue to show up when trying to do any sort of task (such as refreshing project)
4. Using mq patches cause the mercurial status of project to not update regularly (in Navigator pane). I've since moved away from using mq patches in favor of managing multiple local heads.
5. The merge view/window no longer automatically shows up when performing a merge in eclipse. It has to be manually opened, but used to automatically show up - I liken it to performing a search but the search results not automatically showing up unless it's part of your workspace layout already.
These are what I can recall off the top of my head, but I've also shied away from using the plugin since quite a while ago. In general the speed is sometimes frustrating to deal with, when wanting to view history of a file. It feels like it's gotten slower after some time that we've had largefiles extension turned on, though it's unfair to assume so because I'm biased.
The only exception could be lack of history. Largefiles make some tricks with storing the hash of file X in a .hglf/X file and folding it back in the right namespace is tricky and had some errors. Annotate and largefiles are conceptually incompatible; largefiles is intended for big and binary files where annotate wouldn't work no matter what.
My experience is that the main problem with largefiles is the problem it is trying to solve. It is not a good idea to store large files in a VCS - especially not in a DVCS. Storing large files in VCS is last resort. Given a situation where you have to / want to do it anyway, largefiles is a fine solution. It works quite well for us and without significant problems.
One other thing that's been a problem in the past with hg + largefiles (or only started happening since around when we turned on largefiles): Cloning a largefiles repo using "--uncompressed" flag would re-open all the closed named-branches in the newly cloned repo.
But these assumptions break down with binary assets. Changes are not small, they typically change entire files at once. They're also not diffable. As a result, conflicts are not rare, they're common, and impossible to resolve for both parties. That's why locking or "checking out" certain files is a needed feature, so changes are ordered strictly linearly -- a graph structure doesn't work.
I think that really is a problem with the diff tools, not the format itself. This is why both Mercurial and git allow you to pick special diff and merge tools per filename extension.
Depends. For some uses (e.g. a change in one corner of the file, encoded by the same tool with the same parameters), the diff of JPEG would be fine, differences will be restricted to the 8x8 pixel blocks that were touched by the change. Other changes (e.g., changing encoding quality, trimming a row of pixels off the edge of an image) would lead to more complex diffs.
You would never check a .o file into your SCM; an SCM, as the name implies, is for managing source files, not object files. A PSD or a DOCX is also a source file, despite being binary: they're representations of the work-in-progress itself, containing enough data to let you resume editing the document. A JPEG, on the other hand, is a compiled object—something you export from your image editor, not something you open and edit and save again. (Unless you're intentionally going for that recompressed-shitpost look, I suppose.)
When you pull down a source repo—a thing you edit—you should expect to get source. That applies to both your text/code assets, and your image/binary assets.
If, on the other hand, you need some assets to just sit there and be consumed by your project, then those aren't source, and so don't belong in your source repo. Those are likely dependencies, which can be resolved to (a triggered compilation of) the relevant source repo for those assets, or which can be resolved to a linear(!)-versioned binary package containing the compiled objects for those assets.
Which is all to say: PSDs, given an appropriate diff tool, could go in git. Final, "product" JPEGs, on the other hand? Those should be sitting in a gem (or equivalent), which was the tagged continuous-integration result of pointing a buildbot (exportbot?) at the relevant source repo full of PSDs. When you build your project, that gem gets pulled down (hopefully from your own private CDN) and suddenly you have some JPEGs, just like suddenly you have some native module .so files.
Brushes wouldn't be easy though. This could be handled by trivial svg or ps paths which would get rendered in the make process. This way even brush strokes could be diffable. I wounder if it would be possible to convert existing PSD files or Gimp project files into such a format.
How far do you take the decompression? For raster images, you'll probably have to decompress all the way to bitmap because the same image could have multiple completely different binary representations in a format like png.
How do you diff changes? If we have a raster image and one person changes one thing by a small amount, say increases brightness 1%, this could alter every pixel of the image! How would you detect that change and interleave it with something like a contrast adjustment of 1% that could also change every pixel? Sure, the merger would still have to choose which adjustment goes first if the changes aren't independent, but how would they know that's what changed? I.e. how would the diff tool know that the changes are "brightness +1%" and "contrast +1%" and not some other arbitrary number of adjustments?
If you change the top-left pixel of a PNG, for example, between the intra prediction and the DEFLATE compression, the new file can be totally different, and to reconstruct it you either hope the destination is using the exact same libpng with the exact same settings, or you have to find a space-efficient way to write down all the arbitrary encoding decisions the format allows.
If I were to make a contrived analogy, how do you know how to diff random arrays of bytes ? Where do you start, where do you stop ? How do you know that "\n" or "\r\n" is some kind of delimiter ? You put that knowledge in "diff" and in your editor, and git stores the raw array of bytes. It's the same with binary content: git doesn't care that you don't deal with UTF-8 characters, it doesn't care that it isn't bounded by newline characters.
If you take things this way, you start to understand that the "diff" tool you use must be appropriate to the content you have, and it's not the scm's business. Now, how exactly would a diff work for images, I have absolutely no idea.
template<typename T>
T diff(const T& a, const T& b)
{
auto sz = std::min(a.size(), b.size());
T img(sz, 1);
for (auto i = 0; i < sz; ++i) {
img[i] = std::fabs(a[i] - b[i]);
}
return img;
}
// usage
vector<float> a, b;
a.emplace_back(1.0); b.emplace_back(0.5);
a.emplace_back(1.0); b.emplace_back(0.0);
a.emplace_back(0.5); b.emplace_back(0.5);
auto img = diff(a, b);
write_exr("filename.exr", img);
The resulting image ends up with 0.0 black in pixels that are identical and non-zero values in the pixels that differ. When you look at it in an image viewer only the portions that differ will be visible.You often need to crank up the gain when the differences are small.
https://github.com/cameronmcefee/Image-Diff-View-Modes/commi...
Format-specific diffing and merging tools that are aware of git-lfs would probably help, and those can come later.
"Not invented here" much?
I would think bridging the gap in annex to track by file pattern would be easy but a lot of people might prefer not to know how to make annex go. So using simplicity as differentiator.
You could at least read the examples on the git-annex page[1] before passing judgement that the use cases are at all different (they're not). Instead of using a new 'lfs' command that ties you to GitHub, you use an 'annex' command (along with a few others).
Git-annex does just fine "keeping track of larger objects inside your git project efficiently", and is no more divorced from your normal project workflow than GitHub's lfs.
I'm sure this will be much easier to use for the end user like other github products and will "just work" out of the box.
The documentation lays out the workflow: https://help.github.com/articles/configuring-large-file-stor...
As does the website: https://git-lfs.github.com/ (see: "Getting Started")
git annex add large_file
git commit
And with mixed (haven't experimented with that yet), I'm pretty sure you could:
git annex add large_file
git add small_file
git commitThe protocol is open (https://github.com/github/git-lfs/blob/master/docs/api.md) and the client additions are open source. There is a reference server implementation at https://github.com/github/lfs-test-server.
edit: added protocol spec
This isn't about this particular instance (Github's LFS), but in general, a "reference implementation" isn't the same thing as having an open protocol.
Having a reference implementation without a proper specification means that any other implementations have to re-implement the existing reference implementation, including any bugs. The purpose of a specification is to outline undefined behavior as much as it is to outline defined behavior. That is, the specification says, "these are the portions of the program which you may not rely on".
We've seen this happen in some languages in which a particular implementation is either the de facto or de jure standard. Other compilers or interpreters end up having to mimic their bugs when it comes to things like arithmetic overflow/precision errors, because developers have come to rely on the language behaving one way, in the absence of any clear rules telling them otherwise[0].
[0] Not that developers may not rely on things that a specification explicitly tells them not to - there are plenty of examples of this too - but at least then it's possible to say determine either that a particular program will run on any standards-compliant implementation, or that it is implementation-specific.
Would have been nice to have had debate on existing solutions to weigh the pros/cons.
See my other comment here: https://news.ycombinator.com/item?id=9345242
What will be interesting is to see whether GitHub's implementation of LFS allows a "bring your own server" option. Right now the answer seems to be no -- the server knows about all the SHAs, and GitHub's server only supports their own storage endpoint. So you couldn't use, say, S3 to host your Git LFS files.
That's exactly what git-annex does. Except it can host on your own servers, or S3, or Tahoe-LAFS, or rsync.net, etc. And it's free software. And it supports multiple servers for the same repo, so you have redundancy.
Adding an S3 remote is just setting the AWS keys and running a single command: http://git-annex.branchable.com/tips/using_Amazon_S3/
Besides, I'm sure Joey Hess wouldn't refuse the help.
That's exactly how Mercurial's largefiles works too:
http://mercurial.selenic.com/wiki/LargefilesExtension#The_lo...
Also, you don't need any kind of special server. Any hg repo can turn into a largefiles store by just flipping the bit in the repo configuration.
Does this mean that with the free tier I can upload a 1GB file which can be downloaded at most once a month? Even a small 10MB file, which fits comfortably in a git repo, could be downloaded only 100 times a month. Maybe they meant 1TB bandwidth?
I don't think these facts are a problem. They create an open source tool, provide a location to try it out, and a service to pay to use it if you like it and don't want to host yourself. Seems like a fair offer.
git-annex has been great for my photo collection (which is strictly binary files). It lets me keep a partial checkout of photos on my laptop and desktop, while replicating the backup to multiple hosts around the internet.
At work we have a bunch of video themes that are partially XML and INI files and partially JPG and MP4. LFS would work great for us, except we don't use github (we don't have a need for it.) It looks like this is going to be very simple for that kind of workflow.
Just yesterday HN user dangero was looking for this exact sort of thing, large file support in git that didn't add too much complexity to the workflow: https://news.ycombinator.com/item?id=9330125
For example:
git config annex.largefiles "*.mp3 or *.mp4 or *.jpg or largerthan(100kb)"
git annex add .The main fundamental advantage (vs implementation quirks of git) I can see is that these files are only fetched on a git checkout. But (of course) this breaks offline support, and it requires additional user action.
Wouldn't it have been fairly easy to build exactly the same functionality into git itself? "Big" blobs aren't fetched until they are checked-out? This also has the advantage the definition of "big" could depend on your connectivity / disk space / whatever, rather than being set per-repo.
Git-annex renames binaries with their SHA256 hashes, puts them in a .git/annex/ dir, and replaces files in the working dir with symlinks. Git-LFS seems to use small metadata pointer files (SHA256 hash, file size, git-lfs version) instead of symlinks. Not sure whether the files reside in something like the .git/annex/ dir; I'm guessing the do or there wouldn't be those pointer files.
You can clone a repo without having to download the files.
With git-annex you can sync between non-bare and bare repositories without having a central server. Git-LFS seems to have a separate server for binaries. It looks like it may act like a git-annex special remote rather than git-annex's usage of synced/master branch.
Git-annex repos share information about the locations of annex files, how many repos contain a given file, etc. You can trust and un-trust repos. It doesn't look like Git-LFS offers this.
Git-LFS has a REST API. I'm using an old version of git-annex so I can't say if it does (I think it does).
Git-LFS is written in Go, git-annex is Haskell.
Git-LFS is a GitHub project. GitHub will offer object hosting.
Update: clarity, speling, added a bullet point.
I hesitate to say this means git-lfs is not distributed at all, but it seems significantly less distributed than git-annex, which can keep track of files that might be in Glacier, or on an offline drive, or a repo cloned on a nearby computer, and so can be used in a more peer-to-peer fashion when storing and retrieving the large files.
This is why git-annex uses a separate branch for location tracking information, which it can merge in a conflict-free manner.
I've been trying out git-fat on a large repository that has some binaries in it. Since we are already using ssh to access the server (and have the usual ssh-agent setup), it was easy to integrate it with scp.
Unfortunately, the performance for our workload (45K files, 5GB total) with git-fat wasn't much different than plain git.
It seems most of my problems stem from the number of files, rather than the size. If I had a smaller number of very large files, git-fat might be a good solution.
There's a number of other solutions open source out there, some of which are documented in our readme.
Or to put another way, what problems will I run into if I just commit large media files without using this?
Centralized version control systems such as Subversion don't have this problem (or at least, to a lesser extent), because as a user you only download a single revision of each file when you check out the repository.
Extensions like git-media, git-fat and now git-lfs solve this issue by only storing references to large media files inside the Git repository, while storing the actual files elsewhere. With this, you will only download the revision of the large file that you actually need, when you need it. It's sort of a hybrid solution in-between centralized and decentralized version control.
I'm probably missing something since there this, and git-annex, and git-bigstore, and others...
On the other hand game devs at this point are very used to Perforce, and it looks like Perforce is interested in solving this problem from the other side, by adding Git features to Perforce Helix and making it distributed.
Would it be possible to extend git-annex with a command that lets you set one or more extensions? By using git hooks you can probably ensure that the normal git commands work reliably.
What we don't have is the locking. I agree with the people commenting here that locking is a requirement because you can't merge. We need to do that.
A GB doesn't get you very far if you are working with raw audio and video.
Does it make sense to think about storing virtual machines images (.vmdk) in git on GitHub with LFS?