Git partial clone lets you fetch only the large file you need
about.gitlab.com
about.gitlab.com
That would e.g. make it reasonable to check in virtual machine images or iso images into a repository. Extra storage (and by extension, network bandwidth) would be proportional to change size.
git has delta compression for text as an optimization but it’s not used on big binary files and is not even online (only on making a pack). This would provide it online for large files.
Junio posted a patch that did that ages ago, but it was pushed back until after the sha1->sha256 extension.
However, with self synchronizing hashes of the kind used by rsync bup and borg, it doesn't matter - you could have a 1TB file, delete a single byte at position 100 - and you only need to store or transfer one new block (with average size 8KB for rsync, configurable for borg) if you already have a copy of the version before the change.
It's somewhat comparable with diff/patch but not exactly; it's worse in that change granularity is only specified on average; It's better in that it works well on binary files, does not require a specific reference diff (can reference all previous history), and efficiently supports reordering as well small changes - if you divide a 4000 line text file to four 1000-line sections and reorder them 1,2,3,4 -> 3,1,4,2 you will find the diff/patch to be as long as a new copy, whereas a self synchronizing hash decomposition will hardly take any space for the reordered file given the original.
So how do these self-synchronizing hashes work? Like a Merkle Tree? (Ah, okay https://en.wikipedia.org/wiki/Rsync#Determining_which_parts_... )
So rsync uses 8KB for chunk size, so for a file 1GB it has 125 000 chunks. (And if every chunk needs 16 bytes of hash data to send, that's about 2MB, pretty darn efficient, especially if it can spot reorders.) Though according to Wikipedia it only does this if the target file has the same size, so adding new files to ISOs might not work in case of rsync, but still, the possibility is there for diff algos and version control systems.
But it will definitely use hashes when size differs (unless forced to copy whole files, or copying between local file systems)
One obvious example where you could have a lot of common blocks (even following the offset where a change was made) is zip files. The zip format basically compresses each file individually and then concatenates all that together.
Let's say you have a build and it packages the results up as a big zip file. (Java builds often do this. A jar is a special type of zip file.) If you change a few source files and rebuild, and if your build is deterministic (and/or incremental), then the new zip file will contain a lot of the same stuff as the previous version. And if your zip archiver is deterministic (pretty safe assumption), it should produce a zip file that is mostly the same sequences of bytes as the previous zip file, even if there are changed files in the middle.
If you write a .tar.gz archive, then one change in the middle will throw everything off from that point on because it compresses the whole archive instead of individual files. In theory a binary diff can work around this by first undoing the gzip that was done to create each large blobs, then doing a binary diff on that, and then arranging to be able to recreate what gzip did. Obviously that's messy.
Of course, not every file is an archive. Some are filesystems. But any writable filesystem (notably not including ISOs) that is capable of being used on a hard disk will of necessity not rewrite everything. If it did, changing on one file on a filesystem would take hours because the rest of the partition would have to be rewritten.
Another obvious type of big blob is multimedia. I don't know a lot of specifics, but I would think file formats meant for editors would keep changes localized for reducing IO (for example, so that changes in a non-linear video editor don't need to write a giant file), but formats meant for export and delivery might change the whole file since they're aiming for small size.
Similar effect can be achieved with gzip --rsyncable, which IIRC resets the dictionary based on a rolling sum.
Someone was trying to talk me into git subtrees though...
It'd be great if they worked like Python's editable package installations.
I could totally see a .gitmodules.reqs file specified in terms of semver specs against tags, or just listing a branch to check out the HEAD of; resolving to the same .gitmodules file we already have. Not even a breaking change!
Would be nice if git submodules could also point to a branch instead of specific commits. That way, the superproject's state would not be modified every time the branch is updated.
It's interesting to see the wheel reinvented. We used to run a 500gb art sync/200gb code sync with ~2tb back end repo back when I was in gamedev. P4 also has proper locking, it is really the is right tool if you've got large assets that need to be coordinated and versioned.
Only downside of course is that it isn't free.
Another downside is that it consumes insane resources (our servers are in the dozens of TiB of ram, with huge NVMe based storage arrays directly attached)
Another downside is that you have to maintain connection to p4 to do any VCS operations (stashing included).
Another downside is that branches are very "expensive" (often taking days) and are impossible to reconcile. We never re-merge to MAIN.
DVCS is in direct opposition of workflows that include binary files(yes I'm aware that git lfs has locking, it's also centrally orchestrated) because you can't merge almost every binary format.
We were using P4 ~15 years ago for these workflows and rather than understanding what made them work people are just rediscovering the same problems that have already been solved.
My guess is that we'll next see a solution that dynamically caches most downloaded files in a geographic friendly way, heck we may even call it "P4Proxy".
I've seen so much FUD around how git is the "one true workflow" because other solutions "don't scale" when they don't understand the constraints that certain workflows impose. Git/DVCS is great for a lot of things but sometimes you should use the right tool for the job rather than hack something together.
[Edit] These reasons are exactly why you see Unreal supporting P4/SVN out of the box[1] and no mention of git.
[1] https://docs.unrealengine.com/en-US/Engine/UI/SourceControl/...
Locking helps with preventing collisions, but honestly the issue is still always communication. Why are people even touching files they shouldn't be touching?
Meanwhile perforce is a pain for code heavy projects and requiring a central perforce server. Git works great there.
The issue is neither is a silver bullet for the others workflow and needs, and they both suck horribly for mixed code and binary asset workflows.
That's not even considering cost.
Meanwhile film studios generally prefer keeping the considerations separate and using symlinks or URI to their data store and that works really well. But that doesn't work great for remote workflows.
So again, I think you're applying a very p4, game centric view to this. There are lots of different use cases and team structures that none of these version control systems are able to address in their entirety.
Because there's 300+ people working on a project, and it's not feasible to know what every other person is working on or planning on working on.
The file lock (code can be merged too, just like git, so this is only really for binary assets) is a crude communication tool saying "hey I'm using this file".
> Meanwhile perforce is a pain for code heavy projects and requiring a central perforce server. Git works great there.
I think you're applying a git biased view here. I don't think perforce is unsuitable for code heavy projects at all (see workspace views as a prime example), and for the majority of people, having a central server isn't an issue. Most people treat github/gitlab as a centralised server anyway. I've _never_ in a decade of programming heard someone suggest adding an extra remote to git so I can share your changes, it's always been "push it as a separate branch and I'll merge it". If you have a missing internet connection, you're likely not able to share code _anyway_, and with p4 you can always reconcile offline work when you're back online.
I think locking is a fine utility to have but I think a lot of workflows use it to workaround poor communication.
And I think you're constraining your views of git to just your workflow.
I've worked in a lot of scenarios where you need multiple remotes such as having an internal repo and an external one.
And similarly there are lots of scenarios where having a decentralized copy of the report is very useful for being able to work in offline scenarios and compare multiple branches. Things like when commuting on a plane or being in low connection areas.
I don't see how my view is git centric. I'm saying each VCS has very useful areas and equally big rough spots. The problem is that each VCS group believes there's is the only right system.
I will say however that if you think locking is optional then you already don't understand these workflows and why they're so critical. Art/design/animation doesn't care that "they should not have been touching the file" they just care that they have to throw away two days of work because someone made multiple edits to the same package file. I've literally seen multiple teams almost come to blows when this happens.
You can separate code from art/assets. That comes at an integration and iteration cost, it'll drive your designers mad.
There are workflows out there where code is not the first class citizen, in those cases I've seen git shoehorned in and untold pain follows.
I don't think attributing my disagreement with you to not understanding workflows is a fair characterization.
I still think locking , while useful, is only mandatory for work cultures with poor communication. Otherwise, many companies get by with very large workforces who don't hit these issues without having locking.
Separation of code and art assets also don't need to be painful. It's very doable but does require some amount of architectural consideration.
And I very much acknowledge there are projects where code isn't the majority makeup, which is why I say that none of the VCS systems cover mixed projects well or cover all the needs of the others well.
It sounds like we just come from different development cultures. Your solution to lack of locking sounds like a top-down hierarchy that wouldn't be flexible enough to support the teams I've worked with.
Having seen both approaches(and how they break down) I'll take a centralized locking solution over communication mistakes that lead to days of work being lost.
For context I developed the publishing pipelines for the majority of departments in the studio. I have several hundreds of assets being published through my pipelines on a daily basis, if not more, from a variety of departments.
We only hit collisions on a very rare basis, and which were often resolved in an hour or two in the worst case.
We've scaled this from very small teams to large ones, from very scrappy realtime productions to feature length offline rendered films.
I don't doubt that locking helps. I just argue that maybe it's not as critical as people make it out to seem.
At the end of the day humans make mistakes, especially when involving communication. I'd rather have a physical system that prevents breaks instead of requiring cross-team/cross-discipline coordination.
Maybe gamedev is much more coupled than film(we regularly had design, animation, art and code touching the same common core packages). Look at Unreal or any other gamedev pipeline and you'll see a bias for locking source control solutions.
Lack of organizational backup, do you mean cultural from the studio or infrastructure? Both are a problem no matter what solution you pick.
If a team goes AWOL, that's on them. The tooling usually allows for some amount of arbitrary workflow but they can't go completely off the rails. But that's true of p4 too. So I think that scenario would have to be more specific.
And yes people make mistakes, and you need tooling to guide them. Locking is a tool, but it's not the only tool. I feel very much that many workflows use it to hide deeper issues. That's not to say it's not valid, it is, but it's not a panacea either.
Unreal heavily favors perforce and SVN because that's what it was designed around. There's no absolute reason it could not work with other versioning systems and their paradigms if it came to being necessary.
Unity on the other hand is quite happy to work with any version control system, and works quite well with git or perforce.
You again seem to be trying to approach this from the angle of only the system you're familiar with working. But maybe try stepping outside the box and seeing if your workflow isn't a byproduct of your tools.
After all, you were asking git users to look at perforce as the solution. I don't think it's fair to then go ahead and assume that p4 is the only workable solution.
Take AOSP, even Google had to overlay the repo[1] tool to scale past git. It's a hot pile of garbage that won't let you sync all repos to a specific point in time. Not to mention the nature of cross-repo commits are not atomic. Good luck bisecting a breaking change across millions of lines of code and build files.
I've spent over a week chasing down how some homespun tool for storing binary assets side-by-side with git works so I could get a single file into a build.
Last company I was at which was a leader in the Android space just put the whole thing in P4, branch per device and it worked without many major issues. Pulling source took 1/50th the time a repo sync took. Literally an A/B comparison of one tech vs the other. That's before you even start to consider prebuilts.
Like I said, I think we're just going to have to agree to disagree and leave it at that.
> Google had to overlay the repo[1] tool to scale past git
It was created to allow for a forest of git repos to all coexist in a world in which git submodules wasn't suitable yet OR the repos spanned security domains. However, almost all of the shortcomings of submodules have been addressed and so—at least—the team that I lead now is considering migration to it from Repo.
> cross-repo commits are not atomic
Yes, that is a feature. But I think you meant that there's no cross-repo coordinate in the timeline to sync to. However, there is. That's exactly what a Repo tool manifest snapshot is. Our CI system ensured that change that had deps across repos were committed and a Repo manifest snapshot only was taken with all inter-commit deps satisfied.
> Good luck bisecting a breaking change across millions of lines of code and build files
The team that I led implemented this. We simply snapshotted the forest at every T time intervals. For bisection, we walked the snapshots. Once a specific manifest snapshot was identified as the culprit, we further bisected within repos for a specific change.
> Pulling source took 1/50th the time a repo sync took.
Yes, that's what the partial clones (the article you replied to) and sparse checkouts solves. Once these two things are widely available, I don't see any benefits to P4 remaining.
It's been 12 years since the Dream was released and we're still not to the same level of perf/features as just stuffing the whole of AOSP in P4. I get that git has advantages, and it's an awesome tool when used appropriately but the desire to use it to solve every SCM problem under the sun is a bit misguided.
I don't think it's accurate to say they shouldn't be touching the files.
It's possible to make two unrelated changes in the same binary file just as it is possible to make two unrelated changes in a source file (or other merge-friendly text file).
Just as there may be nothing wrong if one person changes the function foo() in a file and someone else changes the function bar() in that same file, there may be nothing wrong if one person opens a CAD file and makes a change to a drawing in one part and another person makes a change to a different part of the same drawing.
In that case, they could coordinate by communicating (even though their tasks are unrelated) but then they're just doing the same thing as locking the files but manually and informally (and probably inconsistently) without the benefits of automation.
Of course there are times when locking catches failures to communicate, but that doesn't mean that that's what locking is for.
Again, I'm not saying locking is an invalid solution. It is. But to me, it's often (but not always), a crutch for a deeper issue.
That binary file should be set up to be modular if it is intended to have areas that multiple users can touch without directly affecting each other.
Git LFS has been just fine for multiterabyte repositories for years.
Now if someone wanted to make an open source P4 replacement that would be a neat thing.
[1] - https://dvc.org/
Git is a distributed VCS, and we should support keeping it that way.
And GitHub's scheme is pretty much a de-facto standard at this point—GitLab's implementation is an exact copy of it, for example:
https://<host>/<org/project>/archive/<ref/branch/tag>.tar.gz
Edit to add: Also, git-archive --remote is actually most of the way there, but it's not an HTTP download, of course. :(One of my biggest remaining pain points is resumable clone/fetch. I find it near impossible to clone large repos (or fetch if there were lots of new commits) over a slow, unstable link, so almost always I end up cloning a copy to a machine closer to the repo, and rsyncing it over to my machine.
> Partial Clone is a new feature of Git that replaces Git LFS and makes working with very large repositories better by teaching Git how to work without downloading every file.
However I’ve heard many studios use Perforce instead. However not being open source is a downside to some, but I don’t really know too much about it personally.
Then if working with a lot of non code files, sounds like some solutions have locking. I guess not two people could edit the same Blender or PSD file at the same time and then merge them later on.
Kinda wouldn’t surprise me if some companies actually run multiple versioning control systems. Code on one system, game assets on another.
You’re more right than you think about multiple versioning systems, although keeping synchronized becomes an issue. Perforce is a bit of a boon for management, as they get a GUI for versioning across a multidisciplinary team.
P4’s GUI/model is also intuitive for non-programming roles to learn and use historically compared to git, so a team with wide skills can ramp up quickly with a unified toolset. A less-technical manager gets a GUI that has versioning across changes from a multidisciplinary team. You can probably guess what inertia that has in a space with higher turnover compared to other industries.
As mentioned, things are changing though. git and GitHub have become a mainstay and are what new programmers likely learn in schools. This has a trickle effect on new projects with smaller teams and results in more investment into git setups. I use git in a AAA context at work, and it’s not uncommon to find sentiments from more seasoned game programmers on git that are similar to HN comments about the latest fad in web frameworks.
And again, you keep blaming Git here, now around a lack of intuitiveness and a lack of GUI. Again, Git has multiple GUIs around to choose from and multiple integrations with almost any editor and IDE you can think of, some meant for beginners and trivial usage.
And no, things are not "changing" and Git is not to be compared with a "web framework fad". Git became the version control system more than 5 years ago.
I think you're correct in that git is the go-to version control software. I'd reach for it as a default tool every time. I do work with older programmers where the majority of their careers have been in Visual C++ and Perforce, and I've definitely heard sentiments seeing it as wheel-reinventing. I don't agree with them, but it's what I've experienced.
The two biggest issues are: - git lfs doesn't automatically identify large files or binary files. So it's very easy for even experienced engineers to have set up lfs but forgotten to track a file or extension
- git exposes too much of its internals. It's really cumbersome for artists even with UI tools.
That's not to fault git as a technology but I think there's a place for an artist friendly layer on top of git, perhaps a very artist centric UI and set of tools and workflow guides
How would a partial checkout help?
Not really. Modules are specced based on zip files and metadata in text files. There's just support for extracting that data from git repos transparently.
Here's a slightly out of date write-up: https://research.swtch.com/vgo-module
git archive --remote=<your-URL> | tar -t
You can get the most recent tree for a repository (no history, just the current state of the repo) with `git clone --depth=1`. That's often good enough for slow connections.
Wrong? There's a --depth option for the git fetch command which allows the user to specify how many commits they want to fetch from the repository