Mercurial with Largefiles: Why it is not a solution for game development
ennoble-studios.com
ennoble-studios.com
- you have lots of large files which are not amenable to diff and change frequently;
- everyone is working within the same company, and usually on the same network;
then a DVCS is unhelpful because you have to spend the disk cost for the repositories on every machine having a full copy of everything that's ever been checked in, regardless of whether they need it or not.
Many games are tens of gigabytes when shipped. It's easy to imagine a process which accumulated hundreds of megabytes of asset changes every single day over a multi-year development process. Then you can imagine having to have expensive terabyte SSDs just to work on it with all your tools.
I'm actually looking at this problem at work for possibly converting a large repository from svn which is a decade old and merely tens of gigabytes. Frankly svn handles it just fine so I'm going to defer the problem until I absolutely have to migrate.
120 GB for the game data repo. This is the repo for the source and art that actually ships.
2.5 TB for the raw art repo. This is where the pre-export raw files go.
Obviously it would be impractical to require that the artists all have the entire repo. We use SVN.
are there that many games which have a lifetime of over 11 years? I only know about one...
Think of League of Legends, Team Fortress 2, EVE online, World of Warcraft. There are hundreds more smaller games.
as said there aren't that many games that have a lifespan of over 11 years.
There are so many games that have this long a lifespan.
Why? Perforce's integration to Unity is quite poor (they have a huge untapped market here), so you end up having to resolve a lot of things slowly in their tools. Git/HG are much faster in my experience at detecting changes and interacting sanely with them. Also the team could never learn why a file is checked out, why can't they commit, etc.
We regularly clean out our largefiles cache on disk, so most of the time everyone just has the latest version of a given binary file on disk. The server of course has every revision, but I want that.
And most important of all: with small tweaks we're able to use Phabricator for all of our task management/documentation workflow. Getting VCS hooks out of the box to let artists say "Adding typewriter model, please review T555" in their commit and having that task automatically get assigned to the reviewer is priceless.
Most of my team doesn't have any idea what VCS is, but they've learned to use TortoiseHg (they call it "the turtle") and Phabricator to organize ourselves.
While Mercurial isn't the only way to get there, its free, its fast, its simple, and it unlocks the power of Phabricator (so does Git+LFS I believe).
So in my experience, I would say hg+largefiles is an excellent solution for game development.
I have been using Bitbucket and Mercurial for all side projects of mine for quite a while. But when you start with game development, you will reach the repository limits quite fast. Textures, meshes, sound, music, concept art and other binary blobs eat alot of storage.
Git-LFS is a bit a pain to setup, because you need to define before checking in, which extensions need to be stored as large file. And then there are check-in hooks, which sometimes did seem unreliable. Visual Studio git integration is also quite limited, but SourceTree did serve me well.
It's quite liberating if you're able to check in code and assets together without taking into account the space needed.
1) https://blogs.msdn.microsoft.com/devops/2015/10/01/announcin...
Free and open source (apache v2) at http://bitkeeper.org
Take-aways: http://seanmiddleditch.com/my-gdc-17-talk-retrospective/
https://twvideo01.ubm-us.net/o1/vault/gdc2017/Presentations/...
Game Development seems to have such a different workflow than most of the stuff I'm familiar with like backend, web dev, and the occasional network programming.
https://gamedev.stackexchange.com/questions/480/version-cont... This link recommends peforce as the standard .
What do people in the industry actually use ?
Is there anything close to a text-based procedural texture format ? Textures could be procedurally generated at startup and transformed into bitmaps. I am aware of kkrieger, but is there anything other than proof of concept ? No one takes voxels seriously anymore...
Are there any issues with this approach I should be aware of, considering that Hg with Largefiles seems to have some, too?
1. Git clones are deep and wide by default, so you end up with local repos that are the size of the total history. (easy to work around)
2. Git narrow clone support isn't the greatest, so you will probably want to fit the "working directory" on one computer. (somewhat painful to work around)
3. Git has to checkout all files every time you change branches which can be slow if you have a large working directory. (very painful)
4. Many operations (status, diff, commit for example) require git to scan the working directory to see what changed. This will be very slow if you have many files in your working directory. (very painful)
Luckily all of these problems can be solved seamlessly with a virtual filesystem approach. For example https://github.com/Microsoft/GVFS
This isn't true. Largefiles aren't stored as deltas, but as complete blobs. Mercurial still reads them in their entirety, so it still uses a lot of memory, but it's not diffing.
> The next problem is everyone collaborating on the project would have to take a huge Pull with the new large files, for every version of the large file they don't yet have [...] if you want to go back to a revision you haven’t pulled yet and the Server is not up you’re out of luck. That means you should get all the commits at some point anyway (because you want all the code versions at your side), so what’s the point?
Largefiles's mechanism doesn't require that you download every version of every large file. That's a key part of its design. If you decide that as a matter of policy that you want to download them all anyway, then no, largefiles won't help much.
> Well, they handle files by placing them outside the repo, and storing only the hash of the file in the repo itself (all bigfile hashes inside one file). This has an unfortunate effect that you won't be able to tell which exact bigfile/largefile has actually been modified when looking in the history – the only thing you'd see in the repo is the cumulative file that holds the hashes of all bigfiles as having a change.
This isn't true either. Largefiles stores the hashes in separate files, and the history machinery is able to interpret the records properly.
I've written a little script to demonstrate the structure of a largefiles repository:
https://bitbucket.org/snippets/twic/7eeAxy
One thing that would be really useful that largefiles doesn't (that i know of) do would be to opt out of downloading some largefiles at all. If i'm checking out an old revision just to read some old code, i don't want to spend ages pulling largefiles that i'm not going to look at. You can do this with Facebook's remotefilelog extension, which lets you make shallow clones, which can omit the large files, but it's awkward:
https://bitbucket.org/facebook/hg-experimental/src/default/r...