Something being autogenerated, or binary, doesn't mean it shouldn't be in version control. If step one of your instructions to build something from version control involve downloading a specific version of something else, then your VCS isn't doing it's job, and you're likely skirting around it to avoid limitations in the tool itself. People still use tools like P4 because they want versioned binary content that belongs in version control, or because they want to handle half a million files, and git chokes.
In my last org, we vendored our entire toolchain, including SDKs. The project setup instructions were:
- Install p4 - Sync, get coffee - Run build, get more coffee.
A disruptive thing like a compiler upgrade just works out of the box in this scenario.
It's a shame that the mantra of "do one thing well" devolves into "only support a few hundred text files on linux" with git.
> then have your build toolchain unzip the archive before it runs
My build toolchain shouldn't have to work around the shortcomings of my environment, IMO.
> et voila you now have this large file versioned in Git.
No, it's on a separate http server that is fetched via git lfs. Subtle, but important difference.
This is a non-issue for images and autogenerated files, since you shouldn't ever be doing a merge on them.
> breaks the concept of D in the DVCS of git.
git-annex is distributed and works well for files that will never be merged (such as images, or autogenerated files)
I think the SHA should be in version control. The file should be reproducibly built [1], then cached on a central server.
This means that a build target like a system image could be satisfied by downloading the complete image and no intermediate files. And a change to one file in one binary will result in only a small number of intermediate files being downloaded or reproducibly built to chain up to the new system image.
This is something that's really lacking in, for example, Git.
Requiring reproducible builds to handle translations or images is a bit much. Also, if it's cached on a central server, that now means you need to be connected to that central server. If you require a connection to said central server, why not just have your source code on said server in the first place, a la p4?
I do agree that NixOS is a great idea, but personally 99% of my problems would be solved if git scaled properly.
> why not just have your source code on said server in the first place, a la p4?
That would be great. A version of git where cloning is almost a no-op, and building is downloading the package assuming you haven't changed anything.
I'm not aware of how p4 allowing this. My recollection of perforce is that I still had most source files locally.
You vendored all your compilers/language runtimes in the source control repo of each project? Including, like, gcc or clang? WTF?
> It's a shame that the mantra of "do one thing well" devolves into "only support a few hundred text files on linux" with git.
Because the Linux kernel source tree and its history can accurately be described as "a few hundred text files".
Yeah, right.
Having local commits intermingled with an upstream code base can make for really hairy upgrades, but I guess every situation is slightly different.
Well we don't put them in git, we put them in perforce because git keels over if you try and stuff 10GB of binaries into it once every few months.
I think the real question is the other way around though, why _not_ use git for versioning when that's what it's supposed to be for? Why do I have to verison some things with git, and others with npm/go build/pip/vcpkg/cargo/whatever?
Yep. Along with paltform SDKs, third party dependencies, precompiled binaries, non-redistributable runtimes, you name it.
Giant PSD or FBX files? 4K Textures? all of it.
Client mappings are the bread and butter of P4 (or Stream views more recently which are not as nice to work with) - you say "I don't want the path containing MacOS" if you don't want it.
> Because the Linux kernel source tree and its history can accurately be described as "a few hundred text files".
I was off by a little bit, it's ~60k. But it's still "only" 60k text files, no matter how important those text files are.
(There are various good reasons why you might not! But "because binary files shouldn't go in version control" is not one of them)
There are plenty of cases when including generated files is appropriate. It has many advantages over not doing that - probably the biggest are
* Code review is much easier because you can see the effect on the output.
* It's easier to find the generated files because they're next to the rest of your code. IDEs like it much more too.
In fact the upsides are so great and the downsides so minimal I would say it should be the default option as long as:
* The generated files are not huge.
* The generated files are always the same.
Even when they are huge it might still be a good idea, but you can put the files in a submodule or LFS. I do that for a project that has a really difficult to install generator so users don't need to install it.
That said, if the autogenerated output is stable, it's fine. After all, in a sense, compiling your code is also a kind of autogenerating and few people will advocate for keeping compiled code in git.