Nowadays, the repo-tarball split is largely unnecessary, but the practice remains. The xz incident became a wake-up call that the Reproducible Build movement focused on binary reproducibility, but have so far ignored tarball reproducibility. Hopefully, this problem will be addressed by the community in the future.
The reproducibility must stretch over all the outputs.
A related idea is the Hermetic Build", where the sameness of all the input is ensured, even including the build tools.
A project can be reproducible when built from the release tarball, without being able to reproduce the release tarball from a git tag.
I think locking the source distribution mechanism to only git would be detrimental in the grand scheme. A source tarball is universal, independent of preferred tooling, easy to hash, sign and verify, and archive. Even git has loopholes in that a tag does not necessarily have to be a part of the master branch, and not everyone uses (or wants to use) github, or a similar online interface for git...
It's also worth considering that source generation tools are often keen to changing their API frequently, while their output is more weaponized against the passage of time. Autotools is significantly better about that these days, but many other tools aren't...
Pulling from Git doesn't really require using GitHub other than connecting to their IP. You don't need an account or any tool other than Git to clone a public repo.
It seems to me that pulling directly from the primary source of truth is a good idea these days where possible.
'Dumb' HTTP involves requesting `info/refs` which lists all the references (branches and tags) available. The client has to find the tag you want, then request the corresponding commit object. That lists the hash of the tree object, which has to be requested, and recursively the tree object references other trees and finally the file blobs.
Any object request could fail because the object isn't stored 'loose' on disk, but is instead in a pack. So the client first downloads `objects/info/http-alternates` to see if it's in a different location, then if that doesn't list anything, it asks for the list of packs `objects/info/packs`. Then it asks for the index of each pack to determine which pack contains the object, then finally downloads that pack.
Because that's very slow, Git also supports 'Smart' HTTP. That moves all the lookups server-side, with a back-and-forth between client and server about what the client has and what it wants, ultimately building a custom pack containing what you asked for. This obviously can't be cached.
In contrast a tarball is one request for one file on disk which is fully cacheable, both by the server hosting the file and any proxies on the path.
This is why things like `bower` and `DefinitelyTyped` got deprecated. Some of them were even hosting the registry itself as a Git repository on GitHub.
I keep seeing this argument, but I don't understand why CI runners can't generate the configure script rather than offloading that onto the end user requiring them to have autotools/automake/autoconf/m4/etc
If you've ever tried to fight with that on an decade old enterprise system, those tarballs with configure scripts in them are useful. They get generated with new tooling on modern distros, but run on ancient systems. Getting the right versions installed can be an absolute pain in the ass.
These days most Linux kernels can do everything we need, and our release tag of QEMU might only be one or two patches not yet in the upstream branch; and in any case, you can always use the most upstream release. So this isn't necessary anymore (and indeed we only include QEMU in the release tarball, not Linux or pvgrub). And you can still "./configure && make && make install" to get a fairly complete system, it will just clone other repositories on your behalf.
On the other hand, I actually checked just last month, and the tarball for Xen 4.18.0, released back in November, was getting 700 downloads a week. Who are all these people downloading the release tarball, rather than using the distro version of Xen? Do they want and need these extra bits inside? We don't know and were hesitant to make any breaking changes.
I think with the xz fiasco, we now have justification to switch entirely to a `git archive` tarball of a specific tag, whether it's inconvenient for people or not.
Making reproducible tarballs from VCS isn't hard - with the advice from https://reproducible-builds.org/docs/archives/ (and requiring gnu tar) it's possible to get byte exact output from both Mercurial or git mirrors, on both Linux and MacOS. Put that in the github CI and it's reasonably difficult to subvert (compare against a local build too at release time).
It would be nice if "git archive" had guarantees about archive format, and if the github "download a tar.gz" matched that, but it doesn't seem to be the case at present.
Removing that blind spot is either Harder Than You'd Think or Easier Than You'd Think, depending on your perspective and expectations. You can find some issues listed here:
https://reproducible-builds.org/docs/
Or the homepage of reproducible-builds.org for a general take on the subject. (I am not associated with that website.)
If they built from the source there wouldn't be this issue.
One primary one is, sustainability. The build tarballs are kept forever. Try that with a remote repo that could vanish in a mere 10 or 20 years.. or tomorrow!
And tar is the most stable archive format in existence, with at least half a century of use.
note: Companies that care, should ask themselves what do you do, if you have a build system which relies upon externals, and a part of that build goes down?
And you have an urgent fix to PROD required?
Hope you can find all the bits, unvarnished, scattered on dev boxes? Cobble them together and hope they build, while your PROD is currently borked?
Or do you hotpatch PROD?
If your build process breaks due to an external repo going MIA, then you're doing it wrong.