FYI: LLVM-project repo has exceeded GitHub upload size limit (2022)
discourse.llvm.org
discourse.llvm.org
Anyway, the problem is that git stores all changes, forever. A better approach would be to clean old commits, or somehow merge them into snapshots of fixed timespans (like, anything older than a year get compressed into monthly changesets)
I've never used it before, but from what I understand, it's very powerful but also very confusing and easy to mess up, and of course with a sort of vague ambiguous name that makes it hard to discover; in other words, it's quintessentially git.
If a 10 year old vulnerability is found in OpenSSL, it be nice to be able to investigate if it was an accident or an act of espionage.
If this is my source code, I want the whole history. I want that 10-year old commit that isn't used in any current branch. A build machine may not need any history: it just wants to check out a particular branch as it is right now, and that works too.
But there is an intermediate case: Let's say that I have an issue with a dependency. I might check out that code and want some history to know what has changed recently, but I don't need that huge zip file that was accidentally checked in and then removed 4 years ago. If it were a consistent problem, perhaps you'd invent some sort of 'shallow' or 'partial' clone, something like this:
https://github.blog/2020-12-21-get-up-to-speed-with-partial-...
Accidentally publish secrets/credentials? Rotate them yes but also remove them from the published history.
Accidentally publish a binary for a build tool without the proper license? Definitely remove it (and add it to your .gitignore so it doesn't happen again!)
You discover a major flaw that bricks certain systems or causes data loss? Retroactively replace the Makefile/configure script/whatever to print out a warning instead of building the bad build.
I'm sure there are others.
In fact, Git's distributed nature actually makes it infamously bad at scaling with repo size - since it requires every "participant" in a repo to have a copy of the entire repo before they can start any work on it.
Introduced: https://devblogs.microsoft.com/devops/introducing-scalar/
Integration into Git: https://github.blog/2022-10-13-the-story-of-scalar/
https://git-scm.com/book/en/v2/Git-Internals-Git-Objects
Pushing from a shallow clone to a remote is more complex, but is supported in modern git versions.
Overall looks like ~750MB of unnecessary file duplication for one install of LLVM.
I get that Windows doesn't have the same niceties when it comes to symbolic links as macOS and Linux, but that seems really suboptimal and wasteful.
https://blogs.windows.com/windowsdeveloper/2016/12/02/symlin...
My applications definitely have modules that only get touched when we do a major UI refresh (about once every 10 years.)
The addage "if it ain't broke, don't fix it" comes with the corollary that if you build it right it won't break...until it does. And then someone has to do some archaeology to get a process back online that hasn't been documented since dot-matrix printers and typewriters, much less git development. I'm also building stuff and making decisions today for machines that have 10, 20, or 30-year expected lifetimes, and the core components should last indefinitely as long as maintenance is performed.
Many businesses aren't built for and aren't compatible with a 2- or 3-year obsolescence cycle.
> As it turns out, we only run repacks on a repository network level, which means that repacks need to consider objects from all forks of a given repository.
> Repacking entire repository networks will always lead to less optimal pack sizes compared to repacking just objects from a single fork. For GitHub, disk space is not the only thing we optimize for, but also performance across forks and client performance.
Repacking the repo locally more than halves the size:
> I tried locally right now to run git repack -a -d -f --depth=250 --window=250 and the size of the .git folder went from 2417MB to 943MB…
But GitHub repacking the repo network reduces the size by only 20%. So presumably GitHub can't aggressively remove things from a repo without affecting how forks work.
This blog post is very old (2015) but there's a section that describes their use of Git alternates to facilitate forks: https://github.blog/2015-09-22-counting-objects/#your-very-o...
git clone https://github.com/llvm/llvm-project
When entering the directory my fancy prompt timed out [WARN] - (starship::utils): Executing command "git" timed out.
But it seemed OKish after `git status` (disk cache?) time git status
...
real 0m0.287s
user 0m0.189s
sys 0m0.392s
The size of `.git/objects` after the clone was 2.5GB.Output of `git-sizer`: https://pastebin.com/T5HRMfg9
Are you really suggesting it's impossible for bad practices to bloat a got repo?
Is that too much under one umbrella? Probably. But it's not just a compiler. It's a monorepo.
It's not so difficult. A team of tens of programmers can reach that size in a couple of decades, just by writing code and textual documentation. No graphic assets required, all it takes is a generation of steady work by a mid-sized team.
This made me laugh
Yes, and for large assets there are extra solutions like Git LFS (which GitHub has support for).
CocoaPods and Homebrew both hit similar issues in the past with huge work trees by hosting all their specs on GitHub, resulting in breaking workflows for their users.
Those projects made changes to their workflow to mitigate this. Since the source here is around 6 months old, have LLVM done something similar?
git filter-branch --index-filter "git rm -rf --cached --ignore-unmatch /path/to/file" HEAD
https://stackoverflow.com/questions/43762338/how-to-remove-f...Or is git-lfs generally avoided for good reason? Or so all the changes and not necessarily large files cause this repo to exceed the limit?
Who are you going to trust to do this? Someone will need to execute this and then do a force push. Are you going to compare it? Are you going through all the history of files removed to see if there is no change to the source code? And what about the files you're actually going to move to git-lfs? How can you prove that they haven't changed in the migration?
Provenance is a thing. https://slsa.dev/provenance/v0.2
I knew that rewrites compromise the history. It is low stakes and I didn't want to allocate a new repo to start anew, just learned my lesson there.
I'm mostly curious about git-lfs as large file viability or it should be avoided and plan for large build artifacts to host elsewhere.
> Repacking entire repository networks will always lead to less optimal pack sizes compared to repacking just objects from a single fork. For GitHub, disk space is not the only thing we optimize for, but also performance across forks and client performance.
So the lesson here is you can DoS existing open source projects somehow by forking them and increasing the forked repo size >2GB?
The issue is that GH doesn't accept too big packs. git by default pack everything into a single pack. Maximum pack size can be specified either in config or as an argument to repack. The way I read the error message a user can push a huge repo by making sure it's packed into a few packs under 2GB limit.
It doesn't seem like there's an easy way to turn this into a DoS. GH would repack the fork network on its own schedule. A user probably can not trigger repacks. The repacks on GH side would probably be smaller then their limit, too. packs are probably scoped to a fork and the server is an active client so it most likely wouldn't return objects from other forks. I don't think it would be easy to DoS GH just by pushing big packs (either under or over the limit).
A shell loop can push N commits at a time, using git's handy syntax for that:
`git push origin p4/master~${x}:master`
Is it still a fundamental limitation, but with a known client-side workaround?
> If you happened to push the entire llvm-project to another new (personal) GitHub repo recently, you might encounter the following error message before the whole process bails out:
> The crux here is that we tried to push the entire repo, which has exceeded GitHub’s upload limit (2GB) at once. This is not a problem for majority of the developers who already had a copy of llvm-project in their separated GitHub repos
If you fork the repo on the Github website, they manage the fork server side, and you never have to push up the entire repo.