Why is it several gigs? Is that really necessary?
Why is it several gigs? Is that really necessary?
A lot of code written by a lot of engineers over a lot of years.
I'm not sure what other answer you're expecting?
I work with a compiler that has just tens of engineers working on it over just a decade or so and even that's a 6 GB repository. No binary assets. Just source code and configuration I think. I really don't think it's that unusual.
> Is that really necessary?
What would you do? Delete history every year or so? I regularly annotate files and see useful history from ten years ago that I need to do my work.
I've seen some things. Binaries, large files, isos, generally just any large file with no lines that somebody just didn't know better and committed, changed a few bytes a few times, and generated huge deltas quickly.
I'm far more surprised you don't know what he was expecting than I am that he asked, considering how many times I've run into this in the past.
https://github.com/oracle/graal
It's not even a mono-repo - this is just part of the project.
Maybe someone's got some tools that let them dig around in the history and find large things or explain why it's so large? I don't think they've been checking ISOs in.
Receiving objects: 100% (981372/981372), 187.08 MiB | 6.39 MiB/s, done.
% du -sh graal 328M graal
% du -sh .git 221M .git
And that includes many binary files that were added and removed over time, but not filter-branched out of existence, including jars, pngs, pdfs :)
(https://stackoverflow.com/questions/10622179/how-to-find-ide...)
The biggest issue is actually the l10n stuff. It’s all text with lines but there is a lot of it and historically the exports were not really ordered so every major update to the translation files scrambled everything.
It turns out folks can write an ungodly amount of code.
that or check-in binaries, logs, database backups or other things that don't belong in VC. TBH it's pretty much always this
That's a lot of code! I think that is at least a little unusual. The monorepos I see are usually large because people gratuitously commit large binary files (and sometimes change them, generating huge changesets).
At one of my work places one of the monorepos was over 50GB... took an hour to clone and many minutes to pull/rebase if I was a few days behind...
I don’t like monorepo approach because of such monstrosity, but the upside is it’s easier to commit/sync changes that span multiple products/components, and easier to merge many commits a minute from thousands of engineers...
Just wait until you have to store deep learning training data.
GitLab and Gitea come with a git-lfs server, but I don't know if there is any standalone server that you could use with git-shell.
It's just too much of a hassle to figure out the security implications of yet another server, and administrating it, setting up the connection between Git and the plugin, etc.