Recommended reading if you're interested in Git's data model, it's pretty easy to understand compared to Git's UI: https://git-scm.com/book/en/v2/Git-Internals-Git-Objects
(And yes I have no idea what Github's backend does, but I can't imagine "duplicate all data for every commit" would be a feasible implementation strategy.)
The Linux kernel repo contains 1.1 million commits and 80k files on master. That would mean a naive total of 88,000,000,000 files being stored as a “full copy of the repository”.
Does this pass the smell test?
And if it isn’t impossible to use then I wouldn’t post about how it is.
22 million commits is a lot though; the author states that the repo was over 8GB "last they checked", and the entire thing looked to be an experiment to determine "Github's or git's breaking point" - I guess they found it.
The Linux kernel repo contains 1.1 million commits and 80k files on master. That would mean a naive total of 88,000,000,000 files being stored as a “full copy of the repository”.
Does this pass the smell test?
Basically, yes it may write snapshots per-file (never per-commit) locally but there is a separate routine to transparently repack the whole thing with deltas.