Generate a Git repository with 2^28 commits—one for every 7-character shorthash
github.com
github.com
A=shasum; C=echo; I='($C "A$s$A$X C$s$C$X I$s$Q$I$Q$X Q$s$k$Q$k$X k$s$Q$k$Q$X X$s$Q$X$Q$X s$s$Q$s$Q$X $I" | $A)'; Q="'"; k='"'; X=';'; s='='; ($C "A$s$A$X C$s$C$X I$s$Q$I$Q$X Q$s$k$Q$k$X k$s$Q$k$Q$X X$s$Q$X$Q$X s$s$Q$s$Q$X $I" | $A)
Note that this is not in any way opening its own source code as a file or through shell magic or something like that, in fact it's not looking at its own code at all. It really generates its own hash computationally out of itself, and would work the same if this was a compiled C program computing (not containing!) the hash of its own source code, even if the source code was thrown away.
This may seem impossible at first, but it's really just a funny variant of a quine (a program that prints its own source). http://www.madore.org/~david/computers/quine.html came up with the idea and explains the concept.
[1] You could trivially change it to output its SHA2 hash instead (which just wasn't as common at the time I did this), but then you'd loose the particularly "nice" hash. That property is entirely unrelated to outputting its own hash, and done the same as the commit hashes linked here.
You can easily adjust quines to output any f(program_contents), including f=some_hash_function.
Also, quines are related to fixed point combinators, the most famous of which is the Y Combinator (in case you didn't know where does the name of this fund/forum come from).
7-8 character shorthashes are perfectly reasonable for tracking a few thousand objects without much chance of collision. There's a tradeoff between the uniqueness guarantee and the friendliness if humans ever need to see the identifier.
Of course, sequential identifiers can be even shorter and friendlier, but they are more troublesome in several ways than hash-derived ones.
It gets a bit dicy once you setup CI to start generating builds/images off every commit to master, and then off every push to every branch. In practice, though, I haven't seen it bite of the projects I've been on. Usually something will break and then you just update the scripts to use N+1 characters.
In theory, a short-sighted script without safeties could do something wacky like deploy an image that's several months/years old. Running the numbers suggests something so catastrophic is quite unlikely.
org/imagename@sha256:<hash>I think their point is that the image which a tag refers to can be changed in the repository, so if you're concerned someone might generate a Git commit with a colliding hash and produce another build with the same hash, there is in fact a much simpler attack scenario: replace the container image in the repository with any different image with the same tag.
The syntax they provided identifies the image by its own hash, so producing a different image with the same hash is _much_ more difficult.
Linked repo has only 2 commits.[0]
[0] https://github.com/not-an-aardvark/every-git-commit-shorthas...
What the hell?!
Any guesses if this would also happen if you tried to push it bit-by-bit? (although you'd of course need reasonably large groups of commits still, to not end up with an impossible number of pushes)
I expect GitHub wouldn’t handle the push well anyway, though, as mentioned.
So 2^28 is 1/4 of a billion (10^9)
> Git can figure out a short, unique abbreviation for your SHA-1 values. If you pass --abbrev-commit to the git log command, the output will use shorter values but keep them unique; it defaults to using seven characters but makes them longer if necessary to keep the SHA-1 unambiguous
https://git-scm.com/book/en/v2/Git-Tools-Revision-Selection
and also
> Git doesn't really truncate anything, internally everything will be handled with the complete value.
https://stackoverflow.com/questions/7128444/how-does-github-...
> Due to the birthday problem, any repository that has at least 19291 commits is likely to have a pair of ambiguous commits somewhere.
Edit: nevermind; I am stupid.
Some choice quotes:
> The Google codebase includes approximately one billion files and has a history of approximately 35 million commits spanning Google's entire 18-year existence.
> Google's codebase is shared by more than 25,000 Google software developers from dozens of offices in countries around the world. On a typical workday, they commit 16,000 changes to the codebase, and another 24,000 changes are committed by automated systems.
Note that this was 5 years ago. If you look at the number of commits over time, it is growing exponentially, about sevenfold over the last 5 years:
https://dl.acm.org/cms/attachment/c219bd10-b97a-402e-bc94-79...
If you assume another sevenfold increase over the past 5 years, then they should have around 250 million commits now.
Decent guess. Rate's gone up; they're at about 382 million now. [1] You can see the current number in the PiperOrigin-RevId: on several public repositories that get changes synced from Google's internal Piper.
I'm not looking at that paper (as I'm about to walk out the door); one thing to check is if the number there includes just the actual commits or also the pending CL numbers (same concept as Perforce). The number I'm quoting includes both.
[1] https://github.com/google/tcmalloc/commit/6f42c40d6ebcf2aeb8...
Note that perforce is not like git. One does not generally clone the whole repo. You create client views that contain the files you want to edit. Everything else stays on the server.
Paths are also how one does tags and branches on old school tools like Subversion (just copy your branches/master to tags/releaseX and then never change it)
This isn't a problem for Git. It accepts arbitrary prefixes so you'll just have to type a few more characters if you want to refer to a commit whose 28-bit prefix (7 hex digits) is not unique.
Which was a git based proof of work, the server would only accept pushes for commits if it had more leading zeros in its hash than the previous commit on that branch.
Discussed here: https://stackoverflow.com/questions/9392365/how-would-git-ha...
It costed 6500 CPU years to create the first collision containing a valid pdf document. Storing both versions into git still works, because git prepends a header before creating the SHA-1 of the blob. I believe because of that it is harder to create a git blob collision than a pdf collision. (I admit I did neither read very carefully nor try to think it through very seriously.)