A new hash algorithm for Git
lwn.net
lwn.net
Rewriting the whole thing including every git repos history seems like throwing the baby out with the bathwater, when you could just add a secondary transparent verification instead. Just seems like there has to be a better way.
Are all the other future commits still valid, or am I going to suddenly get conflicts or garbled text? Depending on where the modification is done, that code might have gone through much more churn -- especially if there are a bunch of sha-256 commits after it (which I can't attack). I don't know enough about how git stores content blobs to answer this.
Second problem: Can I push my replacement commit to another repository (eg, github)? Would even force push work? Do I have to delete branches and re-push my own? If I already have enough permission on the repository to do this, it means I can already push whatever I want -- so does this attack even matter at all?
Assuming that's successful (or I can trick people into using my own repository), what will happen to someone that already has a clone and does a pull? Will they get my change (and will it work or be a pile of conflicts or garbled text)?
Even if only fresh clones will get the changes it could still be quite devastating -- especially if using CI -- but I'm just not clear if this attack is even theoretically possible.
Why? It's not the same as saying 'versions after vX are safe', it's the same as saying 'any unsafety after vX was there before, not introduced since' (both with 'as a result of SHA-1 collision' qualifiers of course).
> Can I push my replacement commit to another repository (eg, github)? Would even force push work?
Implementation dependent I suppose, but I wouldn't have thought so - I don't see why they'd actually check the content when the hash is supposed to indicate whether it differs or not.
> Do I have to delete branches and re-push my own? If I already have enough permission on the repository to do this, it means I can already push whatever I want -- so does this attack even matter at all?
I think an attack would look more like:
1. Create hostile commit that collides with extant commit SHA
2. Infiltrate a package repository, or GitHub, or corporate network, or ...
3. Insert hostile commit in place of real one
Of course it's a problem if 2 & 3 happen alone anyway, but the problem with the collision commit is that it makes it so much less detectable.The downside to the migration would be that all unchanged files would be stored twice (once identified by SHA1, once identified by SHA-256). But you could work around that by hardlinking identical files.
A blob is a "snapshot" of a file. The next version of a file is a completely different blob with no direct relation to the previous.
"Pack files" use delta compression in order to lower the actual size of "similar" blobs.
You could get conflicts if you tried merging or rebasing over the nefarious blob, and the "patch history" (git log -p, which builds the patch view on the fly) would show possibly unexpected complete file replacements.
Presumably the proposed "hash translation store" could use an approach similar to notes, and include the hash translations as objects in the git database (hopefully in a way that could be signed by a tag).
[1] http://alblue.bandlem.com/2011/11/git-tip-of-week-git-notes....
The Git team made the right choice: SHA2-256 is the best choice here; it has been around for 19 years and is still secure, in the sense that there are no known attacks against it.
Both BLAKE[2/3] and SHA-3 (Keccak) have been around for 12 years and are both secure; just as BLAKE2 and BLAKE3 are faster reduced round variants of BLAKE, Keccak/SHA-3 has the official faster reduced round Kangaroo12 and Marsupilami14 variants.
BLAKE is faster when using software to perform the hash; Keccak is faster when using hardware to perform the hash. I prefer the Keccak approach because it gives us more room for improved performance once CPU makers create specialized instructions to run it, while being fast enough in software. And, yes, SHA-3 has the advantage of being the official successor to SHA-2.
SHA-512/256 is a standard peer-reviewed and well-studied way to run SHA-512 with a different initial state and then truncate output to 256 bits.
This is heavy bikeshedding, but SHA-512/256 would be a more conservative choice than SHA-256. Under standard assumptions, SHA-256 is no weaker than SHA-512. The structure is extremely similar to SHA-256, but a collision on intermediate state requires a collision on all 512 bits of state instead of 256.
On most 64-bit CPUs without dedicated hash instructions, SHA-512/256 is faster for messages longer than a couple of blocks, due to processing blocks twice as large in fewer than twice as many operations.
Currently, the latest server and laptop CPUs have SHA-256 hardware acceleration but not SHA-512 acceleration. I'm not sure how many phone CPUs support sha256 but not ARMv8.2-SHA extensions (SHA-512). If it weren't for this difference in hardware acceleration, there would be few reasons to use SHA-256.
That being said, the current difference in hardware acceleration support probably makes SHA-256 the right choice here.
I am aware of the length extension issues, but they are not relevant for Git’s use case.
In terms of support, SHA-512/256 has, as you mentioned, less hardware acceleration support, and it’s also not supported in a lot of mainstream programs like GNU Coreutils. I also know that some companies mandate using SHA2-256 whenever a cryptographic hash is needed.
Git made the right choice with SHA2-256: It’s the most widely supported secure cryptographic hash out there.
Is BLAKE 3 still faster than sha-256 when using the cpu speciliazed instructions? I think most modern desktop CPUs has built-in instructions for SHA256.
I’m guessing when people compare BLAKE 3 to SHA 256 they’re comparing software to software, but this wouldn’t be the case in reality?
I can tell you this much: It is only with Ice Lake, which was released in the last year, that mainstream Intel chips finally got native hi speed SHA-NI support. Coffee Lake and Comet Lake, which are still the CPUs in a lot of new laptops being sold right now, do not support SHA-NI.
type 16 bytes 64 bytes 256 bytes 1024 bytes 8192 bytes 16384 bytes
blake2s256 46720.33k 187461.21k 305314.65k 373840.55k 398207.66k 401528.15k
blake2b512 38423.44k 155318.81k 422325.08k 592401.75k 674843.31k 681743.70k
sha256 84620.44k 279840.47k 723573.76k 1199678.81k 1484693.50k 1510484.65k
sha512 33854.38k 135674.20k 275343.70k 444872.36k 545802.92k 554166.95k
sha3-256 26146.35k 103860.27k 253944.92k 308119.21k 347477.33k 351906.47k
sha3-512 26349.83k 105590.85k 144236.03k 173082.62k 189448.19k 189814.10k
It's possible that Blake3 might be faster than accelerated SHA-256 on large inputs, where Blake3 can maximally leverage its SIMD friendliness. OTOH, Blake3 really pushes the envelope in terms of minimal security margin. Performance isn't everything. SHA-3 is so slow because NIST wanted a failsafe.OpenSSL info:
OpenSSL 1.1.1c 28 May 2019
built on: Tue Aug 20 11:46:33 2019 UTC
options:bn(64,64) rc4(8x,int) des(int) aes(partial) blowfish(ptr)
compiler: gcc -fPIC -pthread -m64 -Wa,--noexecstack -Wall -Wa,--noexecstack -g -O2 -fdebug-prefix-map=/build/openssl-D7S1fy/openssl-1.1.1c=. -fstack-protector-strong -Wformat -Werror=format-security -DOPENSSL_USE_NODELETE -DL_ENDIAN -DOPENSSL_PIC -DOPENSSL_CPUID_OBJ -DOPENSSL_IA32_SSE2 -DOPENSSL_BN_ASM_MONT -DOPENSSL_BN_ASM_MONT5 -DOPENSSL_BN_ASM_GF2m -DSHA1_ASM -DSHA256_ASM -DSHA512_ASM -DKECCAK1600_ASM -DRC4_ASM -DMD5_ASM -DAES_ASM -DVPAES_ASM -DBSAES_ASM -DGHASH_ASM -DECP_NISTZ256_ASM -DX25519_ASM -DPOLY1305_ASM -DNDEBUG -Wdate-time -D_FORTIFY_SOURCE=2
NOTE: /proc/cpuinfo shows sha_ni detection, and the apt-get source of this version of OpenSSL confirms SHA extension support in the source code, but I didn't confirm that it was actually being used at runtime. Blake3 SHA-256
66743 84620 Tiny
534057 1199679 Medium (1024 bytes)
573611 1510485 Largeish (16384 bytes)
This is based on the parent’s numbers with a fudge factor to account for Blake3 being a faster version of blake2s256 (i.e. the 32-bit version of Blake2 which is the only version in Blake3)Of course, this does take in to account that Blake3 has tree hashing and other modes which scale better to multiple cores.
(Edit: update figures; I need to scale up Blake2s256 not Blake2b512)
> git convert-repo --to-hash=sha-256 --frobnicate-blobs --climb-subtrees --liability-waiver=none --use-shovels --carbon-offsets
Is it sarcasm ?
(https://fossil-scm.org/forum/forumpost/50a5bea5fb)
> That's appalling. Fossil's implementation doesn't require a conversion.
“This is a key point, that I want to highlight. I'm sorry that it wasn't made more clear in the LWN posting nor in the HN discussion.
“With Fossil, to begin using the new SHA3 hash algorithm, you just upgrade your fossil binary. No further actions, workflow changes, disruptions, or thought are required on the part of the user.
* “Old check-ins with SHA1 hashes continue to use their SHA1 hash names.”
* “New check-ins automatically get more secure SHA3 hash names.”
* “No repository conversions need to occur”
* “Given a hash prefix, Fossil automatically figures out whether it is dealing with a SHA1 or a SHA3 hash”
* “No human brain-cycles are wasted trying to navigate through a hash-algorithm cut-over.”
“Contrast this to Git, where a repository must be either all-SHA1 or all-SHA2. Hence, to cut-over a repository requires rebuilding the repository and in the process renaming all historical artifacts -- essentially rebasing the entire repository. The historical artifact renaming means that external links to historical check-ins (such as in tickets) are broken. And during the transition period, users have to be constantly aware of whether they are using SHA1 or SHA2 hash names. It is a big mess. It is no wonder, then, that few people have been eager to transition their repositories over to the newer SHA2 format.”
For example, floppy.c could be replaced in a repo with file with the same sha1 hash as long as the last commit that modifies floppy.c used a sha1 hash.
Right?
For example, the manifest of the latest SQLite check-in is see at (https://www.sqlite.org/src/artifact/29a969d6b1709b80). You can see that most of the files have longer SHA3 hashes, but some of the files that have not been touched in three years still carry SHA1 hashes.
An attack like what you describe is possible if you could generate an evil.c file that has the exact same SHA1 hash as the older floppy.c file. Then you could substitute the evil.c artifact in place of the floppy.c artifact, get some unsuspecting victim to clone your modified repository, and cause mischief that way. Note, however, that this is a pre-image attack, which is rather more difficult to pull off than the collision attacks against SHA1, and (to my knowledge) has never been publicly demonstrated. Furthermore, the evil.c file with the same SHA1 hash would need to be valid C code that does something evil while still yielding the same hash (good luck with that!) and Fossil (like Git) has also switched over to Hardened SHA1, making the attack even harder still.
As still more defense, Fossil also maintains a MD5 hash against the entire content of the commit. So, in addition to finding evil.c that compiles, does your evil bidding, has the same hardened-SHA1 hash as floppy.c, you also have to make sure that the entire commit has the same MD5 hash after substituting the text of evil.c in place of floppy.c.
So, no, it is not really practical to hack a Fossil repository as you describe.
The attack may be difficult and unlikely I'm not questioning that, but if I understand correctly then Fossil's migration is straightforward because they did not address the same issues Git chose to.
I think more is at play here.
(1) You can set Fossil to ignore all SHA1 artifacts using the "shun-sha1" hash policy.
(2) The excess complication in the Git migration strategy is likely due to the inability of the underlying Git file formats to handle two different hash algorithms in the same repository at the same time.
But, I could be wrong. Post a rebuttal if you have evidence to the contrary.
But, I could be wrong. Post a rebuttal if you have evidence to the contrary.
It seems unfair to demand a rebuttal when you are the one who made the claim.
According to the article at least, the difficulty stems mainly from their migration strategy, for converting all existing SHA1 hashes.
That's essentially the same difficulty, since the only strategy for doing this that has been historically proven to work seamlessly and painlessly involves being able to handle both hash algorithms in the same repository at the same time.
...and also produce an innocent-looking diff!
I mean, you could stuff a bunch of random bytes into a C comment to force the desired hash in the output using these documented attack techniques, but anyone inspecting the diffs between versions is likely to see such an explosion of noise and call foul.
If you want an analogy, it's like someone saying they've learned to impersonate federal agent identification cards, only it requires that the person carrying the fake ID to have a thousand rainbow-dyed ducks on a leash in tow behind him.
Such attacks are fine when it's dumb software systems doing the checks, but for a source code repository where people do in fact visually check the diffs occasionally?
Well, let's just say that when someone manages to use SHAttered and/or SHAmbles type attacks on Git (or even Fossil) I expect that it won't take a genius detective to see that the repo's been attacked.
Also, if something is replaced in the history how often do people go back and view diffs in old code? Hardly often enough to rely on it being spotted.
Sure, many thousands of people doing blind "git clone && configure && sudo make install" could be burned by a problem like this, but someone would eventually do a diff and see the problem on any project big enough to have those thousands of trusting users in the first place.
I'm not excusing these SHA-1 weaknesses, only pointing out that it won't be trivial to apply them to program source code repos no matter how cheap the attacks get.
For instance, the demonstration case for SHAttered was a pair of PDFs: humans can't reasonably inspect those to find whatever noise had to be stuffed into them to achieve the result.
I also understand that these SHA-1 weaknesses have been used to attack X.509 certificates, but there again you have a case very unlike a software code repo, where the one doing the checking isn't another programmer but a program.
...which will likely contain thousands of bytes of pseudorandom data in order to force the hash collision...
> they cannot raise any alarms
You think a human won't be able to notice that the diff from the last version they tested looks awfully funny? Code that can fool the compiler into producing an evil binary is one thing, but code that can pass a human code review is quite another.
You might be surprised how often that occurs.
I don't do a diff before each third-party DVCS repo pull, but I do diff the code when integrating such third-party code into my projects, if only so I understand what they've done since the last time I updated. Commit messages, ChangeLogs, and release announcements only get you so far.
Back when I was producing binary packages for a popular software distribution, I'd often be forced to diff the code when producing new binaries, since several of the popular binary package distribution systems are based on patches atop pristine upstream source packages. (RPM, DEB, Cygwin packages...)
Each time a binary package creator updates, there's a good chance they've had to diff the versions to work out how to apply their old distro-specific patches atop the new codebase.
Someone's going to notice the first time this happens, and my guess is that it'll happen rather quickly.
If we weren't worried about sha1 collisions in git then we wouldn't switch to a new hash function.
I mean, I see you're expressing concern, but the first major red flag on this went up three years ago, and another big one went up last month. (https://sha-mbles.github.io/)
When we dealt with this same problem over in Fossil land, we ended up needing to wait most of three years for Debian to finally ship a new enough binary that we could switch the default to SHA-3. Fortunately (?) RHEL doesn't ship Fossil, else we'd likely have had to wait even longer.
Atop that same problem, Git's also got tremendously more inertia. Git has to wait out not only the Debian and RHEL stable package policies but also all of that infrastructure tooling they brag on. Every random programmer's editor, merge tool, Git front end... all of that which a project depends on will have to convert over before that one project can move to a post-SHA-1 future.
This is going to be a colossal mess.
It just seems to me that the Fossil maintainers have decided that keeping all old SHA1 hashes is acceptable, while the git maintainers have decided that it is not.
Unless I've misunderstood, this is why it was "so easy" for Fossil to transition to a new hashing algorithm. Not some superiority in the design of Fossil, as implied on the Fossil forums.
It seems like a problem very few people need to worry about and Fossil has made the right trade-offs.
1. Keep in mind that Fossil and Git are both applications of blockchain technology, which in this particular practical case means you must not only forge a single artifact's hash, you must also do it in a way that allows it to fit into the overall blockchain.
2. Fossil's sync protocol purposefully won't apply Dr. Hipp's hypothetical evil.c to an existing Fossil blockchain if presented it. Fossil will say, "I've already got that one, thanks," and move on. Only new or outdated clones could be so-fooled.
Are we saying this now? More like blockchain is an application of git technology.
If you're looking for prior art, ZFS's application of Merkle trees predates both. I think there was some other public use before that, but I can't recall right now.
Then it also has a similar looking-ish migration to SHA3-256.
But the only reason this would be attractive is because then people could keep using the existing prefixes to refer to the whole commit. But of course doing this would be insecure. So for this to make any sense at all, people would need to make good choices on when to use an insecure prefix and when to use the whole hash, because it's security relevant. This seems a bit doubtful to me.
In fact, https://github.com/bradfitz/gitbrute exists.
Also: the prefixing increases the length of the hash (and hence the desire to shorten it) without adding any security.
It may be a better to display SHA-256 commit hashes, but accept SHA-1 hash prefixes for old commits. It may be confusing for git to accept hashes that aren't visible in `git log`, but it's probably for the better.
This solution would basically just make the UI backward-compatible while still requiring the complete modification of the internal to change the hash function.
You'd still risk a collision if you refer to commits using a shortened hash outside of git but something tells me that you don't even need a vulnerability to take advantage of that if you have an attack vector. For instance github seems to use 7 hex digit in short hashes, this could probably be bruteforced relatively easily (be it for SHA-1 or SHA-256). To give you an idea I looked at the current bitcoin difficulty (which AFAIK uses two rounds of SHA-256 internally and works by bruteforcing hashes with a certain number of leaning zeroes) and the hashes look like this: 000000000000000000028048b31e42bd53d3b36da90d1a840ae695ec1a5ee738
The proposed method would have the advantage of keeping existing known abbreviations, which are _already_ less secure than SHA-1, while keeping the security of the second hash.
It also has the disadvantage that the full hash would become excessively large and unwieldy, so pros and cons.
Note that git doesn't concern itself with reversing a hash function. The commit contents are part of a repository, there is no value in guessing the commit contents basing on its hash. Here, the hash function choice is purely about collision resistance.
But yeah, don't do weird things with hashes. Cryptography is hard. Don't invent memecrypto: https://twitter.com/sciresm/status/912082817412063233, it's not going to increase the security. Use a single algorithm if you can. Don't transform the output of a hash function in any way.
Notwithstanding, i still dont like it as an idea.
Not a dealbreaker by far, but still a slight mark against this solution.
I wonder if it would still be practically possible to manipulate the commit id.
Fossil lets you force the project ID on creating the repo, but the capability only exists for special purposes.
git convert-repo --to-hash=sha-256 --frobnicate-blobs --climb-subtrees \
--liability-waiver=none --use-shovels --carbon-offsets
Surely some of those options aren't real...(never fails to amuse me)
Gives me the impression that it’s a construction of the article alone. Unsurprising, given the snark of the options.
How long until a specified length preimage attack can break bittorrent blocks?
I remember a paper published a ~decade ago estimating very short (well funded) ASIC sha1 collisons. Anyone have that ref?
EDIT: Should I have not said preimage? My understanding is bittorrent is broken (by DDoS, not infohash(?)) if you can make a bad block that matches the length and sha1 of a target block.
There are three different attacks
1. Collision, which is practical (expensive but practical) for SHA-1 today, lets somebody make two documents A and B which have the same hash. This is only useful if you can fool people somehow into accepting document B when they think it's document A because of the hash, for example with digital signatures.
2. Pre-image, which is not practical for any hashes you care about including MD5. This lets you find the document A given the hash(A) value. This is very niche, since obviously for large documents by the pigeon hole principle there will be many such pre-images and it's impossible to get the "right" one, for small inputs it can be relevant, sometimes.
3. Second Pre-image, likewise not practical. Given either document A or hash(A) which you could easily determine from document A, this lets you produce a new document A' that is different from A but hash(A') == hash(A). This would be extremely bad, and is what you'd need to attack real world Bittorrent from somebody else.
Often people say "pre-image" meaning strictly second pre-image, it's usually clear from context, and a true pre-image attack as I explained above is only rarely relevant.
Collision would only let bad guys corrupt their own purposefully constructed collision bittorrent, which like, why? So yes, Bittorrent would only really be in serious trouble if there was a second pre-image attack. But on the other hand, don't use broken cryptographic primitives. Attacks only get better, always.
https://en.wikipedia.org/wiki/Merkle%E2%80%93Damg%C3%A5rd_co...
https://www.reddit.com/r/crypto/comments/44p5jc/eli5_why_are...
Fossil uses SHA-3, which has an entirely different construction, which is not at this time known to have a similar weakness. SHA-3 is also much newer, with a much shorter list of known attacks.
Anyway, as hinted above, chosen prefix has nothing to do with the type of hash construction, except in the sense that so far there were lots of Merkle–Damgård hashes and some of them are no longer safe, whereas until recently there weren't many of the Keccak family hashes.
The Wikipedia article is talking about Length Extension, which is a different phenomenon from chosen prefix collision attacks, and if it was a problem in Git (or indeed Fossil) would have doomed them both immediately anyway.
For a generic crypto hash you should use SHA-512/256 (NB this is not offering a choice that slash is part of the name) to avert Length Extension but since the DVCSs already seemingly put the effort in to be safe against it SHA-256 is a perfectly reasonable choice.
It's also not particularly surprising. Just by its length SHA-1 has in its best case 80 bits of collission security and 160 bits of preimage security.
Now its important to understand that attacks usually don't cause full devastation, but they usually make attacks a bit better than optimal.
Attacks in the 60 bit range is what's possible, attacks in the 70 bit range is what's dangerous. It's easy to imagine that relatively small deviation from optimal security gets SHA-1 from 80 into the dangerous territorry (the attacks are in the low 60s range). However getting from 160 bit down to the 60/70 bit range would require massive improvements in attacks.
It's safe to say that SHA-1 is still very far from preimage attacks. Still to be clear I'd still recommend to get rid of it whereever you can. The far bigger risk is that you think you only need preimage security, while you actually need collission security for scenarios you haven't thought about.
Even MD5 still doesn't have a known preimage attack, so... many many years?
Oh wait, perhaps you actually meant preimage as you said rather than I assumed second preimage. OK yes, that isn't ever going to be possible for non-trivial inputs.
https://electriccoin.co/blog/lessons-from-the-history-of-att...
I personally would never allow a repo with two hashing algorithms to exist on my watch.
If you have ever had to use a tool like BFG to prune large objects from a repo you'll see it's not that bad, but it does require users to re-clone.
I would want to use the same process for SHA256 - that is let it be the default for new projects and then convert older projects based on need.
But there needs to be a BFG style conversion tool that spits out an object id map as output.
Here's more info on BFG: https://rtyley.github.io/bfg-repo-cleaner/
The sad thing is that the ARX BLAKEx functions seem to be gaining undeserved amounts of hype. I do not think they are getting comparable scrutiny from researchers, seeing as BLAKEx hashes are ARX, and also changed considerably since the SHA3 contest (so it is far from clear that the scrutiny that BLAKE did receive translates to BLAKE2 or BLAKE3).
The Keccak team published a short and poignant relevant blog post back in 2017 as an answer to that notorious "Maybe skip SHA3" blog post: https://keccak.team/2017/not_arx.html
A HN commenter from 2017 explained ARX's safety downside better than I: https://news.ycombinator.com/item?id=15292103
> The nuance that's being made here is that the public cryptanalytic results we have are from researchers that need to publish. However blackhats (be it government or private) have no such need. Thus, they do not care if the analysis is elegant or neat.
> This means that ARX functions will have less published analysis, but may still be successfully attacked.
> This isn't even a new argument they're making here. It's been well understood that simple cipher designs are better, because they are easier to understand. If you can understand it well, yet not break it, that gives confidence. If you don't understand it, it might break as soon as you do.
https://en.wikipedia.org/wiki/Secure_Hash_Algorithms
Since collision resistance is roughly half the number of bits, it seems unconscionable to me that anything below 256 bit hashes even exist, because 64 bits is crackable but 128 bits effectively never will be. This was well-understood even in the 90s when MD5 and SHA were first published.
Just thinking about this for the first time, I don't buy any argument about storage or performance, since those become less important as time goes on. It feels like Linus made a mistake here, and offloaded the inevitable work of upgrading repositories onto the general public (socialized the cost) which is something that all programmers should work harder to avoid.
Said as an armchair warrior who has never accomplished anything of any importance, I realize.
Reusing the precise collision from the shattered attack is made impossible by initializing the state with anything other than the prefix from the shattered attack. But the cost for mounting such an attack yourself is only 11k USD. However, as git uses the sha1collisiondetection library, such an attack would be detected by current git. Thus, this library is a much better protection than the length encoding.
> this new version would have to contain the desired hostile code, still function as a working floppy driver, and not look like an obfuscated C code contest entry
It's still plausible that one can pull a trick like that to introduce malicious code into the repo, but improbable.
> An attacker would not just have to do that, though; this new version would have to contain the desired hostile code, still function as a working floppy driver, and not look like an obfuscated C code contest entry
The whole idea is that they want to switch away before these things become likely. They are unlikely now, but SHA-1 is only getting weaker as time goes by and more research is done.
The full quote here is even better:
"and not look like an obfuscated C code contest entry (at least not more than it already does)."
git --output-format=sha1 log abac87a^{sha1}..f787cac^{sha256}
> For a Git user interface this is relatively straightforward and conciseNo, it isn't. It's a complete and utter user interface clusterfuck. Just say no to this insanity.
But the more 'improvements' they make to it the more literal that accusation becomes in my head. And what's worse is that I've grown enough callouses now that my response is an eyeroll instead of pain. I use git all the time, but it's terrible and I need something that is better, not just sucks less. And apparently soon, because I don't know when that koolaid is going to start looking good but it's not long now.
Send help.
You forgot to include the end of that sentence, that acknowledges your issue with it:
> , but one can still imagine that users might tire of it relatively quickly.
"Tire of it quickly" and "have an immediate gag reflex" are two completely different categories of negative reaction.
It's hard to see the sunset when you're down in the muck, and eventually 'less bad' starts to look like progress to you. It's a trap and you should be aware of it.
I was hoping I captured that by saying "very rarely". However, if SHA1 collisions can be made willingly, doesn't that mean that one can also willingly make a SHA1 hash that matches with the prefix of an existing SHA256 hash?
When and if someone injects a SHA1 attack into your repository, and the main git CLI throws up its hands and says "hash collision" trying to access it, I'm not seeing major problems here. The git CLI doesn't need to provide convenient commands to interact with attacks that are not practical today. To the extent that these will become practical, I think git should drop the SHA1 lookup after a migration period regardless, and it would not hurt to provide a gitconfig knob to disable SHA1 lookup.
No, the “prefix of an existing SHA256 hash” stops being relevant at that point – that’s just a full preimage attack on SHA1. Isn’t known to be feasible yet.
> I was hoping I captured that by saying "very rarely"
It’s rarer than that. :)
Don't cave in to sky-is-falling bullshit regarding the existing SHA1.
Git is not a crypto system; it's just version control.
We've used version control systems just fine that had no integrity features at all. For isntance you can go into a RCS ,v file and diddle anything you want. Some BSD people are still on CVS, and their world hasn't fallen apart.
The average git command is along the lines of "git ph-nglui --mglw=nafh Cthulhu...R'lyeh -- wgah^nagl fhtagn"
git diff
git commit -p
git rebase -i HEAD~3
The command quoted in my original comment is just this we strip away the SHA256 garbage: git log abac87a..f787cac
(Or maybe it is: git log abac87a^..f787cac^
I cannot guess whether the ^ operator still has the same meaning or whether it is part of this ^{sha...} notation.)The hashes will typically be copy and pasted, so you type just the git log, .. and spaces.
The fixed parts of convoluted git syntax can be hidden behind shell functions and aliases. But notations for referencing objects are not fixed; they will end up as arguments.
This isn't the first ^{...} notation. The manpage gitrevisions(7) also mentions <rev>^{/<text>} for referencing a commit based on a regular expression of its commit message, like
git checkout 'add-search^{/finished query builder}'
Though, this new notation is probably more in-line with the notation <rev>^{<type>}, which lets you disambiguate what you put in <rev> as in deadbeef^{tag}, so that it's not confused with deadbeef^{commit}.EDIT: The article doesn't mention it, but I imagine one interpretation would take precedence and cause git to issue a warning when it's ambiguous. Right now, if I tag a commit with the hash of another commit, its interpretation as a tag takes precedence and I get a warning at the top, "warning: refname '368bc6e' is ambiguous." That would mean you'd only ever write ^{sha256} when the provided part of a sha256 hash is ambiguous with an existing sha1 hash or something else like a tag. That's also vice versa with ^{sha1}.
> 'For a Git user interface this is relatively straightforward and concise'.
It kinda looks like you missed the joke and are now doubling-down on your disagreement.
The author does not think the proposed example is reasonable. You're in agreement.
But unless you're going to take this up with Linus, you're just yelling at your fellow disappointed spectators.
Anyone who compares the git CLI to being driven insane by Elder Gods is not defending the git CLI.
I know he's not defending it.
What I said is that he (kazinator) is inadvertently attacking somebody that's also not defending it (the author).
Well, let's see, the Fossil equivalents are:
1. Do nothing at all for a conversion from the SHA-1 to SHA-3 — yes, 3, not 2 as in Git! — because it's automatic for months now and dead easy going back 3 years now. (https://www.fossil-scm.org/fossil/doc/trunk/www/hashpolicy.w...)
2. "fossil diff"
3. "fossil ci"
4. Why are you rebasing in the first place, again? https://www.fossil-scm.org/fossil/doc/trunk/www/rebaseharm.m...
> Rebasing is the same as lying
And I think, "Holy crud do I not want to be part of this community."
The nice thing about Git is that (within reason) once I understood it, I was able to use it in very flexible ways.
It's really common for different projects I manage to range all over the place from the extreme "commits as literal history" perspective all the way to the "commits as literature/guide" perspective. Sometimes I don't rebase at all, sometimes I rebase a lot. Sometimes I commit everything, all the time, sometimes I refuse to commit any code that isn't a deployable feature. Sometimes I leave branches as historical artifacts, sometimes I don't care about history and I'm just trying to coordinate developers across timelines.
That's not to say that Git isn't opinionated about some things -- nearly all good tools have at least a few strong opinions. But Git passes the (IMO extremely low) bar of not conflating a workflow decision with a moral failing. Over the years as a software engineer, I've learned to be somewhat skeptical of programming/workflow heuristics advertised as rules, and to be very skeptical of heuristics advertised as ideologies.
I really don't understand the perspective of someone who can't think of even one good reason why they would ever want to edit history. You've never accidentally committed a password to repo, or had to respond to a takedown request?
The fact that Fossil preserves history does not prevent you from coordinating with people across timelines. It is rather the whole point of a DVCS.
> conflating a workflow decision with a moral failing
I think it's fairer to say that we don't think a data repository is any place for lies of any sort, even white lies.
> I've learned to be somewhat skeptical of programming/workflow heuristics advertised as rules, and to be very skeptical of heuristics advertised as ideologies.
Sure, flexible tools are often better than inflexible ones, but you also have to consider the cost of the flexibility. Here, it means someone can say "this happened at some point in the past," and it's just plain wrong.
That isn't always an important thing. Most filesystems and databases operate on the same principle, presenting only the current truth, not any past truth.
Yet, we also have snapshotting in DBMSes and filesystems, because it's often very useful to be able to say, "This was the state of the system as of 2020.02.04."
You don't need a snapshotting filesystem for everything, and you don't need Fossil for everything, but it sure is nice to have ready access to both when needed.
> You've never accidentally committed a password to repo, or had to respond to a takedown request?
Fossil has shunning for that: https://fossil-scm.org/fossil/doc/trunk/www/shunning.wiki
And no, shunning is nothing at all like rebase, which should be clear from the article.
Fossil also has the `amend` command: http://fossil-scm.org/fossil/help?cmd=amend
And no, it is also not like rebase, because it only adds to the project history, it never destroys information.
I, too, wish this extreme hyperbole would be just left out of the discussion completely. It is offputting, and I think it's intentionally a bad faith argument, it fails to acknowledge the utility, the design intent, and the context behind rebase, which has been talked about at length by Linus and others.
When rebase is used as designed, according to the golden rule, it's not modifying published history, so it's not "lying". Whether rebase has safety problems is a separate issue from whether it's use as designed amounts to being "dishonest".
I'm all in favor of improved design choices, and if Fossil is making those better design choices, let them stand on their own without intentionally denigrating git and every user of git through utter exaggeration.
When I revise history in Git, even if it's just doing something as simple as removing sensitive information, I often need to replace that information, either through new commits, or by introducing minor edits to surrounding commits. I could add those changes on top of my current HEAD, but then checkouts of old versions would be broken. On the other hand, if I can just replay my commits while inserting extra code, I'll end up with something that's pretty close to my original history, with just the offending information excluded/replaced.
That carries the cost that people will need to force pull my repo, but at least the repo history will still roughly correspond to what development looked like, rather than being out-of-order and mostly impossible to build except for at my current HEAD.
As a followup question, what do you do if the sensitive information you need to exclude is in a commit message? `amend` won't help you, since it's not destroying information. Do you shun that commit and then... what?
It just seems like destroying information isn't enough unless you can also replace it?
> Sure, flexible tools are often better than inflexible ones, but you also have to consider the cost of the flexibility.
I appreciate this -- I like having multiple tools for different purposes. I don't see a problem with having a VC that focuses on auditability, or having one that goes in a radically different direction from Git. Fossil has very interesting ideas, which is why I try to pay it some attention whenever I see it mentioned or linked to.
However, whenever I follow those links and start digging deeper into the philosophy behind its design decisions, inevitably the conversation changes from, "here's our alternative approach to Git" to "what Git does is fundamentally wrong". It's not, "Fossil doesn't have this problem because we eschew rebasing", it's "why would anyone rebase?"
(Nearly) all architectural decisions have good and bad consequences. Sometimes those consequences are imbalanced, so we have heuristics that can say things like, "often X is a bad idea." That's fine.
More harmfully, sometimes people extend heuristics into rules that say, "it's never a good idea to do X". Programming rules are usually wrong.
But programming ideologies the worst, because they say, "there is something mentally or morally wrong with a person who would do X". This is toxic for the reasons that Fossil devs already mention in their documentation:
> programmers should avoid linking their code with their sense of self
Programming ideologies explicitly encourage developers to have egos, because ideology conflates architectural decisions and workflow processes with individual worth. Programming ideologies make it harder for people to grow as programmers, because they tie intellectual growth to fears about being wrong. They're completely toxic.
And is Fossil's documentation promoting an ideology? I'm guessing that you'd disagree with me on this, but my take is that when Fossil's official documentation says things like:
> Honorable writers adjust their narrative to fit history. Rebase adjusts history to fit the narrative.
or
> It is dishonest. It deliberately omits historical information. It causes problems for collaboration. And it has no offsetting benefits.
That's not designing a focused tool to support specific heuristics, or making a case that, "sometimes strict auditability is important". That's just trolling for fights.
No. You start with the ideology based on your local culture and project needs, then you pick the tool that supports your project's needs.
This is why we spend so much time talking about philosophy in the Fossil vs. Git article, particularly this section: https://fossil-scm.org/fossil/doc/trunk/www/fossil-v-git.wik...
Which of the two philosophies matches better with the way your project works? That alone is a pretty good guide to whether you want Fossil or Git. (Or something else!)
A foobar is an assertion that there is a single right or wrong way to look at the world. Not even just a single correct or incorrect way, but a right way -- a proper way. If the problem with a rule is that it overgeneralizes what the world is, the problem with a foobar is that it generalizes what the world ought to be. To the extent that a foobar allows space for deviation or alternate approaches to architecture, it's only with the implicit understanding that those deviations are on some level, a kind of small sin.
Not all foobars are necessarily wrong, but in the world of software, they are particularly dangerous, and should be approached with caution. Under a foobar, a rebase isn't an organization strategy, it's a "white lie". A writer isn't optimizing for a specific audience or purpose, they're "honerable". An agreed-upon set of rules for everyone accessing a repo can be "dishonest".
Different people have different standards for this kind of thing -- but is it really all that weird or abnormal to worry that this kind of language can encourage toxicity in a community, or that it could encourage developers to think of architectural outcomes as personal validations or attacks? To me, that language sounds very foobar, and it makes me nervous about what experience I'm going to have if I adopt Fossil and then start asking questions to the community about how to use it in unconventional ways.
To be clear, there are other pages in Fossil's documentation that are much, much better about this kind of thing (particularly the Fossil vs Git page).
But even on those pages, the thing is: I use Git constantly. I am intimately familiar with its strengths and flaws. I really don't need the documentation to tell me that Git's storage is an "ad-hock pile-of-files", because I've worked with those files before and built 3rd-party tools to manipulate them, and while there are flaws, sometimes being able to do a completely dependency-free read on any OS/platform to find the current HEAD is quite useful.
When I read the docs, I just want to know what makes your software different. You're not going to convince me that actually all of my experiences were wrong, and everything I like about Git is secretly terrible. You might be able to convince me that there are specific problems Git isn't optimized for, and that Fossil can solve them.
When Fossil is talking about Bazaar and Cathedral development, I'm really interested in learning more. When Fossil is taking cheap shots at purposeful design decisions in Git that are actually really good for certain classes of problems, I lose confidence that the docs know what they're talking about.
If I were to attempt to help with the terminology, instead of sorting out the definition of ideology, I might say you’re talking about dogma and wyoung2 was referring to philosophy most directly above, but indirectly using philosophy to justify dogma.
There’s a fairly stark irony here in using words like ‘lie’ and ‘dishonest’ to judge this git workflow while at the same time taking cheap shots... but in the end I suppose the Fossil devs can describe things any way they want, and I don’t have to like it or use Fossil.
If you think your filesystem-based Git repo is easy to manipulate, go poking around in there, and what you'll find is a bespoke one-off pile-of-files database! Given a choice between Git's DB and SQLite, I put more trust into SQLite.
> I just want a C program in /usr/bin that does version control.
...which Git doesn't provide. Git is hundreds of files scattered all over your filesystem, a large number of which aren't C binaries anyway, and of those that are, only one of them is the front-end program sitting in /usr/bin, whereas Fossil can be built to a single static executable in /usr/bin.
And if you can't build Fossil statically on your system, it's likely due to an OS limitation rather than something about Fossil itself, as on RHEL where they've made fully static linking rather difficult in the past few releases.
Getting back to Git, large chunks of Git are written in POSIX shell, Perl, Python, and Tcl/Tk. Almost all of Fossil is written in C, and the rest of the code is embedded within that binary running under built-in interpreters rather than depending on platform interpreters.
This has nice knock-on effects, one of which is that Fossil is truly native on Windows, whereas you have to drag along a Linux portability environment to run Git on Windows. Another is that Fossil plays nicely with chroot/jail/container technology.
> I'm also not interested in version control systems that are dragging along a wiki and bug tracker.
Not a GitHub or GitLab user, then, I'm guessing?
> Not a GitHub or GitLab user, then, I'm guessing?
Absolutely not.
You mean your password is something funny? In that case random.org can help you with password generation.
The point is valid that when we rebase, we are losing history: the context of where that change was originally parented.
However, (1) the history does not matter if the change was parented in some temporary context, like your unpublished changes and (2) the information can be tracked in other ways, such as a Gerrit Change-Id (or something like it) in the commit message.
Regarding (1) the extra parent pointers in a merge commit cause retention of garbage. If we do everything with merge instead of rebase, we will never lose any of the temporary commits. If we prepare an unpublished change through numerous rebase operations, all that temporary crap will stay referenced from the head, waste space and confuse other people with irrelevant information when they try to navigate the history.
It does if it means a big ball o' hackage lands on the public working branch, since it complicates merges, backouts, cherrypicks, and bisects.
Git users can also hide individual commit messages behind one big combined message, losing part of the project's development history and logical progression.
When I pull your repo and build it, and I find that it doesn't build on my system, I don't want to dig through a 500-line merge commit to figure out why you changed this one line from the one that used to build last week, I want the 14-line diff it was part of so I can begin to understand what you were thinking when you committed it. If I later find out that that 14-line change was wrong but the rest of your 500-line merge was fine, I want to be able to back it out with a single command. (In Fossil, it's `fossil merge --backout abcd1234`.)
> confuse other people with irrelevant information when they try to navigate the history.
How much time do you spend navigating the project's history vs looking at the tip of the current branch?
I'd wager that the times you dig back into the history, it's because you are in fact trying to figure out why you got here, which means a trail of detailed breadcrumbs will be more likely helpful than "...and between one week and the next, something changed in commit abcd1234, but we've lost all of its internal context, so we'll be spending next week reconstructing it because Angie's on vacation now."
It's reasonably common for me to start exploring a problem space, stub out a concept, and have a long drawn-out conversation with the compiler that touches many files, before finally reaching a point that is working enough to be interesting.
At that point, I can take a step back and note that actually, not all of those changes have to be made all at once, and I can break that patch up into a bunch of simpler pieces.
In Git, I have two essentially-equivalent choices:
1. stage the commit in small chunks, adding a separate descriptive comment for each.
2. commit everything as a WIP commit so this working state is in the reflog, then break it up into smaller commits with interactive rebase.
In either process, I'm able to get a better comprehension of my own thoughts along the way.
Fossil, refusing both staging would force me to commit the proverbial 500-line blob all at once, which is less helpful to the reviewer trying to discern my thought process.
If rebasing isn't important to your workflow, Fossil probably is a better choice for you than Git. It has a lot of comforts that I really appreciated even eight years ago when I did frequently use Fossil (distributed wiki and tickets are really nice, the web interface serving raw artifacts from any point in time is great for HTML5 game jams, SQLite is a very portable repository format, etc. etc. etc.), and I'm sure it's only improved in that regard. The only reason I use Git for personal projects is because rebase is that helpful to my process.
Nope.
If it were me doing such a thing as you describe, I'd start the work on a feature branch. If I'm working on that repo with other active developers, this lets them see what I'm up to and possibly help; and if not help, then at least be aware about where my head's at, so they can better predict what's likely to land on the shared working branch later.
If I got to a point where only part of the branch needed to be applied, I could cherrypick those individual changes, either down to the parent branch or up to a higher-level feature branch.
All of this happens in public, with the work fully recorded, so someone doesn't have to reconstruct the development history after the fact later.
This mode of development helps keep your project's bus factor above 1.
To be clear, my typical approach is certainly to commit every time I return to a working state. But in more experimental modes, I often reach that the long-way-around and have ended up with multiple semantic changes I wish to break apart for study.
Unless I'm missing something and Fossil has gained the ability to cherry-pick selective lines from a commit.
The connection between the two is that git has a script called git rebase, which has an interactive mode, and that can squash commits.
git merge has squash functionality also (git merge --squash).
Rebase on a local working copy is normally used for cleaning up a string of commits that is messy and/or separating commits that mixed multiple logical changes together.
Local commit history before push is arbitrary. There’s nothing sacred that needs to be preserved about the exact order I typed things into each file, that’s not what I want from a version control system.
Personally, I haven’t really seen use of rebase complicating merges, reverts, cherry picks, or bisects. I can imagine ways it can happen, but I haven’t seen it be a problem in practice. However, I have seen cases where failing to rebase caused problems. Allowing build breakage between two commits is an example where bisect is affected, and squashing the fix into the first commit before push is much preferred.
So, anyway, your example feels totally contrived.
> How much time do you spend navigating the project’s history vs looking at the tip of the current branch?
This is a false dichotomy. I need both. I happen to navigate project history quite a lot, like multiple times per day. In addition to how I got here I usually need to know who changed it, so I can talk to them.
What happens if a Fossil repo that has had SHA3 commits written to it is accessed by old Fossil software before that change was introduced?
If you have an old clone made from before the transition and try to update it, I'm not sure what it says, since I don't have any of those around any more. It has, after all, been three years since Fossil began to move on this problem, so that it's largely a past issue for us now.
This transition time was indeed annoying for us over in Fossil land, but Git's going to have to go through a transition like this, too. The question isn't whether but how long we'll have to wait for it to begin and how long it'll take to complete.
The moment I can't read new repos with an installation of git 1.6 or 1.7, I'm ditching the garbage and finding something else.
Forward and backward compatibility, forever, please!
Beyond about 10 years, you usually end up freezing old binaries in place along with old data in order to continue manipulating it anyway.
SHA-256 is fine. The biggest problem is switching to it...
No. Hashes are really cheap.
This annoys me a bit, because every discussion about hashing goes into endless bikeshedding which hash function to use. The simple truth is: SHA2, SHA3, Blake2/3 are all good enough from both a security and performance perspective that for almost any use case and the advantages and disadvantages are so minor that it really doesn't matter.
Personally, the last time I was in a place where I had to choose which cryptography to use, I used SHA3’s direct predecessor, RadioGatún, because I needed a combined hash + stream cipher and, at the time (late 2007), RadioGatún was the only option.
RadioGatún also benefits from being about as fast as BLAKE2 (it would be faster in hardware, FWIW, having SHA3’s hardware advantages), and is approaching 14 years old without being broken by cryptoanalysis. Also, unlike BLAKE2/3, and like SHA3 and all sponge functions, it’s computationally expensive to “fast forward” in RadioGatún’s XOF (stream cipher, if you will) mode, which is beneficial for things like password hashing. Another nice thing about RadioGatún: It doesn’t have any magic constants in its specification, allowing a useful implementation to fit on my coffee mug, e.g.
#include<stdio.h>//RadioGatun
#include<stdint.h>/*32-bit**/
#define b(z) for(c=0;c<z;c++)
uint32_t c,e[42],f[42],g=19,h
=13,n[45],i,j,k;void m(){j=0;
b(12)f[c+c%3*h]^=e[c+1];b(g){
i=c*7%g;k=e[i++];k^=e[i%g]|~e
[(i+1)%g];j=j+c;n[c]=n[c+g]=k
>>j%32|k<<-j%32;}for(i=39;i--
;f[i+1]=f[i])e[i]=n[i]^n[i+1]
^n[i+4];b(3)e[c+h]^=f[c*h]=f[
c*h+h];*e^=1;}int main(int c,
char**v){char*q=v[--c];for(;;
m()){b(3){for(j=0;j<4;){f[c*h
]^=k=(*q?255&*q:1)<<8*j++;e[c
+16]^=k;if(!*q++){b(18)m();b(
8){j=c;b(4)printf("%02x",(e[1
+j%2]>>8*c)&255);c=j;if(c%2)m
();}return 0&puts("");}}}}}//
If someone asked me which hash algorithm to use, I would suggest SHA-256, unless I think they needed protection from length extension attacks (so SHA-512/256), or needed an XOF (stream cipher-like) construction (so SHAKE256).If performance mattered more than a conservative security margin, BLAKE3 (software performance) or KangarooTwelve (SHA3 variant; excellent hardware performance) would be good choices. If I were to do choose a hash + XOF for use today, I would use KangarooTwelve’s variant with a little larger security margin: MarsupilamiFourteen.
The nice thing about RadioGatún is that it only takes about 2k of compiled code (and can fit in under 600 bytes of source code, as seen in the parent) to pull all this off.
This was the best way to pull it off back in 2007, when RadioGatún was the only secure Extendable-Output Function (XOF) that existed.
Than sha256 will likely be preferable in the long run: It's faster with SHA-NI than blake3.
If you're not developing on a system with sha-ni, get with the program. Zen2 is freeking awesome. :)
SHA-NI was introduced with the Intel Goldmont microarchitecture.
(On AMD the first generation zen have sha-ni, FWIW)
At this time, Intel often experiments with or introduces features that are particularly interesting for embedded usages first on the Atom. For example the already mentioned SHA-NI. Another example are the MOVBE instructions (insanely useful if you handle big-endian data, for example in network packages (I am aware that on older x86 processors, there exists the BSWAP instruction)) - they were first introduced with Atom.
Under "feedback from git people" on https://www.mercurial-scm.org/wiki/SHA1TransitionPlan
We know only collision attacks which is "produce 2 files with the same hash, but you can't control what hash". So you can't target any existing repo. You need to use social engineering to get one of your special files into a repo.
It makes the distributed aspect of git untrustworthy, as previously you knew if you pulled from anywhere and the hash was good, you’d pulled the correct code. With SHA1 being functionally broken that’s no longer necessarily the case.
It is untested, unstable code that can only write to repositories and not read them.
"Much of the work to implement the SHA‑256 transition has been done, but it remains in a relatively unstable state and most of it is not even being actively tested yet. In mid-January, carlson posted the first part of this transition code, which clearly only solves part of the problem:
"First, it contains the pieces necessary to set up repositories and write _but not read_ extensions.objectFormat. In other words, you can create a SHA‑256 repository, but will be unable to read it. "
For smaller projects (like my own), can i move to sha-256 with no expectation of backward compatibility today ?
If you want it to be write-only, sure, go ahead!
Anyone who tries to use a git client more than 5 years old wouldn't be able to pull+push to a new repo. Sounds reasonable to me. Git clients more than a few years old are pretty broken already due to TLS changes.
Keeping around a dual hash system forever sounds like baggage and complexity that outweighs the benefits.
I don't see why changing the hashing algorithm is so problematic, hence the reason why I asked the question. Converting a repository to SHA2 should be straight forward (the only issue is everyone's tooling), you could also run the repositories side-by-side. I'm genuinely interested as I think Git & Bittorrent are quite elegant solutions to complex problems.
Exactly! If you've ever worked in a corporate environment, you know the fun of having to support 10-year-old versions of your favorite cutting-edge software.
Please fork git for this and call it something else, like git6, and ensure that git6 cannot push to git repos.
Color me surprised, dropping the "blockchain" word in the middle of the introduction
Being, as it is, a chain of signed blocks.
It is entirely possible and likely that it is used for didactic purposes as many people are familiar with the blockchain structure and its use of hashes.
And honestly fair enough. The inovative part of bitcoin is not the blockchain but all the economics & game theory going on to create trust in the system
Again it is not a criticism of the article, but it is not a criticism of the criticism either.