Linus' reply on Git and SHA-1 collision
marc.info
marc.info
1) Git doesn't rely on SHA-1 for security. It relies on HTTPS, and a web of trust.
2) Even if git did rely on SHA-1, there's no imminent threat. What happened today was a SHA-1 collision, not a preimage attack. If a collision costs 2^n, a preimage attack costs 2^(2n).
3) Even if someone managed to pull off a preimage attack, creating a "poisonous" version of one your git repository's objects, they'd still have to convince you pull from their repo. This requires trust.
4) Even if you pulled it in, your git client would simply ignore their "poison" object, because it would say, "oh, no thanks, I already have that object". At worst, the code simply wouldn't work. No harm would be done.
When it comes to git, an attacker's time is better spent creating a secret buffer overflow than wasting millions of dollars on a SHA-1 collision.
Imagine the NSA publishing a crypto algorithm and contributes it to openSSL or some hypothetical crypto library using git. If they commit their new algorithm, everyone will be looking at that. They could do something devious like tinker with the way random numbers are generated elsewhere and reduce the possible keyspace of another algorithm to something very small and easy to brute force.
When this keyspace shortening is found out it would be hard or impossible to track back. No amount of inspecting the files that reportedly changed would reveal that the NSA did this.
HTTPS lets you verify that you're fetching changes from, say, a Github server. Just because a repo is hosted by Github doesn't mean that you can trust its contents.
> 2) Even if git did rely on SHA-1, there's no imminent threat. What happened today was a SHA-1 collision, not a preimage attack. If a collision costs 2^n, a preimage attack costs 2^(2n).
This needs investigation, but I suspect some plausible attack vectors exist based on collisions. Say you generate a good file and a bad file with the same SHA. If Github uses some kind of object cache for viewing files through the website, you could probably get the good file into their cache, then open a PR with a commit containing the bad file. The project maintainer would see the cached good file, but when they merge the PR, the bad file would be merged.
I'm not sure if this exact attack would work, but probably something like it would. If not with Github, then perhaps with Bitbucket or GitLab.
Another possible approach would be to send a PR which discreetly introduces the bad file into the project maintainer's object store. (It doesn't have to be in a commit you're asking them to merge; it could be in a separate branch which they likely wouldn't notice.) Once that's done, you send a PR which introduces the good file. If they review the second PR on a website, they'll see the good file, but if they merge manually, they'll get the bad file from their local object store. Even if the maintainer merges through Github, if they deploy from their own machine, they'll deploy the bad file.
> 3) Even if someone managed to pull off a preimage attack, creating a "poisonous" version of one your git repository's objects, they'd still have to convince you pull from their repo. This requires trust.
Not really -- on Github, contributors are often strangers, and project maintainers often fetch from a stranger's repo in order to try something out. They expect that fetching objects is harmless. Even if some project maintainers do look for signs of trust before fetching, they'd probably be easily fooled by fake info in a profile. A determined attacker could even make a large network of fake accounts which star each other's projects and so forth, similar to black hat web rings.
> 4) Even if you pulled it in, your git client would simply ignore their "poison" object, because it would say, "oh, no thanks, I already have that object". At worst, the code simply wouldn't work. No harm would be done.
If an attacker had the resources to perform a preimage attack, one thing they could do is take the latest version of jQuery on the day it's released. They could append some malicious code, then add junk in a comment until they get the desired SHA. Now they just have to get a project maintainer to fetch from their repo ("Check out this feature I added! Just fetch from my fork and run the server."), and then wait for the maintainer to upgrade jQuery.
Face it - it's far more likely that a github account is compromised and that repo you rely on has been amended. And you don't really have a good way of verifying which commits are "safe" whatever that means. At best, commits can be cryptographically signed by ther authors, to prevent this. But if the author goes rogue then all who depend on them are up the creek.
And all this for what?
For applications, such as signing pdf or other documents, SHA-1 should be retired. But the collective crypto community have said that for more than a decade, so I have little sympathy for companies that is affected by this.
But for Git, there is no reason for immediate concern. Should they upgrade to a more secure hash function, yes. Ideally they make it rather straight forward to use different function in the future and perhaps multiple hash functions. I doubt anyone is able to find an input that MD5 and SHA-1 both hash to the same value. It would significantly reduce reliance on a single hash function.
With the attack vectors discussed above, the type would always be blob. I don't think the length field is much help either; see https://news.ycombinator.com/item?id=13720725
In all the examples in your link, while true that you can add "silent data" to all of the examples, they are all examples of structured data. So not only do you have to figure out the collision, you have to do so within the structure of the format you're attacking. This is a lot harder than just prepending the exact bytes you want.
That's not how the attack works. Not at all.
> This is an identical-prefix collision attack, where a given prefix P is extended with two distinct near-collision block pairs such that they collide for any suffix S.
And
> Our example colliding files only differ in two successive random-looking message blocks generated by our attack. We exploit these limited differences to craft two colliding PDF documents containing arbitrary distinct images.
As was mentioned in the original collision thread, concatenating insecure hash functions has fundamental weaknesses: https://www.iacr.org/archive/crypto2004/31520306/multicollis...
Let us assume that we have 2 160bit hash functions, and one is broken to the degree that we can find collisions in constant time. This now means that we can break the combined hash in 2^80 rather than 2^80 + 2^80. The total brute force complexity was not improved, but the reliance on either hash function was.
I thought that if you signed your git commits or tags with GPG, you implicitly relied on the commit checksum, and thus on SHA1.
Is that right?
One day Egor Homakov finds a new hack and finds his way into access to the master branch for library X. As a prank, he force pushes a change to master.
Being aware of such a possibility, instead of setting up my script to pull from master or even a specific tag, perhaps I pull from a specific sha1sum. That way I know I'm getting the exact version I want. That's today. Tomorrow, when preimage SHA-1 attacks are cheap, it will no longer save me.
There is no way to securely deploy a package directly from the internet. The sooner you understand that, the sooner you'll be able to sleep through needless catastrophes like this hypothetical attack -- or leftpad.js.
Besides, some people use Heroku simply to deploy their rails or node apps, and they don't have an artefacts server.
Fetching a file from the internet and verifying it against a hash or signature is the only way to securely deploy a package from the internet. This is exactly one of the use cases theses cryptographic primitives are built for. Don't blame the users when they utilize the tools the way they are supposed to be, blame the tools when they fail when doing so.
And by the way, if you ever need to run `apt-get upgrade` or `apt-get install` on your production server from a public mirror, then you are guilty of "deploying a package directly from the internet", too. (My apologies and congratulations if don't.)
But yeah, you could put an unverified repos to be used
git checkout 0.1.0
git tag -v 0.1.0Tag signatures are mostly worthless now from a crypto point of view -- with the caveat that you can still get some value from them if you still trust sha1 to be secure against second-preimage attacks.
Not only do you likely need to populate your host with packages not from your host. But also, your host will also still be connect to a public net, even if only indirectly (e.g. private net), and hence potentially manipulated.
Not to mention that with pre-built binary packages your deployment speed and repeatability get significantly better, as you don't need to rebuild the artifacts every single time.
Ansible gives larger and easier to exploit attack vectors than git.
And then, with pulling semi-random things from internets you have much bigger problems with your deployment procedure. You should never ever download a git repository with software, instead you should be using package system supplied by your operating system.
2005-me vehemently agrees with you.
Also, from rants I hear about Homebrew and its breaking random libraries on upgrades, I don't think it should be mentioned as it was state-of-the-art or something.
Npm: https://docs.npmjs.com/cli/install
Bundler: https://bundler.io/git.html
Homebrew: https://github.com/Homebrew/brew
You are picking on the useless details. Those are all commonly used package managers in production environments. They often provide software that simply never gets packaged with the OS. They likely always will, because they have more focused design goals.
A better way to argue this would be point out that specific ways to use the package managers better. For example: bundlr supports saving all the required packages offline. This provides the opportunity to do a security review and save the packages locally/internally rather than always trusting the whatever is on the Internet.
OS-supplied packages don't grow on magical trees. If you don't have the necessary software in official repositories (or if it's your software), you can package it yourself. Deployment then becomes a breeze, and you save yourself otherwise completely useless process of recompiling things over and over again.
> You are picking on the useless details.
Quite the contrary. Those details make important difference.
> Those are all commonly used package managers in production environments. They often provide software that simply never gets packaged with the OS.
Apart from Homebrew, which is for workstations (hardly anybody runs macOS servers), none of these "package managers used in production environments" provide you a complete way to rebuild your software. You can be fine for a while if you stay away from modules that are interfaces to C or C++ libraries and from tools from other languages (e.g. I have used Python's Sphinx to document Erlang daemons quite successfully), but once you hit that, deployment starts to be PITA, because you'll need to remember to install all the required libraries, -dev packages, compilers, and what not.
On the other hand, DEB or RPM with artifacts will just automatically pull the required libraries, and its build dependencies give a dedicated and standard place for the necessary build tools.
Your comment supports my opinion that today's programmers usually don't want to be bothered with learning things that have been working for sysadmins for twenty years already.
And I have yet to hear that port and emerge are incapable of doing dependency resolution. They're battle tested systems that work well in production environments.
Second, ports and portage have support for and networks of mirror servers that keep copies of software available through these packaging systems. It's trivial to switch if one of the mirrors goes down. For pip, gems, or npm you need to plan ahead for the problems and deploy your own package cache, from what I know.
Third, I was using Gentoo with one of these "battle tested systems that work well in production" for several years. It was doable, but it wasn't pretty, could lead to breaking software after updating some random deep dependency (if it was recompiled with different flags), and generally required more work and attention than APT would, all that for very little gain (if any gain at all). Oh, and it ended up working with binary packages, after all, I just needed to compile them myself instead of having a half an hour downtime of production MySQL because it needed to get compiled (which could fail, leaving me with no working database installed).
I bring that up to highlight some of the wrong assumptions you make. You make several needless assumptions and use those to draw funny distinctions between things. I am not even sure of the point anymore.
likely any of these systems could be used in a variety of environments for a variety of purposes.
Those are development tools, not deployment ones (even though they are used as such; programmers usually don't bother with learning what sysadmins do, so it's not a surprise).
Well, Moore's law is petering out isn't it? :)
However, the largest change we have seen the last couple of years is the accessibility of computer power. These days an attacker can be concerned mostly by how many CPU/GPU hours he needs to rent from Amazon or equivalent providers to achieve his goal. So the accessibility, convenience, and price of processing power is still improving quite fast.
1) Git doesn't rely on SHA-1 for security. It relies on HTTPS, and a web of trust.
I don't think I've ever (intentionally) used git over HTTPS. I always clone using ssh (which has its own authentication mechanisms) or the git protocol (which is read-only).
2) Even if git did rely on SHA-1, there's no imminent threat. What happened today was a SHA-1 collision, not a preimage attack. If a collision costs 2^n, a preimage attack costs 2^(2n).
Thankfully, the collision attack doesn't apply to git (see above) so the cost ought to be greater than 2^n for a collision.
3) Even if someone managed to pull off a preimage attack, creating a "poisonous" version of one your git repository's objects, they'd still have to convince you pull from their repo. This requires trust.
Find a popular git host (say, Github but if a popular project is on Git Lab or Bitbucket, they will do just as well) and compromise them. Target a recent release for a popular project (say, Rails) and poison a relevant object that gets pulled down by all the downstream maintainers to package the release.
The benefit to straight up compromising a git repo without faking the hash lies in introducing a vulnerability without the maintainers nor developers of the project noticing (or noticing months/years after the fact).
Please note that the shattered-{1,2}.pdf files both have exactly the same length. And even with cleartext it is easy to pad passages so they contain the same amount of bytes. See how the quoted paragraph above has exactly the same number of characters as this one.
>>> len('''> Wow, Linus raises an entirely different issue which is that the PDF-based attack can't and won't work on git at all. Due to length prefixing it is extremely difficult to insert some nonsense into the middle of a git object which is how this attack works on PDFs.''')
264
>>> len('''Please note that the shattered-{1,2}.pdf files both have exactly the same length. And even with cleartext it is easy to pad passages so they contain the same amount of bytes. See how the quoted paragraph above has exactly the same number of characters as this one.''')
264I would say cryptographic integrity definitely counts as "security".
I agree this isn't currently a big deal, but it's probably time to start migrating to a better hash.
> Do we want to migrate to another hash? Yes.
I don't really see how HTTPS is relevant here either, I clone most of my repositories over SSH for instance. And you can use git over plain HTTP too.
Fourth, people use git in creative ways. Linus may think it is a cardinal sin to commit binary blobs in a git repository, but I can't imagine I'm the only one using git as a poor man's backup and file sharing solution.
And last but not least, relying on Sha1 takes effort of constantly asserting its use in Git is still secure. Support request to that end will skyrocket from now on, both of the constructive kind, like the technically concerned coworker ("but isn't git insecure now that Sha1 is broken"), and of the regulatory kind ("if you use Sha1, MD5, … please fill out these extra forms explaining why your process is still eligible for certification with ISO norm foobar").
Since we have to migrate away from Sha1 at some point in the future, I'd like it to be sooner rather than later.
[1] See Wikipedia for a timeline and references: https://en.wikipedia.org/wiki/MD5#History_and_cryptanalysis
If it helps, the git devs recognized that SHA-1 would be replaced at some point and have been preparing to move away from it. It's just a lot of work on basically one volunteer. A non-SHA-1 prototype might show up in a year or two, hopefully.
"A contrived collision on MD5 in 2004 got perfected to a single block collision in 2010 [1]" so they'd have at least years to fix it were they using md5?
He says they'll migrate, but it's no reason to go crazy. If anything, calmness of this sort is what we need more of (this industry, anyway... we go crazy about stuff way too much).
https://www.schneier.com/blog/archives/2005/02/cryptanalysis...
Fix what's broken, no doubt. But stay rational and look at the problem from all perspectives.
git-annex (written by Joey Hess from the email) is a way to manage binary blobs using Git, but IIRC it uses SHA256.
Some security-focused developers sign git tags and/or commits, specifically to have things verifiable end-to-end and not having to trust HTTPS and all the middle men that entails.
Would that not be a case where git relies on SHA1 for security? Someone could replace a tag or commit with a malicious version that verifies fine since it has the same hash the original developer signed.
> 3) Even if someone managed to pull off a preimage attack, creating a "poisonous" version of one your git repository's objects, they'd still have to convince you pull from their repo. This requires trust.
In the case of signed commits/tags, this opens projects up to malicious action by hosting companies and others. Usually signed commits and tags are used specifically to avoid that exposure, because the developers don't trust the infrastructure.
> 4) Even if you pulled it in, your git client would simply ignore their "poison" object, because it would say, "oh, no thanks, I already have that object". At worst, the code simply wouldn't work. No harm would be done.
That only protects existing checkouts that already have fetched that commit. What about new checkouts, or older checkouts that haven't been updated yet?
Not disputing that this SHA1 collision does not signal any immediate emergency, just pointing out that git is used in different ways by different people, and some of those uses very much do depend on git's SHA1 for security.
This is a completely orthogonal, separate layer of security that has nothing to do with this particular issue.
> Would that not be a case where git relies on SHA1 for security? Someone could replace a tag or commit with a malicious version that verifies fine since it has the same hash the original developer signed.
What's your response to this? To me it seems like this would be a serious issue.
Alas, libfoo author has been corrupted by The Adversary and later arranges for same hash to resolve to different code. (currently we're talking of collision not 2nd preimage, but libfoo's author could have been planning it all along, before your audit.)
["libfoo" here is fictional, I'm not refering to any of the projects actually named that.]
Similar things happen when Git hashes are exchanged by any side channel. The ability to use strong hashes as pointers to content inside any other data is the whole beauty of Merkle DAGs. Git submodules just happen to be one such channel that's part of git, but I think it's important to accept that git commit hashes are widely used outside git itself.
P.S. I see git-evtag already covers submodules. Nice.
This is a big lie. When the -S flag is used, git signs the SHA-1 of the commit. Moreover HTTPS does not provide any form of authentication due to the extremely broken CA model. NSA could simply ask any CA to give them a cert for github for example. Not to mention that https would only authenticate that you are talking to the github server, it would say nothing concerning the authenticity of the code.
>they'd still have to convince you pull from their repo. This requires trust.
How about compromising your servers instead? Or maybe simply have NSA asking github to let them modify a commit (which would end up having the same SHA-1 and being signed by you).
>4) Even if you pulled it in, your git client would simply ignore their "poison" object, because it would say, "oh, no thanks, I already have that object". At worst, the code simply wouldn't work. No harm would be done.
https://stackoverflow.com/a/34599081
Didn't Git v2.0 break backwards compatibility? Couldn't they simply move to Sha-2 during that time?
> Didn't Git v2.0 break backwards compatibility?
No, it did not.
> You are _literally_ arguing for the equivalent of "what if a meteorite hit my plane while it was in flight - maybe I should add three inches of high-tension armored steel around the plane, so that my passengers would be protected".
> That's not engineering. That's five-year-olds discussing building their imaginary forts ("I want gun-turrets and a mechanical horse one mile high, and my command center is 5 miles under-ground and totally encased in 5 meters of lead").
> If we want to have any kind of confidence that the hash is reall yunbreakable, we should make it not just longer than 160 bits, we should make sure that it's two or more hashes, and that they are based on totally different principles.
> And we should all digitally sign every single object too, and we should use 4096-bit PGP keys and unguessable passphrases that are at least 20 words in length. And we should then build a bunker 5 miles underground, encased in lead, so that somebody cannot flip a few bits with a ray-gun, and make us believe that the sha1's match when they don't. Oh, and we need to all wear aluminum propeller beanies to make sure that they don't use that ray-gun to make us do the modification _outselves_.
> So please stop with the theoretical sha1 attacks. It is simply NOT TRUE that you can generate an object that looks halfway sane and still gets you the sha1 you want. Even the "breakage" doesn't actually do that. And if it ever _does_ become true, it will quite possibly be thanks to some technology that breaks other hashes too.
> I worry about accidental hashes, and in 160 bits of good hashing, that just isn't an issue.
What all the commenters here doing a hatchet-job on Torvalds are missing is that he's saying "show me the money"/"perfect is the enemy of good". The hatchet-jobbers are too busy doing the usual tittering over his language to actually absorb the point.
Also, "more secure" doesn't necessarily mean "more complex", or "slower". Blake2b for instance is as fast as md5 and has a simple RAX (xor, rot, add) core from Chacha20.
There's also important concern about SHA-512 (which has led to SHA-3 and other options), but I'm not sure it can be put in the same category; at least all of the best attacks today are significantly reduced-round.
https://en.wikipedia.org/wiki/SHA-2#Cryptanalysis_and_valida...
By contrast, Wang's original attack (which we had more than 10 years ago) achieved a 2¹¹ speedup for collisions against full SHA-1 and her second attack later the same year achieved 2¹⁷ speedup.
Evidently the speedup for the final version of Stevens's attack (which worked) was also around 2¹⁷. By contrast, a Moore's law improvement over 10 years would only be expected to make computers around 2⁶ times faster! (At the risk of mixed or mangled metaphors, we might say that mathematical insight during the time period you mention has been about 2¹¹ times more useful for attacking SHA-1 than computer power increases.)
Edit: someone else linked to Valerie Aurora's chart, which I'd forgotten about, which gives a synopsis of the historical status of the most popular hashes. On that chart, SHA-1 from 2005 was one category worse ("Weakened") than SHA-512 is today ("Minor weakness").
"I have no evidence of SHA-1 being broken which is evidence that SHA-1 can't be broken"
I think this is a shockingly good example of how smart people get security questions utterly wrong. The right analogy when it comes to security has to involve some type of adversary, not just random, unmotivated natural phenomena -- as long as we're using aircraft analogies, it's not so much "a meteorite might randomly hit my plane in flight" as "there's angry and armed people shooting at my plane with armour-piercing ammunition".
Indeed, in the real world, putting hundreds of kilograms of armour on aircraft isn't the absurd/childish idea Linus seems to think it is:
https://en.wikipedia.org/wiki/Fairchild_Republic_A-10_Thunde...
And it hunts tanks.
If one of those is lingering, I wouldn't even fight if I was the enemy.
It might not be able to kill a modern main battle tank, but I bet spraying it with rounds like that would fuck it up pretty seriously.
Another story that I've heard is that a B-1 flying at operational altitude (200 ft above ground level, mach 2) was often as effective as dropping munitions.
That very topic is being discussed right now on the Aviation Stack Exchange: http://aviation.stackexchange.com/questions/35771/is-it-corr...
People tend to really underestimate how much harder it is to go faster once you reach 0.90 or so, especially at low altitudes.
But the if a small meteorite (only a few grams) were traveling that slow, 1.5 inches of titanium would do quite a bit.
"At some point, usually between 15 to 20 km (9-12 miles or 48,000-63,000 feet) altitude, the meteoroid remnants will decelerate to the point that the ablation process stops, and visible light is no longer generated. This occurs at a speed of about 2-4 km/sec (4500-9000 mph). From that point onward, the stones will rapidly decelerate further until they are falling at their terminal velocity, which will generally be somewhere between 0.1 and 0.2 km/sec (200 mph to 400 mph). Moving at these rapid speeds, the meteorite(s) will be essentially invisible during this final “dark flight” portion of their fall."
just btw, 25 km/s is almost 100.000 km/h.
https://en.wikipedia.org/wiki/Impact_event#Airbursts
One they get big enough, they slow down from ludicrously fast to still ludicrously fast. The smallest impactor shown in that table enters at 17 km/s, loses 90% of its energy traversing the atmosphere and smacks the ground at nearly 5km/s. You'll probably want to take your titanium armour and stand somewhere else.
The plane was built for survivability. It can withstand an engine being shot off (why the engines are external on "pods"). The cockpit is (I think) surrounded by a 2 inch titanium tub.
I'm not sure the JSF, i.e. F-35 will be capable of taking over this role, as is intended.
The radioactivity is unrelated to its use as bullets, which relies on its high density, but it is radioactive.
(1) P(crash | section hit)
and adding armour to those sections where that quantity is maximized (maybe with some thought to the relative weight of armour needed for each section, but I digress).Let's directly apply Bayes' rule:
(2) P(crash | section hit) = P(section hit | crash) * P(crash) / P(section hit)
The denominator can be further expanded: (3) P(section hit) = P(section hit | crash) * P(crash) + P(section hit | no crash) * P(no crash)
So we can see from the 2nd term that if aircraft regularly comes back with a section that's been hit and yet it hasn't crashed, then that directly reduces (1), meaning that section needs less relatively less protection, all else being equal.Another point in this method's favor is if crashed aircraft frames are too damaged to permit us to identify which sections were damaged. In that case, we can still estimate (1) just by replacing all the P(section hit | crash) terms with a uniform term.
This analysis can be further expanded to the actual amount of damage each section took in a as well. The more damage a section took on surviving aircraft, the less protection it needs.
It's not like a rifle bullet.
"A METHOD OF ESTIMATING PLANE VULNERABILITY BASED ON DAMAGE OF SURVIVORS" BY ABRAHAM WALD (1943) : http://www4.ncsu.edu/~swu6/documents/A_Reprint_Plane_Vulnera...
It seems like a straightforward, almost painfully obvious definition in retrospect now, after having it articulated to me. I think education on security manners is poor and should be a standard topic.
What's interesting here is that, just 5 years ago, the 2017 cost of this very attack (the Stevens attack) was estimated at 2^18.4 = $350k [0]. The collision announced today cost about $100k. Perhaps even less. [1]
Cloud computing is cheaper today than many expected it to be, and it seems like we're only now entering the era of fierce competition. Who knows what the next decade will bring. If your modeled attacker cost is within an order of magnitude or two of the danger zone, beware. Better to have many orders of magnitude of headroom.
Along these lines, I really like what cperciva did with attacker cost modeling in his scrypt paper. See the table on page 14 [2]. I wish more security choices were presented this way, with estimated attacker costs. Taking that even further, I wish the numbers were updated dynamically against present day hardware & compute costs. Maybe even with trendline projections. It's difficult to make good security choices without knowing costs.
[0] https://www.schneier.com/blog/archives/2012/10/when_will_we_...
> it could be incredible bad luck that caused that good-looking patch to be mistakenly matching a dangerous object
It's to communicate an idea to another person in a way that can also convey subtleties, and not just the literal words being conveyed.
In this case, he was trying to convey the idea that the risk is so small and so remote that it really isn't worth spending a lot of time on.
You understood the point, I understood the point, and everyone else understood the point. Which means the analogy was successful.
So please, stop trying to pull the conversation on some tangent so you expound on why smart people should spend more time on the perfect analogy that you approve of.
The purpose of an analogy is to simplify something that's too hard to understand for the person you try to convey your idea to. Sometimes analogies are appropriate, e.g. when you teach something. When you want to convince somebody whose opinion is very different from yours, analogies aren't appropriate. They sound condescending: "because you are not smart enough to understand the real rationale behind my opinion, here is an oversimplified argument based on analogies I made just for you". The question is, did Torvalds have any real argument back then? Maybe he fell back to using analogies for the lack of any real argument.
> In this case, he was trying to convey the idea that the risk is so small and so remote that it really isn't worth spending a lot of time on.
To convey the idea that the risk is very small, one needs to have a proof. Real proofs shouldn't involve analogies. They should use facts and logic.
And to be quite honest abstracting ideas is core to problem solving, and I think it's a bit disingenuous to say that anyone misunderstood what Linus was getting at there.
Sure. But don't confuse abstractions with analogies.
Here is a good proof about both the integers and rational numbers. Every integer and rational is a real number (abstraction). When you add any 2 real numbers, you get the same result regardless of their order (fact). So it must be true that, when you add any 2 integers, you get the same result regardless of their order (correct conclusion #1). It also must be true that, when you add any 2 rationals, you get the same result too (correct conclusion #2).
Here is a bad proof about the integers and rational numbers. Both the integers and rationals are very similar: you can add them, subtract them and so on (analogy). Between every 2 rational numbers there is another rational number (fact). So it must be true that between every 2 integers there is another integer (wrong conclusion).
That's just the nature of an analogy, but that doesn't make it useless or fair to attack the speaker for using an analogy.
And now we're debating on what does and doesn't qualify as an analogy.
Anyone who doesn't understand why the risk was so small doesn't belong in the conversation.
Can you even imagine where our medical field would be if we expected surgeons to talk amongst themselves as if they were speaking to the general public?
It is absolutely acceptable for the speaker to make assumptions about the listeners knowledge, and that doesn't reflect poorly on the speaker.
> The purpose of an analogy is to simplify something that's too hard to understand for the person you try to convey your idea to.
It's to convey an idea. That's it, anything you add to that is your own bias at work.
Linus was trying to get across the scale of just how small the risk was.
The Git devs are not heart surgeons, software development is not a medical field. Again, you are using an analogy to "prove" your point. Can you, please, use a real argument?
>> The purpose of an analogy is to simplify something that's too hard to understand for the person you try to convey your idea to.
> It's to convey an idea. That's it, anything you add to that is your own bias at work.
Do you disagree that using an analogy is oversimplification? If you don't, then that part about "to simplify a complex idea" in my statement should absolutely stay. If you do disagree, then please show why and how an analogy doesn't oversimplify a complex idea.
When someone uses an analogy they're not trying to prove anything, they're trying to get you to see things from a specific perspective, or they're trying to communicate an idea.
You can still walk away from the analogy and disagree with them, but you should have a better understanding of their perspective or their argument.
> Do you disagree that using an analogy is oversimplification?
I think your entire approach to analogies is unnecessarily combative. You view an analogy as someone trying to prove something rather than trying to communicate their position better (or just an idea in general).
And your approach is to point out that the analogy isn't perfect, and therefore you've "disproved" the analogy.
Only analogies are, by their very nature, imperfect. When an analogy is perfect it ceases to be an analogy and becomes the thing being discussed.
It should automatically be understood that there are the analogy isn't perfect and anyone can find flaws in it. That doesn't mean it isn't an effective way to communicate.
An analogy applies a principle to a common setting without loss of specificity. Specifically the dedicated adversary is lost in this abstraction, so it's a bad analogy.
Sure, if he's saying he is powerless against higher powers. But if he's trying to make a quantitative assertion, a qualitative analogy is not the right tool.
Using analogies is okay when you try to explain a difficult idea to somebody who wants to learn about your subject. Using analogies is not okay when you want to make a solid argument, when you want to convince somebody whose stance is very different from yours.
So this is a shockingly good example of how quotes taken out of context are pretty useless in a discussion.
Why is this quoted in support of an argument that Linus used to come across as a lunatic in online correspondence ?
This seems to me like an entirely reasonable way to make it very, very difficult to ever attack, because an attacker would have to be able to generate collisions for both of your hash functions.
(I'm not a crypto expert -- maybe you are, if so and my above comment is totally wrong, can you explain how it's wrong?)
But of course, now is not then. And I think zkms' metaphor "there's angry and armed people shooting at my plane with armour-piercing ammunition" is apt here. The threat is much more sizable now, and so the cost vs. benefit analysis has changed.
Incidentally, one of the requirements of the SHA-3 competition was that it not be related to previous hashes. So while there are no demonstratable attacks against SHA-2, we nonetheless have "two ...hashes, and that they are based on totally different principles."
[1] https://news.ycombinator.com/item?id=13715146
[2] https://www.iacr.org/archive/crypto2004/31520306/multicollis...
This was his point, and it's still true. Generating a specific SHA-1 hash is still not feasible.
Not really. It's not a preimage attack. They spent several hundred dollars to find two random byte strings with the same SHA1 hash. There's still no way to SHA1-collide a specific byte string instead of random junk.
Now to be fair, he also keeps repeating that he hasn't seen the attack yet. Which leads me to question why is this post interesting to HN? Is it to show how Linus aimlessly speculates and gets his guesses wrong?
--
He says:
> pdf's don't have that issue, they have a fixed header and you can fairly arbitrarily add silent data to the middle that just doesn't get shown.
In other words, he expects the PDFs to have the same size because silent data has been arbitrarily added to the middle that doesn't get shown.
Maybe it's possible, but it doesn't seem very likely for source. You'd be more likely to succeed if someone stores binary files in git where people don't have an easy way to audit what's causing the difference. But in that case it would seem you likely have simpler attack vectors.
I think it's interesting because it is another example of his basic attitude toward security, and a lot of people who should know better use Linux believing it is secure because "it hasn't been broken yet."
That being said, his motto is more like "don't freak out", and in the specific git context, I can only agree.
I wasn't thinking of it like that, but now that you mention it, yeah. There's definitely something to learn from the way people make assumptions and are subsequently led by them.
No. It doesn't change anything if the size is in the PDF header. The size of both PDFs are the same, the header of both PDF files is the same on the both "shattered" files now.
What Linus says is that if you tried to put these two PDF files in git, it would not see them as the same, as git calculates the sha1 differently. But Google would be able to produce two PDF files that would, as git sees them, appear to be same just as easy as these that were produced.
P.S. (answer to your answer to this message) Note, You wrote one level above
> If PDF had a header at the beginning of the file that states the file size, then it could be harder to find a collision.
And I argued that it isn't harder, but irrelevant.
From your answer:
> But to generate a collision with a different prefix q one would have to do the expensive computation all over again
Yes. Now read what your claim was again. It's not harder. Exactly as easy as the first time.
Right, but they would have to re-do their enormous calculation. ("This attack required over 9,223,372,036,854,775,808 SHA1 computations.")
Google started with a common prefix p (the PDF header), then computed blocks M11, M12, M21 and M22, such that (p || M11 || M21 || S) and (p || M12 || M22 || S) collide for any suffix S. Given p, M11, M12, M21 and M22, anyone can make colliding PDFs that show different contents quickly. But to generate a collision with a different prefix q, e.g. one including the file size, one would have to do the expensive computation all over again, I think.
Note: I'm not trying to argue that SHA-1 can be made secure with padding. I was just trying to say that the statement "The PDFs have the same size" misses the point.
But it does add an extra constraint, and it prevent attacks that rely on random length filler data. This is not much of an improvement, but an improvement non the less.
shattered-1.pdf and shattered-2.pdf have the same size and sha-1 hash, but git is still able to recognize that those are two different files and creates two different commit hashes for them. So clearly just having the same size and sha-1 hash is not enough to fool git.
The product was cancelled. I always wondered if the patch would be of any use to anyone.
You think? I dunno why you wouldn't try to submit a pull request though.
http://crypto.stackexchange.com/questions/9435/is-truncating...
At best you reduce the brute force complexity, at worst you enable pre-image attacks.
One thing I hate about crypto talk is statements like this So, truncating one of the SHA-2 functions to 160 bits is around 2^20 times stronger when it comes to collision resistance.
Which is all too broad. What if SHA-1 is down to 2^10, is truncated SHA-2 2^30? Does it mean we have proved that no weakness exist in SHA-2? A correct statement would simply be that no known attack exists on truncated SHA-2 yet.No. Moore's Law has been dead for years and will never come back. The benefits we saw in recent years came from people figuring out how to compile code for SIMD processors like GPU's, not faster or cheaper silicon.
> Do we want to migrate to another hash? Yes.
Wouldn't all that time trying to explain away the SHA-1 issues be better spent on developing a safe transition plan? Work on this could have started long ago, and if it would have started, going from SHA-256 to SHA-512 to SHA-3 to ... would be a no-brainer by now.
In the simplest case, ensure that all newly created git repositories work woth SHA-256 by default (or SHA-512, or whatever), and switch back to SHA-1 for old repositories.
In the more advanced case, provide the possibility for existing repositories to have multiple hash values (SHA-1, SHA-256) for every blob/commit, then phasing out client support for old hashes as time goes on. When some SHA-1 collision happens, those who use newer git versions would notice and keep having a consistent repository.
If all those different browsers and web servers were able to coordinate a SSL/TLS hash transition SHA-1 to SHA-256, then a protocol like git with roughly 2 widespread implementations should be able to do that, too.
I read through this thread yesterday, and walked away with the impression that they have started working towards a hash migration (and general cryptoagility) already, albeit not with much priority.
> pdf's don't have that issue, they have a fixed header and you can fairly arbitrarily add silent data to the middle that just doesn't get shown.
This doesn't seem like much of an obstacle, since you can add silent data to all kinds of files, like
- With HTML, JS, etc. you can just add whitespace.
- Some formats like GIF89a have variable-length comments.
- With any media format that uses palettes, you can add extra, unused colors.
- Just about any compression algorithm can be tuned to manipulate the compressed size. E.g. with DEFLATE (which is used by PNG in addition to some archive formats), you can use a suboptimal static coding rather than the correct Huffman tree.
- With most human-readable document formats, you can add zero-width spaces or something similar.
It is much more difficult to change source code in a way that:
1) generates a collision 2) is still valid source code 3) the changes cause a desired effect (like a backdoor) 4) has the same file size
pretty easy with comments.
Given a case where someone with permission to push gets compromised and a malicious actor can pull this sha-1 attack off, aren't there bigger problems at hand? The history will be there and detectable or if they're rewriting history, usually that's pretty noticeable too.
I may be totally missing a situation where this could totally screw someone, but it just seems highly unlikely to me that people will get burned by this unless the stars align and they're totally oblivious to their repo history. So I guess I agree with the "the sky isn't falling" assessment.
If you were able to forge commits with the same SHA-1 you might theoretically be able to rewrite the history without invalidating the signatures, which would be a problem. We're not there yet though, but it's one step closer.
Now suppose they've contributed a commit that collides with another one that puts a backdoor in, and they use this very selectively, MitMing git clones made by high-value targets. If anyone were to compare one of those clones against the real kernel repository then that would burn their operative, sure. But how likely is it that anyone would ever do that?
I have put git commits into scripts I run on automated servers for example, to be sure that every server runs exactly the same copy of the program.
I don't know much about git internals, so forgive me if that is a bad idea, but what does everyone think about it working like this:
If future versions of git were updated to support multiple hash functions with the 'old legacy default' being sha1. In this mode of operation you could add or remove active hashes through a configuration, so that you could perform any integrity checks using possibly more than one hash at the same time (sha1 and sha256). If the performance gets bad, you could turn off the one that you didn't care about.
This way by the time the same problem rolls around with the next hash function being weakened, someone will probably have already added support for various new hash functions. Once old hash functions become outdated you can just remove them from your config like you would remove insecure hash functions from HTTPS configurations or ssh config files. Also, you could namespace commit hashes with sha1 beging the default:
git checkout sha256:7f83b1657ff1fc53b92dc18148a1d...
git checkout sha512:861844d6704e8573fec34d967e20bcfef3...
Enabling/disabling active hash functions would probably an expensive operation, but you wouldn't be doing it every day so it probably wouldn't be a huge problem.
Just some context - git calculates an object's name by his content in the following way. Say we have a blob that represent a file who's content is 'Here be dragons', then the file name would be:
printf "blob 17\0Here be dragons\!\n" | openssl sha1
# => a54eff8e0fa05c40cca0ab3851be5aa8058f20ea
So the object gets stored in '.git/objects/a5/4eff8e0fa05c40cca0ab3851be5aa8058f20ea'For example, if you want to stream the blob over http you can use it to set the content-length. Otherwise, you have to use chunked transfer.
So I could imagine in a large source file, it would be possible to have some malicious code plus some data in comment blocks to make the hash match. That said, the PDF's are 422k, and I think it's a much more difficult attack on more typical, smaller size source files that one would typically check out and build from git. Maybe Xcode .nibs and that sort of tool output could become relatively easy attack vectors, though.
https://git-scm.com/about/info-assurance
These claims are wrong as long as it uses SHA-1. Full stop.
It'd be really nice if git had cryptographic integrity. Not just because it'd prevent some attacks on git repos, but because it'd make git essentially a secure append only log. Which would be interesting, as it'd more or less automatically give some kind of software transparency for many projects.
Like for instance rebases?
If a low signal/noise ratio is still the purpose of information then the thread is less interesting than Linus mail:
- if we add size it will make forgery harder - yes SHA1 should be replaced
What linus is missing is people rewriting history. This will not be a concern for git, but certainly will for any crypto currency relying on SHA1 in a close future. (Hint this transaction belonged to me)
Can you name one that does?
To me it feels like this would just be a small hurdle? But I don't really know this stuff that well. Can someone with more knowledge share their thoughts?
I think Linus also argued that SHA-1 is not a security feature for git (https://youtu.be/4XpnKHJAok8?t=57m44s). Has that been changed?
If you don't have permissions to my repo that already limits the scope of attackers to people who already have repo access. At that points there's tons of abuse avenues open that are simpler.
If someone could fork a repo, submit a pull request and push a sha for an already existing commit and that would get merged and accepted (but not show up in the PR on github) well that would certainly be troubling, but at that point I'm well past my understanding of git internals as to how plausible that kind of attack would be...
1. Go to github, do a git clone 'repo' --mirror 2. cd to the bare repo.git and do a git show-ref and you will see all the pull requests in that repo. If any of those pull requests contained a duplicate hash, then in theory they would be colliding with your existing objects. But since git falls back to whatever was there first, I think it would be very challenging indeed to subvert a repo. You'd essentially have to find somebody who's forked a repo with the intention of submitting a PR, say, on a long running feature branch.
You could then submit your PR based off hashes in their repo before they do, which would probably mean your colliding object would get preference.
Its pretty far fetched, but the vector is non-zero.
For example, let's say it goes (letter##number represents the hashes):
State 1:
Real repo: R3 -> R2 -> R1 -> R0 (note: R3 is the head)
Benign PR: B1 -> B0 -> R1
Malicious PR: M1 -> M0 -> R1 (note: M0 = B0, but the contents are different)
State 2, after merging Benign PR:
Real repo: R4 -> R3 -> R2 -> R1 -> R0, R4 -> B1 -> B0 -> R1
If Malicious PR was merged now, Git would, I imagine, just believe that M1 is a commit that branched off of B0, since that's where the "pointer" is at.
So, yeah, what would this actually accomplish?
Git can implement checking for easily collided data and warn the user, potentially even look to implement the safer hash countermeasures too. The fact that this isn't a second preimage, or that SHA1 isn't used to auth a repo doesn't really factor in to it.
He does say that the sky isn't falling, and there are some steps they can take to mitigate it.
Edit: or do you mean all the posts here in the comments?
Another possibility, but this is a hack to keep key length to 40 chars, would be to change key encoding from hex encoding to base64. In 40 chars you could encode 240 bits instead of 160. It is preferable to get rid of the hard coded 40 char limit. It shouldn't be that hard.
Of course there is probably 20 hardcoded in many places too. So in this case the bit length is an issue too. You are right. Switching to base64 encoding would not solve the hardcoded bit length if any.
Using file hashes as file identifier doesn't look like a good idea as suggested here https://valerieaurora.org/hash.html because hash lifetime are short. The system should support changing the hash every year. The work required to compute hashes increases too.
> Git has opaque data in some places (we hide things in commit objects intentionally...
I think the bigger problem would be external -- tooling and other integrations. I'm guessing if they did move to another algorithm, as part of the migration git would need to re-compute the hash for every single object in all of our repos and migrate all our refs over to the new hashes, so that repos created before and after the change would be indistinguishable. This would mean that every commit hash which appears in plaintext in commit logs, emails, bug trackers, etc. would be wrong. Not to mention 3rd party tools which make the same assumptions about hashes that git itself does. It sounds like a nightmare to me, and one that I would only want to force on the community if absolutely necessary.
Somebody already submitted patch series to (optionally) use it in git in place of SHA-1:
Git sees them as different despite them having the same hash. You can test with:
mkdir shattered && cd shattered
git init
wget https://shattered.it/static/shattered-1.pdf
git add shattered-1.pdf
git commit -am "First shattered pdf"
git status
wget https://shattered.it/static/shattered-2.pdf
sha1sum *
md5sum *
mv shattered-2.pdf shattered-1.pdf
git status
So it doesn't see the files the same.Apologies for those on mobile (please fix this HN!): the commands are: mkdir shattered && cd shattered && git init && wget https://shattered.it/static/shattered-1.pdf && git add shattered-1.pdf && git commit -am "First shattered pdf" && git status && wget https://shattered.it/static/shattered-2.pdf && sha1sum * && md5sum * && mv shattered-2.pdf shattered-1.pdf && git status
EDIT: Ah, of course! git adds a header and takes the sha1sum of the header+content, which breaks the identical SHA1 trick. You can add a footer on and they keep the same SHA1 though. Don't have time to play about with this more just now, but try it with `cat`ing some identical headers and footers onto the pdfs.
EDIT2: Actually, this is discussed more extensively in the other thread which I hadn't read yet. Go there for more details: https://news.ycombinator.com/item?id=13713480