How I tripped over the Debian weak keys vulnerability
hezmatt.org
hezmatt.org
From my IRC logs at the time:
17:23 < luciano> has really an accident. I was needing many primes numbers... 0:-)
17:23 < Sesse> and you got the same numbers every time?
17:25 < luciano> Sesse, not every time :PThanks for sharing.
This is where statistics of "many eyes" and "sunlight as disinfectant" hits home. However bizarrely improbable that anyone would just wander along and find a a bug, people do, because they can. With proprietary/closed code, the probability is zero.
People discover bugs all the time in closed source software...
Ah I see what you mean.
Problem of language then.
"Finding bugs", to me as a coder, means something different from observing the effects of bugs.
Sure, I can see incorrect behaviours in lots of proprietary software. And I can guess what might cause it. But without the source code that's not the same as "finding bugs".
Reporting bugs vs finding them is perhaps what you're looking for, but I'd argue finding is overloaded. It's perfectly valid to kick the tires on some software and find some bugs.
I've found thousands of bugs in closed source software before, and subsequently reported those bugs.
This is a good point of course. Implementation, protocol and runtime bugs are indeed invisible from the source POV. As are lower level Ken Thompson "trusting trust" [0,1] bugs in your tool chain.
Most of all though, many "bugs" are just malicious functions the coders meant to put in there.
[0] https://cybershow.uk/episodes.php?id=2
[1] www.cs.cmu.edu/~rdriley/487/papers/Thompson_1984_ReflectionsonTrustingTrust.pdf
They however didn't find the cause of the bug.
And in many software, open source or proprietary, there are a lot of bugs sitting out in issues tracker for days, weeks, months or even years.
I can think of at least two or three occasions where I discovered an exploitable vuln, didn't have a contact to report it to, decided I didn't want to have the back-and-forth of some non tech company threatening me with a lawsuit, and took no further action (typically for lower severity vulns in less important services).
With open source software you can hop on GitHub and dig deeper into a vuln or other bug and actually report it somewhere.
On top of that, the vast majority of people won't know what to do if they find a bug. In the distant past I was one of those people, and only realized some of the bugs I saw in hindsight. It's been nearly 30 years so the details are mostly gone now, but I remember messing around in Windows and I could make the Microsoft Netmeeting application crash with a buffer overrun error.
Of course I was really new to computers then and understanding that buffer overflows in networked applications were really bad things (and seemingly beyond a lot of people that had been in the industry). Even then attempting to report security issues back in those days would have been far more difficult (hell and even risky in many cases).
So really it is many things that are required. Running into the issue. Having a deep enough understanding of computing to realize the issue is bad. Having a means of reporting the bug to a place where people will look at it. And a security culture that knows when and how to act on the bug report.
1) Observing
2) Understanding as bad
3) Reporting (or fixing)
4) Closing the loop on remedy
Both proprietary and open code runs into problems at (3) because people don't want to hear it, for commercial or ego reasons.
With FOSS at least you get the direct intervention route of simply fixing and publishing the patch, which done responsibly may or not force a maintainer's hand. With proprietary you can piss into the wind and be ghosted, or sued, hence more irresponsible/anonymous disclosure.
Anything that maximises the likelihood of getting from (1) to (4) safely must be a good thing, so I think FOSS yields the better security model.
What I was wondering when I read the same sentence you quoted: how many really serious security bugs like Heartbleed, CVE-2008-0166, or the zx drama are happening without people finding out about it and publishing their findings?
The problem with closed source is there is a third possibility: ignore the problem and save on the cost of fixing it. The responsible disclosure regime we have now is because companies almost always chose this option, ie denied it was a problem and refused to invest to fix it. When the discoverer then released the bug anyway they we so enamoured this this approach tried solving the disclosure problem by suing the researcher.
If you think companies still don't ignore security issues when they are given a choice, you are kidding yourself. The problem compounds because when you do find a bug open source makes it easy to see if you can use it to create a security issue. In proprietary code that's much harder, so I'm 100% certain a fair number of potential security issues don't get patched because it isn't obvious how to exploit them. Nonetheless they are chinks in the armour, so they give the bad guys a excellent set of starting places to start looking.
1) Someone noticed something that was odd. They were able to go look at what was happening, along with source code and see that suspicious things had happened.
2) They were able to reach out to security experts across almost all the main distributions and get additional eyes on it who also confirmed there was a security thing and were immediately able to take actions to deal with the insanity.
3) Disclosure went public, and lots of eyes with expertise across all kinds of aspects of software and security have been able to tease apart what was done, and how, figuring out risks etc.
4) Other suspicious commits across other bits of software, from the same developer, have been tracked down and identified and the work continues in figuring out consequences.
5) Every distribution has become alert to the minutia of the ways that things were compromised around the build archives etc and have been able to start figuring out how to catch other cases and stop this in future. Stuff will start to spread to various open source projects as distributions work on their packaging tooling/processes, helping ensure other projects don't become vulnerable to compromise through the same mechanisms.
If you contrast this with closed source, where unless there is an exploit, it's reports of e.g. "Your software is running slightly slow" like in the xz/openssh case, is unlikely to get much attention, if any. Then once the closed source company finally finds out about it, at best you'll get a very carefully phrased explanation of what happened, revealing just the bare minimum amount of data they think they can possibly get away with. That badly harms the whole industry's ability to avoid repeats.
The result was a terrible vulnerability, but it seems more of a case of spectacularly bad luck of everyone not spotting the issue.
Obligatory Dilbert https://imgur.com/uR4WuQ0 and XKCD: https://xkcd.com/221/
Plus, the upstream OpenSSL code was invoking undefined behaviour. Hence the compiler could have validly made the exact same transformation as the Debian maintainer. At the time this felt academic: surely compilers can't be that mean! Since then I think undefined behaviour is better understood as a thing to avoid entirely.
Then, eight years later, Heartbleed was discovered. And we suddenly all realised how badly maintained OpenSSL was. In their defence, it was pretty much a volunteer job. Thankfully, subsequent funding has improved the situation.
And what the fallout would be.
But now that pension funds get into Bitcoin, one has to wonder what would happen if a giant like Coinbase turns out to be holding their funds in cold storage but with guessable keys.
As Android wallets were the easiest to use wallets and exchanges routinely robbed people, it mattered a lot even back then.
While Coinbase has a quite good protocol for securing keys, what's scary to me is that so many ETFs are trusting it to be the only key holder instead of using multi-signature wallets between multiple entities (like Fidelity and Unchained capital as the other 2).
I know they state that they do it via a custodian and name it like "Coinbase", but is that really all? All they got is an IOU from Coinbase? Maybe they have more specific setups?
Over the last 22 months, Unciphered has been working on a vulnerability which affected BitcoinJS, a popular package for the browser based generation of cryptocurrency wallets, as well as products and projects built from this software. Over a period of years, this vulnerability caused the generation of a significant number of vulnerable cryptocurrency wallets.
Bitcoin isn’t that safe to hold value in, if you might be targeted by someone with brains and money.
Base Numbers:
- 2*128 guesses on average
- public state of the art for ECDSA on an FPGA is 1315tps [0]
- retail price of said fpga $10,000 [1]
- total net income for xilinx from advanced FPGAs FY2022 (936M * 0.74 + 879M * 0.72) = 1325M [2]
Ballpark numbers, we'll assume that attacker can buy 10x every FPGA xilinx made in 2022 for a 50% discount and can run them non stop for zero cost.
We'll also assume they have a bunch of secret math geniuses and have a faster ECDSA implementation that can do 1,000,000,000 tps ( or 1,000,000x SOTA)
- (1325,000,000 / 5,000) * 10 = 2,650,000 FPGAs
- 2,650,000 * 100,000,000,000 = 2,650,000,000,000,000tps
- 2*128 / 2,650,000,000,000,000 = 1.28e+23s
- 1.28e+24 / 60 / 60 / 24 / 365 = 4,080,000,000,000,000 years
- ~4 quadrillion years
- or 4x the time until every planet has been ejected from every star system and the sun has cooled to 5K [3]
But why stop there, lets assume that the attacker can use the entire planets GDP to buy chips and has a 1,000,000,000,000,000x faster ECDSA implementation.
- World GDP (2022): 101.3 trillion
- (101,300,000,000,000 / 5000) = 20,260,000,000
- 100,000,000,000,000,000 * 20,260,000,000 = 2.66419e+28tps
- 2*128 / 2.66419e+28 = 12,772,451,173s
- 12,772,451,173 / 60 / 60 / 24 / 365 = 405 years
So even then, we wouldn't see a single BTC key broken within our lifetime.
(Unless you believe that 3 letter agencies have successfully built a quantum computer that practically implement shore's algorithm, in which case you should probably be more worried about the fact that they can break public key encryption globally)
[0]: https://arxiv.org/pdf/2112.02229.pdf
[1]: https://www.colfaxdirect.com/store/pc/viewPrd.asp?idproduct=...
[2]: https://web.archive.org/web/20211203065624/https://investor....
[3]: https://en.wikipedia.org/wiki/Timeline_of_the_far_future
Then, since you know every possible public key ... it's just a matter of looking up its private key in the index.
HOWEVER, I was off by a few magnitudes on how big the seed numbers get. If you are curious, this is basically the biggest seed number:
904,625,697,166,532,776,746,648,320,380,374,280,100,293,470,930,272,690,489,102,837,043,110,636,675 ...
or 904 trevigintillion (thanks wolfram alpha!)
So, yeah, it's not happening anytime soon, though it doesn't invalidate the premise, just changes the timeline to unachievable with current technology. Technology could improve to the point where it is feasible in a human lifetime.
"Computers made of something other than matter and occupying something other than space"
Edit to add: just found https://keys.lol that does exactly this, but generates a few dozen at a time and doesn't store the results.
Computationally the two are the same: cracking a key generated by a secure algorithm is equivalent to generating all possible keys and checking to see which one matches.
The problem is the the index to the first occurrence of any particular N byte sequence will on average take around N bytes to store so your "compressed" file is about the same size as your input file.
[1] This is not known to be true. A number that has the property that every possible finite sequence of digits in base b occurs in the number's base b expansion is called a "normal number in base b". A number that is normal in base b for all integer bases b ≥ 2 is simply called a "normal number". Almost all real numbers are normal, but it is not known if π is among them.
Well if we're being picky then I'm going to point out that almost all real numbers cannot be written down.
Of the numbers we can write down in a reasonable way, the computable numbers, we've only proven a relative handful of them to be normal in any base, and barely any normal in all bases.
So if you have an arbitrary irrational number where you can ask if it's normal, then the answer is probably "we don't know, we haven't figured out a proof", rather than "almost certainly". Those overwhelming odds about "almost all real numbers" don't apply to your number, because your number is computable.
The cracking algorithm talked about earlier is that you create a key from seed 0, 1, 2, etc. until you find a match, and then you can keep going to look for more matches if you so desire. Or you can go in a different order, because the outcome is the same.
The only change in your version is that you don't look for matches until you've finished going through every seed.
That makes your version strictly slower.
It's not invalid, it's just worse, for the use case of cracking existing keys.
And for the use case of cracking keys made after you do most of your computing, you still wouldn't use the "one giant index" method for precomputing things. It's not efficient.
This sentence is funny to me. Maybe because I'm not a native speaker, but I can't avoid reading it as a big problem with the GH team itself, heh. I was expecting the next sentences to "get into" that.
> I wonder if Luciano hadn’t found it, how long it might have been before it was found
I think perhaps only GH or one of the big cloud providers would have stumbled upon it, few places I can think of that would have thousands and thousands of keys stored from users.
It does / can read both ways, but a comma after "problem" would have removed that doubt.
> Ezra Zygmuntowitz pointed GitHub in my direction (and let me take the time to really get into the problem) with the GitHub team
> Ezra Zygmuntowitz pointed GitHub in my direction (...) with the GitHub team
Somebody wrote a patch which fixes this, and then - even without assistance from an LLM just using normal human incompetence, somebody said "There's another similar copy nearby, we should get rid of that one too" and Debian landed the patch to apply both changes.
The result is now OpenSSL doesn't copy any bytes. It doesn't copy the uninitialized data, which is good, and it also doesn't copy the truly random entropy into the pool. Oops.
To be fair to the Debian patch, reading from uninitialized junk doesn't sprinkle randomness into the desired buffer: it invokes undefined behaviour. The Debian patch was lawful!
But I think you're slightly wrong here, I think Cox is describing three cases where OpenSSL takes a bunch of bytes and says here you go entropy pool, here's some potential entropy to add, if it wasn't really entropy that's fine, stir it in anyway.
A. Buffers you've provided and are asking to pour random data into. OpenSSL thinks this is clever - best case the data is unpredictable and it helps, worst case it's just zeroes and the entropy pool will handle that anyway. Valgrind and similar tools will be very angry when you do this to actual uninitialized buffers, which in C is really easy to write and so likely to be common. This is the case that already had a comment about such problems and is the second of two places Debian removed.
B. Buffers OpenSSL tried to put entropy from the Linux random device into. OpenSSL doesn't bother checking that the randomness went into the buffer, after all if it didn't the buffer is uninitialized... This is the first of the two places Debian removed (in terms of line number). This is where "Debian Weak Keys" come from because now the "random" numbers are not random. In almost any real world scenario this code was unproblematic although very stupid, until it was patched.
C. A Buffer read from a file the user provided with entropy in it. This only ends up with uninitialized data if the file shrinks while OpenSSL is trying to read it, which can happen but would probably be some sort of attack. Again though OpenSSL doesn't care, Debian did not patch this code AFAIK but very few people care about it.
Edited to add: One thing I really like in the discussion below the actual post you linked is the idea that the OpenSSL people are thinking the Debian dev wants to debug their software and are saying it's fine to comment out the randomness in order to help them do debugging not realising the developer actually intends to remove the randomness because they believe it's a bug. When you re-read the mailing list posts with this understanding and forgetting the Debian Weak Keys bugs, it snaps into place - we can't know if that's really what happened of course, but it makes sense.
Why MySQL rather than sqlite?
They were trying to speed up access to ~/.ssh/authorized_keys. That's the kind of situation mysql is designed to shine in. I would have expected that patching OpenSSH to check for ~/.ssh/authorized_keys.db would be less work than patching it to use MySQL.
Just buy a copy of GitHub Enterprise, deobfuscate the files (it's a fun challenge to deobfuscate them, and not all that hard), and look around to see.
It's unfortunately not open source, so we can't share and talk about the code, link to it on GitHub, or anything like that, but if GitHub Enterprise is still using such a patch, it's like the prod github is too.
No? Usernames are optional in SSH URLs.
Which isn't always in sync with their GitHub username. Which might even be "root".