A New Life for Certificate Revocation Lists
letsencrypt.org
letsencrypt.org
Did something happen there or some other significant issue get discovered? I'm curious why the move back to CRLs (albeit improved) vs must-staple. It seemed like a reasonable elegant and straight forward solution that fit the web pretty well.
There isn’t any particular force driving for OCSP stapling: Browsers can’t turn it on until it is ubiquitous. CAs don’t want to enforce must-staple because their users will need enormous amount of help rolling it out. Site operators don’t care if browsers are fetching OCSP (especially since Chrome and Edge don’t)
Maybe they saw with some telemetry that very few website actually enable OCSP stampling and decided to implement a fix that cover all certs and can really be deployed
The stapling telemetry is no longer turned on in Release [2], and even if it were, you have to do special things to look at Release data, but some years back (~2018 maybe?) I remember Release stapling was substantially lower than the more tech-savvy Beta and Nightly populations. Which is pretty normal, as tech-oriented sites are more likely to turn on advanced features.
[1] https://telemetry.mozilla.org/new-pipeline/dist.html#!cumula...
[2] "prerelease" https://probes.telemetry.mozilla.org/?search=stapl&view=deta...
IMO, OCSP stapling is the best overall solution until certificate lifetimes are shorter (< 7 days).
As a server developer I'm worried that the focus on independent CRLs will make it difficult to automate certificates in the face of revocation. Currently, Caddy staples OCSP for all certificates by default, caches the staples, and refreshes them halfway through their lifetime. Works great. Every server should do this. And when an OCSP response is discovered to be Revoked, Caddy automatically replaces the certificate. Works great. Every server should do this.
If every browser is independently going to decide which certificates to distrust, now I am not sure of a good, authoritative way to determine "revoked" and then replace certificates automatically. I'm worried this will hurt the TLS ecosystem unless we answer those questions first.
Actually, let's just shorten certificate lifetimes and be done with it already.
Main blockers to short cert lifetimes:
- CA's uptime determines Web's uptime.
Main solution:
- Multiple redundant ACME CAs. If one goes down, try another. (This is what Caddy already does.)
Mainly because we are not talking about validation performance, and OCSP stapling is an excellent performance fix.
Not to mention all the privacy issues that OCSP stapling really fixes.
However, it is optional, so we will repeat the same mistake we made with optional OCSP stapling.
Only nice clients will implement ARI, but they are the clients that need it the least because if they go to the trouble to support ARI they probably already have friendly netizen programming.
As for revocations, ARI doesn't make much sense to me. If we know a certificate will be revoked soon, we might as well stop trusting it right now. Why continue to trust a certificate that we know is being revoked?
Maybe I'm totally missing the point of ARI.
The worst revocations are mass ones that the site operator didn't request. ARI gives a heads-up to renew early, maybe because all certificates issued using HTTP-01 are getting revoked in two days.
Revocations aren't always about distrusting a specific certificate. There's likely nothing wrong with the certificate, but it needs replacement for ecosystem cleanliness. Regardless, 40M certs are being revoked in a 10 minute window on Saturday (oof, because that's the covenant with the BRs, not because Saturday is somehow not awful!). Reissuing 40M certs may take a dozen hours (1000/sec), even if all clients work optimally and begin immediately after OCSP changes status. During that dozen hours, those certs are all already revoked.
It'd be nice if clients could be told in advance: replace this certificate right away, regardless of its validity period. Get it done early before the crowd forms and replacement requires queuing up.
(Obviously using multiple CAs mitigates the downsides to waiting until revocation)
So yes, if we're being strict like the policy is, an early renewal signal is a red flag that a certificate can't be trusted. There might not be anything wrong with it, but we can no longer be sure.
I mean, there's also nothing wrong with a certificate 2 seconds after it expires. Probably. But we can't be sure. And because of that, we immediately distrust the certificate when it expires. (There might be something wrong with it before it expires too. But that's less likely because less time has passed, so we allow it, I guess.)
I think the vision is nice. I really do. I just think the clients that need it most won't support it.
OCSP stapling is temporary, anyway: just until we get certificate lifetimes short enough to make revocation irrelevant.*
* This is how it should be, but the industry seems to be going the direction of keeping cert lifetimes longer.
As for the alternative not a single browser out there validates OCSP out of the box on the hot path (e.g actually waits for OCSP before proceeding)
However, for strong PKI on a consumer or enterprise encryption/decryption device, e.g. smartcards, there is big costs associated with shortening lifetimes. Not going to fly there.
It's much saner to issue a new, shorter-lived cert. Reduce certificate lifetimes to match whatever you'd use OCSP stapling for. Continue to use ACME. Figure out rotation lifetimes such that I can sleep at night (i.e. they need to be at least 2 days, so that if the rotation fails I've still got time to wake up and start working). Work on building client and server tooling to make it easy to accept rotated certificates (i.e. reload your client certificates on SIGHUP).
The advantage of CRLs isn't that they can be used offline, it's that it allows your security team to burn certificates rather than waiting for the CA to do so. You should always subscribe to your CA's CRL too, but you should have a CRL for your own internal use, too.
This is incorrect conventionally
CRLs are signed by the Originating CA in the conventional trust model.
Most CAs won't give you a certificate with the ability to generate your own CRL sharing the same trust, because it can be weaponized to cause a denial of service.
This then brings up non-conventional trust models, and Validation Authority as Co-equal to CA situations for validity which can be significantly more challenging, and requiring VA certificate insertion into every Relying Party, just to name a few things, not to mention all the security concerns.
https://www.imperialviolet.org/2014/04/19/revchecking.html
Revocation checking is useful for your security team to blacklist site. That's the only useful use.
I don't know much about this stuff, so apologies if this is a silly question:
If you needed to revoke all the certificates, couldn't you just revoke the handful of intermediary certificates and call it a day? I assume you'd want to revoke them anyways if there's a situation severe enough that merits revoking 200 million certificates.
There is no "handful" of intermediate certificates -- there are precisely 4 (for Let's Encrypt [0]) and they are essentially on-line root certificates. And if those certificates aren't even compromised, revoking them would only harm the ecosystem.
They're essentially on-line root certificates because they serve most of the function that such roots would serve if they were allowed. There aren't a bunch more available to replace them, so this means recovery now requires a key ceremony, figure on a week to a month to arrange that.
Whereas if you're able to "just" revoke 10 million end entity certificates you can recover immediately.
We have a set of backup intermediates that can be activated if we had to revoke the active ones for any reason, so the disruption wouldn’t be too high hopefully.
(I work at Let’s Encrypt, but this is my own opinion and not that of my employer)
Couldn't they just use actual units? This says absolutely nothing.
When I was reading the thread the other day about a trees worth of oxygen from the MOXIE experiment, I couldn't help thinking: why not just use a term everyone is familiar with, litres per minute air.
Enough oxygen to sustain an adult at rest for x minutes.
I'm beginning to suspect there's an in-joke with science / tech writers about strained analogies.
And I'm not in.
Which has been going on for longer than I've been alive. The idea (I assume) is to take a large number that's hard to conceive and turn it into something everyone can relate to.
But inevitably, they choose things that few can actually relate to, or things that are so vague/variable to be meaningless. It just adds more confusion all around.
It has to be an intentional joke.
CRLs and OCSP are about individual certificates.
The beauty of Merkle Hash Trees for validation is it permits Statements e.g.
CAx = CA2 and 156 ≤ X < 343 (5)
The statement cj indicates that the certificate with serial number X = 156 issued by CA2 has been revoked, while the certificates with serial numbers from X = 157 to X = 343 (both included) issued by CA2 have not been revoked.
From https://www.researchgate.net/publication/220066804_Certifica...
A CRL is not equivalent to modern OCSP, because modern OCSP can operate as a whitelist or a blacklist. However, CRLs are generally more secure than some of the OCSP variants because strong OCSP is used less (Nonced+authenticated/registered), and CRLs are not subject to replay attack like vanilla OCSP can be.
In the space domain, satellite ground systems have used a combination of guards, one way transmission, and others, although I am in favor of system-specific certificate whitelisting via running OCSP with local VAs, backed up by smart clients, with CRLs on the filesystem and in network shares.
CRL processing can cause timeouts, which is why it is always less preferred than the much more lightweight OCSP by comparison.
Furthermore, CRLs are blacklists, and newer OCSP is not exclusively, and, OCSP permits more flexibility with the various trust models that exist for it.
https://dev.to/coroner/why-do-certificate-revocation-checkin...
Edit: I was partly wrong, this is a good thing because you CAN download the CRLs now (see comments below here for info), whereas previously you couldn't. Your browser still probably won't support a full CRL download, but I could be pleasantly surprised.
Previously, you couldn't do this, because not all CAs published CRLs.
Beginning October 1, you will be able to just download the CRLs, because Apple and Mozilla are requiring it.
It's therefore unclear what your beef is,
Correction: Apple and Mozilla will be able to just download the CRLs. Not me. The link in the post SPECIFICALLY says us common plebes don't get that right.
If you think it's because the URLs will be disclosed in the CCADB, note that the contents of the CCADB are published here: https://www.ccadb.org/resources
Specifically, the CRL URLs can be found in this CSV file: http://ccadb-public.secure.force.com/ccadb/AllCertificateRec...
“Our new CRL URLs will be disclosed only in CCADB, so that the Apple and Mozilla root programs can consume them without exposing them to potentially large download traffic from the rest of the internet at large.”
In the same way that you can technically query the DNS root servers yourself but you don't tend to do that because your computer will query a more downstream DNS server.
Does anyone know what the rate of revocations is?
I can easily imagine a situation where it is high enough to cause a browser update every few hours. That is for every installed browser, because - as they say in the article - Browser-Summarized CRLs are "proprietary, browser-specific CRLs". Moreover there are non-browser clients which we must consider if we are to take this proposal seriously.
------
My quick back of the envelope estimation:
CRL size: 4GiB (they say in the article that it could be easily the size of a movie)
Average Cert Size: 75 bytes (first hit in Google, no idea if reliable number)
Time Span: 825 days (CRLs have only unexpired certs and ones older than 825 should all be expired)
4GiB/(75B/cert)/825days = 69141 certs/day
Sounds way too high to me. Where am I wrong?
[1] https://obj.umiacs.umd.edu/papers_for_stories/crlite_oakland...
It won't be fixed until we have name constraints on CA certificates, and a way to decorate trust anchors with local policy name constraints.
> But because OCSP infrastructure has to be running constantly and can suffer downtime just like any other web service, most browsers treat getting no response at all as equivalent to getting a “not revoked” response. This means that attackers can prevent you from discovering that a certificate has been revoked simply by blocking all of your requests for OCSP information.
This is false. Non-nonce OCSP is inherently cachable, and replayable. That means you can have your own HA setups with HA OCSP clients talking to HA OCSP servers (repeaters & responders) backed up by caching in commercial CDNs, and local caching servers like bluecoats.
Likewise, OCSP stapling helps remove much of the performance and privacy issues, pushing it to the serving webserver.
Beyond this, you can just use squid or localized HA OCSP services, and do some DNS rewriting to support it even more HA.
Nonced OCSP is the rare beast that needs to be online, but there are HA OCSP with smart OCSP clients.
> To help reduce load on a CA’s OCSP services, OCSP responses are valid and can be cached for about a week. But this means that clients don’t retrieve updates very frequently, and often continue to trust certificates for a week after they’re revoked.
Trust Stores are inherently manageable. The lag around revocation completely depends on CRL/OCSP publishing, and client update requests.
> And perhaps worst of all: because your browser makes an OCSP request for every website you visit, a malicious (or legally compelled) CA could track your browsing behavior by keeping track of what sites you request OCSP for.
This is why we advocate OCSP Stapling and use of CDNs for OCSP & CRL cache hits. Furthermore, localized OCSP mentioned above decentralizes this even further.
> So both of the existing solutions don’t really work: CRLs are so inefficient that most browsers don’t check them, and OCSP is so unreliable that most browsers don’t check it. We need something better.
CRLs & OCSP work pretty well when actually supported.
When Diginotar happened I polled every single publicly available commercial CA - strangely, a ton of them were not producing any CRL/OCSP at all, putting clients into a fail-open mode.
Lesson of the story: don't blame a protocol for lazy CAs, bad implementations, or the lack of operational excellence from many vendors.
For Firefox end users, a certificate only gets tested against the filter cascade if it is known to have been included in its creation (by examining the embedded SCT timestamps). If it's not definite that the certificate was used to generate the filter, then Firefox reverts to OCSP.
(I'm one of the authors of CRLite in Firefox: https://insufficient.coffee/2020/12/01/crlite-part-4-infrast... )
# Why is CRLite able to compress so much data?
Bloom filters are probabilistic data structures with an error rate due to data collisions. However, if you know the whole range of data that might be tested against the filter, you can compute all the false positives and build another layer to resolve those. Then you keep going until there are no more false positives. In practice, this happens in 25 to 30 layers, which results in substantial compression.
EDIT: Is there any risk of filter blow up (think 1000's of layers) if a CA did a mass revocation (maybe some root key leak)?[0] https://github.com/mozilla/crlite/wiki#why-is-crlite-able-to...
What does happen in CRLite is that you can't keep shipping the tiny "stash" updates to clients, you have to mint a whole new .mlbf filter file, which is about a megabyte. [Edit:] Then you can resume the "stash" updates from there, but the ecosystem 'shock' requires a regeneration of the filter.
(There was supposed to be a blogpost on the Mozilla blog from the research teams; I don't know if it was ever written.)
Source: a WWDC 2017 talk which unfortunately I can't find online anymore
Mozilla's choice here avoids that problem coming up which means nobody needs to push back when it gets "solved" in this regressive way.
The bad case would be if Let's Encrypt discovers a problem (like a security flaw or implementation error in a validation method, as happened with the TLS-ALPN-01 method before) and concludes that it has to mass-revoke a very large number of affected certificates.
So, no, it's not common. But it's necessary.
[1] https://dl.dod.cyber.mil/wp-content/uploads/pki-pke/pdf/uncl... - Pg 7, under Local Cache
CDNs + Localized OCSP + Tactical OCSP + Smart OCSP Clients + Network caching + OCSP & CRLs on the filesystem, just to name a few (not including delta CRLs and other solutions).
The DoD OCSP Responders are configured to share hash sets with downstream OCSP Responders & Repeaters, which makes promulgation particularly easy.
Looking at a random CRL [2] it's 41 bytes per revoked certificate.
8 million records at 41 bytes per record would be 300+ Megabytes. And a cautious CA might keep revoked certificates in their CRL for more than a year.
So if an event like heartbleed happened again and uncommonly large numbers of certificates needed to be revoked, the gigabyte range is within the bounds of possibility.
[1] https://news.netcraft.com/archives/2014/04/11/heartbleed-cer... [2] http://crl3.digicert.com/Omniroot2025.crl [3] https://www.grc.com/revocation/crlsets.htm
A CA does need to keep the revoked certificate in the CRL until at least the natural expiration date on the certificate, so for the CAs that give certificates out with 5 years out or more expiration dates, they may need to keep CRLs for much longer than just a year just naturally by nature of their expiration dates.
The previous status was the certificates could last up to 825 days, however that policy changed at the end of August 2020, so there are no extant certificates under those rules which expire after this year. And before that the policy was 39 months, but the last such certificate expired in 2021. Before that the policy was 5 years, but that policy changed in 2015 and so such certificates are long expired.