Revoking certain certificates on March 4
community.letsencrypt.org
community.letsencrypt.org
It looks like a nasty and subtle pass-by-reference of a for-range local variable, although I'm having trouble figuring out where the reference is stored: https://github.com/letsencrypt/boulder/blob/542cb6d2e06e756a...
I've spent plenty of time hunting down similar bizarre bugs in Go code as well, where the called function ~implicitly~ takes a pointer to the iteration variable and stores it somewhere. Each iteration of the for loop updates the stack-local in-place, and later reads of the stored reference will not read the original value. It's hard to spot from the actual call site :/
EDIT: This was an explicitly taken `&v` reference, but the same thing can also happen implicitly, if you call a `func (x *T) ...` method on the variable.
Yes you could still do it in rust, but any reviwer of the code would say "why in the world are you doing it this way" because it would be forced into a complex cross call monstrosity.
1. The reference outlives the original value.
2. You can't have multiple mutable references at the same time.
If I understood the issue correctly, only one of the references would be a mutable one, so the way Rust could have caught the bug would instead be the related rule: "you can't have an immutable reference and a mutable reference at the same time".
I'll cut the LE team some slack on this one :) the PR does have tests
// Make a copy of k because it will be reassigned with each loop.
But v is reassigned with each loop too.The real question is why there's so much pass-by-reference in the first place. K looks to be a domain name string -- it's almost certainly faster to copy it than to dereference it everywhere.
I don't program in rust, so my knowledge here is limited to what these words mean in C/C++ - however shouldn't making a copy still require dereferencing the copy?
Suppose you have a struct like this:
struct foo {
struct bar elem
} s;
If you know the address of `s`, you just calculate the address of 'elem' from it and read the contents; a single memory read, all the data together cache-wise. Suppose on the other hand you have a struct like this: struct foo {
struct bar *elemptr;
} s;
If you know the address of `s`, you have to first read `elemptr`, and only then read the value of `elem`. That's an extra memory fetch, and probably from a different part of the memory than `elem` is from. Copying on modern processors is very fast, and the resulting copy will be "hot" in your cache. So conventional wisdom I've heard is that unless `struct bar` is quite large (I've heard people say hundreds of bytes), it's probably faster to just copy the whole structure around than to copy the pointer to it around and dereference it.Caveat: I haven't run the numbers myself, but I've heard it from several independent sources; including, for instance, Apple's book on Swift.
It's fine as long as your entire program fits neatly in cache, but once you exceed the cache size performance goes to hell because you force loads of misses of slightly-older data by constantly copying your working data.
Wether you choose to pass by reference or copy, that should just indicate your desired semantics. The underlying compiler can decide if it's really worth making a copy or not.
Sadly most compilers won't elide out most unnecessary copy operations, even when performance would be improved to do so.
Depends on the situation; the parent was somewhat vague.
As one example, suppose you've previously created a `struct foo`, but you haven't accessed it in a long time; then fetching any of its fields will probably be a cache miss. But if it contains a pointer, fetching data from that pointer will probably be a second cache miss – which can't even be started until the first one completes, because the CPU doesn't know what the pointer is. If `struct foo` contains the data directly, there's a good chance it all fits in a single cache line (typically 64 bytes). But even if it's split across, say, two cache lines, the CPU knows which cache lines need to be accessed, so it can send off the second request without waiting for the first one to complete.
That said:
- Strings are variable length, so they usually can't be stored in place anyway, although some languages/frameworks have "small string optimization", meaning that sufficiently short strings are stored in the space that would otherwise be used for the pointer and length.
- In general, Swift is hardly a good role model when it comes to performance. :p
The original bug report, which was initially diagnosed as only affecting the error messages, not the actual CAA re-checking: https://community.letsencrypt.org/t/rechecking-caa-fails-wit...
Brief discussion on revocation exemption requests: https://bugzilla.mozilla.org/show_bug.cgi?id=1619179
Tomorrow will tell if granting a revocation exemption might have been a good idea in hindsight.
My understanding is that the 90 day lifetime is largely because revocation can be thwarted. Thus the practical difference between 24 hours and one week is meaningful for server admins, but inconsequential if someone is staging an attack.
When something is wrong you can revoke them immediately.
Why leave a potential vulnerability open for more than 24 hours?
I’m just starting to think LE is more aimed at large organizations than people running smaller configurations. Which is fine, thankfully we still have traditional CAs. I just hope we don’t devolve into a monoculture of ACME-only SSL.
We could do with another LE-style service (or two) operated independently (both organisationally and geopolitically).
One of the less-obvious reasons: for "serious" usage where you're also stapling OCSP responses, there's a dependency on the cert vendor's OCSP service. You can cache the OCSP outputs to get through short windows of unavailability, but if the vendor's OCSP goes offline for days or suffers some serious incident, it pays to have multiple vendors on-hand. There was such an incident with GlobalSign back in October 2016 (who's otherwise a pretty decent vendor!), so it is a legitimate concern.
For "serious" use-cases, you basically need redundant live certs from redundant vendors, and not having a second LE-like option means one of those is still a legacy CA for now...
Nevertheless technical people bemoan average users clustering towards centralised web-hosts but forget the reality that hosting a website from your own desktop or a VPS is far from trivial even in 2020!
Actually, revocation is broken. Which is a large part of why LE uses 90 days.
For example, Sectigo has misissued nearly every certificate since 2002, including ~11 million unexpired ones (as December) and decided to just ignore their duty to revoke misissued certifcates [1].
Should the rules be changed? Maybe. However, when you're giving an immense responsibility to CAs then public trust is paramount. Ignoring agreed upon rules whenever you find it convenient does not inspire much confidence.
1. "Yes, it's still good"
2. "No, it's revoked"
3. "There was a network problem so I'm not sure"
Of course bad guys who know you'd get answer 2 can most likely ensure you have answer 3 instead. So the only safe thing to do is treat 2 and 3 the same. If we're not sure this certificate is fine then it's not fine. But in practice answer 3 is common anyway. For some users it may happen essentially all the time. So browser vendors don't like to treat 2 and 3 the same, even though that's the only safe option and that can thwart the effectiveness of revocation.
There's definitely further opportunity for improved tooling here. Perhaps this incident will drive it (Let's Encrypt's sheer volume can help in this way).
https://tools.ietf.org/html/rfc2560
OCSP stapling is a mechanism by which a TLS server can make OCSP requests ahead of time and serve the response in-band. TLS clients get a certificate signed by the CA as usual, as well as a recent OCSP response signed by the CA attesting to its continued validity. OCSP stapling allows TLS clients like browsers to know a certificate's revocation status without having to make an extra request, but it changes nothing for an attacker who stole a certificate since they can simply not use it.
https://tools.ietf.org/html/rfc6066#section-8
OCSP Must Staple is an option that can be included on a certificate stating "I promise to use OCSP stapling". An attacker who stole a "must staple" certificate can either include an OCSP response indicating the certificate is revoked, or they can omit an OCSP response which the TLS client will treat as a hard error.
https://tools.ietf.org/html/rfc7633
In short, RFC 7633 makes certificate revocation work. Web browsers and web servers support this today. If you use Let's Encrypt's `certbot`, pass it `--must-staple`.
You need at least effective monitoring and a good OCSP stapling implementation (IIS is supposedly pretty good at this) or else stapling is sadly going to make life worse for you not better.
1. Firefox remains the only mainstream browser to support OCSP Must Staple.
2. OCSP Must Staple does not cover all threat models: if an attacker gains the ability to temporarily issue certificates for the victim's domain (rather than obtaining the private key of an existing certificate), they can request a certificate without the OCSP Must Staple extension. A more effective method would be something like the Expect-Staple header[1] (in enforce mode).
3. It allows the ecosystem to move significantly faster. In a world where all certificates expire after 3 months, phasing out insecure hash algorithms (in certificates) would no longer take many years.
4. It encourages regular key rotation (even if it's not enforced)
[1]: https://scotthelme.co.uk/designing-a-new-security-header-exp...
A better example would be something like Certificate Transparency. Currently, browsers may require Certificate Transparency for certificates issued after a certain date. A malicious or compromised CA may work around this by backdating certificates. This would be less of an issue with shorter certificate lifetimes.
I'm speechless. I used to pay real money to get certs without half the service I get now for free.
Thanks letsencrypt.
You're lucky, because it took me forever to devise a renewal system that met my requirements of being able to renew certificates for domains that are used across multiple machines, both physical and virtual, many of the virtual machines which share IPv4 behind NAT thus cannot sanely use HTTP-01 renewal. Fine, use DNS-01, but: having to deal with my strict DNSSEC setup where I sign zones at home on a trusted device, keeping my keys off live servers, ensuring keys don't end up on untrusted hardware, and wanting my TLS private keys only on the servers they're being used on—all that suddenly makes my deployment more complex.
What I'd rather do is, since my nameserver runs a cronjob to pull signed zones from my home device (unauthenticated because the signed zonefiles are public information), I wanted to incorporate that into my DNS-01 ACME verification too. But all the existing tools (certbot, dehydrated, acme.sh, et al) seem not to support this style of setup well, where I can batch my DNS challenges, wait for my nameserver to pick them up, and then verify the challenges en-masse an hour or a day later (LE challenges last for seven days, so this is acceptable delay for me).
But I'm stuck with the CNAME/alternate NS approach where I have to run yet another network-facing service from home, for the duration of renewals, because I simply do not have time nor patience to sift through ACME's specification and implement a better tool suited for my needs.
>I get notified by email when there is a problem
LE has a mailing list for critical announcements in the interest of current customers and interested prospectors? Do tell, because all I know about is their blog and the Discourse forum, neither of which are solely for critical announcements.
I think they plan to improve their communication after this mishap.
But you can use an rss reader to subscribe to the incidents category (search incidents.rss in this page), which is very low traffic. It's not “action required” level only, but with only two or three a year it may be suitable.
Wow. That is a truly impressive way to handle a security bug and I know it’s not the first time Let’s Encrypt has responded extremely quickly.
I would love to hear how their engineering practices make this possible.
That is, I'm OK with being woken at 0200 to try to understand and if appropriate fix or recover from a disaster only so long as if I'd suspected this might happen the people expecting me to be awake at 0200 would have given me the resource (money, people, whatever) to fix it. If I feel like I don't have that support, I'll only start looking at your disaster during my working day.
My impression is that ISRG pays a lot of attention to preventing disasters, so if I worked for ISRG (not very practical since they're based on the US West Coast and I live in England) I'd be comfortable taking a call in the middle of the night to fix things.
Basically I keep a phone and laptop on me at all times.
This is in comparison to a friend that works somewhere that always has daily on call incidents that are not actually problems 95% of the time. That would piss me off even if I weren't always on call.
At huge scale there just will be incidents every day, if you're SRE for Google then every day is a bunch of problems - but at that scale those are routine incidents and you can afford the follow-the-sun team (actually I think Google SRE found it was better not to have true follow-the-sun but just two groups) to handle them, leaving only still rare "San Francisco just fell into the Bay" level incidents to actually wake people up.
The workforce is spread across the states. When you're drawn to the mission like I am, late nights here and there don't matter at all. I communicated with my wife and I'm sure others informed their significant others what was going on etc and why Friday, Saturday, Sunday, Monday, and Tuesday would be thrown out of wack. Members of the team put in much more hours than I did and that is truly impressive. It takes all of us with our different specialties to make an accurate and effective response.
Some of the things we do are internal post mortems and find ways to prevent the issue from happening again by either improving alerting/monitoring, writing a runbook, fixing code, and fixing misconceptions about a part of the entire system. We do weekly readings of various RFCs, the Baseline Requirements, and other CAs CP and CPS documents to again better understand our system and Web PKI as a whole. This is an understatement, but we heavily rely on automation. From the moment the call was made to stop issuance, an SRE was ready to run the code that disables the issuance pipeline.
The biggest takeaway is that communication and leadership makes all the difference.
I have to go, there's work to be done.
We have an on-call rotation and a system for getting others notified and online quickly when necessary. We make sure not to bring too many people online so that some people are fresh and can rotate in later if the incident lasts longer.
It's not often that staff have to put in time at night or on weekends, and when it happens we work hard to make sure the problem doesn't happen again.
What's supposed to happen:
For each fqdn in the request
if challenge succeeds (eg dns-01)
check whether caa record exists
if (it doesn't) or (it does and allows issue)
issue certificate
In the step on "check whether caa record exists", instead of using the domain name that is being issued in this loop, it uses the first one it found (or one of them, it's unclear which one). So theoretically, if you wanted a cert for: domain1.example.com, domain2.example.com
and you had a CAA record for domain1 that allowed letsencrypt but then a different CAA record was added between the CAA check on domain1 and the CAA check on domain2 (which wouldn't happen because of the bug) you could get a cert for domain2 that the CAA record said not to issue.There's 3M "missing CAA checking results" in total, of which 2M are dated from 2020 and 1M from last month. FWIW the only certs of mine affected were old certs from 2019-12 which had since already been renewed in Feb, and the renewed certs are not affected?
The largest account has 445k certs revoked, and the most revoked certs from last month (most likely to still be in active use?) is 43k for a single account. I hope your rate-limits are in order if you're going to start reissuing all of those before midnight :/
BTW account number 131 at the top of the file seems to mostly be akamaiedge.net sites :)
If you own both https://www.happy-rainbow-nursery.example/ and https://hardcore.bdsm-videos.example/ you probably go to some lengths to avoid visitors realising the connection. Nothing you're doing is illegal or even unethical - but it's obviously going to cause uncomfortable conversations so why not avoid that altogether. Let's Encrypt aren't doing you a favour if tomorrow a mom at nursery says now she knows why you sound so much like Masked Mistress Martha...
I'm guessing customer IDs are associated with e-mail addresses? This seems like a good case of using different e-mails for ever cert. There are open source tools like anonaddy.com you can host yourself or buy from them (they have a decent free tier).
I feel like this list seriously needs to be pulled. There is some serious lack of oversight here.
They are (on Let's Encrypt's end), if an email address was provided.
It's a 1:n relation, the same email may be used for any number of ACME accounts. Roughly speaking, for most clients, the ACME account maps to a specific ACME client on a specific host. If you run three servers with separate ACME clients, you're probably using three ACME accounts (even if you're using the same email and issuing certificates for the same domain).
Large or custom implementations may reuse the same ACME account across many servers and domains. (Issuance would typically be centralized and operated as a separate system in these scenarios.)
(no SSH, terminal access, etc) and it's from the letsencrypt team (linked in blog post).
for domain in $(cat domains.txt); do printf "$domain :" && curl -XPOST -d "fqdn=$domain" https://unboundtest.com/caaproblem/checkhost; done
What is a CAA? Letsencrypt: please dont use initialisms in customer facing blog posts without using the FULL name at the first use. Makes things more learnable and googleable.
It's a fairly common initialism in CA/TLS world (heh). The DNS record is also named "CAA".
PTR and TXT aren't initials they're just short for "pointer" and "text" neither of which is much help divining what they're actually used for, and presumably AAAA doesn't actually stand for anything at all (?) other than it's four times bigger than the A record.
Had I been given a vote we would be using AAAA and AAAAAA records.
Letsencrypt validates the domain ownership for 30 days, so the bug allows you to issue a certificate within that window, even if you added a CAA record after validation that says "don't allow issue by letsencrypt, or only allow issue by MyCA.example.com".
But if you have everything automated, you're checking for renewal and issuing every day and probably validating as part of that, so unlikely to encounter the bug, unless you validate in one step and then sometime 8h+ later, issue a certificate.
(Will still have a look but less stressfully)
I hope they add support for that soon.
https://community.letsencrypt.org/t/certbot-1-3-0-release/11...
(only certbot-auto users are likely to get this release immediately)
Was this feature something already planned (albeit presumably not for release today) or was it entirely inspired by this problem? I confess if you'd asked me "What extra features does Certbot need?" I would not have listed "Check OCSP to trigger renewal" though in hindsight it's a good idea.
Those do exist, though I don't personally have access to them. When things are calmer, you might ask bmw or jsha on the Let's Encrypt forum for some more information.
> Was this feature something already planned (albeit presumably not for release today) or was it entirely inspired by this problem?
I wrote (incomplete) code for this feature several years ago -- specifically inspired by the idea that certificates might sometimes be revoked unexpectedly -- but I don't think anyone planned to continue working on it until this problem came around.
Check your domain using the linked online tool.
I do get "you didn't renew your certificate" messages on a semi regular basis (domains that have passed out of my control) so I know they have my details.
I assume that by default certbot only checks the expiration date of local certificates against the system clock, it doesn't ping any external resources so it can't be aware that the certificate might have been revoked even though it hasn't expired.
I agree that it would be nice if there was such an option, although I assume that it would increase the server load significantly if certbot connected to letsencrypt's servers at every invocation so maybe that's why they didn't do it.
I think the actual issue here is that the certificates have not been revoked yet. We know that they will be revoked, which is why we have to run with --force-renewal, but there is no process for certbot to know that a certificate, although not revoked, will soon become revoked. I would expect certbot to automatically renew the next time its ran post-revocation.
curl -XPOST -d 'fqdn=example.com' https://unboundtest.com/caaproblem/checkhost
Replace example.com with your Fully Qualified Domain. $ curl -XPOST -d 'fqdn=pop3.mydomain:995' https://unboundtest.com/caaproblem/checkhost
invalid name pop3.mydomain:995
404 page not found
$ openssl s_client -connect pop3.mydomain:995 -showcerts </dev/null 2>/dev/null | openssl x509 -text -noout | grep -A 1 Serial\ Number | tr -d :
Serial Number
0325f31485b9c0f393e27b00e4678e881e3cSuppose on Wednesday you get a cert for example.com and www.example.com, and then on Thursday you realise you also need images.example.com - you use the same ACME account (if you run Certbot this will happen by default if you use the same machine and user account, it silently makes you a free account if you don't have one already) and so Let's Encrypt can see that this account showed proof-of-control for two of these names on Wednesday, so only fresh proof-of-control of images.example.com is needed. Unfortunately this bug means Let's Encrypt forgot to re-check CAA for the old names, and so there's a risk they technically were no longer authorised to issue for these names and shouldn't have given you the Thursday certificate.
Rather than try to argue about whether it's appropriate to disregard this check, Let's Encrypt decided to revoke all ~3 million affected certificates. That's maybe 2-3 days worth of normal issuance in the last 90 days, so lots but hardly "the majority".
cd somewhere-nice
wget https://d4twhgtvn0ff5.cloudfront.net/caa-rechecking-incident... gunzip caa-rechecking-incident-affected-serials.txt.gz
for i in $(cat domains); do (openssl s_client -connect $i:443 -showcerts < /dev/null 2> /dev/null | openssl x509 -text -noout | grep -A 1 Serial\ Number | tr -d : | tail -n1) |tee serials/$i; done
cat serials/* | tr -d " " | sort | uniq > serials.collate
grep $( cat serials.collated | head -c-1 | tr "\n" "|" | sed -e 's/|/\\|/g' ) ../caa-rechecking-incident-affected-serials.txt
It will take a moment and then it may tell you that letsencrypt misspoke when they said they sent emails to everyone whose contact details they have.
I asked (on m.d.s.policy) on the 29th how many issuances were affected, Jacob replied saying they intended to spend that day figuring out the answer, but then there was nothing further from him. The incident doesn't seem drastic enough to prompt urgent answers so I intended to revisit later this week if I heard nothing further.
Now we have a complete list of affected certificates instead, (the answer to my original question is about 3 million)
I was sort of hoping the answer was going to be like five thousand or something manageable. Alas. In hindsight I guess this was to be expected.
So while the number of certificates that should not have been issued due to a blocking CAA record is likely small (or possibly even 0), they have to revoke every cert that could have triggered the bug, as they have no way to travel back in time and find out what the CAA records they didn't check would have been.
If all proof-of-controls were fresh the CAA checks are also fresh for those proof-of-controls so there's no bug. That's why the big list is "only" about three million certificates.
Suppose you own example.com, example.org and example.net and all you do is every 60 days or so you spin up Certbot once to get a certificate for six names, the three domains and the associated www FQDNs - that won't trigger this bug because each time your old proof of control have expired and new fresh ones will be used, triggering fresh CAA checks.
You're right though that it's likely the number of truly mis-issued certificates may be zero because the most common way to have a CAA record deliberately changed to forbid Let's Encrypt after having successfully done a proof-of-control (the scenario that would trigger their bug) is researchers looking for bugs in CAA checking, and of course such researchers would have reported this to Let's Encrypt triggering exactly the same incident but probably at a more friendly time like a Monday morning.
The goal would be to have our automation automatically rotate the certificates when similar issues occur in the future.
for domain in $(cat list-of-domains.txt); do curl -s -X POST -F "fqdn=$x" https://unboundtest.com/caaproblem/checkhost ;done | sed '/is OK./d'
Funny to find out the otherwise awesome traefik has no force-renewal other than messing around with acme.json until it renews
The issue would probably also affect people who have geographically separate certificates e.g. if you have two servers in different regions and decided rather than make things more complicated for key distribution you'll just have them each get their own certificates for the same name - that's totally fine with Let's Encrypt (it doesn't scale, but if you had 500 servers not 2 you'd probably redesign everything) but obviously this test only sees one of those servers and won't check the other certificate.
There's no way to know, given that two (or more) valid certificates exist for a name, and seeing one of them, whether the others are still actively used anywhere.
It would obviously be pretty easy to build a web form where you can type in an FQDN and get told if any certificates matching that name will be revoked, but then you get false positives where it says yes, this certificate for some.name.example will be revoked, you rush to replace your certificate for some.name.example but maybe actually the one that will be revoked is from 20 December 2019, and you already got a newer one which was unaffected in February.
No significant load on their infrastructure, and you'd not have to break the "private keys don't move over _any_ network" rule.
Delegated Credentials is the proposed mechanism to do what you actually want to achieve here. The certificate issuance is mostly the same except there's an OID 1.3.6.1.4.1.44363.44 and the digitalSignature KU is present meaning this certificate is intended to be used with Delegated Credentials and thus can sign things.
Although tightly constrained subCAs (which is roughly what you're describing) might be a way to achieve your goal, it's seen as disproportionately complicated and risky, so hence this proposed much more narrowly defined new feature for TLS.
(Pro tip: you can append .rss to many pages on discourse to get an RSS feed)
Is wading through Discourse threads now the new minimum requirement for using services?
(Do check your domain even if you didn't get an email, since they have not delivered emails to everyone who is affected.)
That's not okay. That's serious negligence.
EDIT: My bad. They have known about this for maybe a week, and they did a blog post 4 days ago. I was basing my thought that they had known for 30 days based on the title of the post and in the URL: https://community.letsencrypt.org/t/2020-02-29-caa-recheckin...
They didn't know the extent of the bug until the 29th of February: https://bugzilla.mozilla.org/show_bug.cgi?id=1619047#c1