Around 293 intermediate CAs in violation of CA/Browser guidelines
mail-archive.com
mail-archive.com
This isn’t a problem because a sub-CA can revoke any certificate from any other sub-CA of the same CA. That would be bad, but, at worst, it’s denial-of-service.
Rather, this is a problem because any sub-CA can effectively reverse the revocation of any other sub-CA, or the CA itself. That’s immensely problematic. Suddenly, the CA has no reliable way fully revoke certificates. Revocation is already somewhat broken as it is, but this gives a lot of entities the ability to deliberately interfere with revocation in ways that they shouldn’t be able to.
The author goes on to explain that revocation of the affected certificates is insufficient, because they could be used to effectively reverse their own revocation at any point in the future. Instead, it must be proven that all copies of the keys have been destroyed. That’s quite an undertaking.
What the author fails to mention is that revocation is already pretty broken. Most major browsers have their own built-in CRL replacements that contain the most important revocations they need to know about. Some browsers, like Firefox, may make additional efforts to ensure that any given certificate hasn’t been revoked; others, like Chrome, don’t. If you’ve ever visited a site that gives you a certificate error in Firefox but not Chrome, that’s likely why.
In the case of browsers, it should be possible for each browser to forcefully revoke affected certificates, but revoking a sub-CA certificate is quite disruptive, so I’d be surprised if that happens within 7 days. The catch is that this technique is really only effective in modern, up-to-date browsers.
In any case, the title is misleading. I don’t see where the author guarantees that this will happen within 7 days. The author claims it should happen within 7 days, but considering that the damage is already done and cannot be fully reversed by revocation, I find it hard to believe enforcing that deadline makes sense here.
Also, given that the underlying cause appears to be ignorance, it would be prudent to take things slow and ensure that this doesn’t happen again. As I said before, the damage is already done—revoking appears to be insufficient here.
If this does actually happen within 7 days, though, I will be thoroughly impressed.
EDIT: Mozilla’s reply: https://news.ycombinator.com/item?id=23748561
https://cabforum.org/wp-content/uploads/CA-Browser-Forum-BR-...
Isn't the main reason you would want to revoke keys because they were disclosed, making it impossible to destroy all copies?
To provide an analogy in the context of PGP keys, if an attacker somehow finds a backup of your revoked and destroyed private key someday, they will have trouble using it because your revocation will be public and on record.
How would this be verified? Presumably the keys are stored on HSMs, but you can I'm not sure how you can prove that you didn't make a backup of the key.
In practice, of course, that doesn't mean every one of them will have done. There's 293 of them, after all.
In the event that a key with a Key Destruction Report shows up again, the responsible party for that key will have shown unacceptable negligence and will potentially be subject to the exclusion of their keys as a valid certificate signer.
A lot of these companies core businesses rely on remaining in a position to sign certificates so it is in their best interest to protect that privilege by following the documentation requirements, and properly destroy their keys. It's effectively a pretty good stick.
Basically on Chrome one of my sites is throwing:
"NETT::ERR_CERTIFICATE_TRANSPARENCY_REQUIRED"
for most users but not all, even though they're all on Chrome. It seems to work fine in other browsers.
https://transparencyreport.google.com/https/certificates
When I check my domain here it seems like I have got the transparency certificate so I shouldn't be getting this error.
Is this related to what you're talking about? I would really appreciate any help. I'm using https://letsencrypt.org/ for the cert.
openssl s_client -connect example.com:443 -servername example.com </dev/null | openssl x509 -noout -text
which should print an SCT extension at the end - my version displays it by numeric identifier "1.3.6.1.4.1.11129.2.4.2" but maybe newer versions display it by name.Alternatively, I think you might able to go to https://www.ssllabs.com/ssltest/ and see if your cert has "Certificate Transparency: Yes", but I'm not sure exactly what that means.
Anyway, I don't think this is related, the question at hand is about OCSP, which is a different mechanism from Certificate Transparency. (Arguably Certificate Transparency is a replacement for revocation in general being flawed in practice for many reasons, but they're different mechanisms.)
If you run SSLLabs against your host name, does it say “Certificate Transparency: Yes” or No?
https://www.ssllabs.com/ssltest/analyze.html?d=your-hostname
It's extremely unlikely to have anything to do with this incident.
You should obtain a copy of the certificate which triggers NET:ERR_CERTIFICATE_TRANSPARENCY_REQUIRED and take a look at that. There's an excellent chance there's something else even more obvious wrong (from your point of view as a human) but Chrome decided to focus on the lack of trustworthy SCTs.
My instinct would be that it's likely a middle box (e.g. "anti-virus software" on a PC can install itself to snoop on all HTTPS sites, or a corporate "data loss prevention" proxy or that sort of thing) and the bogus certificate will likely make that pretty obvious if you examine it.
It's a bit of a heisenbug but it's occasionally reported on the Let's Encrypt forums. It always goes away for the reporters just by waiting a little bit.
It would be really nice if a user who runs into this could generate a Chromium event log which would hopefully include the SCT events (chrome://net-internals).
What would the recourse be here if one of those keys were to be compromised, or even if there was reason to believe one might have been? Would the entire CA-level trust chain need to be distrusted, requiring re-issuance of all certificates on that chain?
The "good" news is that most people haven't really been treating revocation (and OCSP) as a reliable mechanism. The major browsers all have out-of-band mechanisms for revoking known-malicious certs via something equivalent to the software update channel, which bypasses reliance on the CA infrastructure. If there's a large-scale attack, the relevant cert/CA will probably be disabled through that mechanism. And most of the smaller clients don't even bother with revocation checking at all (e.g., I'm pretty sure that on an average Linux system, things like curl or "import requests" do no revocation checking) so there's no point in undoing revocation if you're trying to attack one of those systems.
This is correct. Even where there's some provision for checking, it's usually a mechanism where you can supply a CRL (Certificate Revocation List, a signed and dated document that says which certificates were revoked). CRLs are practical for a small private CA but they make no sense at scale. Let's Encrypt doesn't even have CRLs because they'd be enormous.
To be fair 10-15 years ago there's a good chance that average Linux system has a set of CA roots which hasn't been updated in a decade, and most such clients aren't actually checking even CN let alone SANs so bad guys don't need a google.com certificate (or whatever) they can just get themselves a real certificate for actual-bad-guys.example and the client won't check the name matches anyway.
This is a convenient fiction. CA system never protected anyone against state actors. Never did, never will. Subverting a single CA is enough to compromise entire system. And there are hundreds of them.
Security is always grounded in knowledge and physical control — understanding and exercising your capabilities to preserve them. A blind, deaf and fully paralysed person can't be expected to safeguard their own physical security, and neither can an average user — their TLS security. Especially against state actors. More so, when the parties they have to rely on are commercial enterprises whose entire existence revolves around getting paid to issue certificates.
https://www.mail-archive.com/dev-security-policy@lists.mozil...
> We are concerned that revoking these impacted intermediate certificates within 7 days could cause more damage to the ecosystem than is warranted for this particular problem. Therefore, Mozilla does not plan to hold CAs to the BR requirement to revoke these certificates within 7 days. However, an additional Incident Report for delayed revocation will still be required, as per our documented process[2]. We want to work with CAs to identify a path forward, which includes determining a reasonable timeline and approach to replacing the certificates that incorrectly have the id-kp-OCSPSigning EKU (and performing key destruction for them).
Let's Encrypt represents the state of the art in terms of certificate automation, but last time they had an (impending) mass-revocation event, it turned out that even the ACME protocol / client implementations didn't really have any concept of an automated "this certificate is about to be revoked, please re-issue" process. As a result of that event, certbot at least got support for triggering renewal after revocation: https://github.com/certbot/certbot/issues/1028#issuecomment-... -> https://github.com/certbot/certbot/pull/7829
The Firefox out-of-band revocation mechanism (OneCRL) certainly could revoke all these intermediates but Firefox isn't vulnerable to a problem here, so there's no obvious upside to doing that and it's disruptive.
The CAs can't have it both ways: a BR balloting process that they rely on for moral authority when disputing that the majority of deployed browsers have added new security requirements (like shorter-lived certificates), and BRs that they ignore when they screw up.
If Mozilla isn't the majority browser vendor, who cares what they insist on? And if all the CAs band together and say, sorry losers, we're gonna keep doing things our way, what are the browsers gonna do? Cut all their users off from the internet "because principles"? Apple is playing a dangerous game that I don't think will work out in different circumstances. They can't hide behind "protecting users" if their users end up unable to access the internet securely.
We got into this mess because we wanted organizational independence and distributed trust, without considering what internal conflicts would mean to the end users. I'm going to call it and say that within a decade, you'll have to pick which CA you want to trust at browser install time (though you can guess which CA will be the default on which devices).
For one thing you should consider that several of the companies that make browsers also operate a Certificate Authority. For example Google's GTS is represented in that m.d.s.policy thread by Ryan Hurst (whereas Ryan Sleevi is there mostly choosing not to put on his Chromium Web PKI hat). Microsoft likewise controls trusted roots. Apple does operate its own PKI but is not presently broadly trusted, though if it felt the need I'm sure they have people.
Also the most popular (for this purpose) Certificate Authority is ISRG's Let's Encrypt and they've got no reason not to co-operate. The Web PKI is all they do anyway.
This is often portrayed as though browsers are obliged to either distrust everything instantly or allow CAs to do whatever they like, but neither of those is realistic. The major trust stores all already have imposed constraints short of distrust.
You bring up Apple as an example. Unless I've missed it somewhere Apple never announced their 398 day limit as a matter of issuance policy it's simply a fact that Safari and the Mac operating system won't trust new certificates with longer lifespans after a set date. So if one or more CAs decided not to co-operate, nothing happens at first. Nothing whatsoever.
Then, gradually, a few subscribers buy (or renew) 2 year certificates, and these new certificates don't work in Safari. Some of these subscribers will call customer services at the CA where they purchased the certificate. Why doesn't it work in Safari? How can they fix this?
What does the CA say? "We intentionally sold you a product that won't work. Ha ha ha, it's a funny joke, we have your money and you've got useless garbage" ? Maybe they instead try to blame Apple. Apple will point out that the CA knew this wouldn't work and suggest the subscriber seek a refund.
The subscriber demands a refund. The CA is now losing money and it is seeing negative reputational impact. Somebody threatens to sue. It is not a good day to be the CA.
At the subscriber's offices, an IT person has a brainwave and switches to a provider that isn't deliberately disobeying. The web site is back working. Champagne all round.
To achieve the desired robustness/ reliability the Web PKI is structured in a way that makes any individual Certificate Authority expendable. As a subscriber this means you should plan for your CA going away with at most a few days notice. Most people won't do that. Too bad, individually you're expendable too.
IMO, DANE might make sense if DNSSEC wasn't such a mess, although it is a very similar group of parasitic companies involved in DNS. In general, alternative name systems (such as the GNU Name System) could also potentially replace the certificate system and many name and certificate issues are related. Many of the hardest technical issues around certificates relate to revocation and the demonstrated inability of almost anyone to secure anything.
Other options that make a lot of sense in many ways would have govenments or banks involved in identity in a direct way. This is resisted for a varity of reasons.
I hear this a lot, but in my experience (managing c. 1000 DNS zones all with DNSSEC enabled, using a strictly DNSSEC-validating resolver for >5 years, and having built DNSSEC infrastructure for DNS hosting providers), it is both reasonably well designed and generally quite well implemented. What is the mess that you perceive?
Not GP, but the mess is near-zero clients and few recursive resolvers are actually doing DNSSEC validation in practice, after 20 years of deployment.
It’s like IPv6.
Also I believe most active ZSKs are actually held and managed by the larger DNS providers on their customers’ behalf. This leads to very little assurance improvement over unsigned records, as credentials to update a web form is all that is needed to “sign” records. There are no real key management requirements for ZSKs as there are with browser CAs.
The only additional assurance provided by a DNSSEC response is that there was likely no MITM between the authoritative server and validating resolver. Which is something, but that problem is more easily and completely solved by DoH which adds privacy as well as authenticity.
>For example, consider a certificate like https://crt.sh/?id=2657658699 . This certificate, from HARICA, meets Mozilla's definition of "Technically Constrained" for TLS, in that it lacks the id-kp-serverAuth EKU. However, because it includes the OCSP Signing EKU, this certificate can be used to sign arbitrary OCSP messages for HARICA's Root!
>This also applies to non-technically-constrained sub-CAs. For example, consider this certificate https://crt.sh/?id=21606064 . It was issued by DigiCert to Microsoft, granting Microsoft the ability to provide OCSP responses for any certificate issued by Digicert's Baltimore CyberTrust Root. We know from DigiCert's disclosures that this is independently operated by Microsoft.
So my understanding is this: The CA's have issued certificates/sub-CA certs without the proper extension (or with too many extensions), causing those to be able to sign a OCSP response. And the Online Certificate Status Protocol (OSCP) is used to check the revocation status of certificates with the CA.
So, this would allow e.g Microsoft to generate a fake OCSP response? That would perhaps be useful in some kind of MITM-attack scenario?
While not good, perhaps not an end of the world problem either? However, I wonder how much problem will come for people needing to replace those soon to be revoked sub-CA certs...
But as far as I know, browsers are not failing hard on OCSP failure, if you can mitm the connection possibly you can block OCSP requests too.
Would you trust someone who doesn't take issues seriously because they think they're small or unimportant?
EDIT: reading the full report, it seems that the underlying risk is that if one of the intermediate CAs were to be compromised, even if it was revoked it could theoretically forge an OCSP response that it is still valid (and as a trusted CA issue certs for anything). So the response is very appropriate given the potential impact.
If someone compromises a key, typically, you would want to revoke it. However, if that key also allows to revocation to be reversed, you’re in trouble.
I’ve explained more in a top-level comment: https://news.ycombinator.com/item?id=23747524
As I understand it: the issue isn't the nocheck; it's where the OCSPSigning EKU is. You're supposed to see OCSPSigning on end-entity (CA:NO) certificates; the purpose of the EKU is to delegate a non-CA cert the authority to revoke certificates for its parents. When you see that EKU on a CA:TRUE cert, what you're really seeing expressed is that CA's parent delegating OCSP for the root; ie, the CA is granting its customer the right to control revocation for the whole CA.
What nocheck expresses is: "you can't trust this OCSP Delegated Responder to revoke itself, because that's silly; seek confidence in its validity elsewhere". "Elsewhere" apparently usually means "the fact that this certificate has a very short lifetime", which is feasible for an end-entity cert but not so much for a CA.
My understanding is that nocheck (or, lack of it) is how Ryan spotted these certificates, but isn't really the big problem with them.
“ This is https://misissued.com/batch/138/
A quick inspection among the affected CAs include O fields of: QuoVadis, GlobalSign, Digicert, HARICA, Certinomis, AS Sertifitseeimiskeskus, Actalis, Atos, AC Camerfirma, SECOM, T-Systems, WISeKey, SCEE, and CNNIC.”
> For example, consider this certificate https://crt.sh/?id=21606064 . It was issued by DigiCert to Microsoft, granting Microsoft the ability to provide OCSP responses for any certificate issued by Digicert's Baltimore CyberTrust Root. We know from DigiCert's disclosures that this is independently operated by Microsoft.
Pretty much. The whole business model never really made sense: the relying parties have no relationship with the certificate authorities, while the HTTPS servers are the customers of the CAs.
I think it would make a lot more sense for certificates to be issued by domain owners, esp. since the original idea of tying sites to real-world businesses (e.g. with Dun & Bradstreet numbers) has been reduced to just verifying domain-name ownership.
Edit: I think people misunderstand what I am saying here. What I mean is that I think that when one purchases a subdomain of domain, that domain should just issue a certificate — and that domain should only be allowed to issue certificates for its children. So e.g. if one purchases foo.com, then com issues a certificate for foo.com; if one purchases bar.net, then net issues a certificate for bar.net; if one purchases baz.ac.uk then ac.uk issued a certificate for baz.ac.uk. This is essentially what Let's Encrypt and ACME already do: com has the technical ability to reassign any of its subdomains at any time it wants to, and can get a certificate issued for any of them by reassigning & registering a certificate.
And while we're at it, maybe we could kill ASN.1 with fire?
Edit: if you downvoted for this, you have never tried to debug an ASN.1 BER file.
The problem with that approach is that anyone can create a certificate for any domain; so if I go to "example.com" then it's kinda hard for me to detect if my connection is being MITM'd, especially if this is the first time I'm visiting example.com.
This is why ACME requires a verification that you actually control example.com (via http or dns).
I don't think the CA model is perfect by any means, but I don't think it's completely without value either.
The idea that any CA can issue a valid cert for any domain is the heart of what's wrong with PKI.
Software support is far from universal sadly.
Ironically, the only system PKI had to attempt this, Extended Validation, is opposed by the loudest voices in PKI today. Despite arguably being the only real benefit PKI potentially offered: Notarizing that a domain really belonged to a given real-world entity.
EV had flaws, but it should've been improved, not axed. Security detached from people-understandable real-world entities will never provide real security, because at the end of the day people still need to interact with the system.
You misunderstand what I mean: I advocate that the owner of .com be permitted to mint certificates for foo.com, bar.com, since right now the owner of .com and can point those subdomains to any host he wishes, and then generate a certificate using ACME (because he actually controls every subdomain of .com).
Using DNS providers for certificates is an interesting idea; one I haven't heard before. I can't really think of any downsides of that at the moment.
I thought they meant the .tld registry would issue the certificate, so any registrar could sell you the domain+cert but it would have to come from the registry (ICANN say, for .com).
Can't the DNS data have a hash of the cert to avoid 3rd party certs (unless the 3rd party controls the domain registry entry, but then MitM is a [ahem] dead cert).
That might have worked decently in the early internet but it does seem seriously flawed with the current stakes.
That being said, what's the alternative? TOFU? Web of Trust? Those have massive security implications as well. They have the advantage of putting the user back in control but given that the vast majority of the people using the web today doesn't have a deep understanding of the underlying technology and security model I don't see how this wouldn't end up in a massive catastrophe.
It's a tough problem to solve.
I'd prefer a system backed by DNS, and based on verifying the ownership of domains and the authorized DNS provider for that domain. Presumably, in my example, the only domains Google would be authorized to secure would be domains provided via Google's DNS and domain products.
Um no. Google's four production roots (GTS Root R1 through R4) are essentially dormant. You could (but probably shouldn't) manually distrust these roots with no impact.
Never could get the professor to double back around to th He mote problematic stuff.
Formulating a working alternative is far from trivial.
Understanding where web security is right now is about understanding who is making the decisions (regardless of any claims about committees and processes), and what motivations they have to make the decisions they do.
I think the main alternatives people suggest are
- something involving a distributed ledger, where revocation isn't even an option, so that clearly doesn't make it better than the current system if we're talking about revocation being a mess (we could just amend the current system to get rid of revocation and throw out a whole bunch of technical complexity if we wanted)
- something involving DNS, which also involves trusting a bunch of companies nobody's heard of (sometimes the same companies, in fact?) who are hardly obviously better at operating cryptographic infrastructure than the existing CAs
- a TOFU approach like SSH, which hasn't been demonstrated to scale well beyond the dozen or so machines in your known_hosts file (most large companies are using something other than TOFU even for internal SSH)
I don't think PKI is an objectively good system, it's just difficult to picture a better one. The main flaws with PKI in practice aren't really about the companies nobody's heard of or a web browser being run by an adtech company - the main flaws are that people want a lot of things out of the system, some of which are contradictory, and running cryptography at this level of scale is genuinely hard. The alternatives don't really address those problems.
Right now, how many different companies could issue a microsoft.com cert if compromised or sketchy? Hundreds?
Right now CAs delegate trust to bunches of questionable sites as seen here with poor oversight or security based on business interest. On a DNS-based system, the entities involved are limited to those who actually manage your DNS.
It also removes the agency of browsers to decide who does and doesn't get to play, which is the current system.
Given that most of the problems with the CA system historically have not been active attacks but incompetence, I don't think we win much from moving to a system where we can, in fact, kick TURKTRUST out of the pool to one where the question is whether .tr remains part of the internet or not. If Verisign screws up with .com in any way short of revealing a letter from the FBI saying "Please help us MITM Windows Update," there will be immense pressure to allow Verisign to continue being the .com registry and continue holding the .com signing keys.
For similar reasons, I'm not convinced that moving from "Hundreds of unqualified companies could issue a bad cert, but hopefully they won't" to "One unqualified company could issue a bad cert, but hopefully it won't" is a meaningful benefit. It doesn't reduce the theoretical bounds on the attacks, and again in practice, these hundreds of companies haven't been misissuing. (The present story is about mis-delegating the power to issue revocation/non-revocation responses, which is certainly a problem, but only relevant in practice if there are actual end-entity certs that are misissued in the first place.) So while it certainly feels better to have fewer entities that can sign - and to be clear, I am all for distrusting many if not most of them - I don't think it addresses either the fundamental theoretical problems nor the actual real-world attacks.
The Verisign CA function was sold to Symantec. That name might ring a bell too, because with these CAs set to be distrusted as a result of Symantec's mismanagement the whole business was again sold to DigiCert in 2017.
I think the perverse part of your reasoning is that you think .com is trustworthy now. It's one of the worst run registries. Its popularity with businesses probably tells you more about how scammy most businesses are than whether .com is trustworthy, and not very much about either.
It's a mistake to separate out the certificate signing authority for different attention if it would be (as in DNSSEC) hierarchically constrained. Verisign can already screw up badly enough to cause Microsoft to lose control of microsoft.com or let somebody else have it. They've apparently decided they're comfortable with their capacity to mitigate that risk. Fine.
Certificates are problematic. But that's because the world is problematic. They were much more problematic 10-15 years ago. Google and Mozilla drastically mitigated their problems. They're still imperfect, but anything that expresses web trust across the entire world is always goign to be imperfect, and the webPKI at least has some stakeholders that are both empowered and deeply give a shit about security.
If you use any of the major auto-issue, auto-renew certificate platforms, then you do not need to take any action — either you aren’t affected or they’ll issue a new certificate to you automatically. (Let’s Encrypt, AWS Certificate Manager, Google-Managed SSL Certificates, Heroku Automated Certificate Management)
(This affects the PKI as a whole, because a single unconstrained compromised sub-CA can misissue for any domain - so if you use HTTPS, you're affected).
If your question is whether you should take action, the answer is no, unless you're in some way responsible for an intermediate CA, which is something you'd know about.
Edit: Now getting mac error: https://i.imgur.com/JmdC8Yi.png
[1] Normal LE: https://www.ssllabs.com/ssltest/analyze.html?d=www.mail-arch...
[2] Test site: https://valid-isrgrootx1.letsencrypt.org/
[3] In the developer console, there should be a security tab with a View Details button.
Screenshot: https://i.imgur.com/JmdC8Yi.png
Now I get a MAC error insteaf of cert error
Perhaps the stingrays are acting up this morning ;)
1. Get a nice shiny modern Wireshark
2. Tell Firefox you want it to keep records of the session secrets that secure TLS. Set environment variable SSLKEYLOGFILE=/some/path/to/log/secret.keys
3. Packet capture the session you're interested in
4. Give Wireshark the packet capture (if not captured inside Wireshark itself) and the secret.keys
5. Now Wireshark can show the TLS session and you can see what went wrong in detail. So long as you didn't actually do anything secret you can give all these pieces to somebody else to look at.
6. Otherwise, after your investigation destroy the secret.keys and optionally the packet capture itself.
I've used this level of effort to show a customer that, contrary to what they believed they were not presenting the nice client certificate I'd issued them when connecting. It turned out to be a config difference between their staging and production systems or something. But they were absolutely insistent their software was being turned away despite using a client cert (we used mutual TLS) so it took posting a Wireshark capture proving otherwise to get them to actually investigate.
No action is otherwise required for end users, except for those site operators who manually deploy SSL certificates.
If you manually deploy SSL certificates, then make sure that your contact information with the issuer (e.g. Digicert) is up-to-date; if your certificates are affected, they will contact you and advise you on how to proceed with reissuing and redeploying. Or, you can contact their customer service, reference the mail-archive link, and ask if your certificates are affected.
If your SSL certificates are deployed automatically (Let’s Encrypt, AWS ACM, etc.) then no action is required. Either your certificates are unaffected, or they’ll be updated automatically by the automation.
Q: What’s a simple summary of the problem?
Imagine if root tried to userdel a malicious account, and the non-root user being deleted was able to tell the system “ignore that, I’m still valid”. The system would need to be quarantined and a replacement built without that flaw, since you could not state definitively that you had deleted the malicious user account.)
This issue would allow a non-root certificate (‘intermediate’, ‘subCA’) to undo deletion (‘revocation’) of itself by the root authority, as well as undo deletion of its siblings (other intermediates issued by that root authority). To correct the issue, all non-root certificates issued incorrectly in this manner must be revoked and then destroyed, with proof of destruction. This ensures that the mis-issued intermediates can’t zombie-return someday by undoing their own deletion or the deletion of others.
The issue† --- it's super complicated in its particulars but not in its outline --- is that CAs have been issuing constrained intermediate CAs (CA's that can sign only a subset of certificates) that, because of a misconfiguration, can sign any OCSP message for their root CA. So, for instance, Sleevi points out a HARICA (Greek university CA) cert that is constrained by dint of not having the serverAuth EKU (so it can't sign TLS certs), but because it has the OCSPSigning EKU, can be used to sign OCSP messages.
What happened was this: lots of CAs run by companies in the wild (remember: there are all sorts of constrained CAs running inside companies that you aren't supposed to have to care about, in part because they're constrained) run on Microsoft's CA software. Many of those CAs want to support OCSP. The way you generally set up OCSP is to sign a Delegated Responder certificate, which is an end-entity (non-CA) cert with the OCSPSigning EKU, which says "this certificate can be used to sign OCSP certs for this part of the PKI". But Microsoft's CA is broken: it won't let you sign a cert with the OCSPSigning EKU unless your CA cert also has that EKU (the EKUs must "chain"). So, that's what CAs did, despite the fact that chaining EKUs like that changed the semantics of what the certs were intended to express.
So, as Sleevi put it, basically a bunch of CAs accidentally signed the OCSP equivalent of a bunch of CA:TRUE certificates. They gave their customers the ability to essentially disable revocation for the whole CA.
We can go back and forth on how sound the WebPKI revocation infrastructure is with or without this mistake. But at a minimum we should be able to stipulate that strengthening revocation is an important project for browser vendors, and CAs that accidentally break the revocation infrastructure by misissuing certificates are an impediment to that project.
There's an interesting backstory to this, which is what I actually wanted to write about, and that's CA/B Forum SC31. The CA/B Forum is the standards body that coordinates between browsers and CAs; they maintain the BRs, which are the bylaws for operating a CA that browsers will trust. The politics of CA/B Forum are weird, because CAs and browsers are structurally adversarial: browsers want maximal security regardless of the commercial implications for CAs (as they should). As a result of that weirdness, the browser vendors rely not only on the BRs, but also on their own bylaws for their respective root certificate programs. SC31 is an attempt by Ryan Sleevi to align the CA-approved BRs with the browser-approved root programs.
Of course, the browser root programs, particularly Mozilla and Google's, are the only thing that really matters, because those rules determine whether the browsers themselves will honor a CA's certificates. But the CA/B Forum is dominated by CAs who sort of dispute this fact. So, SC31 asks that the BRs import that (now common) browser root program rule that certs can only live for just over a year. The "one year only" rule was considered previously by the CA/B Forum and failed in a vote (CA customers don't like that rule), so the CA's are unhappy to be relitigating that point. But then, the litigation itself is kind of theatrical, in that browsers already made this decision and mooted the debate. Standards bodies are weird.
Anyways, what makes this funny is that, while the one-year restriction is the real problem that seemed to piss off the CAs in SC31, they also challenged some of its OCSP language. In the course of debating with them about whether SC31 described reality or not, Sleevi started looking at their OCSP issuances. And here we are: the CAs aren't even following their own BRs, and are breaking their bylaws in ways that materially impact TLS security.
I understood literally none of this when I first read Ryan's message. I thought I knew some stuff about the WebPKI, but even with help on Twitter I had trouble with this eldritch stuff. I don't know how people like Sleevi do it, and, apparently, neither do many CA operators.
This probably would have been a boring blog post anyways, but to make Kurt happy, I'll conclude with: "use Fly.io!". :)
† Wrinkle: Sleevi flags these CA certs as missing the ocsp-nocheck attribute, which is presumably how he spotted them. ocsp-nocheck says "this certificate can't be trusted to OCSP revoke itself, because that's silly"; the idea is, certs flagged ocsp-nocheck rely on very short lifetimes rather than OCSP to manage revocation. Very short lifetimes are things you expect on end-entity certificates, which is what OCSP Delegated Responders are, but not so much on CA:TRUE certificates, which are what Sleevi actually found.
I would quibble with the description of CA/B as a standards body. It's a standing meeting. Any CA/B documents including the Baseline Requirements manage relationships only between (potential) parties to CA/B itself, the browsers and public CAs.
The reason the meeting exists anyway is that it sucks for the public CAs if the major root programmes have conflicting rules or interpretations of those rules. The idea of the BRs is to as much as possible agree rules with everybody to avoid such conflicts.
Also I would rate Microsoft and Apple as equally significant with Google and Mozilla and it's at least tempting to add Oracle (because Java has its own root trust programme) too.
In terms of whether they'll throw their weight around we're used to seeing that from Google and Mozilla (e.g. on SHA-1 and on the Blessed Methods) but Apple proves here (on the one year certificates thing) that sleeping giants might wake up and kick over everything you're doing if it displeases them.
If Microsoft were to, for example, have refused to trust ISRG (they took a very long time to actually make any decision) I'm sure that we'd moan about it, but we can't make them, and without being trusted in millions of Windows PCs when their cross signatures expire that would the end of the story for Let's Encrypt right?
The other non-Google, non-Mozilla companies don't have anywhere near this level of public disclosure from what I've seen, and I suspect that they may follow some of Mozilla's groundwork (especially in terms of dishing out punishments in response to incidents).
Google actually doesn't operate transparently either, except in the sense that it chooses to participate in m.d.s.policy. You won't find a public process behind Google's decision to require CT for Symantec's roots before it was mandatory for other roots for example, they just announced the policy as a done deal.
You are not alone in concluding that Mozilla's distrust decisions (I wouldn't characterise them as "punishment") are in practice copied by the other root trust stores. It is entirely possible that Microsoft (for example) has a large team of dedicated experts independently investigating incidents and just coming to coincidentally similar conclusions. After all, the facts won't be different if a Microsoft team investigates them than they are when Mozilla and third parties do so for m.d.s.policy. But it's a hell of a coincidence...
I would note that for initial trust decisions Microsoft in particular does not follow m.d.s.policy. If you run Windows there's an excellent chance that your computer (and thus Internet Explorer, Edge and Chrome on that computer but not Firefox) trusts poorly run Certificate Authorities from a variety of organisations and countries which don't seem very trustworthy.
For example the governments of Sweden, Slovenia and Thailand.
[Edited: This used to mention Venezuela but the Venezuelan government CA was in fact distrusted by Microsoft]
Now maybe Microsoft's team carefully vetted all these dozens of Certificate Authorities that aren't trusted elsewhere and concluded they're doing a great job. In some cases we know they weren't able to satisfy Mozilla (or volunteers contributing to m.d.s.policy) but in other cases they never applied at all. Maybe they're just shy?
So far we can say this doesn't seem to have caused any serious reported problems. So maybe it's fine.
That link says the even number month changes are CA led.
Now of course you certainly have much better insight than I do into what's behind those CA led changes because I'm just a Relying Party with their nose pressed against the window. Maybe that new Root Program manager is encouraging participants to clean stuff up with an implied threat that if they don't Microsoft will. But as an outsider it still looks a lot like the old Microsoft root programme to me. Also Microsoft's "revoke or else" rule still sits badly with me despite its purported use to prevent people scamming Microsoft's customers. But I guess I'm glad to hear you think they've "greatly improved".
The problem with intermediate CAs with the 'OCSP Signing' EKU being able to sign OCSP responses for any sibling certificates issued by the root CA seems to have been recognized for a while. But in this case, Mozilla allowed and ignored the 'OCSP Signing' EKU for vendor interop with existing intermediate CAs. Mozilla products would reject such OCSP responses, but the CA policy on issuing intermediate CA certificates with the 'OCSP Signing' EKU was not fixed.
Specifically, MSADCS (Microsoft software used to issue certificates) requiring the CA certificate to have the 'OCSP Signing' EKU, "by design": https://support.microsoft.com/en-us/help/2962991/you-cannot-...
The situation seems to be that the MSADCS policy on requiring the 'OCSP Signing' EKU on the intermediate CA certificate is incorrect, and clients validating OCSP responses signed by a delegated OCSP responder certificate do not require the 'OCSP Signing' EKUs on the issuing (intermediate/root) CA certificate(s): https://groups.google.com/forum/#!msg/mozilla.dev.security.p...
This is in contrast to e.g. the TLS serverAuth EKU, where these EKUs chain: a technically constrained CA certificate must have a TLS serverAuth EKU in order to issue end-entity certificates with a TLS serverAuth EKU.
The 'OCSP Signing' EKU on the intermediate CA certificate is unnecessary: the intermediate CA does not need the EKU to sign OCSP responses directly, nor to issue a delegated OCSP responder certificate with the 'OCSP Signing' EKU. The unnecessary 'OCSP Signing' EKU on the intermediate CA certificate is harmful, because it may be interpreted as a delegated OCSP responder certificate for the issuing root CA, capable of signing OCSP responses for any certificate issued by the root CA.
The workaround for MSADCS is to use an untrusted CA certificate (e.g. self-signed) with the same public key and subject DN as the intermediate CA to issue the OCSP responder certificate, adding the 'OCSP Signing' EKU to both the OCSP responder certificate and the workaround-CA. The resulting OCSP responder certificate with the 'OCSP Signing' EKU will validate as an OCSP responder certificate for the actual intermediate CA, without including the 'OCSP Signing' EKU in the CA certificate.
This BR policy violation approach seems to be an attempt to highlight the security problems and fix this long-standing issue, limiting the use of the 'OCSP Signing' EKU in order to protect clients from OCSP responses signed by intermediate CA certificates that were not intended to be delegated OCSP responder certificates for their root CAs.
Comodo at this point controlled the CA roots that had belonged to Symantec. Trustico, a Symantec reseller (same sort of relationship to Symantec that your local Ford dealer has to the Ford motor company) asked Comodo to mass revoke thousands of certificates it had sold to third party subscribers as reseller for Symantec.
It's not clear what Trustico hoped to achieve by that, maybe they believed they could get back the cost of the certificates? We don't know the details of the (confidential) contract between Trustico and Symantec or to what extent the contract terms survived transfer to Comodo. Maybe Trustico just wanted to push its customers into new deals, because it was not a Comodo reseller and risked being frozen out.
Anyway, Jeremy Rowley, a Comodo VP asked for a reason to revoke these certificates, and by return he got thousands of private keys. Private Key compromise is a valid reason for revoking certificates, so Rowley confirmed the certificates matched these private keys and Comodo began revoking them.
Trustico are the people who had thousands of private keys. A CA is strictly prohibited from having your private keys (and as we saw, if they are shown them they should revoke your certificates) but of course whether a CA enforces this rule on its resellers (via contract terms) is a matter between the CA and reseller. The whole point of private keys is that they're private. So, in one sense Trustico's customers got what they deserved - do not give your private keys to some reseller or trust them to pick keys for you.
Nothing went wrong at Comodo here. And if you as a subscriber followed good practices you weren't affected either even if you'd bought certificates through Trustico. Only customers who'd gone with Trustico and done something inherently unsafe got burned.