The client authors (of browsers) tend to care about making their software easy to use and avoid errors. And there’s a pretty small set of clients.
TLS servers on the other hand often just serve whatever you put in a config file. A chain that works for them on an up to date Firefox may not work for an older android phone, but the server administrator may not be aware of the trade off they’ve made.
My personal opinion is that the tools for debugging for server authors aren’t good enough yet. I’m one of the authors of the certigo tool mentioned in the article, which was written out of frustration from a lack of other tools. Even then, I think it’s not good enough and I hope this class of issue is one we can illustrate better. SSL Labs probably does the best here.
This doesn't make sense. Given a cert that was issued from an intermediate, there is exactly one intermediate that issued it, and that intermediate must be in the handshake. If that intermediate was issued from another intermediate, that next intermediate must also be in the handshake. And so on. There's no ambiguity here. Your server is required to send a fully-formed cert chain whose final cert is signed by one of the roots the client trusts.
It's entirely possible to run a PKI where the set of intermediates is known by nobody and is unbounded and ephemeral, so long as everyone who is given a cert is given the intermediate chain that issued it. The burden is on the bearer of the cert to include it in the handshake. All problems here have arisen from a misunderstanding that causes people to think the intermediates should be well-known or obtainable publicly. That is simply not true.
> it’s not possible to always know what roots a client trusts
It doesn't matter. You send the chain that issued your cert, not what you expect the client to trust. If the client trusts it, you're good. If not, you're not.
Once the missing intermediate certificate issue had been identified, the fix (to the server) was indeed clear. But it's not stated whether the author is in charge of the server in question (it's only mentioned as “couldn't connect to a particular HTTPS site”); I suspect this was some external web site not under their control.
If your client doesn't trust the CA that signed the intermediate that signed your cert, you'll need a cross signed CA cert that works for your clients, but sometimes it's not possible to make a single chain that supports all the clients you want. And it's difficult to differentiate (and even if you can figure out how to differentiate, you're probably going to be patching your TLS server to do it, because it's not widely supported).
When you consider poor client behavior, you get situations like client A doesn't have your CA, but does have another you can cross chain to; client B has your CA and not that other one and has poor verification logic, so if you include the cross chain, verification will fail; you can't make both clients work with the same chain (unless you switch to a different CA). My personal favorite though is when the client helpfully supplies a cached cross chain, but the certificate expired, so it fails validation... even though you didn't send that cert, and you have a path it would validate if it didn't have the cached certificate.
I'm exceedingly glad I don't have to deal with this anymore. But at least TLS1.3 changed the spec to require(i think) clients to handle out of order or extraneous certs which strongly suggests clients should try all possible paths and pass the validation if any are valid. Still no way to present multiple end entity certificates though (end entity must be the first certificate, you couldn't send two full chains)
Either way, the proper solution is "include all the intermediates," which applies to cross-signing scenarios. The client will find the path from leaf to any root it trusts. If you had some kind of crazy number of alternate paths for a given cert, this might become onerous, but in reality, the number of alternate paths is basically always either zero or one.
Consider the case of an older Android client (not from TFA, but bear with me), using certs from Let's Encrypt: https://letsencrypt.org/certificates/
On an older server install, it could be the server was setup with a manual R3->leaf trust, serving (at the time) both copies of R3. This would've been within TLS spec: https://datatracker.ietf.org/doc/html/rfc8446#section-4.4.2 as you are within rights to omit roots:
> a certificate that specifies a trust anchor MAY be omitted from the chain, provided that supported peers are known to possess any omitted certificates.
Both ISRG Root X1 and DST Root CA X3 are well-known roots, and so might be omitted by an operator.
ACME does its thing, but this (admittedly misconfigured) server only updates the leaf cert.
However, original copy of DST->R3 expires. While we have two intermediates in the served chain, because our client relies on the DST trust root, it can't validate it because the cross-signed ISRG Root X1 is missing. But when browsing to a correctly-configured LE-based server (serving the cross-signed ISRG Root X1), this will prime the intermediate cache and allow validation to succeed.
And from the PoV of the old Android based client, this would look like an intermediate is missing (because it doesn't trust the self-signed ISRG Root X1 but does the cross-signed copy).
My 2c, but its entirely plausible (with public CAs) that this could be the case if one of these installations is out of date.
The big difference is https://datatracker.ietf.org/doc/html/rfc8446#section-4.4.2 calling out that the certificate list is unordered, and suggests that clients SHOULD be prepared to handle that in TLS 1.2 as well.
I don't know that there's anything better, and they were certainly convenient, but so many options, so many ways to mess things up.
The one true list of intermediates is "the intermediates between the leaf certificate and the root".
Cross-signing is a corner case.
Technically, TLS 1.2 cannot support cross-signing due to the requirement that each certificate in the chain certify the immediately preceding one (which is obviously not possible with cross-signing, because the cross-signing intermediates break the chain into two paths).
Practically speaking, it works anyway (and TLS 1.3 fixes the problem with the specification) and you should send any intermediates you have and expect the client to construct a chain to a root that it trusts.
It's absolutely not that hard, but my experience is there's a lot of cargo cult engineering around TLS configuration, as well as a degree of deploying-it-until-it-works-for-me.
My personal low was experiencing issues with a server that, upon investigations, was sending 22 certificates including all of the current- and next-generation CAs roots for our internal PKI as well as all of the issuing and intermediate CAs. The TLS handshake was about 200KB on the wire and would fail for some implementations that gave up trying to put together a trusted path.
They should definitely put warnings about these kinds of faults in the dev console, though. I don't think developers that don't get their certs in order will even bother with opening the dev tools on their website, but it's the best place I can think of.
Google and other partnering websites also have a position to prevent these failures. Many websites decide to include Google's stalking scripts and it should be trivial for Google to put scan the domains these scripts run on and warn the website administrators about their bad configuration.
This poses a problem for split horizon intranet websites, but in those cases you'll probably never run into the issue due to the weird setup anyway.
The Google Webmaster Console actually does warn you about all issues regarding your site's ability to be indexed, including SSL issues.