Building Facebook's Service Encryption Infastructure
code.fb.com
code.fb.com
This isn't even the first time I've heard of an issue at FB being caused by a single bad CPU instruction. Working at a scale where "Problem X is a one-in-a-million edge case" and "Problem X happens several times per day" are synonymous is weird...
1. If you're only 99% automated, you're dead.
2. Anything you can easily imagine going wrong is probably going wrong right now.
3. Everything else that can possibly go wrong will go wrong at some point in the not distant future.
Some of the problems we ran into were pretty fun and challenging. The longest substantial one I'm aware of took a couple of years to figure out.
My team and extended teams managed almost 200,000 network devices, spread all over the world, most of which were Cisco, and most of which were installed in stores. And most of the switch ports were connected to customer facing Point Of Sale devices. Among these are employee facing registers and customer facing card scanners. That is, the devices you interact with whenever scan your card to pay for something.
With that many devices in that many locations in a largely unmanaged environment (the switches would be installed all over the store, often in the ceiling, and many of them experienced extreme temperatures), there were a constant stream of failures. The process to manage these failures was optimized, streamlined and largely automated.
However, it was discovered that switches were failing far more frequently in the northern Midwestern US than elsewhere, and then only in the winter.
So this wasn't a really big operational issue, but it had a substantial cost impact, and the rate was high enough that a lot of the affected stores did notice and were complaining.
Right. Very strange, very mysterious.
So, briefly, the root cause:
Apparently, people in the upper Midwest wear wool to stay warm far more frequently than other cold places, specifically the US northeast. And much of the time, the humidity is quite low. So, you have a lot of people wearing a lot of wool in low humidity air. These people generated a lot of static, which they would all too often discharge while interacting with the customer facing point of sale device. And, all too frequently, that pulse of static would end up flowing all the way back to the switch, often killing it.
I didn't follow the subsequent remediation efforts, so I don't know what if anything was done about that.
I suspect it would not have been received as that big a deal to management or to anyone else. In those halcyon days, we were running into and usually solving all kinds of such edgy, extreme scale problems. It was a lot of work, but a hell of a lot of fun too.
[0] https://softwareengineeringdaily.com/2017/06/16/google-early...
Recently, again at Facebook, we found a few machines with CPUs where ADCX instruction is broken on at least one core. This is especially fatal because OpenSSL and Fizz (our TLS 1.3 impl) use these instructions for AMD64 architecture in RSA implementation.
Infant mortality rate of CPUs--especially single socket designs--is very low compared to DIMMs and SSDs. They tend to develop these issues later in their economic lives. They are also comparatively rare to other module failures.
If you don't mind I wouldn't necessarily agree with the comment about JWT by Yueting. JWT is just a format, querying backend to get a new token is not necessary (this is only how people often use them). I actually built a small PoC that mints new JWTs on client side (in the browser) signing them with a non-exportable key (through Webcrypto).
As for Macaroons I believe they could also be adjusted to resemble CATs as I understood them (with layers for different services). I do have other issues with Macaroons though (https://news.ycombinator.com/item?id=17878845)...
I imagine Google runs a bigger scale operation on top of Borg and their internal service mesh.
Companies at this scale have integrations between all different levels and layers of the 'stack' that make the use of off the shelf software difficult or impossible.
My point was encryption of services is built into K8s/service mesh and wondering how it fares compared to FB's approach.
For example, OSKT is a self-service toolkit that lets users build up access controls for clusters, "role accounts" (user accounts for running application automation), and what not. Users use "krb5_prestash" to indicate what hosts should have what role accounts' credentials (the user must own the hosts and role accounts) and krb5_keytab to get keys for services on hosts they are allowed to run. A nifty trick is to have wildcard DNS A RRs for hosts so that one can have HTTP/${USER}.$(uname -n) principals (and keys for them) on any host the user can login to.
All of this is high-performance and self-service. Users don't need to file JIRA tickets or whatever to get their keys for their services.
Self-service credential provisioning is absolutely essential to successful deployment at scale of any authentication system one uses, whether that be Kerberos or PKIX or DANE or anything else one might find or invent.
[0] https://www.eyrie.org/~eagle/software/wallet/readme.html
[1] https://oskt.secure-endpoints.com/
https://github.com/elric1/Thousands? Does one truly need that many microservices, even if you are the size of Facebook? That is a lot.
I would go a bit further and say most companies should only proceed with a microservice architecture if they have sufficient scale and automation such that decomposing their architecture will result in at least a high double digit number of discrete services.
I agree, scale and automation are two important factors when decomposing architectures. I think it would be really valuable if systems could decompose themselves to some degree, based on scale and other factors, without much of an operator's intervention.
[1] https://medium.com/@michael_87395/benchmarking-istio-linkerd...
[2] https://medium.com/@ihcsim/linkerd-2-0-and-istio-performance...
I'm curious at to what the cause of the latency is. TLS handshakes?
https://github.com/istio/istio.io/pull/4220
More here, which basically suggests, don’t stop Istio from scaling out before 500 rps, it doesn’t like that at all:
https://kinvolk.io/blog/2019/05/performance-benchmark-analys...