How Meta patches Linux at hyperscale
thenewstack.io
thenewstack.io
https://blogs.oracle.com/virtualization/post/ksplice-zero-do...
I think the biggest "big company" blunder while I was there was a public blog post by someone high up at the company decrying open source as bad/wrong while at the same time, Oracle was doing all this other open source stuff.
https://docs.oracle.com/cd/B10463_01/web.904/b10320/apjsvsup...
I believe this was circa 2000, but going on longer than that.
They are also remembered for closing OpenSolaris and shaking people down over the VirtualBox extension pack ( https://www.theregister.com/2019/10/04/oracle_virtualbox_mer... ).
Useful word.
> So, if you’d rather not have downtime with your servers, data centers, and clouds, follow Meta’s example and use live patching. You’ll be glad you did.
Maybe if you're working at Meta's scale it makes sense... But I think most well designed services and applications should be able to get by just fine with a full reboot of any single server. I can't really fathom the complexity of managing millions of servers though.
https://www.usenix.org/conference/osdi23/presentation/grubic
I feel like this should be the opposite... I don't work at Meta scale, but I do work for a CDN with 10s of thousands of servers, and everything we do is based on the idea that some machines will always be going down, because some hardware is going to fail every day just from probability. You have to design everything for failure.
Given that, it shouldn't be hard to take servers out of production for patching and updates.
In other words, a hyperscaler is going to have less incentive to minimize down time than smaller shops.
But you got a much simpler process.
Process ain't free either.
If you each sever spends 7.5 minutes each month rebooting, that is 1% of all your computers wasted. If you have 10,000 servers, that’s worth 100 severs. If you have 1 million servers, that’s worth 10,000 servers. If each server costs $10,000, that’s $100 million dollars of compute capacity. You can see how that amount of lost computer capacity can start to justify spending engineering time on driving down the amount of time servers are rebooting.
7.5min/43,200min = 0.00017361
Where are you getting 1% of all your computers wasted per month???
So the correct numbers would be 7.5 minutes of downtime divided by 43,200 minutes times 1,000,000 servers. That’s 173 servers wasted. That is probably still enough servers wasted to devote some engineering time to increasing utilization.
There's a big, big difference, especially at the scale of hundreds of thousands / millions of servers, between designing such that your architecture can suffer 1% of servers being offline and 10% of servers being offline. If you have 1 million servers, even if you could take 1% of them offline at once (i.e. 10,000 servers), if it takes 5 minutes to reboot, you then need to wait 5 minutes * 100 one-percent-buckets = 500 minutes, or 8.3 hours to do a full patch. When you have critical security updates (like Heartbleed) you simply cannot have unpatched servers exposed to the Internet for that much time. And that's not including the amount of time it takes to actually send reboot/patch commands to 10,000 servers.
The larger the bucket, the more likely a bad patch is noticed by the public, and the more likely that an ordinary traffic spike (for which your extra capacity is ordinarily there for) will overwhelm your servers (since your extra capacity is being used to handle patch rollout and rebooting). Sure, you can plan and add even more capacity to compensate, which makes it take even longer to rollout a patch, and now Finance is knocking at your door wondering if you really need all these servers and maybe you could decommission some of them to save money.
It's a fundamentally difficult problem at hyperscaler scale.
AWS has millions of servers in a single AZ.
Running CentOS 8 Stream currently, going to 9 soonish.
I'd stop right there and fix that, because that's a bullshit reason. Cycling hosts in and out of service is easy unless you're not doing things properly.
The Linux kernel is simply not designed to be live patched and it's a total hack to try to do it, it will never work 100% of the time, always be a source of uncertainty, and always be expensive in terms of engineering work. Disaster will always be looming.
By contrast, fixing their system for taking hosts in and out of service, so that it's extremely robust and reliable would likely pay big dividends in reliability.
My guess would be that this approach is papering over organizational dysfunction. One team can patch all the kernels but one team can't make all the hosts support proper cycling in and out of service. And no one cares to fix it because there's no real incentive to do so. Only cool hacks and new projects are properly rewarded.
So what? At large enough scale organization problems are harder than technical ones: if you can fix the former with the latter, that's still a win.
Framing this like it's a good thing is my only objection.
Nothing will work 100% of the time. If their patching mechanism is thoroughly tested and battle hardened, I think the risk would be acceptable. Once you do the initial kpatch security upgrade, you could even schedule the machine for serivce so that it's not relying on that, limiting your exposure to bugs.
There's a lot of systems where you can easily take down some hosts, but taking down more than N% at a time causes issues. If your fleet is large enough then you are limited by the largest set of hosts where you can only take N% down at a time. Now you could say keep the sets of hosts small or N% large. But that can cause other issues as you typically lose efficiency or zonal outage protection.
A solution to this could be VM live migration or something similar. This breaks down for storage systems where you can't just migrate those disks virtually since they're physical disks or places that don't use VMs.
I agree that it’s the right thing to do but it’s hard.
Throwing FUD and saying disaster is looming because its scary computer magic (out of MIT) was a scare tactic RedHat used to throw around about Oracle/Ksplice until they developed their own (Kpatch), then suddenly their sales team had to backtrack and say actually hot patching is good and can be trusted. I'm not saying it's not risky or dangerous, it's operating in kernel space, but that's why they pay really smart people to be careful when doing it, and not digital equivalent of a plumber who can't do more than glue libraries together.
A better understanding of the underlying technology so it's less magic might assuage your fear of it, but thinking Facebook is so dysfunctional that they haven't already made it easier to reboot is to misunderstand the problem at hand.
Facebook hosts can be robustly cycled. Of course. They've been doing this stuff for years. They've figured it out. That's not the issue.
Scaling up brings about new problems. This article specifically mentions the 45 day rolling restart issue. That's not an issue when you have 1000s of servers. It's one that shows up a couple of orders of magnitude later.
So you either solve the problem with a hack like kernel patching or you work to reduce restart times (drain + shutdown + OS restart + process initialization across every service). Get those restart times down 50% (good luck accomplishing that) and congrats, you're down to maybe a 25 day rolling restart, which is still quite a problem.
I wouldn't expect it to be halved by optimization. I'd expect it to be an order of magnitude faster and take more like 4.5 days.
I wouldn't be surprised (but I would be impressed) if they tried and got it down to a full cycle requiring one working day. That's around 1% of hosts cycling every five minutes.
I'm betting it's moreso that teaching developers to write software that tolerates draining properly (or is even able to communicate draining) is too difficult for them so they work around it.
How many individual teams had software running on your hosts? How many those hosts were stateful, and were fragmented across hundreds or thousands of service groups that had their own fault tolerances and unknown (to infra team) warm-up times. Adding complexity (rolling reboots) to already complex systems is almost never a good idea - at some point, there will be an issue caused by hosts rebooted in the wrong order, or too many hosts of a certain type 2-dependency-levels down being simultaneously offline
> hosts rebooted in the wrong order
Order doesn't matter. Host groups set a threshold for unavailability. Hosts are not rolled unless availability targets are maintainable. Usually this just means the oldest host at any time will get rolled.
Facebook obviously uses hot patching kernel updates to work around a social issue. Instead if you are functionally able to prescribe a set of behaviors that teams must comply to, you can easily do things like rebooting the fleet monthly without impacting availability regardless of the statefulness or fault tolerances.
If I shoot a random host in your pool and it matters to you then you haven't achieved fault tolerance. I'm obviously not proposing shooting an unfair number of hosts to you.
That was not my intention - I genuinely would have appreciated answers to my questions as it be useful to compare the complexity of your setup versus Facebook. As an extreme case: million homogenous, stateless hosts are far less complex to manage compared a million heterogeneous, stateful ones, and very little translates from the former to the latter - in my experience.
> Facebook obviously uses hot patching kernel updates to work around a social issue.
Which I think is reasonable when you have tens of thousands of SDEs.
> I'm obviously not proposing shooting an unfair number of hosts to you.
I agree with you, but I'll go on to say "not shooting an unfair number of hosts" is a hard problem to solve at scale, unless you're willing to make it simple and make humans deal with it by continually draining/undraining services which costs a lot of money without increasing the top line, likely far more money than it cost to get a handful of engineers to write kernel splicing. So beyond it being possibly a social issue, it may be a cost/host utilization issue as well
Even if you can splice there are benefits to limiting uptime, with maintenance reaping the majority.
That's a very realistic/optimistic number (especially as you do want to wait for all services to be running and marked as healthy)
"Oh but you can batch this" sure, but you don't want too much of a big batch that will make your service slow or want to risk shooting yourself in the foot - like rebooting your whole control plane then figuring out it doesn't work like that
(The 45 days is probably an estimate as well, I'm not sure they actually do that server by server)
Taking 45 days is probably more about caution and resolving issues systematically rather than pushing a big button and hoping you don’t cause issues.
I’d expect them to have thousands of microservices - and you only have to find a way to break one to cause big issues.
A reasonable compromise in the real world.
Most orgs don’t need and won’t benefit from emulating Meta for the sake of emulating Meta.
SO. Much. This.
Worked at more than a few places where the stack had almost as many layers as it had engineers "because this is how we did it at FAANG..."
Right, and those places also had a few orders of magnitude more engineers on staff to support it all. We do one hundredth of the things FAANG does and we have less than 50 _total_ people in the company; the simpler the stack, the better.
If the infrastructure exists within your org's distribution of choice to do this, it's basically all upside. On AL2023, you just do:
`sudo dnf install -y kpatch-dnf kpatch-runtime`
`sudo dnf kernel-livepatch -y auto`
`sudo systemctl enable --now kpatch.service`
Super simple, one less thing to worry about.
We got bitten by kpatch a few times before decided to abandon it around 2017.
From kpatch github (2023):
WARNING: Use with caution! Kernel crashes, spontaneous reboots, and data loss may occur!
(I'm joking)
So yeah just scaling. I agree I've never heard the word "hyperscale" before and don't think we need that extra intensifier for a well-understood idea.
1000x what? Today's computers are a 1000x the ones from the 90s, should we call them all hypercomputers? Pretty much any startup can boot 20,000 nodes on aws, are they all hyperstartups hyperscaling?
Seems like marketing nonsense.
So you think we need 4 different words for scaling 1-10, 10-100, 100-1000....?
Cut it out, you know it's just marketing hype. Not everything needs its own word.
Why compare to the past capabilities, wtf?
>Pretty much any startup can boot 20,000 nodes on aws, are they all hyperstartups hyperscaling?
Now think how many nodes can Google, Microsoft, Fb, etc. run in their 10s of datacenters.
It’s not 10,000x harder to patch that 10,000x more machines, but it’s not 1x either. Easily 10-20x harder, if not more.
That's not scale, that's organizational sprawl.
In sci-fi a spaceship travelling at "hyperspeed" was perceived to be so much faster than anything known to man that it would be difficult to comprehend.
The performance and cost of a hypercar compared to the average family car is sometimes difficult to understand too. The average person would have to work (potentially) hundreds of years to afford a €10M hypercar.
"Hyperscale" is so large that even us working in tech have difficulty grasping it because we have nothing tangible to compare it to. A million servers is bind boggling to me even with the "cattle not pets" mindset