Here's a simple example:
Split the user population into 365* groups. Assign them to days of the year. Test new code only the group whose day has come up. Follow them until they stop having problems. Now you can deploy that to a month's worth of groups. All good? Deploy to everyone.
Yes, that means that you can't have more than 365 changes in simultaneous development. Tough.
*Yes, yes, leap years. Take a day off from deploying.
They have to investigate it, revert or fix the bad code and start the deployment process again.
Only that it's not 365 groups, because at the size of FB that would be several million people.
The problem isn't that it's impossible -- it's that it's more expensive than just hiring one more engineer to keep papering over the problems.
All of them. Are your roads more reliable when they're constantly being changed or when they are just being maintained? Is NASA achieving its reliability by constantly changing the designs of their ships, or by reusing the same design over and over?
Arguably, with formal verification, you could ensure large parts of your system are perfectly reliable given simple assumptions.
Yes. But a fixed formally verified system will still be more reliable than a formally verified system being constantly changed.
What was said wasn't that FB couldn't be more reliable. It's that they are already so reliable that only new changes introduce problems. Sure you can still work on minimizing those problems, but that's a different point.