Edit: That blog post does say this: "On May 12, we began a software deployment that introduced a bug that could be triggered by a specific customer configuration under specific circumstances."
The scheduled maintenance on May 12 was this: https://status.fastly.com/incidents/dlsphjqst537
Based on that, it sounds like maybe a configuration change could deploy new cache nodes with ip addresses that a customer hasn't explicitly allowed to talk to their backend:
"When this change is applied, customers may observe additional origin traffic as new cache nodes retrieve content from origin. Please be sure to check that your origin access lists allow the full range of Fastly IP addresses"
I'd go so far as to argue that the specifics of the flaw are immaterial right now. At this stage, the important thing is that they have identified a specific code change that was the proximate cause of the issue, and have a mitigation in place. This is contrasted with more mysterious and hard-to-track-down failures. ("We are working to understand why our systems are down and will post another update in 30 minutes")
What will take time, and the thing which will be interesting, is failure tree analysis. (You might hear the phrase "failure chain" or "root cause" but IMO it's quite rare for things to be so linear). That can help identify opportunities to improve processes at many different levels of the product lifecycle.
Humans are fallible, and there's no way we can write bug-free software, so the solution has to be more robust than "hope that every member of our organization never makes a mistake again"
I can poke plenty of holes in this hypothesis, like fastly likely not deploying configuration to all nodes but only subsets. Looking forward to the deeper post.
I doubt they’d build it on Varnish today, but it’s a bit late now since they allow custom VCL (which has now proven to be a terrible idea) and will have to support that for eternity.
They can run two or more serving stacks side by side though if it comes to that.
And to add to your point, they also have a separate process that speaks QUIC. It’s an interesting tech stack with a lot of technical debt.
Ah, okay. I took a look, and it appears they at least didn't allow varnish modules or inline C. But, still, a fairly hefty anchor for the future.
Yes it does. I recently added WebSockets support to a Varnish instance. See https://varnish-cache.org/docs/trunk/users-guide/vcl-example...
> Fastly is a shared infrastructure. By allowing the use of inline C code, we could potentially give a single user the power to read, write to, or write from everything. As a result, our varnish process (i.e., files on disk, memory of the varnish user's processes) would become unprotected because inline C code opens the potential for users to do things like crash servers, steal data, or run a botnet.
Personally, my hypothesis is that somebody uploaded a configuration for their domain `https://IVCL_{raise(SIGSEGV)}.com` (edit: the preceding URL used to contain a heart emoji between I and VCL, apparently HN prefers ASCII, too) in a way that, rather than converting to Punycode, passed a few bytes that weren't in the 96 legal characters accepted by the VCC compiler and caused some kind of undefined behavior.
[1] https://docs.fastly.com/en/guides/guide-to-vcl#embedding-inl...