One Fastly customer triggered internet meltdown
bbc.co.uk
bbc.co.uk
A quick fix, a clear apology, enough detail to give an idea of what happened, but not so much detail that there might be a mistake they’ll have to clarify or retract. What more are you looking for?
Apart from “not have the bug in the first place” -- and I hope and expect they’ll go into more detail later when they’ve had time for a proper post mortem -- I’d be interested to hear what anyone thinks they could have done better in terms of their immediate firefighting.
They are careful to make clear that the customer did nothing wrong and that the problem was a bug in their software.
Read Fastly's statement. There is nothing about it blaming the customer(s) at all. There is nothing trying to save face.
What is your point here?
Is it necessary to refer to "a customer" at all in this statement? What would be problematic if the above were rewritten as something like:
> Early June 8, a configuration change triggered a bug in our software, which caused 85% of our network to return errors.
The advantage is that you wouldn't get ignorant reporting that "one customer took down the internet". I'm not sure there are disadvantages that net outweigh that.
That’s how autopsies work. You describe the cause and resolution. The cause was a bug in the customers control panel.
They’re not trying to absolve themselves of responsibility.
The important adjective "valid" means it was completely normal/expected input and thus not the fault of the customer.
It's perfectly clear you've come at this with a pre-determined agenda of "I bet fastly, like most other public statements after corporate booboos I've seen, will try to shrug this one off as someone else's fault" after reading the BBCs title and haven't bothered to read it at all until now.
Verbatim from Fastly: https://www.fastly.com/blog/summary-of-june-8-outage
There is no customer blaming here. None at all.
That's literally what happened. They even say it was a valid configuration change, it's very blameless.
Saying "a configuration change" loses critical context. I would have assumed that this was in some sort of deployment update, not something that a customer could trigger. Why would you want less information here?
I'll fully retract my statement. This is 100% the BBC's fault, 0% Fastly's.
Can I make one small suggestion that might help to prevent this kind of misleading reporting in future, though? What if Fastly produce the detailed statement they have, with as much accurate technical detail as possible AND a more general public-facing statement that organisations such as the BBC can use for reporting, that doesn't include such detailed information that can easily be misconstrued?
edit: FWIW I had a very negative initial reaction to the headline as well.
It's not limited to one specific customer (i.e this customer isn't the only customer who could have caused the issue, presumably), but it _was_ something the customer (legitimately) did. It wasn't a server outage. It wasn't a fire. It wasn't a cut cable.
"a customer quite legitimately changing their settings (BBC: one fastly customer) had exposed a bug (BBC: triggered internet meltdown) in a software update issued to customers (fastly admitting, when combined with 'legitimately', that fastly are at fault) in mid-May".
Whatever happened to nuanced opinion, where you can see good and bad in the same entity? Why do some people insist so strongly on absolutes?
Here's some excerpts:
Fastly, the cloud-computing company responsible for the issues, said the bug had been triggered when one of its customers had changed their settings.
Fastly senior engineering executive Nick Rockwell said: "This outage was broad and severe - and we're truly sorry for the impact to our customers and everyone who relies on them."
But a customer quite legitimately changing their settings had exposed a bug in a software update issued to customers in mid-May, causing "85% of our network to return errors", it said.
The headline accurately portrays the story given the limit on headlines.
The wording which was actually used makes it clear that that was not the case.
First real customer walks in and asks where the bathroom is. The bar bursts into flames, killing everyone.
Having done some similar stuff with varnish in the past (ecommerce platform), they’re likely taking changes in the control panel and deploying them to a global config - and someone put something lethal in that somehow passed validation and got published, and did not parse.
But then we still don't know what they fixed, was is the incorrect configuration or the underlying bug? I would expect the former instead of the latter, because it is probably not very difficult or dangerous to change that specific configuration while fixing bugs in the code seems riskier and would probably take more time for testing.
We'll see if they will publish a post-mortem. It has become more or less a normal custom these days (and they are frequently quite interesting).
Once the immediate effects were mitigated, we turned our attention to fixing the bug and communicating with our customers. We created a permanent fix for the bug and began deploying it at 17:25.
So they did both. First reverted the config then later fixed the bug.The problem is that if fastly is the best choice for a company then there's zero incentive for the company to choose another vendor. Everyone acting in their own best interest results in a sub-optimal global outcome.
It's actually one of the major problems with the global, winner-takes-all marketplace that's evolving with the internet.
As a software engineer I live by the ethos that coupling and dependency is bad, but if you unravel the layers you start to realise much of our life is centralised:
Roads, trains, water, electricity, internet
These are quite consolidated and any of these going down would be very disruptive to our lives. Connected software, ie the internet, is still quite new. Being charitable, are these just growing pains in the journey to building out foundational infrastructure?
At home I have emergency power, water and internet. If the trains stop I drive, if the car breaks I take the train.
There is even a competitive advantage in living with the risk, as you have less costs and overhead... Sure, you might have an outage once every x years for a few minutes... But that's obviously the fault of the development team, duh
Data centers offer the highest uptime guarantees at the highest price tiers. People pay more for Toyotas, new or used, because of their reputation. Quality is a product feature. If MBA's want to decide if they can cut corners, there are already upsides and downsides, the calculation is something they need to make.
Quality is a product feature when there is competition, monopolies don't suffer from cutting quality.
I find that people have a tendency to be overly narrow in considering competition and declaring things monopolies. There are alternative ways to get tasks done that avoid relying on (and paying for) low-quality internet services if companies find it necessary.
And we are specifically talking about an industry over-relying on a single provider of a service. If there were a variety of competing services, the entire point would be moot.
> Roads, trains, water, electricity, internet
I guess the difference here is that you're (mostly) talking about physical infra, which by definition must be local to where it's being used. We allow (enforce?) a monopoly on power distribution (and separate distribution from generation) because it doesn't make sense to have every power company run their own lines. But with that monopoly comes regulation.
Digital services are different. The entire value prop is that you can have an infinite number and the marginal cost of "hooking up" a new customer is ~$0. This frequently leads to a natural winner-take-all market.
One way to address this is to add regulation to digital services, saying that they must be up x% of the time or respond to incidents in y minutes or whatever. But another way to address it is to ensure it's easy for new companies to disrupt the incumbents if they are acting poorly. The first still leads to entrenched incumbents who act exactly as poorly as they can get away with. The second actually has a chance of pushing incumbents out, assuming the rules are being enforced. And now you've basically re-discovered the current American antitrust laws.
As far as any individual company's best interests, like anything else in engineering, it's about risk vs. reward.
What's the cost of having a backup CDN (cost of service, cost of extra engineering effort, opportunity cost of building that instead of something else, etc.) vs. the cost of the occasional fastly downtime?
I have to imagine that for most companies the cost of being multi-CDN isn't worth what they lose with a little down time (or four hours of downtime every four years).
This is good reasoning but I don't think it's possible to legislate service level objectives like that.
> But another way to address it is to ensure it's easy for new companies to disrupt the incumbents if they are acting poorly.
I agree but realistically there will be many cases when a company is far better at something than anyone else. I think the only way to avoid global infra single points of failure is competitive bidding and multi-source contracts, plus competitive pressure to force robustness (which already works quite well).
I know plenty of people in Texas who will be buying solar panels and batteries after last winter. I will be doing the same.
> Do you roll your own ISP + telecoms network?
If I could magically get fiber directly to an IX I would gladly be my own ISP. I have confidence I would do as good a job or better than the ISPs I’ve had over the years (yes I realize having hundreds of thousands of customers to service is more difficult than a single home).
I have actually been in the position of having to rely on non-mains power all my life.
It bloody sucks.
Because it seems like a relatively recent development that off-grid solar power solutions have become affordable and mature enough to not suck on average.
... And how is this relevant?
... And how is this relevant?
As to the second half of your original question, solar power is not the only kind of backup power that exists.
Is it better for websites to be unavailable at different times as opposed to all at the same time? This seems to be a really common assumption people make re these occasional cloud take-downs, but I don't really understand why people think it.
Seems to me that in cases like this, everyone operating in their own self interest, by all using the best value service, is actually the best outcome. Everyone suffered the same outage at the same time, which minimised the overall cost of the outage (one resolution, one communication line etc. as opposed to many).
It's the longer term outages that are the problem. That's because we start talking about knock on effects.
It's not really a problem if your supplier (and all others) have a short term issue (assuming you don't run super lean). It may be a headache if your supplier has a longer term issue while you set up another supplier (or use your less desired one) but it's not a disaster. It's a big problem if all suppliers are down for more than a short time.
I'd assume people seeing this as a market failure are talking about it in the "this highlights the problem" kind of way, not the "this event was a true disaster" way.
If a massive vendor shutters or has a long term failure, at least you're in the same boat as a bunch of experts, which is a much better place to be than "my obscure or self-rolled solution is now orphaned / hacked / broken".
The unspoken assumption always seems to be "my self-configured solution will have fewer and/or shorter issues than the massive publicly traded solution that everyone uses" but that seems ... very incorrect.
Also... a reverse proxy / CDN always is a single point of failure. The question is... is it a single point of failure that you personally own. In my opinion shared single points of failure are desirable. It's just obviously more efficient.
Yes. If one site is down, it may hurt my productivity a little bit, and I may have to adjust what I work on. But if the entire internet is down that has drastic impact on my productivity, and depending on what I am working on at the time, may completely block me.
(Slight joke here)
Edit: That blog post does say this: "On May 12, we began a software deployment that introduced a bug that could be triggered by a specific customer configuration under specific circumstances."
The scheduled maintenance on May 12 was this: https://status.fastly.com/incidents/dlsphjqst537
Based on that, it sounds like maybe a configuration change could deploy new cache nodes with ip addresses that a customer hasn't explicitly allowed to talk to their backend:
"When this change is applied, customers may observe additional origin traffic as new cache nodes retrieve content from origin. Please be sure to check that your origin access lists allow the full range of Fastly IP addresses"
I'd go so far as to argue that the specifics of the flaw are immaterial right now. At this stage, the important thing is that they have identified a specific code change that was the proximate cause of the issue, and have a mitigation in place. This is contrasted with more mysterious and hard-to-track-down failures. ("We are working to understand why our systems are down and will post another update in 30 minutes")
What will take time, and the thing which will be interesting, is failure tree analysis. (You might hear the phrase "failure chain" or "root cause" but IMO it's quite rare for things to be so linear). That can help identify opportunities to improve processes at many different levels of the product lifecycle.
Humans are fallible, and there's no way we can write bug-free software, so the solution has to be more robust than "hope that every member of our organization never makes a mistake again"
I can poke plenty of holes in this hypothesis, like fastly likely not deploying configuration to all nodes but only subsets. Looking forward to the deeper post.
I doubt they’d build it on Varnish today, but it’s a bit late now since they allow custom VCL (which has now proven to be a terrible idea) and will have to support that for eternity.
They can run two or more serving stacks side by side though if it comes to that.
And to add to your point, they also have a separate process that speaks QUIC. It’s an interesting tech stack with a lot of technical debt.
Ah, okay. I took a look, and it appears they at least didn't allow varnish modules or inline C. But, still, a fairly hefty anchor for the future.
Yes it does. I recently added WebSockets support to a Varnish instance. See https://varnish-cache.org/docs/trunk/users-guide/vcl-example...
> Fastly is a shared infrastructure. By allowing the use of inline C code, we could potentially give a single user the power to read, write to, or write from everything. As a result, our varnish process (i.e., files on disk, memory of the varnish user's processes) would become unprotected because inline C code opens the potential for users to do things like crash servers, steal data, or run a botnet.
Personally, my hypothesis is that somebody uploaded a configuration for their domain `https://IVCL_{raise(SIGSEGV)}.com` (edit: the preceding URL used to contain a heart emoji between I and VCL, apparently HN prefers ASCII, too) in a way that, rather than converting to Punycode, passed a few bytes that weren't in the 96 legal characters accepted by the VCC compiler and caused some kind of undefined behavior.
[1] https://docs.fastly.com/en/guides/guide-to-vcl#embedding-inl...
If you have a design flaw, you don't hint at onus upon the first person to fall foul of it - and we all know how quick many are reading news that they will run fully upon the title alone (something we can all do).
But we have seen a drive towards click-bait/search-bot friendly to garner hits - style headlines. Even the BBC over the years have IMHO learnt towards such tabloid style headlines more and that is just sad.
If a single bug caused the "internet meltdown", it's fairly likely that the bug was triggered by one person so there's no need to emphasize that part.
Headlines are written by editors, not reporters, to maximize CTR and minimize length.
You can't fit detail and nuance in a headline. The point is to get people to read the article, not inform.
If people are drawing conclusions from reading just the headline, not the article, then you can safely ignore them. There's no point in getting mad about it.
no, the engineers/builders who didn't implement proper safety tolerance broke the bridge.
The straw is not responsible for breaking the camels back.
Why should the BBC talk about a bug in the headline if the source says "valid customer configuration"? They don't write for industries insiders. (Plus that industries is shit to begin with and tries to establish bugs as some kind of force of nature no one can do anything about.)
Humanizing the problem OTOH is (presumably) a narrative that resonates with their audience. Who was this guy? What was he doing? How does he feel about bringing down the Internet? (Is he sorry?) Is he anything like me? Could I accidentally bring down the entire Internet? Should I be worried about that? etc etc.
Not saying modern mass media has no room for technical truth, but I would argue their business model demonstrates over and over again it's this other meta-narrative they're tilting at.
I too wish for a world where headlines aren't terrible, but we currently live in a world of clickbait.
Unless it's an edge use case you're not supporting, don't sell me your cost avoidance on any production systems.
I don't know any dev that would agree with "passing them off as rare to avoid dealing with them" - rather, "it's a low priority, but it needs fixing" or "ok, this is an edge case, but holy crap its a bad one"
This article details a new piece of information, and the headline reflects that.
Customer configuration is an irrelevant detail that should have been left out until a full RCA. What does it matter that it was valid? So an invalid configuration would have meant it was indeed the _customer's_ fault??
A more fair treatment would have been, "a customer pentested us and won".
The way I read it, they were trying to communicate the fact that a customer fiddling with their own configuration brought down large swathes of the internet for everybody else. That absolutely deserves to be in the headline.
Which is why it's clickbait. Sensational title, humdrum article.
I read that phrase everytime something like this happens and yet we all still rely on the same handful of companies.
If we have 12 small cdns have 12 outages in a year (combined, 1/year each), each time bringing down 10,000 websites, is that better than 1 large cdn having one outage during the year bringing down 120,000 websites?
If I'm the website owner I think I prefer the latter, my customers blame the cdn instead of me. If I'm the cdn owner I definitely prefer the latter, more customers to amortize my costs over.
Certainly don't want a situation where all the shops are closed.
It'd be better for the internet as a whole if we don't always pick the most popular (so when your email's CDN goes down you can still communicate on chat, when CNN goes down you can still read BBC). But as an individual I have strong incentives to pick the one everyone else picks, because that's presumably the most stable/documented/lowest cost due to volume.
[0] https://slatestarcodex.com/2014/07/30/meditations-on-moloch/
On the other hand, those handful of companies could be asked to structure their services so that an outage only affects a portion of customers and not all their customers. However! That would be more inefficient for them, and more expensive, and that cost would cause the people and orgs mentioned earlier to just flow towards the company that took those shortcuts.
Some examples:
Social network: you only engage on one if your friends/family/coworkers are on the same network.
Search engine: needs to index "the whole Internet", which is less expensive per user if you have more users
CDN: works best if you have edge nodes everywhere, which is quite capital intensive, which is why you need many customers to distribute it over.
... and so on. We might not like it, but many of these quasi monopolies are based on fundamental economics, not (just) on the greed of the companies.
This one isn't exactly based on fundamental economics given that federated social networks exist. Email has similar network effects and is not centralized.
Which leads me to believe that economics and incentives favor big, centralized social networks.
The trade off here is intrinsic and accepting the risks of big CDNs is the right answer.
We took a technology stack designed to survive nuclear attacks, and turned it into something where a single bug can take down half the services on it. Why? Because on the flip side, a single improvement in the centralized service can automatically cascade to all the businesses using it.
Efficiency is a double-edged sword.
https://www.streamingmediablog.com/2020/05/fastly-amazon-hom...
An example is imgix:
It took them 11 hours to recover from Fastly going down for their claimed 40 minutes.
Does this mean that some companies using Fastly could have major costs because of the increased origin load?
All changes to critical infrastructure should be a gradual rollout (emergencies aside). Instant sounds nice, until it isn't. If this rolled out to one region for the first hour it likely would have been caught and Fastly could press the "stop all rollouts" button.
Big surprise
They frequently have a common caching servers located close by. so maybe every 10 or 100 edge locations have a single cache location.
Edge is a reverse proxy and probably handles ssl handshakes. So if your cache is down, all your edge locations in that area are down
The question in various threads is why not have redundancy -- but the point of a CDN is to have extra servers and capacity and lots of locations to make individual crashes just flow elsewhere.
But if the single customer with a valid-yet-crashable config had lots of traffic all over the world... it'll take everything out at once.
Redundancy of CDN is more expensive, and still requires DNS failover. People do the calculation and usually decide that 30 min of downtime every couple years is worth the saving on vendors and code and hassle. They don't like it, but every site that was down made that decision.
But the only way to avoid that is to give each customer its own private hardware, which seems prohibitively expensive (and may not even prevent all failure sources).
So, I am glad things are OK, but it definitely is not that one customer who is to blame for this outage.
I am sad to report that their clickbait worked on me.
Umbasa!
Now, building the functionality you require can still be unrealistic to build yourself. One thing that springs to mind is DDoS protection.
1Gbit is nothing nowadays and can be saturated in seconds; my purposes do not justify the cost of 10g transit. Even owning 4U when I only need a VPS is overkill. But I like owning a small dusty cube of internet. So there's that.
Sure, but in a lot of cases, that "someone else" is an entire team of experts in their particular niche that can do a better job of the specific task at hand than I ever can hope to.
Is this always the case? No. Is it sometimes the case? Yes.
The internet's original distributed nature just isn't compatible with the sheer scale of billions of active users.
I hold out hope that the next version won't be controlled by corporate entities.
Hope this one gets the treatment.
I wonder here, are they sorry enough to run the company on two different tech and software stacks and data centers? Like people don't buy disk drives (SSDs) not all from the same vendor? How much would they spend for the "sorry".
Reading studies about bugs in the past where developers where told to write some code resulted in different bugs.
It seems like it will be very hard to justify the immense costs for this.
I know a second stack hopefully has different bugs. But is that likely (to a meaningful extent)? Reimplementations often reintroduce old issues (which suggests many people make similar mistakes), plus it’s hard to imagine the first stack not influencing the second in various ways.
How do you go about different software though, have one center running a version or two behind to fallover to?