Fastmail 30 June outage post-mortem
fastmail.com
fastmail.com
I'm also not really charmed with how they try to minimise the importance of the incident by repeating it only affected 3-5% of the customers. That might very well be. But those are real people and real businesses that rely on your services that were unavailable for the whole of the EU workday and a significant part of the US workday. Everyone I know who was affected is a paying customer, none of us have received so much as a communication or apology for it.
For a company that's been on the internet since 1999, the single-homed setup is a little shocking. But fine, it's being addressed. But both the communication during and after the incident don't inspire a ton of confidence.
it's not meant to be charming, it's meant to convey the scale of the outage.
as a customer, even if I'm not effected, there's a world of difference between 3-5% and 99%.
> Thankfully, NYI were willing to lease us those addresses, because IP range reputation is really important in the email world, and those things are hardcoded all over the place — but it has caused us complications due to more complex routing. Over the past year, we’ve been migrating to a new IP range. We’ve been running the two networks concurrently as we build up the trust reputation for the new addresses.
Might be one of those things that they did when they were small, but then got hard to change. Hopefully they will fully own their new addresses.
The thing I'm surprised about is that they only have "single path for traffic out to the internet."
There is a nice tick up in complexity when you go to advertising your address space via BGP to multiple providers.
Edit: After seeing the network diagram I have even more questions. What happens if CF is down? This all seems cobbled together and very prone to failures.
They just added redondancy to in/outbound routes.
Is it x2 or x100 or somewhere in between?
There are a lot of aspects to that, but the cost of doing all of the above is a lot less than not having it and failing to have it at the wrong moment and losing money that way. Each business needs to weigh their risk against how much they want to invest and how much they think they can tolerate in terms of downtime.
I've honestly never had a service with a single outbound path. Most datacenters where you rent colo have two or three providers as part of their network. In the cases where I've had to manage my own networking inside of a datacenter I always pick two providers in case one fails.
> Work is now underway to select a provider for a second transit connection directly into our servers — either via Megaport, or from a service with their own physical presence in 365’s New Jersey datacenter. Once we have this, we will be able to directly control our outbound traffic flow and route around any network with issues.
Having multiple transit options is High Availability 101 level stuff.
That's not the issue. With Cloudflare MagicTransit, packets come in from Cloudflare, and egress normally. They were able to get packets from Cloudflare, but egress wasn't working to all destinations. I wasn't able to communicate with them from my CenturyLink DSL in Seattle, but when I forced a new IP that happened to be in a different /24, because I was seeing some other issues too, the fastmail issues resolved (although timing may be coincidental). Connecting via Verizon and T-Mobile, or a rented server in Seattle also worked. It's kind of a shame they don't provide services with IPv6, because if 5% of IPv4 failed and 5% of IPv6 failed, chances are good that the overall impact to users would be less than 5%, possibly much less, depending on exactly what the underlying issue was (which isn't disclosed); if it was a physical link issue, that's going to affect v4 and v6 traffic that is routed over it, but if it's a BGP announcement issue, those are often separate.
Also calling them an armchair QB? Very mature. Their comment is more correct than yours.
I think the whole point of CF is that it isn't.
Their problem was that they only have two transit providers, and one of them black-holed about 3-5% of the internet. Since it was a routing issue, I'd guess it was either a misconfiguration, or that the traffic is being split across dozens of paths, and one path had a correlated failure.
All the redundancy in the world can’t protect you from some random person digging in the wrong place.
Single or million boxes, if they apply wrong config, it doesn't work.
You dual home because you hedge your bets and hope no 2 ISPs gonna have fuckup at same time
That’s certainly possible in specific cases, but not a very good general principle to rely on. One CF could very well be better than two given ISPs.
A second ISP isn't free, it has significant costs in terms of dollars and complexity. The question is, does CF and another provider have significant benefits to justify the additional costs? For it to make sense you have to believe the redundancy CF provides is significantly lacking (and in a way that adding a second provider addresses). Maybe it's true, but it would be nuts to just assume it and start spending a lot of money.
We pay x10 for power alone
We had not that happen in 10 years
> That’s certainly possible in specific cases, but not a very good general principle to rely on. One CF could very well be better than two given ISPs.
You might think that if you have no idea what are you doing.
The entire internet.
AFAIK it's not like FastMail has a crazy number of network-related outages, so overall it doesn't seem that "prone to failure". As with many things, it's a trade-off with complexity and costs.
Things break. More things will break in the future due to increased complexity, brittle network automation processes and poorly written code. You can mitigate failures to a certain extent, but you can't guarantee 100% uptime, even with a triple redundant system. Every business decision is a compromise among various constraints.
I cannot imagine running a service like that with cobbled together DIA circuits and leased IPs.
The problem here is that there isn't an alternative to Cloudflare.
They say this is the article. None of their DDoS solutions can take the heat except for Cloudflare.
So, if you want resilience in the face of Cloudflare being down, you need to build another Cloudflare. Let me know when you build it. Lots of people will sign up.
What you're talking about there is SPF/DKIM etc, which is an anti-spam measure but not IP reputation :)
https://www.fastmail.help/hc/en-us/articles/1500000278242
Read the section on "slots". They keep two copies in New Jersey, and one in Seattle.
However, based on the post mortem, it sounds like they're not willing to invoke failover at the drop of a hat. They allude to needing complex routing to keep their old good-reputation IP addresses alive. That might have something to do with it. (They were "only" 3-5% down during the outage, which is bad for them, but not unusual by industry standards.)
Also, it’s email. It was literally designed to work in such a way that you can be down for days and still get your email.
Fastmail has multiple offices.
> Also, it’s email. It was literally designed to work in such a way that you can be down for days and still get your email.
Delivery is designed that way, not storage. Once Fastmail tells the sender that a message was delivered to an inbox, Fastmail cannot ask the sender to redeliver it if its inbox storage is lost.
When you are running a service like this, redundancy among transit providers is the most basic, table-stakes thing you can do. It's almost negligent to not have that.
The worst switches I ever used were HPE ones in the old C5000 blade chassis. Absolute turds. Packet loss, constant port failures and complete hangs. HPE’s solution was to tell us to buy new ones.
To be honest, those little blue unmanaged Netgear switches aren’t bad at all. We have dozens of them in our lab at work running 24/7 for like decades and have never had a failure that I remember.
As to the netgear switches, I would figure they'd have hot spares considering the cost savings. I'm not entirely sold on a specific vendor for the end-all-be-all for switching needs, support for most is less than stellar, and the need for hot spares grows every quarter report from the big name vendors as they continue to push for larger margins.
In environments like this, it's less about the vendor specific product, and more about the redundancy setup (which appears they lacked but are transparent regarding it).
Just my .02 cents
edit: didn't realize this was such a "hot take"
Ah yes, security through obscurity.
https://www.bleepingcomputer.com/news/security/fortinet-new-...
When the vendors you're buying from aren't taking security seriously, I suppose you take any necessary step in limiting exposure. I'd also argue that outside of the big boys like Cloudflare, no one else is displaying their topology via their own website.
In an ideal world everyone would share their architecture, stack and so on and we as an industry we could learn between each other and everyone would have a net gain out of this information sharing.
In reality at the time that you share something in good faith you will always have someone trying to exploit it.
One example: I’ve worked in a CV production API to recognise certain documents. More than 900 days with no spikes and only real users in the system.
Then the CTO went to a conference to talk about how our performance was great and made a very large advertisement about our system. End result? 1800% spike, and tons of frauds and adversarial stuff coming.
Not being cynical, but I do not think that we’re entitled to have any disclosure from any private company in that regard.