Network Instability in NYC2 on July 29, 2014
digitalocean.com
digitalocean.com
Moderately interesting that they run everything off a single agg pair. Also that they use mlag instead of routing/mpls/etc for availability.
Key finding is lack of visibility in to the layer 2 availability and performance. Would be interestin to see if they try to ecmp layer 3 or use existing lacp frames for fault detection in the future.
There's much less to go wrong with ECMP at L3. Stateful networking components frighten me.
Though the recap is good, there are red flags that show gaps in DO processes and incident management.
Upgrade of software before the original problem was diagnosed and resolved. This is a big no-no, never introduce a new variable in an existing problem even if the service provider or lab testing shows the chances of failure are minimal. I have worked long enough with technology infrastructure to experience situations where service provider insisted on upgrading software during unrelated incident and made situation worse.
A better approach would have been to fail the network to good switch first and when the bad switch was fixed, upgrade the software on bad switch first then failover to the upgraded switch and upgrade software on good switch. This time you guys got lucky but sooner or later your luck will run out.
1. We had experienced bugs with the currently running release which we were fairly sure would manifest when we removed the damaged core from the network. These were primarily around MAC learning.
2. We had performed testing ourselves in a non-production environment and were already planning to take a network maintenance to update these switches in the next few days.
Given those factors, we judged that introducing another variable in the new version was less risky than proceeding with the defects that we knew about in the existing version.
Every host has an occasional problem. It's how they handle the problem that is important to me. This is a night and day contrast compared to the service I received with other hosts with which I've dealt.
1. Two or three things go wrong at once. Some of the problems are spontaneous, and some were always broken but it took the other problem(s) to uncover it. (Here, the SSD failure was presumably spontaneous, but the "not completely successful" routing engine failover sounds like a case of rarely-used-hence-poorly-tested.)
2. A systemic failure hits all redundant components at once (DDOS, fat-fingered global configuration change, calendar bug, etc.)
This is essentially the risk you face when dealing with new providers.
I guarantee that other providers had to go thought the same set of issues and phases prior to achieving a truly redundant infrastructure.
Would love to hear a follow up on the audit.
Edit: Misread a part of the post. Thought they were doing fail-over on a different network level.
Corrected my post. For some reason I thought failure point was outside of the actual switch appliance.
"We are working very closely with our networking partner to understand the nature of the failure, assess the chances of a repeat event, and to begin planning architectural changes for the future."
"Our initial focus was on verifying the configuration so we initiated a line-by-line configuration review by engineers at our network partner"
Yikes. You get what you pay for.
At our scale, partnering with our vendors is a necessity. We often operate in areas that, while they are technically within specification, are toward the higher end of the range.
I'm not saying an overlay (e.g. VMware NSX née Nicira) is the solution either. I'd suggest rethinking why you need any VLAN anywhere, which is likely the reason for VPC/MLAG. If I'm wrong and you don't put the same VLAN in multiple cabinets, then rip VPC/MLAG out immediately and switch to a pure L3 design. It is so much simpler and consistent with industry best practices.