CenturyLink 911 outage was caused by a single network card sending bad packets
twitter.com
twitter.com
This also happened to coincide with them receiving a notice of a $50/month increase on their bill from CenturyLink, as well as my SO giving them a Roku for the holidays and setting them up with streaming services, which they got lots of practice using in the two whole days the cable TV didn't work.
Guess who's cutting the cord and switching to Internet-only service from Verizon FIOS at a much cheaper rate now?
This is what happens when you treat IT like a cost center and don't provide the necessary funds to tackle technical debt and keep your services up and running: You get huge costly outages revealing your basic incompetence and customers fleeing to superior competitors forever.
The major downside is that their IPv6 support is nonexistent and they have no plans for rolling it out.
Yeah, it's congestion on the transmitter (upstream) end.
(One big problem I've experienced in the past is that any time you configure both IPv6 and IPv4, all programs prefer IPv6. But tunneled IPv6 has much worse performance than the native IPv4, so you really want applications to prefer IPv4 in a tunneled v6 environment.)
Also, I doubt HE's ipv6 tunnel is going to carry anything close to 1Gbps for me for free ;-).
> Typical is 500-600 Down
500/941 vs 500/1000 is 50% vs 53% of advertised. The 3% isn't really significant. I don't think it's bad service for the price, but it certainly isn't true 1gigE.
> but often can reach 900 Mbps up. Typical is 500-600 Down, 900 Up; the highest I've seen down on e.g. fast.com is 900 Mbps
Typical observed speeds down, which is what everyone cares about for residentical internet, are 500-600 Mbps — far below 1Gbps. I mentioned the outliers for color but they're either not useful to me (I don't need 900 Mbps upload) or not representative of real service delivered (I've only observed 900 Mbps once; as I said, 500-600 is much more typical). Off-peak is typically around 700 Mbps down.
So I still don't know what comment or point you were trying to make by claiming GigE TCP can only reach 940 Mbps. It clearly isn't a counterargument to my claim that Centurylink's $80/mo, residential "gigabit" fiber offering does not deliver 1Gbps internet service (nor would I expect it to).
I'd also consider that segment "comprimised" anyway, from a security perspective anything past my modem could be MITM'ing my connecting.
I would guess the vast majority of their revenue comes from extorting old people and farmers with shitty overpriced ADSL that squirrels eat into and take down annually.
I cut cable TV years ago and exclusively stream. I like having the TV on for background noise when home, even though I’m usually not actively watching it the whole time, and you’d be surprised how fast one can chew through data usage. I usually use around 1.5TB a month when you factor in offsite backups.
I can get 400/35 Fiber to the Home with no datacap (no throttling either) for £54 a month (~$70).
UK average fixed-line download speeds in Q4 2017 were about 26 Mbps [1]. The same Q2/Q3 2018 statistic for the United States was over 96 [2].
American broadband is crappy. But in mean technological leadership, it’s ahead of the UK. (At the leading edge, I get 400/35 for $80 in Manhattan.)
Last month apparently I hit a new high for me: 890GB. Lots of devices needing updates/games downloaded...
I am really surprised that with a 7 person household you don't go through your data even faster.
I even run a local caching server for all my Apple devices, which over the last 30 days has saved me about ~50GB from having to be served from origin, with about ~100GB being served to clients.
Although I tend to also be a heavy user of Apple's iTunes and particularly their movies service, so I have a feeling a lot of that cached data is me putting movies on in the background and them being streamed from the cache rather than from Apple directly.
It is surprising how big docker containers get after a while.
https://support.cableone.net/hc/en-us/articles/115009159707-...
Multi-day outages tend to be indicative of a dysfunctional internal organization.
Especially now with cordcutting becoming the norm and traditional cable subscription number dropping, large providers are scrambling to find ways to extract profit.
With the rise of streaming, peering arrangements are more lopsided than ever and ISPs are actually having to pay for their own internet access.
Because of this we're going to start seeing a lot more data caps, neglecting basic network infrastructure, and other tactics to extract as much profit as possible from customers.
Don’t most ISPs cache this sort of content closer to their customers? Netflix at least offers to install hardware that does this.
Cost-cutting didn't seem to have played a role in this outage. They had a bad linecard in their out of band network that was somehow impacting the control plane of their devices. You can talk the pros/cons of how they built out their national Out-Of-Band (OOB) network, but you can't blame this outage on cost cutting.
You're simply bashing a large company that had a bad day. They might deserve your vitriol for plenty of other reasons in other cases, but you can't blame this outage on cost cutting.
Any time something really bad happens that could have been prevented, it's reasonable to ask why it wasn't prevented. This wasn't some black swan event, just a routine hardware failure that they didn't have defenses against.
To be perfectly honest, I don't know what's the best choice for running a management out-of-band network for a network provider of this size. Do you?
I think it's very reasonable to suspect underinvestment and technical debt here. Not only did they have a problem, they had a very hard time a) finding the cause of the problem, and b) mitigating the problem while they were looking for the cause.
We can't know, of course. But I think it would be hard to argue that this is optimal, and that nobody at CenturyLink ever suggested it could be better if they spent a bit more money.
Look, I'm not defending them or their choices. I'm just saying that this more more nuanced than "They suck they should have known better or spent more money." When you run a network as large as this and with as many "legacy" technologies as they do, things aren't very cut and dry.
NTT probably has the best managed global network with their extensive use of SDN. They run a tight ship and everything goes through their automation frameworks. This requires a huge investment in R&D. Even then, it doesn't cover 100% of their network because "there's that T320 for that customer in Chicago, and the stuff up in Michigan for Dorian's T1." Multiply that times 30 or so to consider size differences between NTT and CenturyLink....and yeah, stuff happens.
But your last paragraph specifically describes technical debt as part of the problem. If even the best org (NTT) has technical debt and CenturyLink isn't the best org, then I think it's safe to suspect that a) CenturyLink has significant technical debt, and b) they have underinvested in that technical debt compared with industry best practices.
The issue was a single bad port on a switch about 3 hours away. It took two days of downtime for them to figure it out.
Now the place I work at just deals with having fibre cut 2-3 hours away and losing a days worth of work for it. At least now I don't have to rely on insiders to tell me what actually is going on. But Centurylink does offer a 4g service for backup connections....
I'm not surprised at all with this. Business as usual for Centurylink.
> investigations into the logs, including packet captures, was occurring in tandem, which ultimately identified a suspected card issue in Denver, CO. Field Operations were dispatched to remove the card. Once removed, it did not appear there had been significant improvement; however, the logs were further scrutinized .. to identify that the source packet did originate from this card.
> Support shifted focus to the application of strategic polling filters along with the continued efforts to remove the secondary communication channels between select nodes.
And then
> By 2:30 GMT on December 29, it was confirmed that the impacted IP, Voice, and Ethernet Access services were once again operational. Point-to-point Transport Waves as well as Ethernet Private Lines were still experiencing issues as multiple Optical Carrier Groups (OCG) were still out of service.
And finally
> The CenturyLink network is not at risk of reoccurrence due to the placement of the poling filters and the removal of the secondary communication routes between select nodes.
Looks like the root cause analysis has a way to go. Addendum says:
> The CenturyLink network continued to rebroadcast the invalid packets through the redundant (secondary) communication routes.. These invalid frame packets did not have a source, destination, or expiration and were cleared out of the network via the application of the polling filters and removal of the secondary communication paths between specific nodes. The management card has been sent to the equipment vendor where extensive forensic analysis will occur regarding the underlying cause, how the packets were introduced in this particular manner. The card has not been replaced and will not be until the vendor review is supplied. There is no increased network risk with leaving it unseated. At this time, there is no indication that there was maintenance work on the card, software, or adjacent equipment. The CenturyLink network is not at risk of reoccurrence due to the placement of the poling filters and the removal of the secondary communication routes between select nodes.
Talk about a needle in a hay stack.
We actually ended up replacing them on 2 racks with 24port unmanaged gbit 3com ones we bought from some high street electronics shop one night during an outage... Not a fun memory..
(Of course, those were later replaced with some proper equipment)
I’ve had invalid spanning tree (protocol typically used to prevent loops in networks) packets cause a trunk link to flap as the invalid packet made the switch think its only trunk link (the link to the rest of the network) was part of a loop and shut it down. When the link went down, it could no longer get the bad packets, so after a delay it would enable the trunk again and get the invalid packets again.
I have a very low opinion of STP. I'm also not a fan of the fad to have very flat networks where there are very few routers and instead everything is switches in one gigantic subnet. Packet storms are notoriously difficult to track down on big networks.
That's not a very radical opinion in network engineering circles. No one ever liked it due to it's non forwarding links for loop prevention but it was good enough to work until we discovered it's successor.
Which is?
Impossible, unfortunately, or we wouldn't need STP.
> and moving all redundancy to layer 3 and routing protocols.
Not always possible either, sadly. Works well for new deployments, can difficult to retrofit into older deployments depending on the scale and applications involved.
There are lots of proximate reasons that a failure can avoid all of your sanity checks, but the simplest is when your organization is a little insane. I try to steer the conversation toward avoiding surprises, other times toward ergonomics (don’t rely so much on humans in the moment to do the right thing).
Often people don’t fight me too much on that, but sometimes it’s a near thing. Lots of senior devs are senior because they keep everybody else down, so blaming human error for an outage seems perfectly reasonable to them.
Agreed. I think it’s a case for running a Joy Driven Development operation first, and looking to organizational responsibility first (versus looking for somewhere to cast blame), but practically, that probably too utopian to bear out in practice. No reason not to keep it in mind, though.
lookupByCity(...) {
....
try {
conn = connectionPool.getConnection();
stmt = conn.createStatement();
...
} finally {
if (stmt != null) {
stmt.close();
}
if (conn != null) {
conn.close();
}
}
}
close() can throw, and in the circumstance of the outage it did for stmt, leading to the connection not getting closed and eventually the pool being exhausted with every thread blocked waiting for a connection. It's an interesting chain of failures, arguably the presence of such a chain is the real root cause, rather than the unhandled sql exception.The case you mention above might have been prevented by using checked Java exceptions. Our programming languages and tools could be doing a lot more to catch these problems at compile time or make them impossible by language or API design.
As you said it would be nice if the languages could make these situations impossible.
I think its a bit of a chicken and egg problem because some language issues don't come to light until people are using it and it is too late.
This is one area Rust really excels - the same mechanism for making sure memory gets cleaned up also automatically closes network sockets and file descriptors when they go out of scope. Even in the case of errors it’s impossible to forget to clean up. That entire finally block is unnecessary in rust.
I'm learning Rust coming from Go. It looks cool, but it also concerns me how most data structures in the stdlib use unsafe blocks to defeat the borrow checker. This is not the point of Rust, I would have thought?!
Beyond that, to some degree, it is the point of Rust: limit unsafe things so that you can reason about them more easily. The CPU is inherently not safe, so it has to exist on some level. rust gives you tools to manage this.
This is really interesting and something which bugs me about root cause analysis and it's a neat coincidence that this has been quoted relative to an aviation incident.
In aviation, incidents and accidents are investigated with the understanding that there is never a single cause of an accident. It's known as the swiss cheese model. All the holes in the swiss cheese have to line up for something to go wrong. Even in a seemingly simple "pilot error" accident, there are years of initial and recurrent training factors, ergonomic and human factors and so on which all lead to the event. It's exceedingly rare for a single "root cause" to be the whole story.
Medicine is starting to adopt techniques learned from aviation like checklists, crew resource management and no-blame, swiss-cheese accident investigations. I am hopeful that the software industry will take similar lessons over the next decade or so.
The software industry that programs space craft?
The software industry is vast and not every system involves copious amounts of human decision making. Often the idea of the root system cause, and the root process cause(software construction, operations, etc) cause is separable.
I would say that aviation is almost inverted in the that regard compared to booking systems, banking systems, and most of what software engineers are exposed to. A person can not fly from Dallas to Chicago without many, many human decisions being involved. However, a packet traveling from Dallas to Chicago involves nearly zero new human interactions.
No mention of fixing the design flaw in the system that allows a single piece of malfunctioning hardware to knock out 911 service for millions of users for two days.
Gotta love this.
> No mention of fixing the design flaw in the system that allows a single piece of malfunctioning hardware to knock out 911 service for millions of users for two days.
Yep, they put a bandage on it and called it good.
Just wait until writes Bloomberg that this was because China had put a secret microchip on a network card. Months of entertainment to follow.
Hopefully they're also tracking the design flaws, and yes that's worth following (and asking whether they're planning to do so?), but bear in mind people have limited time and resources, so don't be too hard on them (or they'll be less willing to help and investigate in future).
yes, clearly knocking out 911 service for millions of people isn't a problem. won't someone think of the poor programmers??
This is of course gibberish. A "frame" is ethernet or L2 concept, packets are transport layer. Using the term "Frame packets" in an official RFO is laughable.
A NIC on their management subnet disrupted their entire network? There are so many levels of absurdity to this.
A mangled ethernet frame would be dropped if the CRC was incorrect. A "show int" on a switch would have shown drop counters incrementing. If it was a broadcast storm it also should have been obvious which device was sending an outsized amount of traffic to the all 1's address. Management networks are generally low traffic - ssh and some SNMP. It would have should have been obvious looking at interface graphs by TX on the management network.
Further any modern switch from a major vendor has a storm control setting which disables a port when it goes beyond a certain threshold for either broadcast, multicast or unicast. Even if storm control wasn't enabled it would have been trivial to do so, find the offending port and work backwards from there.
>"A polling filter was applied to adjust the way packets were received in the network equipment"
"Polling filter" is not even an idiomatic network engineering term. I'll assume this means an access list. So it took them 50 hours to apply an ACL? And this required engaging the hardware vendor?
This is a garbage RFO even if its not meant for a technical audience. It sounds like the real RFO is due to incompetence, bad network design and probably a horrid corporate culture shaped by fear, silos and CYA at this company.
OTU also has frames. Optical gear is generally happy to pass along mangled packets :/
> any modern switch from a major vendor has a storm control setting
Look at the switches embedded in optical transport gear. They are pretty rudimentary.
Optical transport gear (L1 networks) are full of impressively clowny behavior.
Yes but STS "frames" and OTN are layer 1 concerns. There would still never be "frame packets." It's just as egregious.
Also do you believe anyone would use DWDM for their management network? Management interfaces seldom require anything more than a few megabits of bandwidth. Burning an entire wavelength for a management network would be pretty crazy.
In the RFO Centurylink also mentions - "A decision was made to isolate a device in San Antonio, TX from the network as it seemed to be broadcasting traffic and consuming capacity." Lightwave gear most certainly does have any concept of broadcasts.
Yes, in the form of the Optical Supervisory Channel (OSC), which is built into DWDM gear and generally implemented as Ethernet over SONET.
The OSC can also carry management traffic for other devices (aka datawire).
It's Ethernet, so it has broadcasts...
In this sense its no different than how a copper ethernet management VLAN should not be able to take down your entire production network.
Optical control plane generally hasn't benefited from the hardening that's happened in the IP world.
Things like CoPP haven't become common practice yet.
OK, but I imagine we can probably both agree that proper network design is orthogonal to the pace of development in optical transmission gear ;)
Don't "bad packets" get dropped at the first switch? Isn't that one of the main benefits of packet based switching?
Was this even an ethernet packet or something else like an optical transport protocol (eg OTN)?
Since checksums are hardware accelerated, the invalid packet probably had a valid checksum applied to it.
Makes sense to have it at the end.
I'd love to see an in-depth technical analysis of the outage.
Don't expect a technical report from CL/L3. We had 60+ mpls/vpls circuits from them and all our reports were very high level.
Source: am neteng
Summary: the specific Symantec disk imaging software was partly loaded via PXE boot, and that machine started to flood the network with bad packets. Switching the computer off for a few seconds didn't help, since the SMPS capacitors still held current - enough to keep the card alive and for the sysadmins to not suspect that computer!
If the whole thing is a single flat logical network (one that could allow bad packets to propagate as we witnessed) that would suggest it is also quite vulnerable to malicious actions.
It is all well and good applying a filter, but that seems like a bandaid fix. Why is equipment even able to talk that has no reason to do so? Seems like they've put convenience over good network governance.
https://catless.ncl.ac.uk/risks/9.62.html#subj2 https://catless.ncl.ac.uk/risks/9.63.html#subj3
Also see datagram, cell and probably others I’m forgetting right now.
No, use frame when your talking about layer 2, packet when your talking about layer 3, and segment for layer 4. Datagram and protocol data unit (PDU) are general terms that can apply to any layer.
A switch forwards frames, while a router routes packets.
https://stackoverflow.com/questions/31446777/difference-betw...
If your network isn't properly configured these things can happen easily.
Absolutely true. However, if you are an ISP, then not correctly configuring your network is... unimpressive.
I have tracked that kind of thing down before. "Line noise adapters", otherwise known as former NICs, can be a pita.
But taking down the whole service, or for a pretty big region?
I am off to read the details!
Or that said infrastructure may exist, but not redundant... I mean, RS232-over-IP or RS232-over-ISDN boxes are no secret sauce, but when their access line is routed over the same thing the box is supposed to remote-manage, then one has problems.
a. There was no error monitoring.
b. That a SPoF existed.
c. That it wasn't found sooner.
The FCC, with their Verizon lackey Ajit Pai, should fine them $100 million bucks to get their attention, but they won't because corporate welfare.