A coding error caused Rogers outage that left millions without service
theglobeandmail.com
theglobeandmail.com
In Canada 911 service is accessible without a SIM card; if your phone doesn't have a SIM, any cell tower should in theory still accept and route the call as normal. However in Roger's case, because the cell towers and their authentication mechanisms were still operational, any 911 call with a Roger's SIM card in it would route through Roger's network and only Roger's network; the one that couldn't service any calls. In essence, your 911 service was completely cut off.
The workaround is to pull the SIM card prior to making the 911 call, but it leaves an interesting question about what you're supposed to do in an eSIM world where pulling the SIM is not possible but you're again in this kind of situation.
To be clear I don't know if this is a real problem or not, but it is an interesting thought either way.
ergo, I keep a spare ejector tool on my keychain for this very reason.
Though at this point we should just move to tool free sim slots like with Sony phones.
See "Erase your eSIM": https://support.apple.com/en-us/HT212780
The actual “but” clause is nobody is going to be able to diagnose the problem just from the terminal device. It’s like when our cable service goes down* the kids shout out “WiFi is down” (while to me it’s working fine).
* this happens a lot because I live in Palo Alto which, despite having deployed fibre close to every address in the city back in the 1980s was never able to actually, you know, deploy IP service.
I live in midtown where the city hasn't gotten around to undergrounding the wires. In theory that means it's easier to put in fiber. In reality it means no fiber and cable internet that get flaky when it rains or if the wind blows.
I looked at using the fiber to connect my house directly into the PAIX (I confounded an ISP years ago). PA was happy to rent me dark fibre which I could light up however I liked. They charged by the distance (I was about 1.5 km from the PAIX). IIRC it was going to be about $25K for the cap ex and then 2K/month to use it (plus my equipment). My wife sensibly said no.
>Staffieri says Rogers has made meaningful progress on a formal agreement between carriers to switch 911 calls to one another’s networks automatically, even in the event of an outage on any single carrier’s network. He says the company is physically separating its wireless and internet services to create an “always on” network that meets a higher standard of reliability.
-https://vancouver.citynews.ca/2022/07/24/rogers-competitors-...
This may not work if Rogers is the strongest signal as the handset may latch to it despite its non-functional state.
And that’s for any handset without a SIM, an originally Rogers handset or not.
For example, on iOS:
Settings > Cellular > Network Selection > Switch off Automatic
Update:Grammar.
I'm no expert but I feel like roaming still require some sort of "authentication" over the original network and obviously Rogers would have failed on that part.
Admittedly, this is neither fast nor easy/intuitive. As pointed out in a sibling comment, the fault lies with Apple and Google for not forcing the 911 call on a different network after it fails.
The fix carriers are discussing sounds like finding a way to re-route 911 calls over other networks on the same tower, or some other failover mechanism. It’s not clear if that would be simpler than having a handset try again on a different network.
This is an interbank network which is essentially a cooperative operated by the country's major financial institutions.
Who would have thought Interac would've only had one ISP? Thankfully, they've stated they're adding at least one additional service provider.
I drove from Winnipeg to Calgary that day (a 12 hour drive) and had to stop in several small towns with TDs to pull cash to fill up.
It was quite bizarre checking Twitter all day to see if it was back yet and seeing it was still down.
Disclaimer - I am Canadian but was born in USSR, came here some 30 years ago.
Why? Avoid blame and I do not speak up when my manager leaves me in a bad position.
My priority in many meetings is merely to not get nailed down on something that I can be whacked with later. Better to avoid all accountability, as Canada doesn't really reward doing a good job.
In Toronto and the surrounding cities, for example, about half of the population are foreign-born.
A significant proportion of the population of many other Canadian urban areas are foreign-born.
About 20% of the overall Canadian population are foreign-born.
Many of these people are from cultures that are quite different to anything resembling "traditional" Canadian culture.
Acquiring Canadian citizenship later in life doesn't necessarily change a person's values, attitudes, and so on.
A lot of Canadian-born individuals have one or both parents who are foreign-born, which also can have an impact on one's values, attitudes, and behaviour.
Even among those with multi-generational ties to Canada, there were already significant cultural/values/attitude/behavioural differences among the various groupings.
Ultimately, in any given interaction with somebody in Canada today, there's a good chance you're dealing with somebody whose ties to "traditional" Canadian culture are limited, or even non-existent.
Maybe some kind of a relatively cohesive "Canadian" culture or identity existed at one point, several decades ago, but I don't think that's the case any longer, especially in the urban areas.
I consider myself one even though I am foreign born. But maybe since I do not pour maple syrup on my eggs'n bacon I am a fake.
>"About 20% of the overall Canadian population are foreign-born."
For reasons unknown most of company owners/reps I've done business with were WASPs. Count me "lucky".
Also what you said does not explain why is it so different in the US.
What was most surprising was how long it took some store clerks to flip their point of sale systems into queuing mode so that people could keep buying things. There was a period where we just went from shop to shop trying to buy the same thing until we found one that didn't decline the card.
It was amazing to think that one provider could go out and that would render me completely useless in my own city. I’ve outsourced memory to my phone for so long that I only know places I frequent.
He ended up using a TD. I failed prairie hospitality.
It sounds like their control/management plane (with the user's database) was dependent on their data plane. So a data plane outage was more challenging to mitigate than it should have been in a decoupled architecture. Good lesson for any architecture.
Lots of things went wrong here, to name a few:
- lack of rigour in their process that allowed a critical modification without understanding its effects in context
- lack of monitoring and metrics to alert them immediately of the problem
- lack of emergency rollback capability to revert back to the known-good config
- lack of independent business-continuity comms channels. if you're a Rogers CTO or senior responsible adult and you don't have a backup Telus or Bell SIM card (and vice versa), you learned why you need one.
From their own documentation, in mid-2015 (kinda late, but better than never), the providers issued eachother eachother's SIMs to use in these situations. Who qualified exactly is a mystery.
I'm impressed they went to these lengths to get services up!
[source - 20 years of NetSec and NetEng experience]
And all through the day I kept thinking to myself "I bet someone pushed an update to prod, causing this. And I'm glad that this time it wasn't me."
> But, at 4:43 a.m. on July 8, a piece of code was introduced that deleted a routing filter. In telecom networks, packets of data are guided and directed by devices called routers, and filters prevent those routers from becoming overwhelmed, by limiting the number of possible routes that are presented to them.
> Deleting the filter caused all possible routes to the internet to pass through the routers, resulting in several of the devices exceeding their memory and processing capacities. This caused the core network to shut down.
Lesson no 1: Do not design your system to have a single point of failure.
I mean, if you just pushed a config change and the whole network goes kaput, take a look at the config change before you start suspecting hackers.
This is why some hospitals still use the old pager systems to contact people in the city. One hospital-owned antenna on a battery can coordinate a lot of people. I don’t know what the equivalent to that would be in this case though.
It still works, you know?
Also, pagerduty works over wifi...
even when the internet is down?
I would imagine using it for running a nurse-doc comms would constitute professional use.
Of course I'm sure this does run the line of if this is emergency or not.
May whatever deity you worship help the instructor teaching atmospheric propagation to some of the staff I’ve worked with in the past.
Yes sure, there's repeaters, and other details to know about, but tx/rx on a single frequency isn't exactly the rocket science that sad hams make it out to be.
(Many hospitals do also have ham radio operators as part of the disaster plans and operate during their drills.)
INAL, contact ARRL for farther questions.
They probably did it from home/overseas. Can't check what you did after you dropped the country itself offline.
And the Big Failures are always around this spot in the stack. Things like routing topology control (also top-level DNS configuration, another famous fat-finger lever) are "single points of failure" more or less by definition.
"In the Rogers network, one IP routing manufacturer uses a design that limits the number of routes that are presented by the Distribution Routers to the core routers. The other IP routing vendor relies on controls at its core routers. The impact of these differences in equipment design and protocols are at the heart of the outage that Rogers experienced."
I read this as: we were halfway through a big project to replace 6509s with new Juniper switches but that takes a few years, so in the meantime we sometimes make configuration changes to each router that have different behaviors.
Seems like an innocent and unavoidable mistake.
My opinion: You know how Juniper routers require you to commit a change? That should be enforced by a dedicated firmware and there should be multiple rollback confirmations (optional). So you commit a change and confirm, an hour later you have to confirm again else a hard reset and rollback takes place with a protocol in place where the rollback can be delayed when other devices in connected segments are also doing a rollback.
There is no one thing you could do to avoid catastrophic failure. The best thing you can do is break your system up entirely and simply never do any change which can affect the entire system. Which is still very hard to do.
Imagine you have a route that crosses all of Canada. For each end of Canada you want at least 3 separate networks that can each carry 1/2 of the traffic. On failure of one network, somebody has to take that failed network's traffic and send half to each of the two remaining networks. But if somebody fucks up and sends that traffic to one of the networks, that network will be flooded and inoperable, and the first network also will be inoperable, leaving only one network with 1/3 the country's traffic online. Whatever customers designed their own stuff to depend on the first or second network is also screwed. This is just one of a million different scenarios that needs complex controls to prevent a (mostly) catastrophic failure.
There is nobody using formal proofs to show the design makes sense, so it's just humans winging it and hoping their contingencies work out.
I don't get this. If a switch can handle only x number of routes, it won't get "overwhelmed" if you try to add more, it would just refuse to accept any more routes. Why would the network go down completely? These things are designed with extremely robust software, you simply start getting error logs saying that the device has run out of memory, but the operation continues. Most of the heavy lifting is done in hardware chips, so the software just can't program new routes, but the routing itself should continue unaffected.
It not how it worked in old Cisco routers - when more and more routes are added at some point it runs out of fast memory (TCAM) and becomes either too slow or not operational at all. Unless you manually configure route number limit it would install more routes into FIB than it can handle. So one have to monitor TCAM utilization and based on this configure various limits to prevent its exhaustion.
May be TCAM exhaustion is still a problem even in modern Cisco equipment - I've moved to another field and don't know much about it.
https://www.theglobeandmail.com/business/article-how-a-codin...
The Outage The configuration change deleted a routing filter and allowed for all possible routes to the Internet to pass through the routers. As a result, the routers immediately began propagating abnormally high volumes of routes throughout the core network. Certain network routing equipment became flooded, exceeded their capacity levels and were then unable to route traffic, causing the common core network to stop processing traffic. As a result, the Rogers network lost connectivity to the Internet for all incoming and outgoing traffic for both the wireless and wireline networks for our consumer and business customers.
The Recovery To resolve the outage, the Rogers Network Team assembled in and around our Network Operations Centre (“NOC”) and re-established access to the IP network. They then started the detailed process of determining the source of the outage, leading to identifying the three Distribution Routers as the cause. Once determined, the team then began the process of restarting all the Internet Gateway, Core and Distribution Routers in a controlled manner to establish connectivity to our wireless (including 9-1-1), enterprise and cable networks which deliver voice, video and data connectivity to our customers. Service was slowly restored, starting in the afternoon and continuing over the evening. Although Rogers continued to experience some instability issues over the weekend that did impact some customers, the network had effectively recovered by Friday night.
what was the root cause of the outage (including what processes, procedures or safeguards failed to prevent the outage, such as planned redundancy or patch upgrade validation procedures);
Like many large Telecommunications Services Providers (“TSPs”), Rogers uses a common core network, essentially one IP network infrastructure, that supports all wireless, wireline and enterprise services. The common core is the brain of the network that receives, processes, transmits and connects all Internet, voice, data and TV traffic for our customers.
Again, similar to other TSPs around the world, Rogers uses a mixed vendor core network consisting of IP routing equipment from multiple tier one manufacturers. This is a common industry practice as different manufacturers have different strengths in routing equipment for Internet gateway, core and distribution routing. Specifically, the two IP routing vendors Rogers uses have their own design and approaches to managing routing traffic and to protect their equipment from being overwhelmed. In the Rogers network, one IP routing manufacturer uses a design that limits the number of routes that are presented by the Distribution Routers to the core routers. The other IP routing vendor relies on controls at its core routers. The impact of these differences in equipment design and protocols are at the heart of the outage that Rogers experienced.
The Rogers outage on July 8, 2022, was unprecedented. As discussed in the previous response, it resulted during a routing configuration change to three Distribution Routers in our common core network. Unfortunately, the configuration change deleted a routing filter and allowed for all possible routes to the Internet to be distributed; the routers then propagated abnormally high volumes of routes throughout the core network. Certain network routing equipment became flooded, exceeded their memory and processing capacity and were then unable to route and process traffic, causing the common core network to shut down. As a result, the Rogers network lost connectivity internally and to the Internet for all incoming and outgoing traffic for both the wireless and wireline networks for our consumer and business customers.
how did the outage impact Rogers’ own staff and their ability to determine the cause of the outage and restore services;
At the early stage of the outage, many Rogers’ network employees were impacted and could not connect to our IT and network systems. This impeded initial triage and restoration efforts as teams needed to travel to centralized locations where management network access was established. To complicate matters further, the loss of access to our VPN system to our core network nodes affected our timely ability to begin identifying the trouble and, hence, delayed the restoral efforts.
Despite these hurdles, our preestablished business continuity plans enabled staff to converge at specific rally points. Those equipped with emergency SIMs on alternate carriers that enabled our teams to switch carriers and assist in the initial coordination efforts. Further, we rapidly relocated our employees to two of our main offices in the GTA (# #). The critical network employees were able to gain physical access to our network equipment. Other essential employees were able to use alternate SIM cards, as per our “Alternate Carrier SIM Card Program” (described in Rogers(CRTC)11July2022-1.xiii below). Other employees were able to work from # #. Together, these groups were able to establish the necessary team to identify the cause of the outage and recover the network.
what contingencies, if any, did Rogers have in place to ensure that its staff could communicate with each other particularly in the early hours of the outage;
On July 17th, 2015, the Canadian Telecom Resiliency Working Group (“CTRWG”), formerly called Canadian Telecom Emergency Preparedness Association, established reciprocal agreements between Rogers and Bell, and between Rogers and TELUS, to exchange alternate carrier SIM cards in support of Business Continuity. This is to allow TSPs to communicate within their organizations in the event of loss of their respective networks. Bell, Rogers and TELUS took the lead to provide SIM cards to all CTRWG members.
# #.
When it was realized that Rogers entire core network was offline, employees started swapping out our Rogers SIM cards with our alternate carrier SIM Cards. This previously established contingency plan allowed us to begin communicating within our organization in the early hours of the outage and to start restoring services.
i feel like there's a lot of passing the buck in the video
It is ineffective. Their remit includes "architecture" but the Rogers network architecture is a clear failure.
There was no political will to evaluate the results of such a group.
As a Canadian I feel that the current arrangement gives Canada most of the perks of America with very few of the downsides.
[0] https://www.pcmag.com/news/comcast-is-americas-most-hated-co...
Who hurt you?
Let's say this was a mistake. When they tested the config in a pre-prod environment, they would have noticed the redistribution of those routes. If they were pasting the config into a router instead of sending it via scp, maybe there's a mouse buffer/paste error, but you don't do that on a core device anymore. Judging by the response time, they weren't doing this onsite at the console either and their OOB access didn't work because it probably used the bunged cell network. Not negligent, but certainly risky.
I have personally made routing mistakes that caused peer resets, like mistyping a prefix on a static route that overrode the route to a peer interface over an exchange, and I have also overloaded circuits because of misguided path prepends. I haven't had enable privs on a core router in probably 20 years, but this wholesale redistribution is caused when you take one peers routes and announce them to another with you as the origin, known as "announcing the internet." As I remember, this isn't a single filter.
Not to muddy it, and it doesn't help my attack theory much, and please call this out as bullshit with a correction if you recognize it to be as such, but someone who used to work there told me something to the effect of their architecture had combined some of the older ASNs from previous Rogers acquisitions and that was part of the problem, whereby they essentially have the old AS's speaking iBGP to each other internally within Rogers, with only the single Rogers AS facing the internet as a peer, sort of the way the some of the old mega ASNs. It implies there were different policy domains with different levels of maturity inside the main AS, and those internal peers don't filter routes they recieved from the core. When you have a network that has grown organically over the last couple of decades, this happening at least once is inevitable. Legacy policies just redistributed what they were sent, and between flapping, dampening penalties, transiting and backhauling mobile voice traffic over IP networks, and backup and OOB reachability, this one of those perfect storms where for years the network both worked magically, and then one day failed just as magically.
Undoubtedly, there are probably some senior engineers sitting around saying, "I told you this would happen, and we've been trying to tell you for years," but when you know this stuff, it's your responsibility to also be persuasive.
I was very attached to the attack theory for a bunch of reasons, but the conceptual chain of events above makes more sense to me, and it is a lot more like a rubber gasket on a fuel tank expanding too rapidly an unexpectedly cool day. Obvious in hindsight and simple to explain, but invisible up until the moment of ignition.
When the human waste hits the circular cooler device, that's when we find out how well an organization is built and managed.
Who is surprised here it took them this long to recover? The necessary no-blame culture, temporary removal of decision barriers, all-hands-on-deck to get back to normal, all the things the lucky of us take for granted are guaranteed to be missing here.
"The typical Rogers Communications Systems Administrator salary is $67,329" (https://www.glassdoor.ca/Salary/Rogers-Communications-System...)
"The average system administrator salary in Canada is $71,981 per year or $36.91 per hour. Entry-level positions start at $61,425 per year, while most experienced workers make up to $94,400 per year." (https://ca.talent.com/salary?job=system+administrator)
https://www.ontario.ca/document/industries-and-jobs-exemptio...
But I'm sure Rogers would never abuse any of those.
Ottawa is cheap at around $965,000.
And even if you are into charity, it's usually more efficient for you as an individual to make more money and then donate to whatever GiweWell lists at the top of their list of most bang-for-your-back charities. Instead of working in a soup kitchen, or a Canadian bank.
For example, mobile phone connectivity is still very regulated, but in many countries competition seems to be much fiercer there than for home broadband.
I am curious. What are some of those aspects?
https://en.wikipedia.org/wiki/Impact_of_the_privatisation_of...
About the last: the (partial) privatisation of British Rail was a mess. Far from a textbook case. But still ridership numbers that had been in seeming terminal decline picked up in absolute terms, and rail's relative share of the overall transportation market increases total.
https://www.econlib.org/library/Enc1/AirlineDeregulation.htm...
The impact of deregulation of (American) airlines has been tremendous. Flying is now cheaper than ever. But to address my point: flying is still very strictly regulated, just less so than before (especially in the areas of pricing and competition).
https://www.econlib.org/library/Enc/TruckingDeregulation.htm...
https://www.econlib.org/library/Enc/SurfaceFreightTransporta...
Despite common complaints on the Internet, the US is actually still really, really good at running railroads. It's just that their area of competence is freight rail. Passenger rail in the US is anemic.
Guess which part is in government hands, and which one is comparatively free market and less regulated? However, neither area is completely free market or completely state controlled.
Another example: a few years ago Germany legalized long distance busses. The old ban was a hangover from the Nazi era.
https://www.dw.com/en/regulations-eased-on-long-distance-bus... is an overview article from just before the legalization of busses came into effect. I leave it as an exercise for the reader to find some sources from afterwards. Google Translate might be helpful, if you don't read German.
And, of course, legalization and deregulation here just mean different (and a bit less) regulation. There's still plenty of rules. We are talking about Germany, after all.
Working purely for money and being emotionally detached from your work is one of the bad consequence of capitalism -- in fact it's the most hellish consequence of capitalism that most people working software development jobs are likely to personally experience. I don't understand why people working at adtech companies talk about being emotionally checked out all day as a healthy way to adapt to the system.
I can tell you that jobs sucked compared to comparatively capitalist West Germany.
(Basically, capitalism might or might not be bad; but incremental improvements to it have a much better track record that trying to switch to a completely different system.)
I used to work for Google for a while. The gig was pretty cushy, but I could have still rattled off lots of complaints, if you asked me.
Almost any programmer job is almost infinitely better than what most normal people have to put up with. And better paid.
Kaiser arguably did more innovation than OneMedical, as far as I can tell.
These are sleepy places where you go to work 15 hours a week and on demanding days call in sick.
But besides there being no money in it, between regulatory hurdles, fragmented ecosystems and lack of incentives for individual doctors or hospitals to be good at sharing or maintaining these types of records, I could give the software away for free and still fail to gain traction.
In all of these spaces I don’t see how it progresses without government regulation and/or incentives, because the incumbents have literally no reason to improve.
It's just tech debt. If you've ever been to China, it's crazy how well WeChat works. There's just nothing like it in the EU or the US because we have 20 different decades old things for everything that people can't or won't replace.
I ruled out the USA because of the CFAA. I do not like second guessing everything I do online whether it violates a badly written federal law. Of course, United States v. Elcom Ltd. didn't make me much happier. Three Felonies A Day came out later and it just strengthens all this. We also should mention the almost complete lack of social net and the insane health care system.
There are numerous arguments that by now the United States is a failed state https://www.thenation.com/article/world/trump-usa-election/ https://eand.co/how-america-collapsed-and-became-a-fourth-wo... etc.
Some of the eHealth stuff uses active FTP connections, so this isn't even an exaggeration.
I would have. I call in sick during prod failures at one of my gov jobs.
In jobs with no real rewards for excellence but decent job security, a crisis is a reason to get out the door.
We have this where I work and I have mixed feelings about it. On one hand, it's certainly freeing that we don't get fired for mistakes. On another, it's strange to feel that the people who objectively caused the headache for everyone else go without serious attention to remediation.
IMO, we don't have to name names, but Root Cause Analysis should include the decision tree that the primary motivator chose. No blame culture also hampers a lot of insider-threat analysis.
Of course, letting people run rampant and acting everything is fine is not great, but there's a good chance it might be an organizational problem and, if not, a HR problem (which is not necessarily related to a specific incident).
> MO, we don't have to name names, but Root Cause Analysis should include the decision tree that the primary motivator chose.
I definitely agree with this! I actually have seen cases where people have asked for revisions to a root cause analysis not due to a lack of naming individuals but a lack of sufficient explanation for what the anonymous individual actually did to cause the outage, which I think is fair feedback to give.
> No blame culture also hampers a lot of insider-threat analysis.
Yeah, if no-blame culture is being used as a shield for maliciousness or incompetence, I think that could have a harmful effect on things. I guess I always just interpreted "no-blame" as actually meaning "no blame for honest mistakes", but I guess that's too long to be catchy.
I have twenty years of experience in the framework we work with and I can't roll a production hotfix without QA approval.
What this comes down to is: if prod goes down, I should be able to focus on getting it back up and not worrying whether I have a job tomorrow. That's not conducive to success.
if you do retros and the same thing keeps coming up and it's always coming from the same person, then the manager will have to solve that
They have no independent management infrastructure. Their management network rides logically on top of their main physical network. So when the latter went down you probably had people going to the physical locations of various DCs. But then you have to coördinate between them, and if your employees use your own cell phone service for communication, and that's down…
So the first couple of hours was figuring out what went wrong, then it was having people scramble to locations to plug in serial consoles, then it was buying a bunch of SIM card from other companies, and finally folks managed to wrap their heads around the big picture and come up with a way to bootstrap the network.
Sounds like some employees already had those for such an event, which is impressive. From the article:
> some employees started swapping out their SIM cards for Bell or Telus SIM cards that they had received back in 2015 as part of an emergency contingency plan established between the wireless carriers.
Even shopify is getting bled dry