"Out of Band" network management is not trivial
utcc.utoronto.ca
utcc.utoronto.ca
That record has not been maintained in the digital era.
The long distance system did not originally need the Bedminster, NJ network control center to operate. Bedminster sent routing updates periodically to the regional centers, but they could fall back to static routing if necessary. There was, by design, no single point of failure. Not even close. That was a basic design criterion in telecom prior to electronic switching. The system was designed to have less capacity but still keep running if parts of it went down.
Most modern day telcos that I have seen still have multiple power/line cards/uplinks in place and designed for redundancy. However the new systems can also just do so much more and are so more flexible that they can be configured out of existence just as easily!
Some of this as well is just poor software, on some of the big carrier grade routers you can configure many things but the combination of things that you can figure may also just cause things to not work correctly, or even worse pull down the entire chassis, I don't have immediate experience on how good the early 2000s software was, but I would take a guess and say that configurability/flexibility has had a serious cost on reliability of the network
Unreliability is unreliability even of it comes through software and we ahould treat broken software as broken, not as "just a software error".
I view programming languages and software dev tools as fundamentally about the management of complexity.
The evolution of software development tools reflects a continuous effort to manage the inherent complexity of building and maintaining software systems. From the earliest programming languages and punch cards to simple text editors, and now to sophisticated IDEs, version control systems, and automated testing frameworks, each advancement enables better ways to manage and simplify the overall development process. These tools help developers handle dependencies, abstract away low-level details, and facilitate collaboration, ultimately enabling the creation of more robust, scalable, and maintainable software.
Notable figures in software development emphasize this focus on complexity management. Fred Brooks, in "The Mythical Man-Month," highlights the inherent complexity in software development and the necessity of tools and methodologies to manage it effectively. Eric S. Raymond, in "The Cathedral and the Bazaar," discusses how collaborative tools and practices, especially in open-source projects, help manage complexity. Grady Booch's "Object-Oriented Analysis and Design" underscores how object-oriented principles and tools promote modularity and reuse, aiding complexity management.
Martin Fowler, known for advocating the use of design patterns in his book "Patterns of Enterprise Application Architecture," emphasizes how patterns provide proven solutions to recurring problems, thereby managing complexity more effectively. By using patterns, developers can reuse solutions, improve communication through a common vocabulary, and enhance system flexibility. Fowler also advocates for continuous improvement and refactoring in his book "Refactoring" as essential practices for managing and reducing software complexity. Similarly, Robert C. Martin, in "Clean Code," stresses the importance of writing clean, readable, and maintainable code to manage complexity and ensure long-term software health.
While speed and efficiency are important, the core value of software development lies in the ability to manage complexity, ensuring that systems are robust, scalable, and maintainable over time.
I don't think there's any sort of expectation that it would definitely go one way though, as you say. It's all business and legal constraints providing engineering constraints to build against.
https://en.wikipedia.org/wiki/AT%26T_Building_(Indianapolis)
They didn't have downtime logs, but that doesn't mean that the rapid growth of telephone demand didn't outpace the Bell System's capacity to provide adequate service. The company struggled to balance expansion with maintaining service quality, leading to intermittent service issues .
Bell System faced significant public dissatisfaction due to poor service quality. This was compounded by internal issues such as poor employee morale and fierce competition.
So mobile phones would try to make a connection to the tower just enough to connect but not be able to do anything, like call 9-1-1 without trying to fail-over to other mobile networks. Devices showed zero bars, but field test mode would show some handshake succeeding.
(The CTO was roaming out-of-country, had zero bars and thought nothing of it... how they had no idea an enterprise-risking update was scheduled, we'll never know)
Supposedly you could remove your SIM card (who carries that tool doohickey with them at all times?), or disable that eSIM, but you'd have to know that you can do that. Unsure if you'd still be at the mercy of Rogers being the most powerful signal and still failing to get your 9-1-1 call through.
Rogers claimed to have no ability to power down the towers without a truck-roll (which is how another aspect where widespread OOB could have come in handy).
Various stories of radio stations (which Rogers also owns a lot of) not being able to connect the studio to the transmitter, so some tech went with an mp3 player to play pre-recorded "evergreen" content. Others just went off-air.
https://www.theregister.com/2022/07/25/canadian_isp_rogers_o...
In sane handsets (ones where the battery is still removable), that tool was and still is a fingernail, which most have on their person.
I believe the innovation of the need for a special SIM eject tool was bestowed upon us by the same fruit company that gave us floppy and optical drives without manual eject buttons over 30 years ago.
Meanwhile “PCs” had a functional but ugly rectangular opening the size of the entire drive, and the drive had its own bezel, and imperfect alignment between drive and case looked a bit ugly but had no effect on function.
(I admit I’m suspicious that Apple’s approach was a cost optimization, not an aesthetic optimization.)
If the emergency call doesn’t go through, try the call over a different network. This would also mitigate problems we see from time to time where emergency calls don’t work because the uplink to the emergency call center was impacted either physically or by a bad software update.
I didn’t think this failure mode was even possible.
How does one know ahead of time if any particular change is "enterprise-risking"? It appeared to be a fairly routine set of changes that were going just fine:
> The report summary says that in the weeks leading up to the outage, Rogers was undergoing a seven-phase process to upgrade its network. The outage occurred during the sixth's phase of the upgrade.
* https://www.cbc.ca/news/politics/rogers-outage-human-error-s...
It turns out that they self-DoSed certain components:
> Staff at Rogers caused the shutdown, the report says, by removing a control filter that directed information to its appropriate destination.
> Without the filter in place, a flood of information was sent into Rogers' core network, overloading and crashing the system within minutes of the control filter being removed.
* Ibid
> In a letter to the CRTC, Rogers stated that the deletion of a routing filter on its distribution routers caused all possible routes to the internet to pass through the routers, exceeding the capacity of the routers on its core network.
* https://en.wikipedia.org/wiki/2022_Rogers_Communications_out...
> Rogers staff removed the Access Control ListFootnote 5 policy filter from the configuration of the distribution routers. This consequently resulted in a flood of IP routing information into the core network routers, which triggered the outage. The core network routers allow Rogers wireline and wireless customers to access services such as voice and data. The flood of IP routing data from the distribution routers into the core routers exceeded their capacity to process the informationFootnote 6. The core routers crashed within minutes from the time the policy filter was removed from the distribution routers configuration.
* https://crtc.gc.ca/eng/publications/reports/xona2024.htm
These types of things happen:
> In October, Facebook suffered a historic outage when their automation software mistakenly withdrew the anycasted BGP routes handling its authoritative DNS rendering its services unusable. Last month, Cloudflare suffered a 30-minute outage when they pushed a configuration mistake in their automation software which also caused BGP routes to be withdrawn.
* https://www.kentik.com/blog/a-deeper-dive-into-the-rogers-ou...
I've seen outages caused by a single bad router advertisement that caused global crashes due to route poisoning interacting with a vendor bug. RPKI enforcement caused massive congestion on transit links. Route leaks have DoSed entire countries (https://www.internetsociety.org/blog/2017/08/google-leaked-p...). Even something as simple as a peer removing rules for clearing ToS bits resulted in a month of 20+ engineers trying to figure out why an engineering director was sporadically being throttled to ~200kbps when trying to access Google properties.
Running a large-scale production network is hard.
edit: in case it is not obvious: I agree entirely with you -- the routine config changes that do risk the enterprise are often very hard to identify ahead of time.
> "this configuration change was the sixth phase of a seven-phase network upgrade process that had begun weeks earlier. Before this sixth phase configuration update, the previous configuration updates were completed successfully without any issue. Rogers had initially assessed the risk of this seven-phased process as “High.”
> However, as changes in prior phases were completed successfully, the risk assessment algorithm downgraded the risk level for the sixth phase of the configuration change to “Low” risk"
> Downgrading the risk assessment to “Low” for changing the Access Control List filter in a routing policy contravenes industry norms, which require high scrutiny for such configuration changes, including laboratory testing before deploying in the production network.
Overall, the lack of detail of the (regulator forced) post-mortem makes it impossible for the public to decide.
It's a Canadian telecom: They'll release detail when it makes them look good, and hide it if it makes them look bad.
Note: the publicly available post-mortem makes it difficult for the public to decide.
Per a news article:
> Xona Partners' findings were contained in the executive summary of the review report,[0] released this month. The CRTC says the full report contains sensitive information and will be released in redacted form at a later, unspecified, date.
* https://www.cbc.ca/news/politics/rogers-outage-human-error-s...
If each of the six phases was a distinct set of config changes, then they really shouldn't have been bundled as part of the same network upgrade with the same risk assessment. But, charitably, I assumed that this was a progressive rollout in some form (my guess was different device roles, e.g. peering devices vs backbone routers). Should these device roles have been qualified separately via lab testing and more? Certainly. Were they? I have no idea.
Do I think there are systemic issues with how Rogers runs their network? Almost certainly. But from my perspective, the report (which was created by an external third-party) places too much blame on the downgrade of risk assessment as opposed to other underlying issues.
(As you can see, there is a lot of guesswork on my behalf, precisely because, as you mention, there isn't enough information in the executive summary to fill in these gaps.)
Edit: I bet it had diesel generators when it was in service with AT&T to boot.
Listing removed a couple weeks ago.
That's where AT&T screwed up in Nashville when their DC got bombed. They relied on natural gas generators for their electrical backup. No diesel tank farm. Big fire = fire department shuts down natural gas as wide as deemed necessary and everything slowly dies as the UPS batteries die.
They also didn't have roll-up generator electrical feed points, so they had to figure out how to wire those up once they could get access again, delaying recovery.
https://old.reddit.com/r/sysadmin/comments/kk3j0m/nashville_...
I've seen some power outages in california, and noticed that comcast/xfinity had these generator trailers rolled up next to telephone poles, probably powering the low voltage network infrastructure below the power lines.
20 to 25 years ago I visited a telecom switch center in Paris, the one under the Tuileries garden next to the Louvre. They had a huge and empty diesel generators room. They had all been replaced by a small turbine (not sure it's the right English term), just the same as what's used to power an helicopter. It was in a relatively small soundproof box, with a special vent for the exhaust, kind of lost on the side of a huge underground room.
As the guy in charge explained to us, it was much more compact and convenient. The big risk was in getting it started, this was the tricky part. Once started it was extremely reliable.
That's the right English word yes. And that's pretty cool!
Unfortunately the CRTC is run by former execs/management of Bell, Telus, and Rogers, and our anti-competition bureau doesn’t seem to understand their purpose when they consistently allow these 3 to buy up and any all small competitors that gain even a regional market share.
Meanwhile their service is mediocre and overpriced, which they’ll chalk up to geographical challenges of operating in Canada while all offering the exact same plans at the exact same prices, buying sports teams, and paying a reliable dividend.
I'm sure those 2 compete very hard with each other with that level of co-dependency.
Not 100% out of band for a telco though, unless they made sure to use a competitors lines.
However, as I understand it, at least for commercial use, the phone company provides some kind of box that has battery-backing so it can provide phone service for a certain duration in case of emergency.
However, I’m only familiar with the emergency phone call use case, for which voice is enough. I’m not familiar with any legal obligation to provide data service, so I guess that if you need that, it’s up to you to negotiate SLAs or have multiple providers.
[1] Wave division multiplexing sends multiple signals over the same fiber by using different wavelengths for different channels. Each wavelength is sometimes referred to as a lambda.
2. The risk that you and your competitor unknowingly share a common dependency, like utility lines; if the common dependency fails then both you and your OOBM are offline.
The whole point of paying for and maintaining an OOBM is to manage and compensate for the risks of disruption to your main infrastructure. Why would you knowingly add risks you can't control for on top of a framework meant to help you manage risk? It misses the point of why you have the OOBM in the first place.
> If your OOB network is your only way of managing things, you not only have to build a separate network, you have to make sure it is fully redundant, because otherwise you've created a single point of failure for (some) management.
I'm not sure I necessarily agree with that. You can set up the network in such a way that you can route over the main network as a backup if your OOB network was down but the main network was up. Obviously, it's not quite as simple as sticking a patch cable between the two networks, but it can be close - you have a machine that's always on your OOB network, and it has an additional port that either configures itself over DHCP or has a hard-coded IP for the main net. But the important thing is that you never have that patched in, except for emergencies like your OOB network cable being severed but you still have access to the main network. If that does happen, you plug it in temporarily and use that machine as a proxy. There's no real reason for extra redundancy in the OOB, because if your main uplink is also severed, there's not really much you're going to be usefully configuring anyway!
You can cross-connect your out of band network to an in-band version of it (give it a VLAN tag, carry it across your regular infrastructure as a backup to its dedicated OOB links, have each location connect the VLAN to the dedicated OOB switches), but this gets increasingly complex as your OOB network itself gets complex (and you still need redundant OOB switches). As part of the complexity, this increases the chances an in-band failure affects your OOB network. For instance, if your OOB network is routed (because it's large), and you use your in-band routers as backup routing to the dedicated OOB routers, and you have an issue where the in-band routers start exporting a zillion routes to everyone they talk to (hi Rogers), you could crash your OOB network routers from the route flood. Oops. You can also do things like mis-configure switches and cross over VLANs, so that the VLAN'd version of your OOB network is suddenly being flooded with another VLAN's traffic.
(I am the author of the original article.)
Because of how I use it, I was only considering the management port as being for management, and it's separated for security. In the example in the article, there was a management network that was entirely separate from the main network, with a different provider etc. I guess you may have a direct premises-to-premises connection, but I was assuming it'd just be a backup internet connection with a VPN on top of that, so in theory and management network can connect to any other management network, unless its own uplink is severed. Of course, you need ISPs that ultimately have different upstreams.
In the situation that your management network uplink is down, I'd presume that was because of a temporary fault with that ISP, which is different to the provider for your main network uplink. You'd have to be pretty unlucky for that also to be down too. Sure, I can foresee a hypothetical situation where you completely trash the routes of your main network and then by some freak incident your management uplink is also severed. But I think the odds are low, because your aim should be to always have the main network working correctly anyway. If you maintain 99.9% uptime on your main network and your management uplink from another provider is also 99.9%, the likelihood of both being down is 0.0001%.
I'd also never, ever, ever, want a VLAN-based management network, unless that VLAN only exists on your internal routers and is separated up again into individual nets before it goes outside the server rooms. Otherwise, you've completely lost any security benefit of using an isolated network. OTOH, maintaining a parallel backup network on a VLAN that's completely independent to the management network, but which can be easily patched it by someone at that site if you need them to, isn't necessarily a bad thing.
But anyway, these are just my opinions, and it's been a long time since I was last responsible for maintaining a properly large network, so your experience is almost definitely going to be more useful and current than mine.
From there the worst that can happen generally is that the packets spiral the wrong way around the continent
I see you haven't met Google's production backbone network(s)... We intentionally didn't connect the Middle East and India (due to a combination of geopolitics and concerns around routing instability), so any traffic between the two would go the long way around the world, incurring a 200+ ms RTT penalty.
(There was a ThousandEyes report back in 2018 that gave us a black eye. See pages 20-22 of https://pc.nanog.org/static/published/meetings/NANOG75/1909/...)
Agreed entirely on your point that if you're buying multiple redundant links, you're responsible for making sure that they're actually relying on different underlying fiber spans.
That's the point. If you want a reliable separate path, you must test it, and you must be prepared to spend time and money on fixing it. The tests include calling up the engineering manager for the separate path and verifying that it has not been "re-groomed" into sharing a path with your primary -- monthly or quarterly, depending on your risk tolerance.
Operations work does not end because the world keeps changing.
they have incentive and capability to get that traffic off the shared fate should it occur (even if that extends up to starlink serving one of their IP transit providers for OOB). that's why i question the wisdom of being overly concerned with starlink's particular paths.
In practice I don’t know how rapidly they’re able to route around damage to the ground network that could be shared, though.
That means Starlink can, in fact, guarantee communications during outages (so long as the Starlink network itself isn't down). You just need to have Starlink service at both the send and receive sides and the communication effectively acts as a direct link.
What would this look like in practice? Management interfaces like BSPs don't have a great security track record.
It would protect you against things like DDoS attacks, and you can even assign dedicated (prioritized) access for these management ports.
It’s an economical decision I suppose.
Thank you to Chris and to whoever posts his articles here.
Also this sentence makes me question his IQ:
"Some people have gone so far as to suggest that out of band network management is an obvious thing that everyone should have"
Yes Chris, Rogers, the monopoly telco company of Canada, should have OOB network! They can afford it.
Talking about the challenges of OOB is great, but the point the blog post is wrong and dumb.
The report says "Rogers had a management network that relied on the Rogers IP core network". They had no OOB network. They didn't even try.
This is a a symptom of Rogers status as a monopoly, negligence on the behalf of Rogers, and negligence on the behalf of the government who should have regulated OOB into existence. This is some serious clown car shit.
One of the advantages that competitor networks provides is redundancy. Canada doesn't have that, so their networks will remain weak. This will probably happen again some day.
Yes OOB is hard, but not even trying and then throwing up your hands and defending the negligent is stupid.
Makes one realize the reliability of RS232 in today’s day and age.
In telco operations, every MoP should have: an unambiguous linear sequence of steps, a procedure to verify that the desired result has been achieved, and a backout plan if things do go bad. This is drilled into you at every telco I ever worked at. Rogers' cardinal sin on the day of the outage was that they didn't have a backout plan at each step of the MoP.
More structurally, networks have a dependency graph that you ignore at your peril. X depends on Y depend on Z, and so on. And yes, loops are quite possible! OOB management is an attempt to add new links to the graph that only get used in a crisis. These kind of pull-it-out-when-you-need-it solutions are fine, but have a tendency to fail just when you need them. For one, they don't get exercised enough, and two, they may have their own dependencies on the graph that are not realized until too late.
So, what would this Internet rando prescribe? First order of business is to enumerate the dependency graph. I would wager that BGP, DNS, and the identity system are at or near the very top. Notice the deadly embrace of DNS and ID: if DNS is down, ID fails.
Next, study the failure modes of the elements. In the Rogers outage, a lack of route filters crashed a core router. That's a vague word, "crashed". Are we talking core dumps and SEGVs? Are we talking response times that skyrocketed, leading to peers timing out? Rogers really need to understand that. Typically in telco networks when nodes get "congested" like this there are escape valves built into the control plane protocol, eg a response that says "please back off and retry in rand(300)". They need to have a conversation with Cisco/Juniper etc and their router gurus about this.
Finally, the telco industry (or what's left of it) needs to do some introspection about the direction it is pulling vendors. For the last 15 years, telcos have been convinced that if only that can ingest some of that sweet, sweet cloud juice, their software costs will drop, they can slash operations costs, and watch the share price go brrr. Problem is, replacing legacy systems with ones cobbled together by vendors from a patchwork of kubernetes and prayers is guaranteed not to lead to the level of reliability that telcos and their regulators expect. If I'm a Rogers' operations manager and my network dies, I don't want to hear that some dude in India has to spend the next week picking through a service mesh and experimenting with multus to decide if turning if off and on again is gonna work.
> All the OOB in the world will not help you if you cannot reach the management entity (eg IP-enabled PSU, terminal server, etc).
In _healthy_ OOB situations, all of the adjacent OOB infrastructure should be reachable, even if the entire core IP network is completely tanked. The only scenario where this would not apply in my eyes would be a power outage that whacks an entire site including the OOB gear. But in that scenario OOB doesn't help you.
> Next, study the failure modes of the elements. In the Rogers outage, a lack of route filters crashed a core router. That's a vague word, "crashed". Are we talking core dumps and SEGVs? Are we talking response times that skyrocketed, leading to peers timing out? Rogers really need to understand that. Typically in telco networks when nodes get "congested" like this there are escape valves built into the control plane protocol, eg a response that says "please back off and retry in rand(300)". They need to have a conversation with Cisco/Juniper etc and their router gurus about this.
Typically the "crash" is memory exhaustion due to incorrectly configured filtering between either routing protocols, or someone blasting a BGP peer with a large number of unexpected routes. As a former support engineer for BIGCO-ROUTER-COMPANY (either C.. or J..), I can't tell the number of times I've seen people melt down a large sized network due to either exceeding a defined prefix limit (limiting number of routes allowed), or accidentally nuking an ACL controlling route-redistribution, and either cratering all connectivity (no routes), or dump all routes unrestrictedly (no filter), with the latter resulting in memory exhaustion. Luckily, everyone these days working with big routers are culturally conditioned to do change-commit confirmation - if you make a change that blows the box up and isolates it, it will automatically revert the change after a defined period of time.
> Finally, the telco industry (or what's left of it) needs to do some introspection about the direction it is pulling vendors. For the last 15 years, telcos have been convinced that if only that can ingest some of that sweet, sweet cloud juice, their software costs will drop, they can slash operations costs, and watch the share price go brrr. Problem is, replacing legacy systems with ones cobbled together by vendors from a patchwork of kubernetes and prayers is guaranteed not to lead to the level of reliability that telcos and their regulators expect. If I'm a Rogers' operations manager and my network dies, I don't want to hear that some dude in India has to spend the next week picking through a service mesh and experimenting with multus to decide if turning if off and on again is gonna work.
I think your perception of the quality of a K8 telco stack is a bit off to be candid. They are not cobbling together random stacks from unvetted vendors/sources. Nearly every telco K8 stack these days is using an off the shelf K8 vendor, and off the shelf K8-compatible services on top, again from (reputable) vendors.
At the end of the day this was a failure of culture and management. The technology is a side conversation.