WAN router IP address change blamed for global Microsoft 365 outage
theregister.com
theregister.com
From this it sounds like they might have changed the primary loopback IP, which by default is the "router-id" for various routing protocols, causing the entire network to have to reconverge. You can override the default router-id with an explicit address that does not depend on lo0 but lots of networks don't do that.
It's extremely uncommon to change the primary loopback address. It's less uncommon to add an additional one but as the article says that syntax varies by vendor: Juniper will add as additional by default, Cisco and Arista will replace the existing primary one (IPv4) unless you include the "secondary" keyword...
This seems like one of the events in which they changed IP on Route Reflector routers that were pretty busy, which would cause reconvergence and CPU spikes for all routers that it had sessions with. Also, there was a lot of volatility, as part of which re-advertisements were happening continuously. They also attempted rollback, which caused reverse operation, which triggered reconvergence. The other scenario is doing this change on the SDN controller, which affected all other routers.
More details: https://www.thousandeyes.com/blog/microsoft-outage-analysis-... https://www.thousandeyes.com/resources/na-microsoft-outage-a...
Network engineers and the people handling network ops always amaze me.
A crash affecting both sides of a "resilient" virtual chassis I had to work with took off a major broadcast last year (it was a last minute favour I was doing, and I rerouted to a tertiary route in a couple of minutes).
Meanwhile I ran a rather large event going out to some hundred million listeners via two crappy £300 switches which were completely independent of each other, into two independent routers, running via two separate systems (one on a UPS, one on mains). If one of them broke the other one was completely independent and the broadcast would have continued just fine.
As far as I am concerned, that is far better than a virtual chassis.
The kind of bugs that I’ve read about in errata notes over the years is wild and truly unpredictable.
Using multiple connectivity vendors doesn’t guarantee path diversity. Demanding fibre maps and ensuring that your connectivity has separate points of entry into the building, doesn’t cross outside the building, and validating with your DC provider that your cross connects aren’t crossing either, guaranteed path diversity / redundancy.
One of the problems that I've seen in practice that with the degree of virtualization at play that it has at the same time become much more easy to in principle be guaranteed 100% independence and in practice it has become much harder to verify that this is the case because of all of the abstraction layers underneath the topology. One of my customers specializes in software that allows one to make such guarantees and this is a non-trivial problem, to put it mildly, especially when the situation becomes more dynamic due to outages from various causes.
Sometimes of course you have to make judgement calls. From one location near Slough I have a BT EAD2 back to my building a few miles away. I know the route into my building, I can see the cables with my own eyes going in different directions. BT tell me which exchanges those cables goto, and provide me with a map into the field at a 1000:1 scale showing the cables coming in down a shared path. Sure it's possible BT are lying, but it's unlikely. Only use that location sporadically, and when I do it's a managed location, so I can accept the risk of a digger on the ground.
Another location in Norfolk, two BTNet lines, going to two different exchanges. They meet at the edge of the farm and go up the same trunk. That's fine, I can physically control the single point of failure there too, although if peering between BT and my network fails then I'm screwed, but I have a separate pinnacom circuit in a crunch.
Now obviously some failure become far harder to mitigate. A failure of the Thames Barrier would cause a hell of a lot of problems in Docklands, I'm not sure if any circuits in/out of places like telehouse, sovhouse, etc will remain. Cross that bridge etc. Whether my electricity provider will remain with a loss of the internet is another matter, so then it comes down to how much oil there in in the generators, and the generators of any repeaters on the routes of my network.
However the much easier to avoid is the problem of some shitty stacked switch the salesman says will always work.
If you’re buying SDN WAN solutions, you get what you get.
If you’re buying specific paths, you get what you pay for.
Of course I had one failure in Delhi which the provider blamed on 5 separate fibre cuts. Long distance circuits can run via areas where they can sustain multiple cuts across large amounts of area (regional flooding is a good one), and fixing isn't instant. This can be mittigated a little, but you still end up with circuit issues -- I had two fibre runs into Shetland the other month. Frist one was cut, c'est la vie. Second one was cut, had to use a very limited RF link. There's only so much you can do.
On the other hand I've just been given a BT Openreach plan which lists any pinch points of a new RO2 EAD install, I can see the closest the two get during transport is about 400m (aside from the end point of course, and experience has taught me I can trust it.
After that it went on different paths to three different buildings, which from each of those was then routed independently.
We take physical resilience seriously, as it isn't network engineers that do that part of the infrastructure. Enterprise network engineers then throw it all away by stacking their switches into a single point of logical failure.
(Still had a non-IP backup, but sometimes that breaks too -- just in different ways than the IP)
I worked on systems and platforms at the time, and we were more cynical even about vendors we liked.
The first time I saw something like that put into practice was when an experiment in the oil and gas industry that was scheduled to run for years delivered their network design. On the runtime cost of the experiment the extra network wasn't a big deal, but a service interruption would have been and would have caused them to have to restart the whole thing from scratch. It's more than a decade ago and I forgot what the exact context was but the whole thing was fascinating from a redundancy perspective as well as the degree of thinking that had gone into the risk assessment. Those guys really knew their business. Also the amount of data that experiment was expected to generated was off the scale. Multiple petabytes, which at the time (a decade ago or so) was a non trivial amount of data.
In fact it could easily 2x the cost for the same level of quality, which is why it's almost unheard of for cloud.
The fact that it's more difficult and complex to have two separate teams manage two separate networks means it's more prone to error and misconfiguration. The reason it's not done has nothing do with financial costs but rather because it makes no sense, for the very fact I just mentioned.
Like two independent internets spanning from your server to my laptop.
Two completely isolated end to end transports.
That's what OP meant when they said you could make it more reliable by having a redundant network. It's just prohibitively expensive.
Then if one internet goes down in any way I talk to you over the other. That's a fairly straightforward fallback algorithm to implement.
..oh wait. see what I did? ahhAHAHA
https://blog.cloudflare.com/cloudflare-outage-on-july-17-202...
In Microsoft's case, the remediation is not to put in place higher level systems to safely accomplish the goal of the command. Instead:
- "We have blocked highly impactful commands from getting executed on the devices (Completed)"
- "We will require all command execution on the devices to follow safe change guidelines (Estimated completion: February 2023)"
Requiring commands to follow guidelines sounds suspiciously like they're requiring network ops not to break things.
Take this for example, looks like the problem was an unplanned recalculation of routing tables. That's not going to be the case on a small scale test network, and rolling back won't help, indeed in this case it likely would cause more problems.
There isn’t a concept of a transaction or a rollback. You just enter a command, press enter and it’s live.
To counter this we’d write all the commands we planned on executing and peer review it. Nothing was to be done “on the fly” (at least in theory)
In short, coming from a developer perspective with ample version controls and gated releases… networking is a very wild ride.
Yeah, Cisco gear is bonkers.
Mikrotik has "Safe Mode", which undoes all commands since you entered "Safe Mode" if the connection that created the shell gets interrupted. It has saved my bacon on several occasions, but there are several obvious situations in which you can get yourself locked out.
Juniper gear has "commit confirmed $NUMBER_OF_MINUTES", which will roll back everything since your last commit if you don't do a "commit" within $NUMBER_OF_MINUTES. It will also, apply all of the changes you've staged all at once (and do configuration sanity checking before it performs the commit).
I do have no idea how Juniper's rollback works when multiple users are doing simultaneous config editing... maybe don't do that?
You get a warning
Users currently editing the configuration:
bob termainal p0...."
But the failure here is actually sshing to a network switch in the first place.Some cisco kit has restconf which is better for automation, but it's buggy.
It’s been a long time since I’ve touched IOS-XE (Cisco enterprise gear) but Cisco IOS-XR, Junos, Arista EOS and the Nokia SRs all support some combination of configuration transactions with rollback and commit confirm on a timer
This definitely doesn’t stop you shooting yourself in the foot, similar to how you can still push broken config to a k8s controller, but it’s some level of protection for certain types of changes.
This hasn't been true for a very long time. Juniper router's have rollbacks, commits and revisions:
https://www.juniper.net/documentation/us/en/software/junos/c...
and
https://www.juniper.net/documentation/us/en/software/junos/c...
Cisco has similar:
https://www.cisco.com/c/en/us/td/docs/ios/ios_xe/fundamental...
At some point network vendors switched manuals from engineers documenting features whitebox to educated techs documenting features blackbox.
There's a clear transition for docs produced after 2008, prior to which more care went into tech notes and interpreting technologies -- after you're lucky to even get a complete set of steps and caveats without having to cross-reference bugs, release notes, old-manuals, new-manuals, draft manuals, reference manuals, licensing manuals, the inevitable errors that appear in the logs, and of course the configuration guide where this should all be in the first place.
In short, yes, this.
As a dev who has worked at one of the major networking vendors, I can assure you that is the not the case. You’d be surprised by how major bugs are handled internally, especially if the bug affects “important” customers.
Ansible/Napalm is a thing in NetOps in some places. Some folks use Eve-ng / GNS3 to spin up virtual networks to test config changes, and it may be possible to do CI/CD changes if you track things in a repo.
Juniper JunOS has auto-rollback if you don't confirm the change after "x" minutes:
* https://www.juniper.net/documentation/us/en/software/junos/c...
So if you did something that causes breakage and disconnection from the router, you (ideally) don't have to do anything but wait it out.
Compare this to my cisco/foundry/other experience where I would delay changes until I was in the office (physically colocated with main routers) or calling people to be onsite for what was 99% of the time an innocuous change. The stress of it led to me deferring changes or just skipping them entirely which led to more issues/stress/etc.
I'm really not sure there is a single software feature which improved my life as much as "commit confirmed"
The problem is that the state in routing tables isn't stored in a single location, it's dynamically built over time. Breaking a single router in the wrong way can break the state, and there's no rollback of that state
It's possible, it depends on what the nature of the change is. If you use super short commit confirmed intervals (commit confirmed 1) then yes you can cause a situation where you revert a "good" commit and cause a second disturbance. You need to intelligently reason about commit confirmed times to consider this when you're making such changes.
And virtual switches aren't the same as physical switches in any case, they have different bugs, different features, different responsiveness.
"Debugging Under Fire: Keep your Head when Systems have Lost their Mind" (Bryan Cantrill, GOTO 2017)
I remember voicing at our team meeting "boy, they must be panicking at CloudFlare."
Cloudflare works so spectacularly we just wrote it off as a one time thing.
There’s no way it’s DNS
It was DNS
Or at least I think so, trying to figure out a packet loss issue on a virtual machine, for windows xp image.
switchport trunk allowed vlan (add) xxx
Can’t imagine how many outages where caused by the missing « add » command.
no access-list 101 permit something
so long access-list 101!