This.
The higher SLA system that I've created, was for a military project.
Physical network layout: I did choose a double port star topology, this is, every HP5300 modular swith, was connected to each other switch, with two "teamed" ports.
Usually if you only do the connections, you got a network loop. But with STP and VRRP in the HP 5300, I got an "always on" network.
The network did expand to +50 wifi access points with another proprietary wifi controller which also did have the capacity to gracefully failover connections from the APs, on network splits.
Servers behind the switches, where replicated in each segment. So you could at any time turn off (in order, or cutting the power switch suddenly) any rack (switch + servers) and the system did continue to work flawlessly. This was 5 identical racks/switches in star topology.
Acceptance tests did include to literally cut cables, literally turn off the UPS and RACK power, etc.
We could upgrade any firmware (the HP5300 cabinet, their hot swapable modules, or the servers BIOS/network-card/hard-disk firmware), without any service loss.
I'm happy with that result.
I've to say, all my other projects where I've work (+15 years), didn't have resources, neither did give any importance, to the needs of network firmware upgrades or downtime (because of a bomb?). In most cases, it was not because technical issues or handicaps, it was because of management.
Some projects did listen to me, and did contemplate the issue and planned it as a "maintenance window", or as what today is called "immutable infrastructure": prepare the new one, stop the service, replace, bring up the service.
I never did upgrade a switch/router firmware at $job, without having a backup switch ready and pre-configured, in case something went wrong. And preserve the backup one for a prudential time.
In my "always on" military project, firmware did need to pass acceptance tests in environments equal to production, before go to any production environment.
Edit: remove duplicated info