Cisco’s Network Bugs Are Front and Center in Bankruptcy Fight
bloomberg.com
bloomberg.com
Failure to do so is simply incompetence.
Maybe the customer didn't read the fine print of their service guarantees, and that's on them. I would hope that doesn't happen often-- it would be very silly if service guarantees fell apart every time some piece of shoddy equipment (purchased and operated by the service provider) turns out to be at fault.
The higher SLA system that I've created, was for a military project.
Physical network layout: I did choose a double port star topology, this is, every HP5300 modular swith, was connected to each other switch, with two "teamed" ports.
Usually if you only do the connections, you got a network loop. But with STP and VRRP in the HP 5300, I got an "always on" network.
The network did expand to +50 wifi access points with another proprietary wifi controller which also did have the capacity to gracefully failover connections from the APs, on network splits.
Servers behind the switches, where replicated in each segment. So you could at any time turn off (in order, or cutting the power switch suddenly) any rack (switch + servers) and the system did continue to work flawlessly. This was 5 identical racks/switches in star topology.
Acceptance tests did include to literally cut cables, literally turn off the UPS and RACK power, etc.
We could upgrade any firmware (the HP5300 cabinet, their hot swapable modules, or the servers BIOS/network-card/hard-disk firmware), without any service loss.
I'm happy with that result.
I've to say, all my other projects where I've work (+15 years), didn't have resources, neither did give any importance, to the needs of network firmware upgrades or downtime (because of a bomb?). In most cases, it was not because technical issues or handicaps, it was because of management.
Some projects did listen to me, and did contemplate the issue and planned it as a "maintenance window", or as what today is called "immutable infrastructure": prepare the new one, stop the service, replace, bring up the service.
I never did upgrade a switch/router firmware at $job, without having a backup switch ready and pre-configured, in case something went wrong. And preserve the backup one for a prudential time.
In my "always on" military project, firmware did need to pass acceptance tests in environments equal to production, before go to any production environment.
Edit: remove duplicated info
This web company promising such little downtime was stupid. And now they are bankrupt. Good.
They initially thought I was kidding but then we did the fail tree on a white board. It was an interesting experience for them, understanding what it took to get what they took for granted.
My "fail tree" (and I'll put it in quotes as unique to my conception of them) analysis consists of identifying a failure, the system response to the failure, a time to fix for the failed system, and a guess at the uncertainty on the fix. So for example "switch hardware failure" is a failure, with a fix time that varies based on "replacement part on hand" to "order/ship/install (replace) the entire switch". The first order is failure/fix tree with callouts of down time. The second order is mitigation/cost with mitigation strategies and their cost resulting in a new call out of potential downtime, and the third order is mitigation accelerators and their cost (which shorten recovery to non-degraded mode) which affect cost and possible down time.
Much of that you can do on paper, but sometimes you will have to run experiments to see how long things take to fix.
> According to Machine Zone, the hosting service couldn’t make it a month without an outage lasting almost an hour. Another in August of that year was traced to faulty cables and cooling fans, according to the publisher.
Cisco Nexus devices log alerts for low memory/resource situations to syslog, the defaults being 85% minor, 90% severe, and 95% critical. Were they not reading their logs?
Management would be a huge pain in the ass... but it would be doable.
With the right BGP and ospf design you can absolutely meet five nines availability for an end user customer perspective. We have some places that are approaching six nines.
This article reads like they had 1+0 everything and ran into a nasty iOS bug. Running out of ram is amateur hour as well.
When we need to do customer service impacting maintenance that will totally take their segment off the net, the hit can be from 15 seconds to a couple of minutes. And that is in a case where a colo customer is single homed to a single aggregation switch like one of our 48-port 10GbE aristas.
Don't misunderstand me, I'm not defending Peakweb, simply saying that running a five nines network is hard - it typically requires deep experience in diverse problem domains.
It definitely requires ccie level knowledge and at least 10-15 years experience, plus advanced Linux/BSD server admin skills to really do five nines right. It is indeed expensive and requires enthusiastic cooperation from non technical management responsible for budgets.
I worked my way up from a level 1 NOC type position, so I'd like to think that after 20 years I have a good understanding of all the possible OSI layer 1 failure modes (and things you can fuck up in configuration at layers 2 and 3), yet the things some other partner and competitor regional ISPs do continue to surprise me.
Nobody here are saying - as far as I can tell - that they shouldn't have done better. But that'd require them to actually have people with sufficient experience and the budgets and buy-in.
I don't know the details of the countersuit, but that one exists at all is pretty telling. It suggests one of many things: it's possible that the contract did not guarantee any amount of uptime; or maybe Peak Web was not adequately informed of the load that the game was going to take on their servers and therefore they want to argue that the level of downtime was reasonable given that they weren't told to expect those kinds of loads; or maybe the contract had a termination procedure and Machine Zone decided to violate that procedure and just drop the company -- which is not something you can just do.
I mean, lawyers will argue anything for cash, but it sounds like Peak Web isn't exactly rolling in the cash they'd need to do this on a whim. I don't know what the chances are that the lawyers in question are working on contingency, but it seems plausible.
If one is using standard protocols (bgp, ospf, etc) then mixing and matching doesn't really seem to be a problem.
So it's not just attackers, but being susceptible to the same thing triggering the same bug by accident at the same time, or manufacturing defects across a whole batch or model, or affecting a part that's used all over the place.
This is why Erlang has hot loaded code, since around 1986. Those who xxx are doomed to repeat yyy....
Also, the bug in question caused memory exhaustion, which isn't necessarily easy to fix by hot-swapping code without some kind of restart.