How Verizon fixed a recent backbone provider issue
blog.edgecast.com
blog.edgecast.com
Edgecast is owned by Verizon, so cynical-me is having a hard time believing this is not a plant.
We see similar posts pop up once in a while with other such situations (CloudFlare springs to mind). It just happens to be Verizon in this case.
I think there are a few that are. I see very little good in a company like Philip Morris. They manufacture a delivery system for an addictive drug that, as a side effect, results in countless cancer deaths.
I couldn't possibly see myself working for a company like that.
OTOH under the right circumstances I'd work for Verizon. Their evil is venial compared to merchants of death. But a man's gotta eat, and that's why we prostitute ourselves out to these evil behemoths. And why we rationalize that we're "just trying to do the best job we possibly can".
This is a pretty good write-up of what happens behind the scenes when debugging a network issue.
That blog post would have been written something like this:
Verizon: Those sure are nice packets you're asking us to deliver to our mutual customers. It'd be a shame if something happened to them at the interface between our networks.
Netflix: Ouch. Ouch. Ouch. Thank you for bending us over and screwing us. Here, take some of our money. Please make the pain stop.
That's your fucking job, what do you want, a medal?
Seriously, the level of crap that verizon, Level 3 and AT&T put out is immense. I was trying to get a 100meg line in downtown redwood city. (this was q4 2013) at first I was told that the exchange was full. After that they said I could have bonded T3. After much screaming and shouting they decided that they would do me a favour and provision a fibre line.(bear in mind this was part of a large global account, with MPLS and other such niceties)
In the end it took 3 months of epic hassle (this was without way leave) just to get to the point of connection. after another 6 weeks of pointless meandering, I had fibre.
However because of the level of skill at the NOX it was another 3 fucking weeks to get it lit properly. (15% packet loss is not acceptable by the way)
The worst part of this is the cost, $4500 a month for a steaming pile of shite, backed by people to thick to open doors effectively.
In london this is how it went down: Phone up $provider, I want a 1 gig line please.
$provider: sure, that might be up to 90 days, pending legals and survey
Me: ok
$provider: (week later) survey is done, line should be lit in a month
$provider: oh and thats £1500 a month
In conclusion, just fuck right off, get off your arse and fucking do what we all pay you for, provide some fucking bandwidth.
edit: my past experience with CenturyLink was always great -- they provided us with extremely detailed documents of the provisioned line with loss, bandwidth tests, etc.
Funny enough, I'm a direct client of EdgeCast and I believe I witnessed a case of file corruption on their layer last week.
There are lots of built-in monitoring functions for the hardware interfaces. For example, the laser driver for the transmitter often drops in power before failure, so there is a hardware function built-in on the receiver side to report the avg laser power. If the power drops below a certain threshold, an alarm goes off on the monitoring software.
The author is pointing out an interesting failure mode not automatically caught. As a shot in the dark: In the receiver, the hardware is spec'd to receive 31 zeros in a row (iirc, which btw goes way back to early SONET specs). If the detector (eg) is degrading in a way that causes a particular string of zeros (or some other fixed pattern) to fail, you would see it show up much like the test that was run.
What's interesting is that the errors did not light up alarms on the monitoring. Perhaps crc errors have a threshold, and the bit pattern causing the failure was showing up statistically below the threshold.
Anyway, I'm sure there's some hardware failure-analysis guy out there who would love to look harder at that link.
If you want 10 Gbit/sec line rate, one of the ways to do that is to simply forego any and all verification that packets are not corrupted and shove it out of an ethernet port as fast as possible.
I recently ran into this at work, we had some issues with corrupted packets, and we ended up tracing it down to a Twinax cable that was failing. But every single switch in the path to the server happily forwarded the corrupt packets.
Luckily Cisco has some internal counters that show issues like that, and after tracing it down multiple switches/routers we found the culprit and fixed the issue!
Edgecast is a great CDN and company, I'm curious to see how this new ownership plays out.