What caused today's Internet hiccup
bgpmon.net
bgpmon.net
and that doc goes ahead to explain how to increase the limit at the cost of space for IPv6. Worse: The sample code (which everybody is going to paste) doubles the space for IPv4 at the cost of nearly all the IPv6 space, even though we should soon cross the threshold when we're going to see more IPv6 growth than IPv4 growth.
I agree that it would be better not to harm IPv6 growth. But if the alternative is to break the IPv4 internet then the choice seems obvious.
Although it could be argued that "workaround" is a sub-set of "solution" ;)
[1] http://blog.level3.com/global-connectivity/verizons-accident...
> So whatever happened internally at Verizon caused aggregation for these prefixes to fail which resulted in the introduction of thousands of new /24 routes into the global routing table. This caused the routing table to temporarily reach 515,000 prefixes and that caused issues for older Cisco routers.
The connectivity issue was caused by the routing table to grow too big for old routers used all over the place.
The routing table grew too big because of a mistake by Verizon which might or might not have been running old routers.
I question the assertion that many people in charge of Cisco routers doing BGP are in the habit of pasting in configuration changes without thinking about them.
Obviously this only finds routing level problems. We can send a /17 to you just fine, but if you're having an IGP problem and sending every byte of it to null, well, from the BGP perspective that's just fine. Much as if you insist on sending us RFC1918 traffic we'll drop that route and traffic for you just fine, just like we had to eat your 0/0 route you're trying to get us to advertise to the entire internet. I think my head still has a flat spot from hitting it on the desk arguing with people.
Its been a decade since I did that stuff professionally at a regional ISP and I really don't miss it. Not much, anyway.
Note that there’s no good exact opinion about the One True Size of the Internet — every provider we talk to has a slightly different guess. The peak of the distribution today (the consensus) is actually only about 502,000 routes, but recognizably valid answers can range from 497,000 to 511,000, and a few have straggled across the 512,000 line already.
[1] http://www.renesys.com/2014/08/internet-512k-global-routes/
It's interesting how they explain that since there's no true consensus for the actual size of the routing table, the "event" of crossing the 512k barrier has frankly already begun ... and, so far, hasn't been catastrophic, nor likely to be.
In particular, the telling error message is:
%MLSCEF-DFC4-7-FIB_EXCEPTION: FIB TCAM exception, Some entries will be software switched
1) there is a 1 gig inband channel to CPU. the traffic prioritization referred to (selective packet discard) really doesn't matter at the point in which the inband channel is saturated. when you start punting large amounts of traffic to CPU, it will take down your routing protocols, kill ARP, etc, even with selective packet discard, due to this saturation. the only way to prevent this is CoPP in the fast path.
2) inband traffic is interrupt driven on this platform. high inband traffic, by itself, will cause the CPU to spike and drop protocols (missed hellos, etc).
the result will be, without a doubt, an outage. packet loss and high latency would be the best case scenario, but only on a box that doesn't carry much traffic (typically not the case for anything taking full tables).
as a side note, these boxes should have started alarming well before overrunning the TCAM (iirc, it begins at 97% utilization), so operators should have had notice to implement the necessary TCAM carving changes.
If anyone recalls Anonymous threatening to 'take down the internet' by DDOSing DNS servers, who knew they could have done it much more simply by dumping 100K BGP paths into the network.