BGP 768K day, and whether it will cause internet outages
blog.thousandeyes.com
blog.thousandeyes.com
And here's the past week. I suspect the big dip is where things actually broke...a little higher than 768k: http://www.cidr-report.org/cgi-bin/plota?file=%2Fvar%2Fdata%...
I believe there's another hardcoded hurdle at 1M IPV4 routes with some existing routers, like the ASR1001.
Guess IPV6 adoption isn't slowing IPV4 growth much.
The article somehow manages to avoid discerning between control plane and forwarding plane which is a key concept for this issue.
Of course in practice things are not that simple. Network operators have a habit of de-aggregating their BGP advertisements for various reasons, including traffic engineering.
It's an interesting problem from a economic perspective. Using BGP for traffic engineering consumes resources from every network participant, but there is no formal mechanism for ensuring efficient allocation of those resources. So far this hasn't been big deal because routing table capacity has been able to stay ahead of demand without too much trouble. It will be interesting to see what happens when/if the supply of routing table entries falls short of demand. Will there be a fee per advertisement? How would that even work? Will operators seen as over-consumers start seeing their de-aggregated advertisements dropped?
ISPs that were previously announcing /16 to /20 sized CIDR ranges to their peers and upstreams are breaking them down into /22, /23 and /24 sized pieces used for smaller downstream customers, and announcing those in a deaggregated fashion.
I predict that the ultimate end state of ipv4 will be a huge number of /22 to /24 pieces all over the world being announced, and further FIB growth requirement for big core routers.
You might say "okay, but the v6 table is not really big right now, so adjust the balance to 900k v4 and 100k v6 routes". But in reality on these ancient platforms each v6 route takes up a great deal more RAM than a v4 route.
If you have a router with 1 million FIB capacity, the time to replace it was five years ago. If you still have one running now, time to hit the panic button and replace it urgently with something like a Juniper MX80, MX104, MX204, etc.
> Engineers and network administrators scrambled to apply emergency firmware patches to set it to a new upper limit. In many cases, that upper limit was 768k entries.
Is there some technical reason for the emergency patch not to have increased the limit to a much higher and future-proof threshold?
In 2014 it shouldn't have been hard to predict that the new 768k would have been hit in just a few years.
In my day to day I’ll work on tables with billions of rows with no issue.