Juniper bug takes down core Internet routers around the world
silicon.com
silicon.com
I've seen many environments where it's a multi-year process to roll out new code, particularly in service providers. I wouldn't classify asking a service provider to change all their code in a matter of months a reasonable request. The amount of testing they do before rolling code is very time consuming. It's not uncommon for service providers to run on 3-4-5 year old code.
I see your point, Juniper missed a serious bug in their testing. But you can't hold them completely accountable when they've already announced and released a fix that corrects the issue.
For systems (like core routers) that are simultaneously too critical and too available to permit timely maintenance cycles, the only solution is to not have any bugs ever.
if something is too critical to be taken offline, it should have a hot standby, right?
This is one of the reasons I'm hesitant to embrace the various NoSQL options at this time - often there's only one implementation of an API, and it's tied to that code.
Compare this to the various message queueing options that all support STOMP or AMPQ, or programming languages that have multiple implementations.
Networking needs to define a format spec for routing and switching, and then have vendors meet the spec. Fortunately we should be getting something like this with software defined networking projects like OpenFlow.
If a router is reset couldn't that cause some large changes in the BGP tables and subsequent route flapping as routers go on and off line. Perhaps large changes like that trigger the juniper bug if there is one.
Please check out the IETF (www.ietf.org) -- This is exactly how it works.
But BGP has no security, is complicated from an implementation standpoint, and you are right, there is a bit of a software duoculture. Juniper and Cisco. That's it.
This has happened before... too bad the routers didn't crash BEFORE propagating the bad BGP updates. :-)