30 karma · joined May 10, 2016
This seems like one of the events in which they changed IP on Route Reflector routers that were pretty busy, which would cause reconvergence and CPU spikes for all routers that it had sessions with. Also, there was a lot of volatility, as part of which re-advertisements were happening continuously. They also attempted rollback, which caused reverse operation, which triggered reconvergence. The other scenario is doing this change on the SDN controller, which affected all other routers.
More details: https://www.thousandeyes.com/blog/microsoft-outage-analysis-... https://www.thousandeyes.com/resources/na-microsoft-outage-a...
However, "is a bit sad in a way" part of sentence is interesting one. Edge services hosted within AWS/Cloudflare/Akamai improved customer experience significantly given that waiting for trans-Atlantic or trans-Pacific latencies is not thing any more.
I was just about to ask about differences with Tailscale, which is solving a lot of the challenges outlined in the post, but you answered it:
"First off, Netmaker is super fast because it can use kernel WireGuard. There are some other WireGuard-based solutions like Tailscale, but they use userspace WireGuard, which is much, much slower."
Good luck!