Was MPLS Traffic Engineering Worthwhile?
systemsapproach.substack.com
systemsapproach.substack.com
Bandwidth engineering, like Bruce said, was religion for some people and anathema for others. It saved at least one company from bankruptcy, letting them run the network they had much hotter than IP would have.
FRR, though, was an absolute game-changer. Local protection to failure and sub-50ms (often sub-10ms) protection switching in an IP network? Immensely valuable. Link connectivity protection was easy to deploy and I will argue made a massive difference in an Internet network's ability to provide real-time services in the face of failure. Node protection and bandwidth protection? Less so. But MPLS-TE-based FRR was the first local protection mechanism that didn't require you to have a specific topology and which could cover 100% of all connectivity failures.
I'm surprised Bruce didn't call out this distinction, because they really are two separate things using the same underlying tech.
https://www.usenix.org/system/files/conference/nsdi18/nsdi18...
I guess it makes sense why the article says how it is central in Google's / Microsoft's network architecture to achieve performance gains.
(Though, after 3 years of QA'ing the stuff, I gave up on MPLS and moved to other technologies. I was just not able to understand it to the deep level I wanted to.)
or if you'd prefer wiki - https://en.wikipedia.org/wiki/Multiprotocol_Label_Switching
From my head:
LDP = Label Distribution Protocol
RSVP = ReSource Reservation Protocol
PCE = Path Computation Element
BGP = Border Gateway Protocol
L2VPN = Layer 2 Virtual Private Network
IGP = Interior Gateway Protocol
MPLS = Multi Protocol Label Switching
TE = Traffic Engineering
FRR = Fast ReRoute
SR = Segment Routing
SP = Service Provider
LDP also enabled L2VPN and was an IGP-bound mechanism to enable MPLS VPN, which was always the killer app for MPLS.
LDP by itself? Meh. But it opened up a whole lot of other possibilities.
running a PCE also adds a massive failure domain compared to the decentralised nature of RSVP.
Does RSVP have a million bells and whistles? yes. Segment routing also has a billion different bells and whistles, and those are far less mature.
While we can argue that this is a niche use case, it would still be interesting to hear of any instances where this type of service is addressed using SR.
Bruce also co-wrote one of the earliest books on MPLS with the original author of MPLS [2].
[1]Was MPLS Necessary?
https://news.ycombinator.com/item?id=32988417
[2]MPLS: Technology and Applications:
I've also written some Verilog code for adding and removing headers on packets on FPGAs. Once you have to support VLANs in hardware, you're already half way towards inserting an arbitrary number of bytes into a packet as higher data rates in FPGAs mean wide busses with all the corner cases of packets being aligned at various offsets into a 256 bit or 512 bit data path. Figuring out how to implement an ethernet CRC32 block that meets timing in a Kintex 7 was certainly fun!
We don't have just switching labels and these other types may not be consumed, leaving the egress PE popping multiple labels. This is fine and the datapath is designed to make these sorts of packet manipulations, it's just harder to do that at full speed as people get further from the nice, simple original MPLS cases. When using multiple layers of encapsulation, customers don't make the same performance assumptions.
Things are worse with SRv6 where the need to read much further into packets is approaching the pain point. Most (all?) network processors have a limit of what you can manipulate at full speed. There's another painful case but I don't think I can discuss that.
With MPLS, you get simple fixed length headers with single integer values is quite easy to push/pop and look up in hardware, and you can have software to generate the policies and manage hierarchies, placing the complexities and hard work at the MPLS edges, and having simple/dumb (aka. cheap) devices in the core.
Also Andrej Karpathy was a schoolkid when MPLS was developed ;)
(My background: I’ve inherited RSVP-TE deployments in a number of continent-wide, and global networks — which has involved driving standards to improve its scalability, and subsequently driving segment routing in the industry and production deployments.)
The issue one has at any kind of scale is that it is non-trivial to acquire capacity that can be coherent in terms of different optimisation dimensions for your network. For example, a network I worked on could acquire limited capacity on Europe to India cable systems, alternate routes were significantly different latency - but there was significant EU-India demand in the network. What were the options for placing this traffic on the network? IGP weights - sure - but this means there is no selective placement of that traffic (i.e., everyone has to take the same route), which one might not be able to support commercially. Looking beyond that, there are limited options _other_ than MPLS-TE based on RSVP-TE. Path Computation Engine (PCE) support, even when it emerged, was RSVP-TE-centric. So, commercially, those networks didn’t have a choice to deploy traffic engineering — and it wasn’t for want of trying. Significant cable system deployments have been driven on routes such as the one that I mention above — so there was capital to be deployed to fix the problem, it’s just that building such systems take years to be built. Did their architects want to deploy RSVP-TE? Pre-SDN (and SDN-in-the-WAN like B4 and SWAN) what option did they have to meet that business requirement? I would postulate very little (at least at the time that I was engaged in these discussions there were no clear alternatives). In fact, I would postulate that the existence of TE in B4 and SWAN shows that there is value in traffic engineering practically. Greenfield/ground-up systems still implemented.
RSVP-TE itself though was not well thought through. The systems design discussion that I think is very interesting here is considering the lessons that we can learn from such a technology. Distributed state in the network, that causes large amounts of signalling following failures, and requires midpoints to be aware of all demands that traverse them and admit them is fragile by its very nature. The scaling analysis that was done during the architectural work (RFC5439 for example) did not think of the RSVP-TE distributed system’s different points of dynamism — it concentrated on steady-state cost, but we’ve demonstrated time and again in production (over many, many years) that practically the system’s scaling was to do with the cost and scaling of dynamic resignalling following events rather than steady-state utilisation (I’ve presented effusively on this, see https://research.google/pubs/pub45800/ and https://youtu.be/NtED7CUHLNE).
Rather than raising the question, from a systems design perspective, as to whether MPLS-TE was a religious mistake — let’s raise the one around how complex distributed systems that have multiple vendors of their equipment (i.e., not ecosystems that are controlled by one party) can iterate on solving business problems without religiously filling in the gaps of protocols that don’t work. In my view, this would be a huge step forward for the networking industry.
Finally, let’s not fall into the trap that we over-generalise. Higher scale networks within a limited geography (terrestrial UK, US etc.) may not have the same business considerations — and therefore may not need the same approach to traffic placement. There’s, as always, a set of trade-offs here.
I work for one of the vendors that makes the boxes that run B4 and SWAN. You are absolutely right.
Throw away because I don't want be identified.