Reimagining the future of routers
medium.com
medium.com
What he is suggesting will drastically impact the latency of packets as well as the throughput. Following his analogy of cache hierarchies in a regular computer, some prefix lookups are going to be 'downgraded' to the main memory and will take maybe 100x the time.
If a router is forwarding at 10gbps it has ~51 nanoseconds in worst case (64 byte packets). So this is the overhead you can probably expect at each router as your traffic traverses the net. I'm 17 hops from my old university's network. That's just under a microsecond of lookup times. Not bad thanks to TCAM or other specialized lookup chips.
If all of these routers adopted this, my lowly traffic would have to read from main memory on each route lookup because I would never be in the top tier for bandwidth usage even if I was connected all day long. A full routing table is >600k addresses now[1] so each lookup may have to reference main memory several times as it walks a trie. Let's assume about 10 times to be generous which (at 100ns a reference) comes to about 1 microsecond. This is just for 1 packet.
As you start to pile on the thousands of other small connections (we're talking about provider network routers) that would be put into this steerage class, there is going to be contention and queuing that could easily push it into 1ms-10ms depending on bursts.
So if my whole path adopted these routers, I could experience jitter of ~170ms or worse. Gross.
This is essentially turning into a crappy QoS system where the biggest bandwidth hogs get the good service and everything else gets garbage.
Thumbs down from me.
> I would never be in the top tier for bandwidth usage even if I was connected all day long
Maybe, but it probably doesn't depend on you unless you are your own AS. Routing tables do not contain individual IP addresses so if you are a "normal" ISP consumer chances are you will be in a block that is in the top 10%.
I remember the days when ARIN would give out AS's to basically anyone. Before all the cable companies merged, I spent my teenage years developing close relationships to my ISP. One of the high level techs told me to throw up BGP on a FreeBSD box and got on the ARIN mailing list with me to get a /30. That's the type of AS that'd fall into the 90 percent category (along with small businesses that who, for whatever reason, have their own AS). Even on my own AS, I'd rarely see latency above 50 ms domestically via BGP or OSPF.
My home IP falls into a slash /22 block I can see in the global routing table. It's highly unlikely that this particular 1000ish IP addresses are responsible for enough traffic on the internet to make top 10%. Top 10% is going to basically be a bunch of hosting providers and CDN prefixes.
If routers become cheaper and smarter the network will grow, you might get more direct routes (less hops), or simply better load balancing (based on latency and-or throughput, maybe even indicated by the size of the flow, so your SSH session will be low latency low jitter but also if you start to download a file via a subchannel it might get shifted to different routers), and with better balancing flows will get distributed not just considering packet drops, but maybe cache hits too (latency basically).
Anyhow, upkeep costs are very important, because they limit the size and efficiency of the network. (High barriers to entry limit growth, high upkeep encourages centralization, which is bad from a reliability and fault tolerance aspect.)
on modern software forwarding core (fd.io/VPP and DPDK) you can forward with a sub 100us latency. so your total latency ends up being roughly 1.7ms "slower".
have a look at this: https://www.youtube.com/watch?v=T66BTHnENY8
yes, but current COTS hardware, via userspace networking can very easily saturate a 10g interface at 64b packet-size / core.
imho, it is the convergence of a large number of advances in COTS hardware over the last couple of years, that has changed the field to mostly a software problem rather than a combination of hardware + software one. when you have proprietary hardware, software tends to be more or less proprietary, which is where most of the networking vendors are currently.
a large number of projects are actively attacking this space e.g. snabb(https://github.com/lukego/snabb), routebricks (http://routebricks.org/) etc.
This is on bare metal though, so might be different experience if you are running stuff on vm's where without sr-iov you would have multiple layers of indirection + context switches before someone can actually do something meaningful to the pdu...
edit: slight clarification added
Citation needed.
http://www.ieee802.org/3/10G_study/public/speed_adhoc/email/...
What are you talking about?
I think the worst-case, steady-state, scenario is actually equal time sharing, and the whole value is that while you aren't talking, other packets flow faster. That's not a doomsday scenario. Noone is out anything. (If this equal-sharing isn't fast enough, your upstream links/peers are over-subscribed.)
Also, your cache/memory analogy disrgards the possibility of simplifying and/or distributing the routing tables.
Route consolidation is one solution. Having no table at all is ideal. My little underpowered home router has 1 entry for it's upstream gateway. There's no reason to keep billions of records in a table in any central router either, given there are only going to be so many uplinks.
The more scalable solution is to distribute the table(s) in parallel to other peers. If the routing tables really are too big and can't be consolidated in core routers, the answer is simply more routers, geographically distributed to best meet demand, not fatter ones in the core, with higher latency, higher costs, and all the downsides of centralization like survaillance.
QoS is a solution in search of a problem; networks scale fine when we scale them out, not up.
Only if you assume a terrible code base. It's very easy to make routing protocols as modules because they just maintain pokey old slow routing information bases that can live in main memory that don't have to react on the nanosecond scale.
I've worked on modules for OSPF on a vendor router and if the customer isn't using OSPF, that daemon and its code are never even executed. No "eternal bug hell".
This whole blog is basically just pitching major feature-gaps as a feature to prepare us for some MVP I expect to see from him in the coming months that only supports BGP and ethernet or something like that.
BTW Snabb is cool, but VPP is more feature complete.
Consider 1 million people watching netflix from the perspective of a transit provider. If you're just looking at bandwidth you can obviously prioritize lookups to netflix servers. But then you have 1 million streams to different client IP addresses throughout the Internet. Each on its own will be a small fraction of the bandwidth, so are you going to punish them all? Not much gain from lookup hierarchies there.
Given that he mentioned Amazon, I'm surprised to see that there wasn't more in this essay regarding Amazon (and Googles) efforts to build their own routers. Also, a number of networking companies have started off by discarding all the legacy networking functions, and starting afresh (Juniper). It would be interesting to review the field and see who else is doing this, particularly in the last 5 years, and what their success has been.
Also - surprised that SDN only gets a brief mention in the conclusion - I thought, reading the essay, that's the direction he was going, and then it was over.
The router, and the dynamic control-plane, as basic forwarding paradigm of the Internet, remains undisputed. However, it gets challenged using new concepts like SDN and NFV, which promise much faster network adoption, automated control, reduced time-to-revenue, which all are good business solutions. In order for router designs to be competitive to those challenges, requires to re-imagine how router hardware and software get engineered.
So, "Re-imagining the future of routers"
(All: suggesting a good title is the best way to complain about a bad one.)
I think the author may have been living in a hole. This is not a new idea. Datacenter routing in many of the big companies is already being done with SDN ala OpenFlow or some other custom protocol.
Right now you can buy a whitebox 'switch' with 40gbps interfaces and load various operating systems that enable different management styles (e.g. OpenFlow control like Google http://opennetsummit.org/archives/apr12/hoelzle-tue-openflow...).
The router has already been 're-thought', it's currently just hiding under the term 'whitebox datacenter switch'.
Going back to Linux, fact is using Linux on a bare metal switch today is like using Slackware in 1995. The need for a more turn-key Linux distro, with better apps (better routing, telemetry, ...etc) is needed to compliment white box hardware today. For a long time the HW has been the limiting factor since most merchant PFE platforms supported <50k FIB entries. Now with 1M+ (on chip) around the corner, SW quality and scale will need to improve accordingly.
I didn't say it was OpenFlow everywhere. I meant custom software in some way (i.e. not an off the shelf Juniper/Cisco/whatever). OpenFlow has pretty narrow use cases once you dump reactive flows so it's definitely rare in the wild.
For other customers using whitebox switches with different OS's booted, look at whoever is paying companies like Cumulus/Big Switch.
Also see Project Calico for containers, they use the linux kernel as FIB, but obviously they could just as well use OpenFlow to program switches for bare metal machines (which then might run containers).
This is very different than a TOR switch or core-router running OpenFlow. Only a few places are running more than a couple thousand VMs in a single SDN, and its nothing like what is being discussed above.
But if you need the big guns there are a ton of OpenStack Neutron plugins, OpenDaylight being one of them (networking-odl) and there are the classical/traditional vendor ones (Cisco, Arista, and so on).
I think the future is very much means this jungle of APIs integrating in whatever tangled ways, and eventually the reliable, robust and sane ones will prevail.
Well handling the routes themselves can be offloaded to regular servers. The issue is indeed managing the forwarding tables. "Something something aggregation, non contiguous bitmasks, NDA" runs wildly for the door. :)
There are plenty of boards that will hold 750Mbs of RAM from the usual folks - DELL, HP and Supermicro. So why is there no viable option? Also this is a hardware concern. How does choice of open source solution - Vyatta, Qugga or Bird matter?
What is the issue?
Full tables in FIB.
This is all very useful for in-building traffic until you get the traffic to the edge of your datacenter/colo/hosting environment network and need to exchange traffic with other ISPs... At which point you need a serious chassis based router with redundancy and full layer 3 capabilities (example: Juniper MX960, Cisco ASR9006 or 9010).
whitebox datacenter switches are a great way to get a lot of 10, 40 and 100 Gbps layer 2 ethernet switch capacity within a datacenter cheaply, they're NOT routers and should not be mistaken with them. They are things that you connect your hypervisor hardware platforms to (example: a whole shitload of 1U servers or facebook/OCP type servers each with dual socket, 16-core xeons in them), is a switch.
Not _that_ kind of router. The type of router connected to you cable modem is a box that performs NAT, not IP routing.
Remember, IPv6 is a necessity in the future if we want to continue to allocate addresses without running out, while better routing methods are simply an upgrade. So I'm not sure there will be as strong as a push either.
Even though the use of hostnames is convenient (and encouraged in IPv6) it's already a headache for knowledgeable people to set up, so imagine non tech guys (you're father or your grandma).
Unless we come up with a simple and easy process for this we're still miles away from a full IPv6 world...
[citation needed]
Also, the statement seems designed to encourage the reader to accept the converse: that somehow a microservice architecture will render technical debt impossible. High-grade bovine excrement, that...
Microservices are just "low coupling, high cohesion" reinvented for the web-heads.
https://en.wikipedia.org/wiki/Coupling_(computer_programming...
The premise is asking to remove a piece of functionality from software:
"Not possible? — Reason it is not possible is because things have been constructed as a monolithic system, mostly by just compiling a new feature. Most often, a given feature is intimately linked to the underlying infrastructure (like an in-memory database, or some event queue processor), and, removing it out of the code base, may get to an effort as large as originally developing the feature. In most cases there is no dedication on how to clean things up later. Every added feature will make future feature additions harder. If you just make the number of supported software features large enough, you can extrapolate, that at some point, this will become unmaintainable. External measure of such a condition, is too hard it is to get functionality into a particular main line release. If you are already using software which can never get “de-featured”, I have bad news — you are doomed to spend your life in the “eternal bug hell.” Availability goes down, operational cost goes up, and your vendor cannot possibly fix it. Time to change vendors is the only way out."
"Therefore i co-founded rtbrick.com where those Hyper-scale design principles are followed, to build the next generation distributed routing and forwarding platform with unbounded scale on your choice of open hardware."
And the comments:
"We are building a routing/system stack which both runs on vanilla ubuntu 14.04 as well as open network linux. The nice thing about our system is that it does not make any locality assumptions. — You can run the BGP control-plane distributed over several compute nodes and the IS-IS control running on different nodes. Yet the whole thing acts as a coherent system and can drive a set of bare-metal switches (e.g. A Dell Z9100)."
I imagine if this provided potential gains that it'd already be a known technique, but I can't seem to find any information about it one way or the other.
"Yet, most routers still support a 100ms+ buffer depth for 100GB/s circuits. Just do the math. You need 1.25 GB DDR4 RAM for each 100GB/s port in a given router."
What is the math? It's not clear at all how he arrived at that calculation. That seems like quite an important detail to omit in your first supporting paragraph. Just saying "just do the math" when its not clear what that is is a bit ridiculous.
buffer size = throughput * latency
1.25 GB = (100 Gb/s / 8) * 0.1 sAnd why is he capitalizing like it's 1390 AD?
But yeah, it's still a router. It's not "carrier grade" big expensive router in the words of Dave Temkin, but it's still a router and will probably smoke the market for lower-end Juniper MX's and Cisco ASR's.
Uses IS-IS to distribute the link-state db. And allows you to utilize an entire mesh, not just a sub-tree.
Ugh this term bothers me. Just call it routing! That hails from a day when routers sucked so much wind the marketing team at Cisco had to invent a new term for the stuff that did it fast.
Yes, I'd very much like to see TRILL gaining more widespread usage and attention, but fundamentally it's equivalent to IP-IP encapsulation plus IS-IS for link-state and iBGP for the outer IP layer, just with fancy terminology (RBridges and so on) and standardized administration/operational semantics. (Which is a good thing of course.)
I run several spanning-tree free networks and will never go back.
It's time to build your own router