Google opens Falcon, a reliable low-latency hardware transport, to the ecosystem
cloud.google.com
cloud.google.com
I've been working with Ethernet devices a lot lately, using the network as a communication bus, essentially. I find that there's a lot of complexity that we simply don't need: ARP, DHCP, DNS... So many points of failure. We know all the devices on our LAN and their unique MAC addresses, and could do everything we need to addressing-wise at Layer 2. But everything's built on Layer 3 and up, so we're effectively working backward to map devices to IP addresses and vice versa. It's unsatisfying.
2. The Ethernet header has less entropy for ECMP than a UDP/IP header. Maybe you could add entropy somewhere but ASICs may not support it.
3. You're breaking compatibility with... everything. Maybe Google could afford this but no one else could.
2. Use some bytes outside of the header?
3. I get the impression this needs hardware to really use well anyway.
2. The point is to find some bytes that are constant for a given logical stream of related packets. Taking bytes outside of the header means taking bytes from the payload, that is by definition not deterministic. That's why everything identifies flows using the IPs + ports + protocol.
Okay, but my point there was saying "Google has millions of servers" isn't relevant, we're not looking at the entire company.
Even with a few addresses per VM, how many racks do you need to put into the same shared-compute mass? One data center is the upper limit, but it doesn't have to be the entire data center.
- let’s put 500 vms on a single one. with 128c256t CPUs it’s easy
- say you can fit 30 of those in a single rack (the common rack is 42U) due to power constraints
- and place 10 of those racks
That’s 500 x 30 x 10 = 150000 nodes to address. With 10 racks you already blow past the scaling limits of the common datacenter switch when it comes to MAC addresses. Here are the limits for Cisco’s Nexus 9000 series, a very common datacenter switch: https://www.cisco.com/c/en/us/td/docs/switches/datacenter/ne...
I've seen computers at moderately sized LAN parties (talking 100 nodes, far from the large or even massive events) that were literally crippled by the broadcast traffic. At some point the flooding and layer 2 discovery (ARP) would do the same as well.
Limiting the broadcast domains with layer 3 really makes the Internet possible. Sure, you can have less overhead and simply do layer 2 only, and really it is completely possible. It's just such a rare use case that it in practice isn't important enough to actually do.
This is a long read but you might find it enjoyable:
https://www.researchgate.net/publication/350711603_Adding_is...
Maybe for a LAN that's fine, but for connecting devices between broadcast domains, it just doesn't cut it.
I'm certain there are reasons IP came to live alongside/on top of MAC, but saying you can't do multi-hop routing with it just isn't true. If all the technologies of the Internet were reset tomorrow, how might you design the perfect layer 2 addressing and routing system?
It just isn’t suitable for this.
AB:33:C6:C6:19:74
I used a MAC address generator to get those two, but I think two is enough to make the discussion. Current reality aside, would you be able to identify those with binary math as being on the same network device, different network devices, across the world? MAC addresses on physical NICs are provided by the manufacturer, sure you can adjust them but I think that leaves the good-faith portion of this discussion.
So if you wanted to have those to communicate no matter what you would have to have a network device state: "I'm network device A, I have this device 0C:F9:31:D2:DB:51" then another state: "I'm network device B, I have this device AB:33:C6:C6:19:74". Then whenever 0C:F9:31:D2:DB:51 wants to talk with AB:33:C6:C6:19:74 it's network device will have to just send it to the next upstream network device or if there are multiple network devices that could be upstream you could send it to them all which is just not great for security whatsoever or you now have to do a recursive lookup for whatever n devices might yet be upstream and wait for a response to see if one of those has it. Overall trying to send ethernet frames globally without an IP network sounds like not a great idea.
Still, there's doesn't seem to be any reason you couldn't just say "device 1 gets MAC 00:00:00:00:00:01" and "device 2 gets 00:00:00:00:00:02" and the gateway controller gets :::00 and there's a special address on :::FF that can be used to talk to everyone...
Is that it? Is that all there is to IP? A loose pattern for reducing search scope, a couple reserved addresses for special cases, and a balance between address bitsize and total number of unique addresses (without requiring additional routing complexity)?
It all seems so... simple
> It all seems so... simple
Because you haven’t even thought through basic use cases.
Then you realize doing some action ends up being O(n^2) so you add some workaround in your switch and cache some things. And you know what they say about cache invalidation. And vendor A implemented it wrong in 1993 so you have a special case for their systems. And then you want to handle abuse cases. And authentication. And you're competing against the whole rest of the world and your thing isn't enough better.
The reason we don't is because at the time IP was introduced, there were many alternative physical layers in active use. And while Ethernet is near ubiquitous now, what we learnt from that was that it is unreasonable to assume that all your data will go over the same physical layer. And so you need a standard addressing format that will work elsewhere too.
Nothing stops you from stripping it back locally and using MAC addresses for everything internal to you, and ditching IP, and "just" gateway to/from IP. Lots of people did gateway between different protocols before IP became the dominant choice.
But you won't get everyone else to change because it'd require new firewall and new routers, and all kinds of software rewrites, and you can see how long the IPv6 transition has taken, so you'd still need to wrap and unwrap TCP/IP and find a way to address IP for everything that isn't 100% local, and even for lots of local-only stuff unless you want to rewrite everything.
There would be potential ways. E.g. you could certainly use a few bits to say "this is external" and then have some convention to pack an IPv4 address into the MAC or let an IPv6 address overflow into the data, and use that to make gatewaying and routing to external networks easier, while everything else just relies on the MAC. But you'd still need a protocol header for other things too, and then the question is how much benefit you would gain from ditching pretty much just ARP, which isn't exactly complex, a lookup table, and replacing the IPs in the header with just a destination MAC. Because the rest of the complexity is still there.
And you can gain most of the benefit of that by getting an IPv6 EUI64 address [1]. They'll work with "normal" IP equipment, and you can optimize in your own software by having the IP stack ditch ARP lookups when they see a local EUI64 address. Whether that optimisation actually makes a difference is another question.
[1] https://community.cisco.com/t5/networking-knowledge-base/und...
The table required for the whole Internet is large, but not gigabytes.
You can't route by MAC-address because it's effectively random. You'd have to store the port number for every device separately. This works fine at LAN scale, but not for the whole Internet.
Not that I see any advantages to the approach but it's almost workable(?), if a little silly, at internet scale:
If every device had a 64byte ID, guesstimating 10billion people * 100 devices/head gets us a 'measly' 64TB of storage. Double that to include routing info gets us to ~128TB. A bit much to be practical, but not entirely insane either.
Question: how does dns lookup differ from MAC lookup. Why is domain name lookup feasible, but not MAC?
a central lookup database for mac addresses (which could be distributed by having separate servers for a segment of the address space) doesn't make much sense because the distance of a server to the location of the device is to great and would make updates expensive.
so the router has to remember each address used. but at least it would not have to store all addresses in existence. actually, i think the storage needs are similar to those for NAT. well, except backbone routers which have to store a lot more.
the actual problem is the initial discovery of a MAC address. where does the routing information for a MAC address come from?
you need some peer finding protocols like DHT, and those are slower.
and i think the self assigning protocol in link-local could even go a step further. instead of hard coding a subnet, it could detect the subnet by copying the one from its nearest neighbor. so start with a random address, talk to neighbor to learn the subnet (and netmask) in use and switch to a new address within that subnet. then possibly run DHCP and update the address again. for static addresses DHCP could identify hosts by its cryptographic host key (like the one for SSH)
when two subnets join one of them may have to adjust its prefix. more complex, but still possible.
subnet prefixes could still be assigned to organizations to avoid overlap on a global level.
i am sure i am missing some details but i think in general this could work.
It works on small scales. We can stitch together a few LANs with ethernet switches. The switches initially forward everything to all ports, but learn where the MACs are so as to send frames only to ports where the destination MAC is known to be.
Ethernet switching won't scale to anywhere near the complexity of the Internet.
There were different teams/universities working on what today we would call LAN and WAN. I forget the details and history (I'm sure someone here, who was involved, could chime in, hah) and might have this wrong, but the result is LAN networking is MAC based while WAN networking is IP based.
It's one of those accidents of history that things are just the way they are and many don't question it. I run into it a lot describing basic networking concepts or early cisco material when people ask _why_ both MACs and IP addresses exist and its just... not always the correct time to explain those details to them.
- "Directly connected/visibile" means node X can contact node Y simply by throwing something on the medium (wire, radio, etc.) and doesn't have to knowingly send to a middleman (router).
When Ethernet was invented in the early 80's there were a lot more L2 technologies. Most are uncommon now (Frame Link DLCIs I think fall in this category, and PPP/dialup was common at one time - no MACs there) except for one: I don't think the cellular network uses MAC addresses at all. I could be wrong with newer 4G/5G stuff which overlaps with Wi-Fi in various places.
Because else you need ARP, IP, UDP, TCP, etc.
That said, it's not very hard to directly talk "ethernet" using raw sockets. Here's an example if you're interested: https://gist.github.com/austinmarton/1922600.
Just write a program to send and receive Ethernet frames. I believe it doesn’t have to contain IP.
Yeah, well, you've basically described IPv6.
My cursory digging indicates that the secret sauce is to grant large IPv6 prefixes and delegate routing to the prefix. An informative-looking Reddit comment says there are 100k IPv6 prefixes (as of Oct. 2020), and each active route takes 1 KiB. [1]
So, IPv6 differs significantly from MAC addresses because you only need to track prefixes.
[1]: https://old.reddit.com/r/networking/comments/j9twgq/how_much...
I believe there are other broadcast protocols that runs on Ethernet for the same reason.
Fibre Channel-over-Ethernet used to be a thing too, but I haven't seen it for a while. Perhaps latency wasn't as much of an issue as people thought and it lost to iSCSI.
I don't think it necessarily is bad idea to run protocols directly on the data link layer. The fewer parts the better. It's just that somewhere someone probably wants to route it, and the more general usage tends to win. So it's always going to be a niche market where latency is really important.
Actually it's interesting that google didn't choose any of these, for their high bandwidth storage needs. They have the money to do their own thing, but why should they?
I don't understand enough about niche high-performance interconnects to know if CXL is a viable alternative for Infiniband where Ethernet-based solutions have too much latency.
Stuff like InfiniBand, HPE Slingshot, Atos BXI, ... There is a consortium that's building a specification for those kinds of things: https://ultraethernet.org/
But still, you would think that some of those lessons could be learned before replacing it. AKA FC routes IP as one of its many protocols on top of the lower levels providing far more service guarantees than one normally gets with ethernet. Much of the QoS/latency/etc metrics were designed into FC from the beginning as a use on storage area networks (SANs). It just never took off as a IP transport because it cost 10x as much as ehernet, including a decade ago when these same groups tried to dump it on an ethernet MAC only to discover that it requires special switches which were $$$$ because "enterprise markup" defeating the whole point of cheap ethernet phy's. See FCoE..
And yet today, there is NVMEoF on FC, which is what one runs when its important that someone scp'ing a file on your network doesn't cause your database queries to slow down.
What I don't get is why OCP doesn't just actually build some of these adapters/etc with a "we won't be greedy" take and sell them not only to the hyperscalers but on the open market. That way someone could actually build say, a FC adapter that has a price similar to an ethernet adapter.
What is "Hardware transport", what is "the ecosystem"? And then there is dozens of random products and technologies that I've never heard of...
This sounds more like a humble brag, than an article trying to inform people about technologies that might actually be useful to them.
As we all know, interest at Google will now wane since you can't get promotions out of this any more.
"The ecosystem" is The Open Compute Project [1], a trade association which mostly puts together quasi-standards that provide targets so computer manufacturers can produce bleeding-edge gear with some hope that it will be interoperable. An example OCP production is the newer 21" racks that are starting to appear in datacenters.
The same happens with:
- databases
- encryption
- operating systems
- frameworks
- programming languages
I've seen this so many times by now it stopped being funny.
As for Google's 'Falcon' project: it smacks of NIH to me, but maybe their use cases are specific enough that none of the off-the-shelf bits were usable.
But none of these are things you can buy today. Well, there's InfiniBand, but if you're wed to Ethernet..
This works for them at their scale with quite a lot of internal tooling specifically to make it work.
However many people defend it as _the_ solution whereas it is a way with trade offs like most solutions.
Granted, Facebook have written their own VCS and Microsoft heavily modified Git to make it usable with monorepos (but only on Windows).
Unfortunately stock Git is bad at multirepos and monorepos. When you have "hundreds of people working full time on a project" scale, stock Git doesn't have a good answer.
The only thing blocking these from becoming standard is that it means userland has direct control of hardware.
It's possible for different vendors to be on different points on the wheel of reincarnation at the same time. https://www.computerhope.com/jargon/w/wor.htm
It's just not realistic to take all the switches, routers, and other garbage you've got in between points in the network off the rack/ceiling/wall/pole because the hardware can't support some protocol.
Good evidence for this is the rollout of fiber, which has been happening neighborhood by neighborhood and house by house for a decade.
E.g. instead of your little server setting firewall rules locally, it tells the router what traffic to allow. That router, in turn, tells upstream about its needs, and so on. Or a server reports its load, and the routers do active load balancing. The hardware wires are still there, as always.
Re. hardware acceleration, I think the earliest form of this was moving the checksum computation [1] from the CPU to the network device, even though the networking device didn't really know about the protocol it was doing the checksum for.
In both, it's just about parallelizing workloads, the same we do with microservices at the upper levels of the stack. Natural progression of distributed systems, with fancy names attached.
[1] https://en.wikipedia.org/wiki/Transmission_Control_Protocol#...
Their SDN implementations are also hardware accelerated.
I may be in over my head since I’m not an HPC/datacenter expert, but not sure I understand how you’d use this on the software side. Maybe someone is aware of specific examples? (beyond the vague “HPC/AI”)
edit: as another comment mentioned, the diagram shows it’s on top of UDP/IP, so it’s mostly an alternative to TCP/IP
> Fine-grained hardware-assisted round-trip time (RTT) measurements with flexible, per-flow hardware-enforced traffic shaping, and fast and accurate packet retransmissions, are combined with multipath-capable and PSP-encrypted Falcon connections ... flexible ordering semantics and graceful error handling ... hardware and software are co-designed to work together to help achieve the desired attributes of high message rate, low latency, and high bandwidth
So like QUIC, but designed for low latency. Maybe. There is no indication of how they achieve it if it is, nor is there a link to further details. The bulk of the article is literally name dropping. Protocol names, FAANG company names, standards organisation names. It reads like C-suite bait. "Come join us boys - all the big guys already have. So it's a sure winner."
For an example of what this means, try setting your MTU above the limit, and watch the raw traffic.
* - it's been years since I cared about the formal definition, my apologies if I got it wrong.
A “lossy” protocol is one that doesn’t attempt to compensate for that. In most cases, but not all, that means a higher level protocol will need to ensure that every bit of data has made it through. (An example protocol that might not care is one for watching broadcast TV on the Internet… if you miss a few seconds it’s not a big deal).
The IT ecosystem has fragmented into mutually incompatible cliques. You are either in the Google ecosystem, the Amazon ecosystem, or some other one, but there are no more truly open and industry-wide standards.
Look at WebAuthN: it enables a mobile device from "any" vendor to sign on to web pages without a password. Great! Can I transfer secrets from an Apple iPhone to a Google Android phone? Yes? No? Hello? Anyone there?
I just got a new camera. It can take HDR still images, which look astonishingly good. Can I send that to an Apple device? Sure! Can I send it to a Google device? Err... not without transcoding it first... on a Microsoft Windows box. Can I send it to a mailing list of people with mixed-vendor devices? Ha-ha... no.
This is the best argument I've seen for splitting up the FAANGs + Microsoft + NVIDIA. Once they get to this behemoth trillion-dollar scale, they become nations onto themselves and no longer need to cooperate, no longer need to use any open standards at all, and can start dictating and pushing third parties around.
Another random example is HTTP/3, which is basically the "What's best for Google" protocol.
Or gRPC, which is "What Google needs in their data centre".
And now Falcon, which is "The transport Google needs for their workloads".
Does it work for anyone else? I don't know, but it's a certainty that Google doesn't care and never will, because they don't need to.
This is one excellent example of the reason that increased/renewed anti-trust actions by the FTC are necessary.
The industry has always been this way. Back in the day there were many many processor ISAs that are now consolidated. There was been many networking standards that consolidated (IPX/SPX anyone?) New things often diverge because of new requirements not out of spite. There is a push and pull between standardization and innovation. Doesn't make it particularly unhealthy unless you can point to specific metrics and compare trends throughout the long arc of time.
For explicitness Google uses stubby which shares a lot of interface level commonality with gRPC but there's differences at a runtime level. Nobody is slinging json or soap around Google data centers.
I suppose when some people advocate for "standards" they just mean the shitty systems prevalent in the industry they happen to have been used to.
Japanese Bullet Train technology
HTTP/2 is that. The version 3 is... quite great, and not really originary from there.
And gRPC is some self-contained thing that doesn't bother anybody.
I do really agree with your point. But those examples are bad, and yet another non-standard cloud service isn't also a good example.
What strings you ask?
The top hits Google captured from their data centre egress.
E.g.: https://www.ietf.org/archive/id/draft-ietf-quic-qpack-20.htm...
Notice that "content-encoding" includes br (Brotli), a Google compression algorithm that essentially only they were using at the time.
Very soon, we will have providers/companies/champions/fighters that keep building the middleware transports to connect their behemoths.
Btw, someone somewhere popped up a better term — AGAMEMNON (Apple, Google, Amazon, Microsoft, Ebay, Meta, Nvidia, OpenAI, Netflix)
They haven't said much lately about sending content updates to their CDN nodes, but I think the throughput requirements on that isn't as high.
Edit: It’s yet another meta protocol built on top of TCP/UDP.
For example, people sometimes refer to TCP as a "Layer 4" protocol even though (a) TCP predates the invention of Layer 4 and (b) TCP is a square peg that does not exactly fit into the round hole that is Layer 4.
Imagine how bad the rest of it was.
I think the idea is that Falcon assumes the underlying network is semi-crappy and works around that (e.g. Falcon assumes that packets arrive out of order then it puts them back in order).
What are you basing that on? Nothing in the article implies that as far as I can see.
The diagram is confusing since it upside down to layering direction, but the article is clear. Falcon is a hardware transport protocol, replacing Ethernet. Like Ethernet, it runs on top of physical transport like RDMA. And IP runs on top of Falcon and Ethernet.
It's not. For example, RDMA can run on top of TCP (iWarp). NVMe can also run on TCP. Now replace TCP with Falcon.
The point is: RDMA is just a term for a computer reading memory on another computer without involving the operating system. The interconnect hardware and physical transport must support RDMA but that doesn't mean it is RDMA, as there are several different implementations of RDMA, and each kind of hardware supports a different subset of those implementations.