Is it time to replace TCP in data centers?
blog.ipspace.net
blog.ipspace.net
For those who don't know Ousterhout created the first log structured filesystem and created TCL (used heavily in hardware verification but also forming the backbone of the some of the first large web servers: aolserver). I was actually surprised to find out he co-founded a company with the current CTO of Cloudflare. https://en.wikipedia.org/wiki/John_Ousterhout
He has both a candidate replacement as well as benchmarks showing 40% performance improvements with grpc on homa compared to grpc on TCP. https://github.com/PlatformLab/grpc_homa
With that in mind I think nobody will replace TCP and I doubt anything not IP compatible will be able to get off the ground. His argument is essentially that for low latency RPC protocols TCP is a bad choice.
We've already seen people build a number of similar systems on UDP including HTTP replacements that have delivered value for clients doing lots of parallel requests on the WAN.
I think many big tech companies are essentially already choosing to bypass TCP. I recall facebook doing alot of work with memcache on udp. I can't find any public docs on whether or not Google's internal RPC uses TCP.
I wouldn't be surprised at all if in the near future something like grpc/capnproto/twirp/etc had an in-datacenter TCP-free fast path. It would be cool if it was built on Homa.
Presentation https://youtu.be/bmSAYlu0NcY
But if I were to make a list of actual reasons too much money is getting burned on AWS bills, I suspect TCP inefficiencies wouldn't make the top 10. And worse, some of the things that do make the list, like bad managers, rushed schedules, under-trained developers, changing technical fashions, and "enterprise" development culture, all are huge barriers to replacing something so deep in the stack.
I am willing to wager that most organizations’ microservices spend the majority of their CPU usage doing serialization and deserialization of JSON.
Being able to use tools like tcpdump to debug applications is important to fast problem resolving.
Unless everything you do is "Web scale" and development costs are insignificant, simple paradims like stream oriented text protocols will have their place.
Bizarrely, this take just a couple of seconds to transfer over 10 GbE so many devs simply don’t notice or chalk it up to “needing more capacity”.
Yes, yes, it’s the stingy sysops hoarding the precious compute that’s to blame…
We could track expected response size, but then every feature launch triggers a bunch of alerts which either causes the expenditure of social capital, or results in alert fatigue which causes us to miss real problems, or both.
This is a place where telemetry does particularly well. I don't need to be vigilant to regressions every CD cycle, every day, or even every week. Most times if I catch a problem in a couple of weeks, and I can trace it back to the source, that's sufficient to keep the wheels on.
I'd say in majority of cases the service was made "too small".
If you waste majority of CPU to serialize/deserialize/send to network you should probably just "do the job" right there and then (aside from loadbalancers and such for obvious reasons)
For your typical website backend between frontend and db, you are doing some async conversion of JSON to db call back to JSON. For an HTTP microservice you are also typically converting some JSON request body to a JSON response body with some kind of I/O call in between.
So that’s a roundabout way of saying I think that the case where majority of CPU is spent on SerDe is more common than you think. And it’s not necessarily a problem if the effort to improve is not worth the savings.
I work on Pyroscope, which is a continuous profiling platform and so I see a lot of profiles from various organizations.
If you want to save the world some CPU cycles I would look into optimizing deserialization. And it’s not just JSON, binary formats like protobuf are not much better.
It comes down to the overhead associated with allocation and tracking (GC) of many many small objects which is unfortunately very common in modern systems.
Even if you're naive and don't care about performance - which is a common sentiment for modern developers who have spent the last decade working for companies where the cost of AWS didn't matter - chains of transformations like this are a good place to switch to a format, any format, less atrociously expensive than JSON.
The example above, btw, comes from a very large unicorn that burns _2 complete cores per outstanding request_on a continuous basis_. To someone who lived in the dotcom, that's so outrageous it's comical, and of course they have years of negative cashflow because of their insane costs.
Anyway, for people who pick relatively more sane application languages, yeah deserialization is pretty much all their CPU does. It’s just such a shame because it really is a godawful format, just like HTTP/1, its basically only benefit is that it’s easily human-readable.
HTTP isn't low latency. Spending all that effort porting everything to a discreet message based network, only to have HTTP semantics running over the top is a massive own goal.
as for WAN, thats a whole different kettle of fish. You need to deal with huge latencies, whole integer percentage of packet loss, and lots of other wierd shit.
In that instance you need to create a protocol to tailor to your needs. Are you after bandwidth efficiency, or raw speed? or do you need to minimise latency? All three things require completely different layouts and tradeoffs.
I used to work for a company that shipped TBs from NZ & Aus to LA via london. We used to use aspera, but thats expensive. So we made a tool that used hundreds of TCP connections to max out our network link.
For specialist applications, I can see a need for something like homa. But for 95% of datacenter traffic, its just not worth the effort.
That paper is also fundamentally flawed, as the OP has rightly pointed out. Just because Ousterhout is clever, doesn't make him right.
The key thing is that HTTP is not designed to be low latency. If latency is important to you, then you need to make your protocol bespoke for that application.
HTTP/3 has a much better TLS connection process, which means that the cost of creating a connection is much lower. QUIC is much more configurable in terms of per steam or connection flowcontrol.
Yes, homa has request/reply semantics, but not in the way that HTTP needs. Homa is optimised for low latency small messages, in a small hop, transparent network. HTTP is a file access protocol with a general data channel hammered in. Sure it'll work in a DC type network, but its also got to deal with lossy, high latency networks.
Within the data center? AWS uses SRD with custom built network cards [0]. I'd be surprised if Microsoft, Facebook, and Google aren't doing something similar.
[0] https://ieeexplore.ieee.org/document/9167399 / https://archive.is/qZGdC
Edit: seriously though, does anyone know what they are up to? We get like 30 megabits from cache to CI.
I've been assuming eventually gRPC over HTTP3[1] would get some traction from big camps. But if Google is using their own internal transport, it seems highly unlikely that these main authors of gRPC are ever going to care.
This blog post talks about how many of the features of HOMA are already present in existing specs like QUIC/HTTP3. I'd expect many of the wins in Ousterhout's HOMA benchmarks could probably be replicated elsewhere, with better transports.
In short, the browser failed to make basic modern http usable by anyone so there's no hope of getting off the janky ugly grpc-web trashfire & making grpc from the client just work, like it always should have been able to do.
Pathetic pathetic showing by the browsers here. Beyond neglect, how they mishandled http2 and http2 and gave webdevs access to almost none of what the most exciting new thing on the www was. Huge huge miss. Incredible bullshit.
Some day we'll be running http3 over webtransport & fixing their wrongs. Thia is such a bullshit long workaround, such a pathetic way to handle the browser getting nowhere with new http standards, end running around their utte unmoving complete inability to advance at all. But we'll start to actually use http intensively again, soon, in spite of the browser & standards community being such ridiculous & farcical impedances against using http2 and http3 at all. What a tragic shit show of uselessness it's been, trying ro actually enjoy the new http standards; resistance on all fronts to real usage.
grpc-web will require very little work to adopt http/3 and the benefit is pretty obvious. The hard part will be if there’s http/1 servers in the way that have an impedance mismatch. Still, I don’t see why grpc which only requires a JS library and matching deployed server which is all open source and doesn’t really require buyin from multiple stakeholders will struggle here.
Your emotional reaction to this seems out of place as there’s no grand conspiracy here. It seemed like a plausible idea. It just never panned out enough to be worth it. Certainly the server push as an API wouldn’t buy you that much in terms of performance because you can emulate it via long poll / websockets unless I’m missing something?
There was a couple very brief moments during fetch's addition of progress where push got a brief bit of attention from big enough names that it seemed like maybe after years there's be some real chance of using http2 push interestingly on the web, but that moment flickered out & died. In general there has been a callous treatment for Push, with blinders on, thinking only of tbe narrowest desires & uses. It's been a completely squandered technological capability that was never opened for use in any interesting form, and it's a shame this sad small vision of http2 push & it's so called failure obstructs us from even considering how many more interesting & powerful uses it could have had.
My understanding is grpc requires push and trailers support, and that the browser still has no designs on offering either capability to developers for use. It's been some years since I've looked, but http3 in the browser once again seems to give developers absolutely no new capabilities, even though http added new stuff like Push & Trailers nearly a decade ago.
My emotional response is because there is such a small & narrow vision, choking how we might be using http & growing the web. Techniques like long-poll & websocket exist, but there's such a clearer better fitting match for sending http resources as they are generated, Push. The lack of browser exposure of new (decade old) http capabilities is pushing us towards a stupid point where we end up running http3-over-webtransport, and it's absurd & enmisersting to see such a lack of follow-through in the deepest most core heart of the web being given a chance to get used, to do the amazing things it could be doing... as opposed to radically non-web non-resourceful hacks like websockets. This harkens back to the HyBi mailing list, and the sad inability for the web & our exhange of resources to be more bidirectional & asynchronous, and that's not a decade od stagnation, it's now two decades of stagnation, stagnation that we almost got a chance to improve past, were it not for the sad silly limited pretense that Early Hints gives us even a thousandth of what Push gave us in terms of capabilities.
I don't know if Homa has multihoming but QUIC and SCTP have and if Homa doesn't then I think it is a huge step backwards.
Netdev 0x16 - Keynote: It's time to replace TCP in the datacenter: https://www.youtube.com/watch?v=o2HBHckrdQc
Netdev is an amazing conference btw, it is insane that we get access to the trailblazing being shown off at conferences for free on places like YT.
The TCP/IP stack works because I don't need to care about what environment the two processes that need to communicate are running in. They could both be on my local machine, or in my home network, or communicating over the internet, or some random intranet, in a data center, across continents, on any OS or any kind of device...it simply does not matter. "Just do a ground-up rewrite of your entire software stack and you'll get a guaranteed 5% efficiency gain" isn't the bullet proof argument that people who come up with these alternatives seem to think.
So, while the actual RPC protocol has its own issues, the one thing they got right was the ability for the portmapper to indicate UDP vs TCP as the transport on a service basis. There have been a few improved generic RPC mechanisms, and I don't really understand why some of these places feel the need to "replace TCP" when really what they need is a more formalized RPC mechanism that can set/detect the datagram reliability and pick varying levels of protocol retry/etc as needed.
The point of these abstractions is that they are insurance. We pay taxes on best case scenarios all the time in order to avoid or clamp worst case scenarios. When industries start chasing that last 5% by gambling on removing resiliency, that usually ends poorly for the rest of us. See also train lines in the US.
Because datacenter is not single LAN segment ? You need to route it
How? Datacenters aren't a single link, you need to route packets.
> The point of these abstractions is that they are insurance. We pay taxes on best case scenarios all the time in order to avoid or clamp worst case scenarios. When industries start chasing that last 5% by gambling on removing resiliency, that usually ends poorly for the rest of us.
I'm not sure what this means. There are other transport level protocols already, UDP is in fairly regular use. Is your argument that TCP offers us some insurance that Homa will not?
I'd suspect that the tcp bits are hidden in the rpc layer anyway, be it grpc or whatever
Additional M.O. protocol cool, but replace TCP with it?
I wouldn't have any trouble ignoring middlebox software (or adding a TCP fallback with a big warning that something suspicious is interfering with the connection) but Windows and macOS still lack proper SCTP support, Linux' SCTP support has some performance issues and usermode raw sockets will probably need to bypass several OS sandboxes to be viable.
That said, in server to server connections SCTP can probably be used just fine.
IPv6 definitely got off on a rocky start but none of that has mattered for decades now. If IPv6 deployment is still such a pain, that's because of either your choices of the choices of your preferred vendors.
Until you remember that you have datacenters that need to talk to each other and then this strategy doesn't work.
And DC applications are far more easy to switch over than general internet, less middleboxes to screw you over, generally better performing networks, and more tightly controlled hosts.
Honestly, it wouldn't be all that farfetched to have AWS implement QUIC over SRD for squeezing the last perf drops out.
Most applications don't care about TCP any more than they care about IP.
This new protocol the sort of thing Osterholt is known for. He came up with "log" file systems, which are optimized for the case where disk writes outnumber reads. Reading then requires more seeks on mechanical drives. If you're mostly writing logs that are seldom read, it's a win. If you're doing more reading than writing, it's a lose.
Most of the gains of QUIC are in exchanging data with the first round trip, which you can do with TCP fast open as long as middleboxes don't mess it up; different congestion control, which you could do in TCP if you control the OS on the sending side as long as the middle doesn't mess it up, which shouldn't happen in a data center; and maybe some benefits for connection switching, which should hopefully not be necessary in a data center (although I've certainly seen cases it would help!); multi-path TCP also addresses this, but has probably worse CPU efficiency problems: if a logical flow is coming in on multiple NIC queues, there's going to be cross-cpu communication on the data structures.
Unless it's changed, QUIC also simply won't handshake if the path MTU is too small, so you don't get to have the excitement of path MTU discovery failures; but you really shouldn't have that problem in a datacenter.
You can make the server think they are using a perfect TCP connection even if the protocol is different. And the opposite, the server/app can use a new protocol like json-stream-whatever and the smartnic can proxy it onto a TCP session.
It turns out that cheap single purpose NICs were great for standardizing on thr lowest common denominator (tcp/udp) but now that it’s clear that the network protocol is the bottleneck, we can imagine and implement new protocols while maintaining backward compatibility.
I wouldn't go this route because by the time the data gets to the DPU you have already burned the cycles and paid the price of TCP. Offloading anything is tricky in general because the protocol for managing the offload needs to be leaner than the thing being offloaded.
But otherwise, yes, it addresses many of these problems.
I even remember that there have been Linux kernel security vulnerabilities in those protocol somehow exploitable even when you don't actively use them, so... I don't compile my own kernels often these days, but when I do, I usually say N to any and all protocols I'm not likely to be using in my remaining life.
In the worst case, you just build on top of UDP and don't care that you're wasting four bytes.
1) It's not TCP
2) It's a PITA on a 'trusted' network because there is not "encryption off" mode
For internal datacentre traffic, it depends on the value of the data. But you will need to do application level encryption.
For links between datacentres, encrypting links based on rules is a sensible and not horrifically challenging thing to do. If you're big enough to have to worry about the volume of encrypted traffic, then you have rules based bandwidth priorities.
for virtually everyone else, just use wireguard or some other VPN with encryption.
- Bandwidth sharing (“fair” scheduling) - Sender-driven congestion control
See: https://arxiv.org/pdf/2210.00714.pdf (pdf download)
Why did it not work out?
I’d also love to hear more about the concept.
Software won, it seems.
This was about the time that grid engines promised mainframe computing without the cost. Only that it required people to understand how to parallelise their workloads. That and modify programmes to run in on many machines at once.
1) cost
2) software support
3) cost
* Infiniband hardware didn't get cheaper as fast as RAM + disk
* Infiniband hardware didn't get faster as quickly as locally-attached RAM + disk
With costs so high, adoption is reduced.
The vendors would've been smart to give away one part of it.
Our prod environment doesn't use it because millions of servers would be cost-prohibitive. Corp uses it here and there for experimental clusters.
The OSI model isn't, and has never been, the be-all end-all of networking models, and it's bizarre to me that people continue to believe that the OSI is the only valid model to follow.
Maybe you can explain why you believe that the OSI model is canonical?
Probably not? Maybe? One would have to either integrate support for new protocol into every server, switch, router, IoT oh and of course all the 3rd party clouds and clients one is speaking to. OR everything leaving the datacenter would have to be dual-stack and/or funneled through some WAN optimizer/gateway/proxy device that can translate new protocol into TCP/IP creating a single point of success bottle-neck. Dual-stack brings up some security issues that this protocol will have to address not to mention more cabling complexity.
I think the best place to start this conversation would be with architects at Microsoft, IBM/Redhat, maybe even Meta since they acquired several kernel developers and let them see how cost effective the gains are. If the big players buy into this and they have kernel developers that can integrate seamless support into Linux and Microsoft to start with and a few big players try it out then maybe it would share the same market cap as Infiniband. Let them deploy a proof of concept pod for free for a year and see what they can do with it. I think this would have to be successful first before other vendors start adding support for new protocols. If the plan is to have one vendor to rule them all then it will not succeed as there would be no competition and it would be too expensive for mass adoption. This would end up being another proprietary thing that IBM or some other big company acquires and sits on.
At least that is my opinion based on my experience deploying SAN switches, proprietary memory inter-connect buses, proprietary storage and storage clustering, proprietary mini-mainframes and server clusters. Speed improvements will impress technical people but businesses ultimately look into TCO/ROI and reliability. Complexity is a factor in reliability and the ability to hire people to support said new thing. If anything I have seen datacenters going the opposite direction; that is, keeping things as generic as possible and using open source solutions to scale first to their vertical limits and then horizontally. Ceph is a great example of this.
linux will even let you do it from userspace (with net_admin cap)
there are some exceptions: like if your endpoint is some a crappy cloud provider (e.g. azure) that provides you something that looks like an ethernet network but really is a glorified TCP/unicast UDP proxy
[1]: ignoring NAT and IGMP/MLD snooping (not that anyone does those outside of internal networks... right?)
not perfect but not unworkable either
Had Infiniband been developed in a way that avoided all of the patent encumbrances it would have replaced IP networking in large data centers (certainly at the Rack level, and likely at the Cluster level as well).