Tesla’s TTPoE at Hot Chips 2024: Replacing TCP for Low Latency Applications
chipsandcheese.com
chipsandcheese.com
With FPGAs using commodity SFP or Ethernet PHY's you can certainly build stuff that runs circles around traditional Ethernet and associated overhead from protocols like IP.
As an afterthought, they publish this for marketing/engineering pull for people who like to optimize (do engineering) for specific situations, while supporting the ROI and keeping cost down.
OP's TCP offload engines do not require any retooling at all, if you consider stuff like buying IP switches not retooling. All you need to do is basically buy a network card.
Also, if you read the article you'll eventually stumble upon the bar chart where they claim that the one-way write latency of TTPoE is only 0.7E-6 seconds faster than existing tech like infiniband, and it's also the only somewhat hard data they show. Does that justify the investment of developing their whole software+hardware stack?
I'm sure the project was fun and should look great on a CV, but overall it doesn't look like it passes the smell test.
(Disclaimer: I work at Tesla, not related to this group, opinions on public info only)
How many ≥10 Gbps chipsets that you'd find in a typical server do not have offload nowadays?
Further, once you're in the ≥50 Gbps card range you can often get ROCE, which helps with things like latency.
Every system has a cost and tradeoffs. Just because someone took an unusual path doesn't mean that they were wrong. And the larger and more specialized their use case is, the less likely that a generic solution is the best match.
Tesla's own charts show ROCE achieving also one-way write latencies in the single-digit microsecond range. If that doesn't qualify as magic bullet, what does that say about TTPoE?
So, they need to compare it to Infiniband, not TCP, and definitely not software TCP. And they need to explain how/if it works with standard huge capacity switches (which is at least a reason to prefer TCP over Infiniband).
There could be reasons to build this, AWS have something, but for Tesla to build their own stinks of NIH bad.
they are purpose building hardware for their specific application. debugging corner cases and making this robust is going to take them a decade. given that nobody else is interested in this non-standard solution, they dont have the benefit of the community debugging it, and improving on it in open-source.
appears to me to be a vanity effort as is the whole dojo project.
Usually, this is line-rate, but if the other side is slow for whatever reason (say the consumer is not draining data), you wouldn't want the sender to continue sending data.
If you also have N hosts sending data to 1 host, you would need some way of distributing the bandwidth among the N hosts. That's another scenario where the credit system comes. Think of it as an admission control for packets so as to guarantee that no packets are lost. Congestion control is a looser form of admission control that tolerates lossy networks, by retransmitting packets should they be lost.
Or I guess even RoCE
Are there other technologies they could have used?
Also, the 80us is supposed to be the worst case, where typical is supposed to be <10us. Again not knowing anything about infiniband, what’s the typical perf? I tried to google but the people who are talking about it are in the know in ways I’m not.
Thanks!
It is definitely possible to go much lower than 80usec on Ethernet. But obviously it depends on the scale, utilisation etc.
At the sizes of GPU clusters we're talking about these days - 32K and up - things get tricky.
The main alternative to Infiniband used in the industry is RoCE - Meta has written a lot about it [0].
There's several reasons to avoid Infiniband, such as cost, availability, vendor lock in, lack of experience etc.
Those are some of the reasons why many players are trying hard to make Ethernet work, see Ultra Ethernet [1].
[0] https://engineering.fb.com/2024/08/05/data-center-engineerin...
IIRC, Arista started off focusing on the financial market with low latency.
There's fairly well regarded in a general sense nowadays (at least /r/networking often has folks recommending them as a vendor).
"Measuring the latency of a 4ns switch":
* https://www.arista.com/assets/data/pdf/Latency-4ns-Switch-So...
There's no analogue in the Infiniband world to dirt cheap 1GbE RJ-45 switches though.
And both price tags will make Elon's "someone's scamming me with a 'you're an enterprise customer' surcharge" sense tingle. The price tags for anything enterprise networking related are seriously inflated, and I would not be surprised if just making your own NICs and switches is cheaper once you hit a certain deployment size.
I'm having trouble feeding things at 400GB/s (not a typo, it's gigabyte/s) per H100 box.
For 10 boxes ideally you want 4TB/s...
> > an IB switch is in the same ballpark as an ethernet switch with the same port speed
> And both price tags will make Elon's "someone's scamming me with a 'you're an enterprise customer' surcharge" sense tingle.
In the context of Tesla doing their own protocol and not-high-end NICs.
The hardest things would've been the DDR4 and PCIE interface. But as they're using standard interfaces, and last generation. I'm sure they got a good discount on all that IP and it didn't cost them hardly any man hours to integrate. And Tesla might've even already had the licenses and IP setup as they make other ASICs.
I didn't do a budget or anything, but at even 10Ks of units, I could see how this could save money. Or at least not loose money. Assuming a comparable IB network card is ~$1000, which I also didn't price.
And there could be other potential cost offsetting features, like power savings.
I mean yeah, but thats why you have negotiators. List price is what suckers pay.
As soon as you start to buy in job lots, or the total price comes to >$500k then stuff becomes a lot cheaper all of a sudden (within reason)
Having said that Infiniband is an arse to deploy, but not as much as custom networking protocol on custom silicon.
> Ideas you’ll never hear at Google or meta
You'd be surprised. Google has a very strong tradition of "not-invented here" which extends to some of our production networking gear as well.
To be fair, at the time, some of this was justified because the available devices on the market couldn't support our use cases back then.
Per section 3.2 of the 2013 B4 paper [0]:
Even so, the main reason we chose to build our own hardware
was that no existing platform could support an SDN deployment,
i.e., one that could export low-level control over switch forwarding
behavior. Any extra costs from using custom switch hardware are
more than repaid by the efficiency gains available from supporting
novel services such as centralized TE.
https://cseweb.ucsd.edu/~vahdat/papers/b4-sigcomm13.pdfAdding to sibling comment about Google, Meta[1] built 2 large-scale production training clusters for science: one with Infiniband, the other one with a custom RDMA over RoCE fabric.
> Custom designing much of our own hardware, software, and network fabrics allows us to optimize the end-to-end experience for our AI researchers while ensuring our data centers operate efficiently.
> With this in mind, we built one cluster with a remote direct memory access (RDMA) over converged Ethernet (RoCE) network fabric solution based on the Arista 7800 with Wedge400 and Minipack2 OCP rack switches.
Google, Meta and Netflix are among the most obsessive on optimizing their infrastructure - it's bold to assume they haven't looked at their COTS network gear and thought "hmmm..."
1. https://engineering.fb.com/2024/03/12/data-center-engineerin...
There's also other differences, such as port counts. AFAICT Spectrum switches at 400Gbps have up to 128 ports whereas equivalent Infiniband NDR Quantum only have 64 [0].
When building clusters of 32K+ GPUs the network cost, power, transceivers etc start to add up.
[0] https://www.semianalysis.com/p/100000-h100-clusters-power-ne...
With the SN5600 for Ethernet (Spectrum-X), which is 64 physical ports, you're running each port at 8x100G-PAM4).
The graph at the end shows they measured (one way) latency at 1.3 microseconds (compared with 2.0 for IB).
But for AI training, where you're simply shuffling around large stacks of matrices, my guess is latency constraints weaken.
And supply chain independence. I've heard that some GPU clouds are delayed because their Infiniband hardware was delayed due to the Israel–Hamas war. Optimally you probably want to avoid critical hardware that's being manufactured in a high risk of disruption zone.
Infiniband suppliers charge crazy prices due to having little competition. It might actually be cheaper for them to design their own than to pay the Infiniband tax.
But now I am curious with the distribution of observed window sizes is in the wild.
Edit: I'd bet the simpler protocol is more vulnerable to various spoofing attacks though.
Edit2: Lol I hope the frame IDs are for illustrative purposes only - https://chipsandcheese.com/2024/08/27/teslas-ttpoe-at-hot-ch...
Such ideas are, however, worth revisiting when the workload is unique enough (in this case, it is), and the performance gains are so big enough...
Multiple parties communicate at the same time? Lower number priority electrically could pull the voltage low, dominating the transmission.
That way, priority messages always get through with no overhead or central communication required.
The technical issue is that you would need global arbitration to ensure that the _goodput_ (useful bytes delivered) is optimal. With training across 32k GPUs and more these days, global arbitration to ensure the correct packets are prioritised is going to be very difficult. If you are sending more traffic than the receiver's link capacity, packets _will_ get dropped, and it's suboptimal to transmit those dropped packets into the network as they waste link capacity elsewhere (upstream) within the network.
This is a protocol between compute nodes in a data center, it's layer 2 so there is no way to reach this over the internet.
But, point taken.
> We proceeded without DCQCN for our 400G deployments. At this time, we have had over a year of experience with just PFC for flow control, without any other transport-level congestion control. We have observed stable performance and lack of persistent congestion for training collectives.
https://engineering.fb.com/2024/08/05/data-center-engineerin...
[0] https://github.com/NousResearch/DisTrO/blob/main/A_Prelimina...
There are ICs you can buy off the shelf for electronic routing and switching of these interfaces.
Once they’ve built this cluster, maybe they can make money renting out compute or something.
Or Elon is using the resources of one company to do work for another company (?):
* https://en.wikipedia.org/wiki/XAI_(company)
* https://electrek.co/2024/04/03/elon-musk-xai-poaches-enginee...
I mean in TCP it's not allowed (Even though, super-theoretically, it's not completely forbidden) to carry a payload in the initial TCP SYN. If you're so latency-obsessed to create your own protocol, that's the first thing I'd address.
If you don’t need it all the time, why bother?
What's disappointing is that it's impossible to do a new protocol on the Internet because of all the middleware boxes that drop packets that aren't IMCP or TCP or UDP.
Google also does not depend on NVIDIA, thx to TPUs. Rents NVIDIA GPUs to external customers - sure, it's a nice side business, but internally TPUs are king and there's no dependency on NVIDIA for that.
On my side, I would like to point that the today HN thread ([1]) that discusses a paper GameNGen ([2]) that runs Doom with diffusion models was trained on TPUs.
I don't see a dependency on NVIDIA there.
If there's a more specific rebuttal to my original statement, please, don't hesitate to state it.
1. https://news.ycombinator.com/item?id=41375548
Deepmind says otherwise. training is most likley all on NVIDIA still. Same for Apple.
The difference is, nobody knows for sure with Apple, because they are a secrecy cult.