Tesla bulks up its GPU-powered AI super
hpcwire.com
hpcwire.com
How is an A100 worth "only" 17 tera-flops?
There are multiple home gpus and even mobile chips rated at tera-flops. And the A100 really is a gigantic chip.
I realize all about bandwidth, and full vs half precision, and inflated marketing vs actual benchmark.
But can a 10-node-iphone-cluster really match an A100 for pure flops?!
https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Cent...
Only if the data is local, small, easily synchronizable, doesn't require fp32 precision and a dozen other caveats.
Can a 10 iphone cluster match an A100 in calculating digits of pi? Yes. Can it do backpropagation in large machine learning models? The gradients won't even fit.
It's optimized for TF32, INT8, bfloat16, and other similar data types that the tensor cores like.
Double precision is important for a lot of simulation codes because of numerical issues. For Deep learning applications it's basically unused, single precision or half precision formats dominate training. The same numerical issues don't show up, so this is a lot more efficient.
Additionally a lot of the 'AI' flops are achieved through matrix cores, fitting Deep learning the workload well. This brings an A100 from only 17 to ~200 tflops.
> But can a 10-node-iphone-cluster really match an A100 for pure flops?!
The iphone number is probably fp32 if not lower precision. However the combined SoCs of 10 iphones have tripple the transistor count and a larger chip area among them. So the comparison is not totally off.
Typically these are things built by national laboratories, university research labs...
We also learn stuff such as ATT developing video phones since the 1960's but they were too early and it never go anywhere until decades later when ATT was a lesser relevant player in the telecom industry. And the Post Office foresaw the coming trend of online bill pay and email eating into their mail volume since the early 90's, yet when their they had projects that would enable senders and receivers to received mail scanned to a USPS account or pay bank statements through a USPS account for those early bill pay adopters, Postal Management did nothing with it.
If all of my local pizza parlours were installing (well, renting) AI/ML rigs, it would be a different story.
1: https://www.top500.org/statistics/list/ (select the "segments" category)
In fact if you look through the historic top500 lists a lot of the space outside the top10 is occupied by industry clusters. Classically oil and gas has been one of the big consumers. In recent years there are more and more cloud clusters in there as well.
So this Tesla cluster isn't really anything unusual.
There are some companies that don't have centralized high-performance compute resources so each department goes out and builds or buys and their own performance compute resources and strings it along. Many universities do this too where physics has a compute cluster and economics and bio-sciences have another cluster and sometimes they are underpowered and if all the departments grouped together and pooled their resources they would have adept system administrators and developers and hardware to support their computing needs but it doesn't always happen with poor governance so they end up with stuff that appears to be built from someone's basement and when that person leaves, retires, graduates there's so much poor documentation none know what to make of anything.
Why would they? So their competitors can copy their business plan?
IBM has had large clusters for decades.
Any large bank or trading company would have clusters that would be at home in that list.
You don't want to know what the Apple cluster is like: they have to train for Siri, in multiple languages, etc. I be that's several orders of magnitude larger than Tesla.
AWS itself is probably several orders larger than most university clusters, probably larger than most country's combined computational capacity.
Meanwhile, FSD is installed on only 7% of Tesla's cars, down from over 50% in 2019. https://www.torquenews.com/1083/teslas-low-fsd-take-rate-off...
Thus a 70 petaflop ultracomputer is quite a big investment to support an increasingly unpopular car option that's inessential and mostly promotional, ergo, a geek tickler.
Tesla is well known to be one of the most vertically integrated companies in the industry. The are well known to use AI for their factories as well. The run lots of simulations for crash testing as well.
What products are they are buying? They do probably buy somethings but unless you have evidence that they buy most things, I not gone believe you without some kind of source.
> Meanwhile, FSD is installed on only 7% of Tesla's cars
You realize it can be bought in later as well right? If quality improves lots of people have to option to buy it.
> that's inessential and mostly promotional, ergo, a geek tickler.
That's your opinion, an opinion that Tesla clearly doesn't share.
Original site I found the video at: https://www.supplychaintoday.com/tesla-artificial-intelligen....
More info: https://www.google.com/search?q=tesla+using+ai+in+factorys&r...
They are obviously not doing this (or not doing it well) as there are several spots on the state highway I repeatedly drive through where my car always slows down from 100KMH to 50KMH for and I have to repeatedly override with the accelerator to stop the person behind going in to me.
Maybe we'll learn something new from the Hot Chips event mentioned in TFA.
New chips are always fun to see and get discussed, but there are simple economic / mass design issues at play here. The Tesla custom chip just won't be made at the volume of NVidia or AMD.
And Intel's attempt to join the market seems like the better approach (even if it fails). You need to get the video game crowd to fund the research and development these days.
Video gamers want high speed SIMD for faster graphics. Without this source of money (and millions of chips sold per year), it looks difficult to sustainably upgrade an architecture.
--------
If not the PC market, then at least try to get the console market or phone market or some other consumer electronics fad wrapped up in things.
I don't think Tesla ever aimed to produce a chip for the public, it's purely to be used in their own products, so they don't have to be competitive with NVIDIA in flops/$ or worry about keeping up with NVIDIA production rate.
The chip only needs to be better, for the specific task they care about, for a specific metric they care about.
Considering that NVIDIA has to make a card that works for everyone, they have to make a huge number of tradeoffs, such as supporting many different number formats going from FP64 to INT4, their cores have to support a ton of different instructions, they need to keep a large amount of space on the chip for large memory buffers to support training of models with billions of parameters, they need to support GPUs interconnect for clustering etc. etc.
Given the importance of real estate on a chip, those tradeoff have large implications.
Tesla has 10x less constraints:
* For the training chip they only support their custom 8 and 16 bits format iic, so they can redistribute cores in consequence. They can ditch many instructions that they'll never use. They can size the memories to be exactly what they need, which is probably less than 80GB needed for large transformers, and use the freed space to add more cores etc.
* For the inference chip, they can additionally ditch all the interconnects as they just need one per car and reduce memory drastically freeing up even more space for cores.
If NVIDIA didn't have to balance so many different use cases on such a small area, then yes it would probably make no sense to try to do your own thing. But since they have to be general purpose, if you target a very specific use case you can be far less knowledgeable than them in GPU architecture and still come out ahead for your use case.
Specificity is the same reason why it still possible to write CUDA kernels for some basic operation that will run 3x faster than the official NVIDIA implementation that was done by some amazing engineers: they have to make their operations efficient for every input size, you don't.
-----
EDIT: And ~7000 chips seems too small for a custom-run of chips. Any "supercomputer" Tesla builds has to be much much larger than this before a custom-chip project likely becomes cost-efficient.
It takes so much time to build that the D1 chip is seemingly going to be obsolete before the supercomputer is even built. Obviously, the D2 needs to be announced soon to remain competitive. But another run of chips is even a few hundred-million bucks down the drain again.
At some point, you have to wonder if the custom chip was worth it at all. A run on a high-end process like 7nm isn't cheap at all, and being forced to make another run at 5nm or 4nm or 3nm as technology improves only shows how difficult the hardware business is.
Will Tesla be able to deliver its chips internally before later NVidia chips (like Hopper) get delivered? Or whatever other AI companies are making chips?
Not paying large margin to Nvidia might also matters on the economics.
Sometimes the first version of something isn't strictly speaking worth it but its still important for the company and what it wants to do.
Its certainty risky, but Tesla has just decided that they are vertical integration is important for them. So far this strategy is working for them but any individual project can of course not work.
[0] https://www.anandtech.com/show/16792/nvidia-unveils-pcie-ver...
I think a better question is: could the heat be re-used, for example, to warm water in bathrooms? Capturing waste heat from computing centers with heat pumps seems like a very doable approach and a practice that is both responsible and profitable.
Do you always take internet videos made by literal competitors over the state licensed regulators?
1. No clear indication that it was even on FSD. Only TACC was probably engaged.
2. Video was posted in 480p so it's hard to see what's on the screen.
2. YouTube Comments were turned off.
3. Bottom half of the screen was cut off on the initial video (most glaring omission IMO). Which would show if the warning that the accelerator pedal was being held down and that the vehicle would not stop. It looks like this: https://i.imgur.com/G3l9Rp1.png
4. Additional footage that was lasted released (also in 480p it's 2022, why?) CONFIRMED this: https://i.imgur.com/1PMbG9f.png
These are just a few examples. But it's hard to believe that the video was not made in good faith.
Yay, let's further our efforts into hooking people and showing them ads! /s