Maybe we'll learn something new from the Hot Chips event mentioned in TFA.
Maybe we'll learn something new from the Hot Chips event mentioned in TFA.
New chips are always fun to see and get discussed, but there are simple economic / mass design issues at play here. The Tesla custom chip just won't be made at the volume of NVidia or AMD.
And Intel's attempt to join the market seems like the better approach (even if it fails). You need to get the video game crowd to fund the research and development these days.
Video gamers want high speed SIMD for faster graphics. Without this source of money (and millions of chips sold per year), it looks difficult to sustainably upgrade an architecture.
--------
If not the PC market, then at least try to get the console market or phone market or some other consumer electronics fad wrapped up in things.
I don't think Tesla ever aimed to produce a chip for the public, it's purely to be used in their own products, so they don't have to be competitive with NVIDIA in flops/$ or worry about keeping up with NVIDIA production rate.
The chip only needs to be better, for the specific task they care about, for a specific metric they care about.
Considering that NVIDIA has to make a card that works for everyone, they have to make a huge number of tradeoffs, such as supporting many different number formats going from FP64 to INT4, their cores have to support a ton of different instructions, they need to keep a large amount of space on the chip for large memory buffers to support training of models with billions of parameters, they need to support GPUs interconnect for clustering etc. etc.
Given the importance of real estate on a chip, those tradeoff have large implications.
Tesla has 10x less constraints:
* For the training chip they only support their custom 8 and 16 bits format iic, so they can redistribute cores in consequence. They can ditch many instructions that they'll never use. They can size the memories to be exactly what they need, which is probably less than 80GB needed for large transformers, and use the freed space to add more cores etc.
* For the inference chip, they can additionally ditch all the interconnects as they just need one per car and reduce memory drastically freeing up even more space for cores.
If NVIDIA didn't have to balance so many different use cases on such a small area, then yes it would probably make no sense to try to do your own thing. But since they have to be general purpose, if you target a very specific use case you can be far less knowledgeable than them in GPU architecture and still come out ahead for your use case.
Specificity is the same reason why it still possible to write CUDA kernels for some basic operation that will run 3x faster than the official NVIDIA implementation that was done by some amazing engineers: they have to make their operations efficient for every input size, you don't.
-----
EDIT: And ~7000 chips seems too small for a custom-run of chips. Any "supercomputer" Tesla builds has to be much much larger than this before a custom-chip project likely becomes cost-efficient.
It takes so much time to build that the D1 chip is seemingly going to be obsolete before the supercomputer is even built. Obviously, the D2 needs to be announced soon to remain competitive. But another run of chips is even a few hundred-million bucks down the drain again.
At some point, you have to wonder if the custom chip was worth it at all. A run on a high-end process like 7nm isn't cheap at all, and being forced to make another run at 5nm or 4nm or 3nm as technology improves only shows how difficult the hardware business is.
Will Tesla be able to deliver its chips internally before later NVidia chips (like Hopper) get delivered? Or whatever other AI companies are making chips?
Not paying large margin to Nvidia might also matters on the economics.
Sometimes the first version of something isn't strictly speaking worth it but its still important for the company and what it wants to do.
Its certainty risky, but Tesla has just decided that they are vertical integration is important for them. So far this strategy is working for them but any individual project can of course not work.