Tesla Dojo Custom AI Supercomputer at HC34
servethehome.com
servethehome.com
It still requires many experts, both to write the XLA to hardware translation, and ML engineers who know how to write TF python that executes quickly.
(note: Google has transitioned many projects to Jax, which also writes to XLA, as TF ended up being a bit of a pig with wings)
Can you say more about this?
But I think the writing was on the wall when the folks building ML Pathways hit some performance problems and realized that Jax made it much, much easier for them to express the computations they wanted and see them run quickly on TPUs (DeepMind had also concluded this). Once Jeff saw that Jax was making stuff that ran faster than TF (for his pet projects) the writing was on the wall.
I wouldn't be surprised if other approaches like genetic programming will make a comeback one day.
If there are any papers out there arguing for/against neural networks to stay in the king's seat forever, I would love to see them.
With "genetic programming" I was referring to applying genetic algorithms to code. So that the result of the evolution is a program that performs well doing a task like driving a car.
Nice research on this has been done by John R. Koza. But nothing new came out in this area for quite a while.
Why would they?
If you assume that your objective is kinda smooth genetic algorithms and related are guaranteed to suck. And most "real" world things of interest can be assumed to be pretty smooth
Your smoothness argument applies to the search approach. I would say "smooth" is not enough to describe the search space of real world solutions. I would rather call it fractal like. There are many smooth areas in a visualization of the Mandelbrot set. But the most interesting things happen in areas that are not smooth but follow some delicate logic. You find tiny local minima there that are hard to reach via gradient descent.
Genetic algorithms have a lot going for them in this type of search space. To pick one factor: GA's distribute the search power so that areas of the search space that are more promising get more resources. Similar to the optimal approach to deal with so called "multi-armed bandit" problems, GA's divide the search space in an optimal way, computation wise.
To achieve parity, GPs and GAs would need demonstrations on the level of what was seen about 10 years ago in NNs, then enough people would get involved to start doing more tech dev to make the hardware do GAs really, really, really fast.
A hosted Jupyter notebook in a sandboxed VM able to send jobs to this new silicon is something that might be possible to set up by a small team in a few months, and could turn into a billion dollar business.
As a bonus, Tesla can use revenue from that to grow this supercomputer, while using any spare/unsold capacity for themselves.
I think they said that during the first AI Day. Here's a 19 minute supercut of AI Day: https://www.youtube.com/watch?v=keWEE9FwS9o
But I think that was the wrong strategic move - they should have opened it up, together with some 'Tesla AI' demo models, in a colab environment. They can hire new employees to do that - it is separate work from that involved in making the self driving car, and will not block or interfere.
The only reason I think they might not is that they don't want to step on Googles toes - there are very close links between Musk companies and Google, and a direct competitor to Googles TPU product might hurt that relationship more than it generates in revenue.
The bottleneck isn't employees, it's chips.
Yet, building a basic colab-like interface, using opensource tools, shouldn't take more than a few people a few months.
Obviously, they might not want to launch that before their hardware is ready, but their hardware has been in production for over a year now.
Clearly something hasn't gone to plan.
There is approximately zero chance that this is a consideration.
My guess would be that, like a lot of things under the Elon Musk umbrella, what they claim they are able to do in theory diverges far from reality. We've seen similar slideshows.
[1] https://cloud.google.com/blog/products/ai-machine-learning/g...
https://blog.google/technology/ai/introducing-pathways-next-...
If you listened to their talk their silicon comes without support for access control and multi tenancy. Letting random people run on this will be an unlimited disaster.
Here's the car. Drive it yourself.
The product has been 'Fools Self Driving' demoware for years.
My comment already agreed with it, and it plays right in to the entire point of you having to still pay attention and drive the car yourself. Hence why I said it is 'Fools Self Driving'
Is this you as well? [0]
Yes [0] was me as well. Apparently the poster was not being sarcastic.
I was manually driving, but the car was constantly freaking out because it thought I was running off the road. :lol Yes, it can be a hard problem.
OTOH, my car is running regular AP and will not engage when no lane lines are available. Nevertheless, I've managed to trick it into engaging in a few cases by engaging while the lines are there and keeping it on when they disappear. It does this close to flawlessly in my neighborhood, which is actually pretty surprising.
If cracks in asphalt fooling the car's AI is a "hard problem", than full self-driving is doomed.
TBH, getting a computer to navigate a parking lot is one of the hardest self-driving problems. That doesn't make them all doomed.
The deadline for "Coast to Coast Full Self Driving" is over 5 years old now. https://techcrunch.com/2016/10/19/musk-targeting-coast-to-co...
An ASIC can strip out features they don't need and save some space. But a good chunk of modern GPUs are memory-controllers, registers, and SIMD-cores. And modern GPUs (both AMD's MI250x and NVidia's A100) have 16-bit matrix multiplication units (aka: Tensor cores). Once we factor in the process disadvantage, I'm not sure if the D1 will be as competitive as they hope.
Tesla's hope is that their D1 chip has more 16-bit matrix multiplication cores than the NVidia / AMD designs. But A100 is quite solid, and NVidia Hopper has been announced at HC34 (aka: NVidia's next generation).
https://www.nvidia.com/en-us/technologies/hopper-architectur...
-------
Most of this presentation on the Tesla Dojo is about the interconnect system. Alas, NVidia's on like the 4th (or was it 5th?) generation of NVlink, available from their DGX servers (and I'm sure a Hopper version will come out soon).
AMD's not far behind, also with a lot of good presentations this year from HC34 that points out how AMD's "Frontier" Supercomputer has huge bandwidth. In particular, each MI250X GPU is a twin-chiplet design (two GPUs per... GPU), with 5 high-speed links to connect to other GPUs in a high-speed fashion. There's a reason why Frontier is the #1 supercomputer of the world right now, in both absolute Double-precision FLOPs, and in Green Double-precision FLOPS-per-watt.
NVidia's Hopper will be hitting 4nm. AMD's MI250x is 5nm. That means the D1 chip has less than 1/2 the transistors at the same area compared to NVidia.
> but I’d imagine ASIC based supercomputers (like Google TPU) are the way to go forward.
Only if you keep up with the process shrinks. 7nm is getting long in the tooth now. All eyes are looking forward to 5nm, 4nm, and even 3nm designs (now that Apple is the customer of TSMC's 3nm node).
-----------
That being said, if the 7nm node is cheaper, maybe this exercise was still cost-effective for Tesla. As the newer nodes obsolete the older nodes, the older nodes become more cost-effective.
Cost-efficiency is less popular / less cool, but still an effective business plan.
> but I’d imagine ASIC based supercomputers (like Google TPU) are the way to go forward.
The issue is that it probably costs hundreds-of-millions of dollars to design something like the D1. Sure, the mass production of the chip afterwards will be incredible, but chips have stupidly high startup costs (masking, engineering, research, etc. etc.)
GPUs on the other hand, are more general purpose and are applicable to more situations. So you can sell the GPU to more customers and spread out the R&D costs. In particular, GPUs capture the attention of the video game crowd, who will fund high-end GPU research just to play video games.
Much like how Intel's laptops allow servers to share the R&D effort, so too does NVidia's consumer GPUs share research/development costs with their high end A100 cards.
While Tesla is claiming less than 400 Tensor-FLOPS on D1.
So yeah, the claims of NVidia's GH100 / Hopper GPU are an order of magnitude faster than the D1. Which is no surprise, because when your transistors are less than 1/2 the size of the competition, you can easily have 2x the performance in a embarrassingly parallel problem.
--------
Note that the A100, released in 2020, offers 312 TFlops of 16-bit Tensor matrix-multiplication operations per second. Meaning D1's chip is barely competitive against the 2-year-old NVidia A100, let alone the next-generation Hopper.
And note that NVidia's server-GPUs (like A100 or GH100) already come in prepackaged supercomputer formats with extremely-high speed data-links between them. See the DGX line of NVidia supercomputers. https://www.nvidia.com/en-us/data-center/dgx-station-a100/
--------
You can't beat physics. Smaller transistors use less power, while more transistors offer more parallelism. A process-node advantage is huge.
But anyway, you dodged my entire point about cost-to-performance ratio by looking just at performance. If NVidia is insisting on pocketing all the performance advantages of the process shrink as profit, then it still make sense for Tesla to do this.
"Dodged" ?? Unless you have the exact numbers for the amount of dollars Tesla has spent on mask-costs, chip engineers, and software developers, we're all taking a guess on that.
But we all know that such an engineering effort is in the 100-million+ project size or more. Maybe even $Billion+ size.
All of our estimates will vary, and the people who work inside of Tesla would never tell us this number. But even in the middle hundreds-of-millions, it seems rather difficult for Tesla to recoup costs.
-------
Especially compared to say... using an AMD MI250X or Google's TPUs or something. (Its not like NVidia is the only option, they're just the most complete and braindead option. But AMD MI250x have tensor cores as well that are competitive to A100, albeit missing the software engineering of the CUDA ecosystem)
------
Ex: A quickie search: https://news.ycombinator.com/item?id=26959053
> For 7 nm, it costs more than $271 million for design alone (EDA, verification, synthesis, layout, sign-off, etc) [1], and that’s a cheaper one. Industry reports say $650-810 million for a big 5 nm chip.
How many chips does Tesla need to make before this is economically viable? And for what? They seemingly aren't even outperforming the A100 or MI250x, let alone the next-generation GH100.
What's your estimate on the cost of an all-custom 7nm chip with no compiler-infrastructure, no software support and 100% all manual software built from the ground up with no previous ecosystem?
Modern GPUs, even with some process advantage (and no shipments her, H100) strive to be general purpose processors too much, and that caps their peak deep learning performance.
Every serious player will have to make or buy their own TPU.
Most training SoCs are focused on building something the size of an NVIDIA GPU, but designed for ML versus general purpose GPU HPC compute (FP64) plus ML. Often those accelerators today have a few types of models they are optimized for. NVIDIA is the baseline and so the competitors are looking for areas where they can get a large boost at a lower cost with something about the size of a A100/H100.
Cerebras is perhaps the biggest exception with its WSE-2, a wafer size chip. Having the wafer size chip means that Cerebras does not need to go into higher-latency and higher-power off-package interconnect as frequently because its chip is 50x larger. In turn, Cerebras drives performance and cost savings by not needing NVLink4 NVSwitches / InfiniBand.
Tesla's Dojo Tile is 25 chips roughly equivalent in size to a NVIDIA GPU in a single package with die-to-die communication facilitated by the base tile and then built for scale up units. Tesla also has focused on the interconnect and pipeline feeding the D1s and Tiles.
Ultimately, I think that it takes something beyond a "solution X saves 30% over NVIDIA in these workloads in performance/ $" to survive. NVIDIA has a massive software ecosystem and can handle more types of tasks versus some of the other AI accelerators. That goes beyond just the training and also to other parts of the data prep and movement pipeline. NVIDIA extracts high margins from this work so that is why some effectively are competing with "it costs less and on some problems can be faster" architectures but what Tesla, Cerebras, Google, and a few others have another level of differentiation.
Nothing is perfect, nor was that explanation, but just a high-level view of why the technology featured is impactful.
That said, they do make progress and the real reason why it's slow is probably because the problem is actually really hard and nobody has found the fast way to the solution yet.
Dude just straight up lies. He may not realize he's lying, but he lies constantly.
Edit, because I'm getting downvoted. Here's some examples: self driving, battery swaps, robotic snake chargers, cybertruck windows, Bitcoin won't be converted to fiat, starlink speeds will improve, he will sell his home, first mars mission 2024, Tesla solar roofs, power packs at every supercharger, brake pads on Tesla cars will never need to be replaced, the gigafactory will be 100% renewable powered by 2020, fixing the Flint water crisis, making bricks from boring company waste, founding a media credibility organization, and probably a TON more.
Oh man, I forgot he took money to fly tourists to the moon.
There are so many lies he's told. So many.
I just want a Toyota Corolla or Camry-style car, for a decent price, that is electric. The i3 was kinda close to what I wanted in spirit, but they had to make it look all "tech" and "future-y".
https://www.theverge.com/2017/2/17/14652026/spacex-red-drago...
Starlink satellites were supposed to have laser interlinks from the beginning and enable low latency multiplayer gaming around the world. The Starlink latency is pretty bad for gaming right now and is variable.
The battery swap stations were a bigger scam than you might remember: California gave them almost a billion for "delivering" them.
Source?
“In 2013, California revised its Zero Emissions Vehicle credit system so that long-range ZEVs that were able to charge 80 percent of their range in under fifteen minutes earned almost twice as many credits as those that didn’t. [...] By demonstrating battery swap on just one vehicle, Tesla nearly doubled the ZEV credits earned by its entire fleet even if none of them actually used the swap capability. [...] For the 2015 to 2017 model years, CARB created a new rule requiring Tesla to actually document a certain number of battery swaps to prove that the capability was actually being used. This development just so happened to coincide with Tesla’s half-hearted “pilot program” and its subsequent decision that its customers had no interest in fast, convenient battery swaps. [...] By exploiting CARB’s fast-refueling rules, Tesla appears to have earned as much as $100 million in additional revenue by demonstrating and hyping a system it seems to never have intended to commercially deploy.”
So far with this Fools Self Driving deception:
Claim 1: 'Full Level 5 Autonomy by end of 2019' (2019) [0]
Reality: As of this year 2022, it is still Level 2.
Claim 2: '1 Million robo-taxis on the road by the end of 2020' (2019) [1]
Reality: As of this year 2022, ZERO Tesla robo-taxis on the road.
Claim 3: 'Tesla's FSD tech will have Level 5 autonomy by the end of 2021' (2021) [2]
Reality: As of this year 2022, it is admittedly Level 2.
Claim 4: "I would be shocked if Tesla does not achieve FSD that is safer than human drivers this year" (2022) [3]
Reality: Clearly it still isn't any safer 8 months ago. [4]
So even if we have given them more time since those claims in 2019, nothing has changed.Not only it was admitted to be Level 2, [5] and still requires the full attention of the driver having their eyes on the road, they have ever continuously raised the prices on a system that clearly doesn't work as advertised in order to get customers to FOMO into purchasing it.
A complete scam of a contraption that can only be perfectly described as a 'Fools Self Driving' system.
[0] https://www.motortrend.com/news/tesla-autonomous-driving-lev...
[1] https://www.engadget.com/2019-04-22-tesla-elon-musk-self-dri...
[2] https://www.cnet.com/roadshow/news/elon-musk-full-self-drivi...
[3] https://electrek.co/2022/01/31/elon-musk-tesla-full-self-dri...
[4] https://news.ycombinator.com/item?id=32401444
[5] https://www.news18.com/news/auto/teslas-full-self-driving-cl...
I do not understand how they continue to get away with it.
You have to know what your saying is untrue for something to be a lie. Otherwise it is simply called being wrong.
https://jalopnik.com/elon-musk-promises-full-self-driving-ne...
And there’s incredibly small chances this system will actually work given the limited sensor set in the car, and especially with the removal of the radar (which is not them “cleaning up the data”, it is 100% entirely due to supply chain issues).
It’s a great car, and I applaud them for making various major advances in the automotive space including the push for EVs and OTA updates. But FSD is a pipe dream and won’t be anywhere near full self driving in the next decade. And other manufacturers are quickly catching up to the current public feature set.
Most of these are either things that are simple changes in strategy, mostly good choice regards to internal investment and roadmap. Others are research projects that were never promised to be products. I really don't understand how anybody can be angry about any of those.
Some of these are products, but apparently they are not as perfect as they could be according to you. This is pretty absurd definition of 'lie'.
> first mars mission 2024 > Oh man, I forgot he took money to fly tourists to the moon.
This is maybe the most ridiculous thing I have ever read. They are experiencing delay on one of the most difficult engineering projects ever.
I guess we should execute him for daring to not having flown people around the moon.
In aerospace even simple LEO orbit rockets often launch late, its normal and costumers know that its a probability. Contracts do handle these cases.
Most of this boils down to Musk saying 'this is our current plan' and then people get angry when plans change. I don't understand why. Did you personally sign a contract with Tesla for a battery swap station or something?
I for one really like getting updates on what SpaceX/Tesla is planning or working on. But I don't get upset when they change it. When I see a snake charger I don't go 'oh they promises this and it will be in my garage in 3 month'. They have absolutely no obligation to me.
I've found that a surprisingly large number of people are really committed to the idea that battery swapping is the only way that EVs could work. These are personally offended that Tesla abandoned it.
But doing battery swapping well is much more capital intensive than their supercharger strategy was. They could never have afforded it on their own anyway.
The whole point of such grants is to figure out if its a commercially viable solution or not, and it wasn't. Or at least not the best one.
I'm actually glad they they've decided to slow down iteration on the self-driving stack - in the early days the self-driving would change behavior from patch to patch - it was extremely difficult to get an idea of what it would be thinking. It's obviously not the product that we all wished it was - but... does any product in the world hold up to consumer imagination?
Apple promises that the newest iPhone is a "new superpower". Am I supposed to believe that? Does anyone believe that? It never fails to crack me up that the folks most upset about Tesla's self-driving promise tend to not be customers and vow that they never will be. What in the world is so upsetting about a company over promising and under delivering? Are these people just constantly boiling with rage? Have they never experienced being a consumer before?
These days Tesla wants to con people out of $15,000:
https://www.thedrive.com/tech/even-tesla-fans-think-fsds-pri...
Why should anyone "consume" this "product" when it simply doesn't work as claimed and doesn't meet any of Tesla's self declared deadlines?
I wouldn't pay $6k for EAP or $15k for "FSD" though.
I’m just not furious is all - and not nearly as furious as folks who… didn’t buy the product. Me being disappointed in a cool toy is hardly a problem worth discussing.
I almost bought a tesla a year ago based on the idea FSD was going to work at some point in the future but after Elon announced some last minute hardware changes to my model's radar (removing it), and I read about all the steps you should perform at pickup, I went and bought a toyota instead.
Still, it is good to see powerful hardware being put to good use for once. I understand it is silly to feel sorry for inanimate objects, but it saddens me to see silicon squandered by some manchild at a national lab working on a dead-end vanity project when it could be crunching numbers for Tesla or—better yet—mining Bitcoin.