Tesla Project Dojo Overview
perspectives.mvdirona.com
perspectives.mvdirona.com
An A100 is an ExaFLOP GPU, if you count INT8 as FLOP.
Then you have the 10 different flavors of FP8, FP16, etc. and even the multiple flavors of FP32.
What people usually count as "ExaFLOP" is FP64, but this hardware can't even do FP32, much less FP64.
So yeah, an ExaFLOP != an ExaFLOP anymore, because these types of announcements never say "An ExaFLOP of what".
Everybody's brain has an infinite througput of FP0 (0 bit FP format). That doesn't mean that your brain can compute faster than Dojo, even though infinite ExaFLOPs >>>> 1 ExaFLOP. The reason is that we are counting different things.
So what's the throughput of this hardware in actual FP64 ExaFLOPs ? Zero, it doesn't support them. Not a very impressive marketing statement though. But what this means is that this hardware would be extremely bad at, e.g., solving linear systems of equations.
I thought Tesla was pretty clear about which floating point formats they were talking about. When they first introduced the D1 they put up a slide with FLOPS measurements for FP16 and FP32. The impressive numbers (exaFLOP in ten cabinets for example) are based on FP16, of course.
I just don't see why volumetric efficiency matters at all for this computer.
All that matters is usable flop per dollar. And if you put it in a cold place with very low electricity costs (eg. iceland), the electrical efficiency doesn't matter much either. All you care about is how cheap can you make high speed transistors.
The Dojo design looks like it really wasn't optimising for what matters at all.
The reason they gave for doing things like "knocking out the walls of cabinets" was bandwidth, not space.
It turns out that the highest bandwidth solutions are also quite compact. One giant piece of silicon instead of many smaller pieces of silicon is great because the different parts of the giant piece of silicon can talk to eachother very quickly with very high bandwidth. Short wires are better than long wires, because you have less noise so you can fit in more signal (also less latency).
It's also just not clear to me how you imagine making things bigger would make them cheaper.
This was also mentioned in the AI Day presentation, and listed as a major factor in the design of Dojo.
It's really Deep Learning 101.
Furthermore, it's a non-sequitur of an argument.
Dojo is generic neural network training chip. It can be used to speed up any and all Deep Learning problems. Self driving is one such problem.
During AI Day Tesla also presented, in depth, the architecture of the Deep Learning network they use.
If you have a critique of their approach or can show how a competitor's approach is better, then do share with the class. If you don't it's just lazy "Tesla bad because they solved a problem that no-one else solved"
They're not just going to bloat the dataset with straight driving.
The guy from comma.ai says the same thing.
2 - it's not true that "no-one else solved" this problem. If we define this problem as: have a self-driving vehicle which safely navigates its environment and does not endanger its passengers or the people outside. See the Waymo whitepapers: waymo.com/safety/performance-data
Ultimately, even if we agree that what Tesla is working on is orders of magnitude better than what's currently available from them, that might not cut it:
Let's assume that the current obstacle detection is 99% accurate... if the newer versions/improved models are 99.9%, 99.99%, etc. accurate...
we'll still have a 0.1, 0.01, 0.001%, etc of chance of an obstacle not being recognized. Tesla cars are a small fraction of those on the road, and there've been several deaths already. If million of Teslas will be sold, those small percentages will still mean a significant number of accidents and deaths caused by Tesla. If instead a Lidar reliably detects all obstacles, well in advance of the vehicle approaching it (and the car will stop/disable self driving if the Lidar becomes unoperable, e.g. due to bad weather)... it would be irresponsible to persuade huge swaths of the population to let a machine drive the vehicle, without providing the extra safety that a Lidar enables.
To date, Tesla has sold about 2 million cars.
Not sure how many of those support HW3 "FSD capable" computer, but I'd guess 1,5 million.
The FSD software package so far has been rolled out as a beta only, to a few selected users. It is based on vision entriely and aims to properly identify everything on the road. The solid obstacles have all been hit by the previous auto pilot version. In the new cars currently sold, the radar has been replaced by a pure vision implementation close to the FSD software, so it remains to be seen which level of safety can be achieved with that.
Nobody uses radar for lane keeping. It can't see the lane lines.
Pretty much every manufacture now has one and they always use forward facing camera and radar in combination. Of course the radar can't see lanes, but its essential for seen if there are cars ahead and how fast they are.
All of these solution suffer from the same problem of the radar not being good to detect a stationary object. That is where you have to do more with the camera data or add some form of lidar.
I think some 1 car (Audi maybe) use solid state lidar. Tesla has switched to vision only for the most part. No company yet uses active lidar as far as I know.
I can't imagine this thought hasn't crossed some minds in their boardroom.
So I think getting in the fab business for Tesla would mean trying to be a direct competitor with TSMC, Intel, etc. by making chips for a large number of customers.
That seems way down the list of Tesla's priorities and best uses of dropping hundreds of billions of dollars.
Intel reportedly produces roughly 10 million wafers per year [1]. Roughly 70 million new cars are sold per year [2], with Tesla currently accounting for 0.5 million of those [3] with roughly 50% YOY growth numbers and a plan on continuing those for multiple years [4].
Cars/wafer is a pretty unclear number, at 5 Tesla is 1% of Intel's market (today), at 25 0.2% of Intel's market. It will take quite awhile for Tesla to hit intel's scale - but the automotive market as a whole might actually be pretty close to it if all new cars start including giant computers.
[1] 884k/month * 12 months/year, from https://www.eenewseurope.com/news/top-five-chip-makers-domin...
[2] https://www.statista.com/statistics/200002/international-car...
[3] https://backlinko.com/tesla-stats (or see [4] but then you have to add up quarters yourself)
[4] https://tesla-cdn.thron.com/static/ZBOUYO_TSLA_Q2_2021_Updat...
There are already multiple competing companies whose sole job is to make the best fabs at high production rates.
Tesla doesn't even make most of their batteries in-house, which is way more core to their business of EV manufacturing. They partner with Panasonic, LG, CATL, etc. because they are in the business of building out manufacturing capacity for battery cells.
You could s/chip/battery/ and make the same incredulous statement 5 years ago
Or upholstery, or charging infrastructure, or casting equipment, etc.
Tesla had the capital to go as vertical as they want.
The other advantage with vertical integration through the manufacturing piece is turn around time on new designs. It is a lot faster/easier to spin new test lots for speculative designs when everyone works for the same org.
The semi industry is really far from that.
Tesla is at least a decade away from the shit Elon is tweeting about.
Plan your next car purchases accordingly.
Sincerely, a Model Y owner
Tesla is not like a regular manufacturer where you can walk around a vehicle and determine if you want to accept it. You have to actually test things - even the things that we've been doing right since the Model T. These things are built in tents by staff that is treated quite poorly, and you should be skeptical in proportion to that information.
Here is a good example: https://github.com/polymorphic/tesla-model-y-checklist
Although I noticed it's missing "check that the roof and trunk are in fact water proof".
Getting any of these things fixed after you take delivery is a huge PITA.
And, tbh, hang around some forums for the company that has been "doing right since the Model T" and you'll see some of the same things. Roof alignment in particular has been a little bit of a problem for the Mach E.
With both companies, there are some horror stories after delivery too.
With both companies, the vast majority of cars are great and end up with really happy customers.
Check list make sense for all manufactures.
> These things are built in tents by staff that is treated quite poorly, and you should be skeptical in proportion to that information.
People are obsessed with this tent. Its just a simple stable structure that you can build quickly. It changes nothing about the manufacturing line being in a building.
The tend idea was set up when a guy with 30 years of manufacturing experience from BMW and other German car makers. They knew what they were doing, and its actually the cars from the tend that improved the quality problems they had early on.
> Getting any of these things fixed after you take delivery is a huge PITA.
Most of these things can be fixed by the mobile service, they can do it while your are not even there. Often the fix it while you work.
But then instead of the presentation ending with "And that's why we're confident in our 5-10 year roadmap to Level 5 Autonomy" they end with Elon tweeting that some beta version is going to roll out to customers in 2-4 weeks and the optimism all comes crashing back down again.
Sounds rather counterintuitive to me. I feel far safer driving on my regular commute because I know where the potholes are, where the blind junctions are despite the missing signs. I know exactly how fast I can go around the roundabout built with a ridiculous camber in wet weather that sends many people in to the ditch every year. etc.
Knowing all this, I can bias my attention to the cars/people around me rather than the environment.
It takes at least decade to build up something like Google Maps, but a lot of the data is already out-of-date.
We are searching for an algorithm which turns camera data into a vector model, right?
This sounds very parallelizable.
Have a lot of computers travel the search space individually. When one finds a algorithm better than the current best solution, have it call the others "Hey guys, this algo works better then what we got so far. Everyone iterate on this one now!".
EG using a managed Database as opposed to rolling your own saves on patching, management etc...
Our neurons don't have a global clock, it doesn't seem to be a problem for us. My intuition is that as long as the input is changing continuously and not through random presentation, and at a rate where big changes happen at a lower order of magnitude than the average compute rate, it wouldn't matter all that much in terms of accuracy.
Would we have been more effective if we built other things like this? 400,000 people were involved in the Apollo program. Would it have been better if we had just one person with a really big brain? How about the Linux kernel and Wikipedia?
And why should I figure out a better algorithm? Shouldn't that be the job of the one guy with the biggest brain who comes up with everything?
Couldn't evolution have settled on bigger brains if they are an advantage? Why all that slow interpersonal communication if it is more efficient to have the combined thought processes inside a single brain?
https://en.wikipedia.org/wiki/Amdahl%27s_law
Some parts of an algorithm are parallelizable, but only up to a certain limit; and that's where the endgame bottleneck is: the last remaining 'in series' parts and the communication overhead for the parallelized parts.
Note that their Dojo system has absolutely massive networking equipment: that's for having a better endgame at Amdahl's law.
----
On another note, one can modify the algorithm itself so that is has more inherent parallelism. For example DNA sequencing: you match short reads, them send them off to be reconciled together. The way to have more parallelism is to generate more matches -by lowering the threshold for matching-. Having more matches means more communication costs, but better parallelism.
For gradient descent, one of the tricks is learning in batches of input data: the neural network doesn't learn from the freshest point of view, but it can be done in parallel (also the way to aggregate new knowledge is to do an average of the weights, which must limit the amount of things learned)
Biological brains are limited by other factors. Human brain size in particular is limited by factors like needing to fit through a human birth canal without wrecking an upright walking gate ([0]) and being able to be powered using a hunter-gather diet.
[0] - loosing the upright walking/running gate would have limited certain ecological options (like persistence hunting). Instead, human gestation is shorter than it should be given our body size when you compare it to other mammals.
The example my father gave when I was a teen was: “nine women can’t make one baby in one month”.
> And why should I figure out a better algorithm?
They said “if”; if you do, you could get very, very rich.
"To get rid of the dependency on the radar sensor for the pilot, we generated over 10 billion labels across two and a half million clips. And so to do that, we had to skill our huge offline neural networks and our simulation engine across 1000s of GPUs, and just a little bit shy of 20,000 CPU cores. On top of that, we also included over 2000, actual autopilot full self driving computers in the loop with our simulation engine. And that's our smallest compute cluster.
So I'd like to give you some idea of what it takes to take our neural networks and move them in the car. And so the the two main constraints that we're working on there here are mostly latency and framerate, which are very important for safety, but also to get proper estimates of acceleration and velocity all of our surroundings.
And so the meat of the problem really is around the AI compiler that we write and extend here within the group that essentially maps the compute operations from a pytorch model, to a set of dedicated and accelerated pieces of hardware. And we do that by figuring out a schedule that's optimized for throughput while working on very severe SRAM constraints.
And so by the way, we're not doing that just on one engine, but across two engines on the autopilot computer. And the way we use those engines here at Tesla is such that, at any given time, only one of them will actually output control commands to the vehicle, while the other one is used as an extension of compute. But those rules are interchangeable, both on the hardware and software level.
So how do we very quickly together as a group to this AI development cycles? Well, first, we have been scaling our capacity to evaluate our software neural network dramatically over the past few years. And today, we are running over a million evaluations per week on any code change that the team is producing. And those evaluations run on over 3000 actual full self driving computers that are hooked up together in a dedicated cluster."