An inside look at the custom CPUs in Tesla's Dojo Supercomputer
semianalysis.com
semianalysis.com
> Their current training cluster would be the 5th largest supercomputer if Tesla stopped all real workloads, ran Linpack, and submitted it to the Top500 list.
which is trivial to do, and pretty much a must when bringing up the cluster to make sure its working properly, so much that most clusters do this on every maintainance, along with another bunch of benchmarks;
and that they say this:
> cost equivalent versus Nvidia GPU, Tesla claims they can achieve 4x the performance, 1.3x higher performance per watt, and 5x smaller footprint.
but have no MLPerf results, tells you everything you need to know about it.
The list of long-term hype-only AI-hardware companies with billions of dollars of VC investment and literally nothing to show is incredible and keeps growing.
Every MLPerf round, the list of companies that want to submit is "huge", and 1 week before the deadline, 99.999% of them have been saying "we'll submit next round" for years.
It's as-if people would spend billions on creating an F1 team, and then notice during pre-season training that the car can't even finish a lap. And then fail to even start a lap on every race of the season. And then do this again, year after year, for a decade. Burning billions and billions...
And if it is tailored for AI, it might not even do 32-bit float, or only at a fraction of the "AI FLOPS".
You have to connect thousands of cables, have hundreds of nodes, with thousands of components, everything interconnected, and if you connect one wrong, the computer outputs incorrect results. You have to routinely update the software, and if a software upgrade introduces a 20% perf regression (which happens), then your 10 MWh cluster starts burning 2MWh for nothing. Or maybe your cooling system sucks, and after a minute of running at full capacity, you need to throttle your cluster to 0.1% of the peak to keep it cool enough that it runs "something".
That's why all systems in the top500 i've been involved with (15 or so) run these benchmarks as an integration tests on every single cluster maintenance (node updates, servicing, OS updates, etc.).
Submitting these results to the Top500 costs you nothing... if your cluster actually works. When you submit to the Top500, they ask for access so that they can re-run them themselves, which happens typically during / after the next maintenance to avoid impacting any users.
If they haven't submitted, 100% sure their cluster does not deliver what they say it should deliver on paper. Maybe it delivers 1% of it, or 0.01% of it (seen both cases in real life). If they haven't fixed it, then maybe it can't be fixed.
HPL, MLPerf, Spec, Stream, OSU.... these are not "benchmarks for bragging rights", these are tests that show that your system works.
Why run someone else's benchmark and not your own application to test performance? And what's the point of submitting to Top500? Why do you care how your system ranks? What's the business or technical purpose in that?
It would be much more difficult to adapt some of your own applications as a test. If the result are bad, how would you even know if it's the application's fault or a problem with the cluster? LINPACK on the other hand is well understood and there has been tons of work to make sure that it uses all the power your cluster can deliver.
If you are claiming something is the best, then without any comparable metrics across systems we are in the Twilight zone.
They are a publicly traded company.
I'm a tesla investor and deserve to know.
If they are right and their hardware is 10x better than the competitor, then I'm happy. If they are wrong and they could have get 10x better outcomes by spending 100x less money, I'll be very pissed off.
> They probably don't want to waste time and energy running a benchmark to compete in someone else's top N list.
It costs no time. It's the first thing anybody does when building any cluster. It's part of accepting the materials from your suppliers to check that they didn't sell you shit.
I'm not an expert in securities law, but I don't believe being an investor entitles you to demand arbitrary proprietary information you want from a company. Otherwise people would buy one Apple share and find out what specs the new iPhone will have.
Tesla is saying their hardware is 10x better than the competition.
I don't believe it. They don't publish any numbers, and everybody else does.
You believe them. Good for you.
I don't. I think not doing that is extremely suspicious, because if their cluster can be turned on, they have the results. So the reality of this is 100% certain that their results just suck, and they don't publish them to save face.
You believe their shit is so good, they didn't even test it while installing it. Good for you.
You think there is a masterplan to keep the results secret. Good for you.
I think their results are just horrible, and that's why they don't show them.
With any serious company, this wouldn't even be a discussion, because they would just show the verified facts about what their cluster can do or shut up about it if its really so secret.
You don't see the NSA or the DOD bragging around about being 100x better than "the competition".
To me they look forced to sell that they are doing something about these clusters to... drive hype up and get the stocks back up, and since their numbers are actually horrible, this is what we get.
My experience is on Summit & am now working with Frontier setup.
In the end if someone just tells you their supercomputer "would" be in the TOP500, they just tell you that it is definitely not in the TOP500. There might be good reasons for that, but still, it's like claiming that you would win Olympic gold if only you'd bother to participate.
Not clear what you're disputing here?
The article, that assumes "FLOPS on paper == FLOPS in practice".
So it is literally unbelievable that Tesla not just stands up a cluster, but created their own hardware to do so, and didn't run any quotable benchmark and only has the theoretic FLOPS numbers for marketing.
They haven't gotten to this point yet.
They have a single tile (A node). It wouldn't surprise me if the tile they have is just a prototype as well.
You have to have a cluster before you can start running cluster benchmarks.
That line is referring to their current Nvidia A100 powered supercomputer which they set up with 5,760 A100 GPUs recently this year [0].
Read the previous line of the line which you’ve shared from the post:
> Tesla has been expanding the size of their GPU clusters for years. Their current training cluster would be the 5th largest supercomputer if Tesla stopped all real workloads, ran Linpack, and submitted it to the Top500 list.
[0] https://blogs.nvidia.com/blog/2021/06/22/tesla-av-training-s...
But I don’t see them not doing that test as an issue as they made it quite clear that their entire system is tailor made, like an ASIC, to focus on neural nets and specific, relevant compute pipelines to what they care about. I would imagine dropping a generic, broad based ML benchmarking tool will not only perform suboptimally but also not be representative of what they’re trying to do. It’s not meant to be a general purpose ML super computer, it’s supposed to be a super computer to solve a narrowish niche of problems.
- Communication speed is one of the biggest bottlenecks to large models, so their bandwidth of 4TBps is very smart.
- They claim a 1.3x perf/watt improvement, which is not really that great for ASICs compared to GPUs. Perf/watt is probably the most important number in datacenters.
- They only use SRAM, no DRAM. This is a huge mistake, which limits their model size. You can only fit a ~10GB model inside a single tile, versus 80GB models for a single A100 GPU.
- Software / compiler stack is as or more important than the hardware itself, because it dictates how much real performance you can squeeze out of the chips. I think Tesla will need to heavily focus on this area before getting anywhere close to real-world GPU performance.
Overall, I imagine the project will have similar pitfalls to Cerebras.
Siemens H100: https://pbs.twimg.com/media/Ee56Q2bWkAA8vBE?format=jpg&name=... https://pbs.twimg.com/media/Ee56Q2ZX0AEqAWo?format=jpg&name=...
How is the author declaring Tesla has one upped everyone in AI hardware and software while the article has exactly zero references to TPUs?
Creating a lab beast is one thing, making it useful and economical in production is another.
Their software only needs to be useful to a small group - the autonomy team. I did not get the impression that Tesla plans to sell supercomputers, but that they are building these supercomputers for themselves to train and deploy AI networks.
You said "Even if the hardware existed, if you can't program it, it won't succeed." But it seems like the key customer - Tesla's internal Autopilot team - can already program it. So I just don't see any problem. They may choose to sell these systems (note that they have been shipping their Gen1 custom chip for years now and it is not for sale to the public), but the real plan for revenue is to succeed at AI and profit there. For that they need incredible computing power internally and in a portable platform for edge deployment, but they do not need to sell their chips as a general purpose compute platform.
So all in all I think the chip is very good and I think they are on a path to success. Whatever the article says to hype it up doesn’t change the value of what Tesla has done in my eyes.
The machine-learning parts are custom but the rest are probably similar to what you would find in a smartphone chip.
A multi-billion transistor CPU is an insanely complex product, there are multiple layers coming together to make it possible. The base layer is the foundry, here TSMC, which does the manufacturing and provides reference design kits. They contain for example the blueprints for transistors and other electrical components. For a digital chips like processors, they are usually used as they come from the foundry.
Next is a layer of software which enables the chip designs, provided by the big EDA companies. This is a rather huge layer, enabling the basic designs on the one side, but also including all kind of simulation and verification tools. And quite some engineering knowledge comes along with it.
So if you want to start designing your own CPU, you still need good engineers who know what they are doing, but large parts of the whole "stack" can and have to bought from the vendors listed above. This enables the quick entries of companies into the market, who were not traditional chip design houses.
> For those counting at home, there are 12 total tiles per cabinet
So 180kW per cabinet. I'd love to see what the cooling system looks like!
If this is as good as they say, they should spin this off and sell these. Probably won't.
i think the most challenging aspect of the idea is interacting with the world, picking up and handling various objects.
its obvious from the presentation that fsd is very good at placing itself in space and mapping out its environment as well as devising routes even when accounting for things like other moving objects. boston dynamics has built machines that take the input of a joystick and translate that into four-legged locomotion. so if you just took FSD and bolted that on top of a spot or atlas, you would have a machine that gets you a very long way towards tesla bot. something that can walk and navigate by itself. so the question is whether or not tesla will be able to recreate what boston dynamics has done and whether or not they could do that on the time-scale that was insinuated by musk. the last time i checked, boston dynamics does not use any NN in their robots.
and then theres the question of interacting with the world. i think its a fair assumption based on ai day that fsd would be able to create a very rich and accurate map of its surroundings. so that half of the question is taken care of. but how could they get it to grasp a drill and manipulate it in an intelligent way? the implication is that they would train a NN to provide that signal. but how are you going to train that? they have thousands of cars out in the world generating tons of driving data. where are they going to get data for performing these tasks? are you going to train it on every task that might be asked of it? its hypothetically possible but thats hard data to generate and label. it would take a really long time at best, and even then it would be an experiment, just as likely to fail as succeed.
it seems to me that this must be a stunt where the intention is to build something that can walk around but not manipulate things or do anything useful. i would change my mind if musk came out with an explanation of how hes going to train interaction.
Sorry to point out but this has nothing to do with the article which is about semiconductors.
The last commercial anthropomorphic was the Willow Garage PR2 back in 2010. It weighed 600 pounds, and had a wheeled base. Each arm had a max payload of 4 pounds. It cost $250,000. The company went bankrupt because there wasn't anything you could do with it.
The tesla bot is supposed to be bipedal, only weigh 125 pounds, and have a "arm extend lift" of 10 lbs. Is that per arm, or both together? Even BD Atlas weighs 196 pounds, and it doesn't have hands!
Like, any one of their numbers would be exceptional. All together, and you are definitely sacrificing something. Either onboard compute is minimal, or it has a battery life measured in dozens of minutes. Something.
They claim a deadlift of 150 lbs. I don't want to say "impossible!", but... difficult? Just holding 150 lbs with any of the robotic hands you can buy on the market today would be hard or impossible.
None of the five finger hands today are as compact as shown in the concept renders. If it does ship with a five finger hand, (stupid, pointlessly expensive) the forearms would be far bulkier. An easy bet is that the first gen will ship with a three finger gripper.
That is, if it even ships at all. BD Handle is a much better design for a humanoid-ish industrial robot. It's going to be simpler, faster and lighter for the same payload as this hypothetical Tesla bot.
https://old.reddit.com/r/robotics/comments/p7t14o/tesla_reve...
(The robot picks up a screwdriver, and a screw. Then the screw slips out of its fingers. Now what? It lines up the screw and the screwdriver. It applies torque. The head of the screw strips out. Now what? The pain of robotics is that you need to cover each and every little error case, because if you don't, the damn thing doesn't work, because it has no brain! This is why every industrial robot is massively overbuilt, and it's environment and fixturing is carefully simplified and fenced off, because error handling is such a pain in the real world, where a dropped item bounces away and hides under a bench or in an orientation where your gripper can't pick it up. 1 in 1000 is too high of an error rate. 1 in 10000 is too much. It has to function perfectly, every time, every grasp.)
And the maintenance costs! Mechanical humanoid hands are terrible end effectors. All little moving parts and lousy tolerances. They would need constant repair and replacement. It couldn't possibly be cheaper than a human in 2021 or 2030.
For industrial applications you could probably have some sort of novel power system, like a tether from the ceiling or special floor that delivers power through the feet.
https://old.reddit.com/r/robotics/comments/p7t14o/tesla_reve...
For those that want to push downward, the present system highly leverages their ability to do so simply through the process or habit of frequent downvoting.
If there is any perception that there was a time when there was very little detectable smell by comparison, it would be good to assess the average downvote-per-user rate then, and compare it to today's figure.
Then measure each user along this scale, perhaps including a time component, or relative to activity in some way.
Allowing for a reasonable standard deviation, it might be better if frequent downvoters past a certain range had the weight of each downvote normalized and see what happens.
This could possibly also be tuned to achieve a target level of discourse relative to a previously-considered-desirable data point in time.
Alternatively, users alone appear theoretically able to overcome the issue if there was a widespread concerted or random effort to frequently upvote the comments or postings seen descending, whether fully deserved or not, keeping them at least neutral without having a negative effect on the commenter's rating.
Mathematically a small uptick in "compensatory upvoting" habits among average users could bring the target way up as long as the overly-frequent downvoters are in the vast minority.
Then when there is true downward consensus it will still always drop through, but those who participate mainly to downvote will have less negative impact.
The only thing worse than the "nattering nabobs of negativity" are the non-nattering nabobs of even worse negativity.
He knows the world idolizes him as a slightly eccentric genius engineer with a heart of gold, and I think there's probably some truth in that, even. Nevertheless, if you are smart enough to be a good engineer, it doesn't take much to realize that you can greatly expand your aura by occasionally trolling people with something stupid just to see if you can get away with it. The fans will adore him more, the critics will pounce, but the net result is more exposure, and more exposure gets investment money. Rinse, repeat.
Personally, I think one of the smarter touches is to release/leak information that looks bad shortly before releasing information saying that the problem was overcome. This constantly reinforces the impression that anything Musk-related is always overcoming impossible odds. "Andrej didn't think we could do it, what do you think now, Andrej?" "The Starlink terminal costs $2400, but just a few months later, it costs $1000!"
I'm sure FSD will be here by Christmas.
It pretty much continues to double: https://en.m.wikipedia.org/wiki/Transistor_count#/media/File...
javascript
Or more likely they'll spend the compute power on the latest JavaScript framework....
[1]https://www.tweaktown.com/news/81229/teslas-insane-new-dojo-...
[2] https://www.tweaktown.com/image.php?image=https://static.twe...