Tesla Packs 50B Transistors onto D1 Dojo Chip
tomshardware.com
tomshardware.com
Die Size is 645mm^2 on a 7nm. This is important because we know the reticle limit which is around ~800mm^2.
The Nvidia AI Chip has 54 billion transistors with a die size of 826 mm2 on 7nm.
I recently saw a Ted Talk, If Content is King, then Context is God. I think it capture everything that is wrong in today's society.
One basic thing I didn't see in the body was power consumption though, anyone know more details on that?
Artificial intelligence (AI) has
seen a broad adoption over the past
couple of years.
And continues: At Tesla, who as many know is a
company that works on electric
and autonomous vehicles, AI has
a massive value to every aspect
of the company's work.
Who is writing like this? And why?What would Tom's Hardware lose if they left out this type of cheap fillwords?
Should I also start writing like this?
Is this type of "reader hostile writing" a new thing or have newspapers always written like this?
These are not rhetorical questions. I am honestly confused.
I understand that you consider the writing objectively "hostile", but the simpler explanation is that you're just not the audience.
That's also just wrong. During the recent "Tesla AI Day", when asked during Q/A, Elon Musk specifically mentioned that they intentionally use machine learning only for very few cases:
Q: "Is Tesla using machine learning within its manufacturing, design
or any other engineering processes?"
Elon: "I discourage use of machine learning, because it's really
difficult. Unless you have to use machine learning, don't do it. It's
usually a red flag when somebody is saying 'We wanna use machine
learning to solve this task'. I'm like: That sounds like bullshit.
99.9% of the time you don't need it."
https://www.youtube.com/watch?v=j0z4FweCy4M&t=9307sHonest question, how much is chip design a factor separate to fab process?
This chip is largely memory and multipliers, both of which are pretty dense.
Fab processes improve over time to have higher density and lower defect rate (which allows bigger chips while getting acceptable yield). So it's not surprising to see a chip on the same node but shipping a year or 2 later (than Ampere) having more transistors.
I think in 25% of cases it will not get them significantly more performance vs Nvidia.
There is a 50% chance that they can outperform off the shelf chips by a significant amount to make it maybe worth it. (This is pretty likely because dedicated hardware tends to outperform general hardware).
However, there is maybe a 25% risk buying Nvidia doesn't get them there soon.
So building their own chips de-risks the worst case, and it's probably not that much more expensive (at Tesla scale). So seems like a pretty good bet to me.
For one, Google TPU. Another: Cerebras wafer scale AI. AMD MI100. Etc etc.
Even if they screwed the pooch with Nvidia, there are plenty of competitors in this space.
Now Tesla has to build its own software stack for large scale distributed learning, which might be harder than the chip design.
Is Tesla really the kind of company that wants to carry the expensive loadstone of training and inference software + hardware?
It's not like PyTorch is gonna run on this thing unless they create a fork. And a huge advantage of things like NVidia are NVlink / NVswitch. Both hardware, and software, that efficiently distributes data at 600GBps across your GPU clusters.
Yes. They are very much that kind of company. Tesla has been pushing vertical integration and that is very much Elon Musk whole approach for most of his companies.
Doing your own battery manufacturing and even supply chain is considerably more expensive and complex compared to making a chip and getting some software developers.
However they are working on having their own battery and battery factory design. They are even working on their own cathode manufacturing plants.
They have a very large 'pilot' plant in California to test a completely redesigned battery manufacturing system. They are actively building additionally battery factories in Austin and Berlin and have a equipment on order for these factories. For the Battery factory in Berlin they have received funding from the European Commission for this.
Tesla will not be fully vertically integrated but they will provide a large parts of their own cells and likely a increasing share over the coming years.
There's no realistic plan for them to get off of Panasonic's cells. They have a purchase agreement, and Panasonic owns half the Nevada Gigafactory.
They have to wait for Panasonic to sell their half of the Gigafactory back to them, during which they'd not have any cell production going on. They basically need to build another US-factory before they even have a chance to become vertically integrated in the USA.
Panasonic in-fact just added another line in Nevada.
These cells are required for Model Y/3. The new cells that they produce themselves will be for some Model Y, Cybertruck and Semi.
I'm not sure why you insist on disagreeing with a simple fact, Tesla is building its own battery plants. They have one battery factory in the US and are already building a second one. This is literally a fact.
1) You hope to make it available to external customers and turn it into a scale business. Your internal use is just the first customer. You need your own hardware so you can keep costs down at large scales.
2). Someone's pet project was to design their own ML ASIC and they thought it would look very good on their CV, and the CEO took the bait.
Hopefully it's the first case!
For all that Elon Musk values openness in some areas (e.g. putting Hyperloop into public domain), he prefers keeping stuff as vertically integrated as possible for everything he deems to be essential to the business, for maximum control.
Tesla has stakes in lithium mining operations, SpaceX has their own metallurgy team and IIRC also a foundry, and while they are using a Liebherr crane at the moment they are thinking about building their own. And for that, it makes sense - SpaceX is only one of Liebherr's customers while SpaceX depends on a crane that fits their needs - so either they get Liebherr to customize their crane or they build their own.
Does Tesla need tens of thousands of these things?
The "tiles of tiles" chip architecture seems like an Elon-obvious, let's just scale what we have approach. Do their neural networks map to that multiscale tiling well?
The WSE2 is much larger obviously, but I would also think it can result in a large performance boost given everything is on a single chip.
It says TSMC 7nm - is that DUV or EUVL?
[1]: https://en.wikichip.org/wiki/7_nm_lithography_process#TSMC
We couldn’t put 50B transistors on a square inch in the 1960s, though. We can now. https://en.wikipedia.org/wiki/Transistor_count lists several larger designs.
So, the engineering is impressive, but not spectacular.
Also, this being a grid of interconnected CPUs means the design is simpler than a single design filling the entire die would be. It’s ‘just’ repeating the same design over and over (possibly with some small variations near the edge)
Of course looking at it without knowledge of the state of the art it is astounding that we can even think of constructing machines with 50 billion working parts