M1 Ultra About 3x Bigger Than AMD's Ryzen CPUs
tomshardware.com
tomshardware.com
Also, let’s not forget that this heat-spreader includes the RAM under it.
Here’s the funny bit: the silicon die for M1 Ultra is actually more than 4x bigger than Ryzen 5000 series. M1 Max was 432 mm2; implying 864 mm2 of silicon. Ryzen 5000 has a 84 mm2 CCD and 124 mm2 IO die.
This was a major reason that AMD was able to price Ryzen so aggressively even at the high end - a high-end CPU is just made of more of the same small dies used in low-end CPUs, albeit requiring a better bin, instead of having to make a much larger single die.
The constant defect rate isn't typically a correct assumption across different products (e.g. different CPU or GPU dies) - different feature designs will have differing defect rates. Maybe Apple was able to design the M1 Ultra's features so that defect rate is very very low, though - I don't really know much about that silicon.
After all, Cerebras made a functional wafer sized chip.
Yes, but the Cerebras wafer-chip is a repeating pattern of logic, memory, and interconnect tiles. Each pattern is well under the reticle limit of the process node they're using.
And how many Cerebras wafers have shipped? I would imagine that they can afford to spend more engineering effort in the late manufacturing phase, fusing off bad bits, maybe salvaging some other bits... but enabling such late-stage functional-unit binning requires a larger mesh topography.
Apple is using a design that's at the limit of economic feasibility at the mass-market sales volume they can support. They can amortize the engineering expense down among the single most popular retail item on the planet.
Cerebras is an engineering marvel, but I wouldn't call it a mass-market item.
The Mac Studio more closely resembles high-end IBM POWER mainframe modules.
Stunning that you can just go out and buy one. And you don't need a data center to feed it.
https://www.anandtech.com/show/16626/cerebras-unveils-wafer-...
I'm not saying it's a good idea, because they'd have done it if it was. They clearly know what they're doing.
I'm just saying that it's not an absolute blocker for a design you'd otherwise be connecting over something like EMIB anyway.
Oh, I see: you are wondering why Apple M1 Ultra, as two chips linked by EMIB, couldn't just be the two chips adjacent on the wafer.
See the difference? Apple knows what their yields are; if they could be sure to get two matched SoC M1 Max chips right next to one another on the wafer, they could punch that pair out and bypass the whole EMIB thing.
I don't know enough about testing and dicing 100-billion-transistor chips to know how expensive that would be.
I don’t know if that means it’s physically one due though.
https://patents.google.com/patent/US20210217702A1/en?oq=2021...
Not entirely unlike what Intel used to do with Xeons before AMD re-entered the picture and what Nvidia kinda does with the likes of the A100.
AMD's next round of APUs on 5 nm will offer a better die area comparison.
Someone correct me if I’m wrong!
Marketing dept has trumped physics.
It is kind of awesome to look at though.
https://www.macrumors.com/2022/03/17/m1-ultra-nvidia-rtx-309...
> Apple's M1 Ultra is essentially two M1 Max chips connected together, and as The Verge highlighted in its full Mac Studio review, Apple has managed to successfully get double the M1 Max performance out of the M1 Ultra, which is a notable feat that other chip makers cannot match.
Unless you're talking about power per watt? GPU workloads are generally inherently parallelizable so performance per watt is actually what generally matters, rather than raw FLOPS.
No, I am not in agreement with you, the RTX 3090 is clearly better at being a GPU.
better is subjective, it depends on your needs.
sometimes performance per watt is more important especially when you're mobile or trying to be conservative w/ electricity.
There is an obvious implicit context when you're comparing performance to an RTX 3090, and it is not mobility or your electric bill.
Personally the M1 Ultra would be better imo as a desktop GPU than RTX 3090 in certain contexts. Putting it in a van, traveling with it, etc.
I'll probably wait until the next go around of them, but I'm loving nearing desktop performance w/ reduced loads.
Considering the 3090 had over double the performance, the lower power parts may well be both faster and more power efficient. You'd have to do the comparison to know, if that's what you care more about.
The FPS was closer than double. The flawed benchmark the 3090 had double.
I'm not sure how the lower power RTX variants compare, maybe that's a closer comparison, has there been a review against those?
No? Certainly not on desktop computers.
But it's almost completely unimportant for always-plugged-in-to-wall-socket desktop chips. Because even the power consumption of stuff like the 3090 is still pretty insignificant on your power bill compared to up front cost, and relative to the power consumption of all the other stuff in your house.
Remember even the 3090 consumes very little power when idle - it's only pulling the rated ~300w when you're maxing it out playing heavy games or rendering or whatever.
RTX3090: 215,034
M1 Ultra: 102,156
The CPU and GPU tests finishes too fast for the system to ramp up to the full power load.
That's not to say that M1 Ultra is equal or better than 3090, just that Geekbench isn't the right test for this.
You can see more detailed analysis here: https://www.youtube.com/watch?v=pJ7WN3yome4
MacOS has nice UI? Great, I will use it. Anything computationally intensive? - Why bother? I would use only Linux on the back end. It’s very easy to couple mid-range Apple product with high-end Linux workstation at your desk. I don’t need to run compute-intensive Deep Learning load while I am at the coffee shop.
Well they're putting it in a desktop and claiming desktop performance...
If Apple is going to put up graphs saying they beat i9 and 3090RTX machines, it's worth investigating even if I plan to put Linux on it.
In this instance it seems like the reality doesn't match the hype as well as in the laptop space, but if it had I would surly consider the option.
> overstretched or overblown analogy / reductio ad absurdum
The m1 max drives its displays directly, the m1 via a starlink tripple 4k displaylink hub.
The problems, in my case, are specific to monitors that take a good few seconds to come back from standby. Macos seems to be impatient and considers them dead. I am down to just 1 screen with this problem, and will replace it at some point.
So, we may have to wait to see what Apple says or if they'll unlock the full performance in future software updates. They did do this in the past with M1 hardware a few times.
Same result here from another channel where it appears GPU isn't at full power atm: https://www.youtube.com/watch?v=0CUoHwMtRsE&t=1172s
>Apple has built a wide enough GPU that they can keep clockspeeds nice and low on the voltage/frequency curve, which keeps overall power consumption down. The RTX 3090, by contrast, is designed to chase performance with no regard to power consumption, allowing NVIDIA to get great performance out of it, but only by riding high on the voltage frequency curve.
https://www.anandtech.com/show/17306/apple-announces-m1-ultr...
This strategy is something that Anandtech's been calling out since the iPhone 5s chip.
>Brian and I have long been hinting at the sort of ridiculous frequency/voltage combinations mobile SoC vendors have been shipping at for nothing more than marketing purposes. I remember ARM telling me the ideal target for a Cortex A15 core in a smartphone was 1.2GHz. Samsung’s Exynos 5410 stuck four Cortex A15s in a phone with a max clock of 1.6GHz. The 5420 increases that to 1.7GHz. The problem with frequency scaling alone is that it typically comes at the price of higher voltage. There’s a quadratic relationship between voltage and power consumption, so it’s quite possibly one of the worst ways to get more performance.
> This strategy is something that Anandtech's been calling out since the iPhone 5s chip.
Not sure that applies to M1 Ultra that's clearly marketed as a desktop SoC, not a mobile device or laptop SoC. Optimizing for energy efficiency or battery life is not the same as optimizing for performance first, which is what M1 Ultra is designed for. Not to mention, there's a reason Windows and Linux have multiple power profiles that changes how schedulers and power loads work. macOS has "high power mode" for m1 on laptops and yet strangely, it is not even available on the Mac Studio.
Keep in mind that Apple showed a graph that clearly shows the GPU was at 100-110w. Why do that if they won't run it at that wattage or even talk about 3090 in the first place? Why ruin their reputation over a silly little thing? Everyone on the planet will easily debunk that.
Also, why even bother adding a high-quality cooling system that the Mac Studio clearly don't need since both CPU/GPU are going to be capped at lower power wattage?
We'll find out eventually one way or another.
The M1 Ultra is using the same CPU cores and GPU cores as an iPhone, just more of them.
It's long been their strategy across the board. Throw silicon die area at tons of execution units and run the chip at a lower clock speed for power efficiency.
For example, the 2012 A5X iPad chip had about the same die area as 4 core Ivy Bridge, which was huge for a mobile device SOC.
https://www.anandtech.com/show/6330/the-iphone-5-review/4
Since they are selling the whole widget, they aren't as dependent on minimizing die area to maximize their profit margin.
Not only that, but it's running them at basically the same clock speeds as the rest of the M1 family, and even as the iPhone. This suggests that they're running at not far above a best efficiency point, rather than way past the elbow of the voltage curve as is traditional for desktop chips.
Or do you think they are mostly using HVt cells?
That effect is actually so bad that I have previously experienced thermal runaway due to it (more power => higher temperature => more power => …)
GPUs and CPUs are not simply "pump more power in it to go faster hurr hurr". It is extremely likely that the M1's maximum performance is attained already. Otherwise, the M1 ultra would have pumped a little bit more wattage and gotten better single core perfs, but guess what, it doesn't.
I agree, they're not a simple "throw more power and it'll go faster" problem. However, that doesn't mean they do not go faster if you increase clock speed and give it a little more power.
After all, overclocking would be utterly pointless over the past two decades and Intel/AMD wouldn't benefit from the so-called turbo boost technology. (I know M1 doesn't have turbo boost).
Just to be clear, yes, there's a limit to everything of course. There's a balance where certain higher clocks would cause a bottleneck in the rest of the system since this is a SoC.
In this case, I used the 100-120w number because Apple used that number in their GPU graph showing the GPU running at ~105w, why bother showing that if it can't accomplish it? Here's the graph: https://images.anandtech.com/doci/17306/Apple-M1-Ultra-gpu-p...
Note that Apple shows 60w for their CPU graph (https://images.anandtech.com/doci/17306/Apple-M1-Ultra-cpu-p...), which did actually match up and does show the double CPU pref that Apple stated.
I didn't pull that number out of nowhere. If Apple used 60w in that graph, then I wouldn't be here.
> Otherwise, the M1 ultra would have pumped a little bit more wattage and gotten better single core perfs, but guess what, it doesn't.
That's because of two reasons:
1. There is only so much data wecan "process" within the same core, more power does not change the data itself, we can only increase the clock speed to finish it faster or widen the amount of data that can fit in the same core. Also of note that Intel Turbo Boost does actually help with single-core thread pref by boosting the single core speed beyond the base line. 2. M1 max clock speed is set to 3.2Ghz (single core), it can't go faster than this (set by Apple). Throwing more wattage here wouldn't change anything but throwing more power to allow all 16 P cores to hit 3.2 does improve performance as long as it isn't overheating. I don't think M1 Ultra does hit 3.2 on all 16 cores (it might have hit 60w max first; I might be wrong but I saw mostly just 3.0ghz in Max video but I have to go back to review it).
Again, this is not about single-core performance, this is about GPU performance, which is extremely parallelizable and more comparable to the multi-core performance instead.
(Although with the m1 line it actually makes a bit of sense for once - it is useful to the end user to set things up so that fans /never spin up/ and the laptop never gets hot to the touch. Whereas with their previous intel models they had terible cooling and still had all of those problems, while not actually achieving a size win over a lot of competitors "thin and light" models that actually had competent thermal design)
Better != Faster or Better != More powerful
Better means closer to fulfil the goals and often the goals are inversely proportional(i.e. cheaper, faster, cooler are the goals but faster meaning less cheaper and hotter).
Apple might be guilty of misrepresentation of raw computational power but with M1 the experience of using computers has become significantly much better.
Mercedes, unlike BMW with Mini and Audi with the VW parts bin has to outsource the low end stuff.
To be picky about this analogy, top speed in a car is an attribute that bears little on everyday utility. In fact, past the 110km/h here in Australia for instance, the only place you can enjoy the extra speed in on a track. %99.999+ of people won't be able to take advantage of anything for entire life of their vehicle.
On the contrary, instant acceleration is something every driver takes advantage of numerous times per day, when changing lanes, correcting for others mistakes, etc. Fuel cost savings and low maintenance saves every driver money. These are the attributes to pay attention to in a car in general, not the top speed.
Point is to pay attention to characteristics which are of utility towards the use case. Obviously, if you are getting a car to go as fast as you possible can on a track, you want a car with higher max speeds so there is that.
This is also why high quality TVs are harder to manufacture than high quality phone displays. You have a lot more waste when you need to throw out/recycle a TV screen compared to a phone screen. And both are considered bad when they have just one bad pixel.
Is this assumption true?
With smaller chips, you have a smaller "grid" for your dart to land on. So the total number of failures (darts) being the same, you still end up with less usable silicon, since the bigger grid means you throw away a lot more surface area with a failure.
Take a look at this picture of a failure map, if you would like some visuals: https://ars.els-cdn.com/content/image/1-s2.0-S09521976120008...
no. intel used to make extra cores for their 128 core cpus to account for the defects, and it was not rare to receive a cpu with more than 128 cores because it was a good one.
AMD uses an IO die manufactured on an older process at another manufacturer (14 or 16nm Global foundries) than their core chiplets (7nm TSMC). I think they even used the same IO die across multiple generations of EPYC/Ryzen, but I'm not sure.
This is also why the approach is an "Apple-only" one.
You're suggesting Apple devices cost between two and three times as much as the competition? I'd like to see that competition, so please show us those $280 Mac Mini and iPhone 13 Mini killers. Or, conversely, show us a superior device for the same price. Remember to match point for point all features of whatever Apple device you somehow believe costs 2-3x as much as it's third party clone.
Also, arguably, the IO Die is the SoC, the chiplets are external to the SoC even though they all live on the same interposer.
(which is a good thing. all-in-one for specialized stuff is fine, but I don't want it to become the norm - upgrading one component is cheaper than upgrading multiple, and tends to support longer hardware cycles)
RAM aside which is Apple's actual performance differentiator, all of these eat the M1 Ultra alive in their respective categories.
(at 864mm2 (2 x 432mm2 M1 Max's), that's a density of ~130MTr/mm2, which is in the ballpark for TSMC's max density for N5 (est. 170MTr/mm2) - N7 max density for reference is ~90MTr/mm2)
So no, they don't even bother to trim it off.
The memory isn't on die. It's just regular lpddr5 but on package. So that's not a yield concern.
As it is, this is a kind of clickbaity article
Soon it will be 6x, then more. Zen 5 small cores are supposed to be just zen 4 cores iirc.
I wonder if we will start seeing more physically large chips in the future because of this?
No, it's not logical, they still could have used a common form factor like M.2.
> CNVi or CNVio ("Connectivity Integration", Intel Integrated Connectivity I/O interface) is a proprietary connectivity interface by Intel for Wi-Fi and Bluetooth radios to lower costs and simplify their wireless modules. In CNVi, the network adapter's large and usually expensive functional blocks (MAC components, memory, processor and associated logic/firmware) are moved inside the CPU and chipset (Platform Controller Hub). Only the signal processor, analog and Radio frequency (RF) functions are left on an external upgradeable CRF (Companion RF) module which, as of 2019 comes in M.2 form factor (M.2 2230 and 1216 Soldered Down). Therefore, CNVi requires chipset and Intel CPU support. Otherwise the Wi-Fi + Bluetooth module has to be the traditional M.2 PCIe form factor.
>if you could somehow link up multiple GPUs with a ridiculous amount die-to-die bandwidth – enough to replicate their internal bandwidth – then you might just be able to use them together in a single task. This has made combining multiple GPUs in a transparent fashion something of a holy grail of multi-GPU design. It’s a problem that multiple companies have been working on for over a decade, and it would seem that Apple is charting new ground by being the first company to pull it off.
https://www.anandtech.com/show/17306/apple-announces-m1-ultr...
How is it not relevant in this thread?
Frankly, it's annoying, and pretty far from intellectually stimulating. What is there to discuss anyways? Do we all need to pat the world's largest tech conglomerate on the back for doing the same thing as other companies did 6+ years ago, with worse performance results and misleading marketing material to boot?
None of those technologies make multiple GPU dies look like a single physical GPU to software.
As covered by the Anandtech quote you dislike so much:
> This has made combining multiple GPUs in a transparent fashion something of a holy grail of multi-GPU design. It’s a problem that multiple companies have been working on for over a decade, and it would seem that Apple is charting new ground by being the first company to pull it off.
However, that article does go into additional detail.
>Unlike multi-die/multi-chip CPU configurations, which have been commonplace in workstations for decades, multi-die GPU configurations are a far different beast. The amount of internal bandwidth GPUs consume, which for high-end parts is well over 1TB/second, has always made linking them up technologically prohibitive. As a result, in a traditional multi-GPU system (such as the Mac Pro), each GPU is presented as a separate device to the system, and it’s up to software vendors to find innovative ways to use them together. In practice, this has meant having multiple GPUs work on different tasks, as the lack of bandwidth meant they can’t effectively work together on a single graphics task.
https://www.anandtech.com/show/17306/apple-announces-m1-ultr...