2022 Mac Studio (20-core M1 Ultra) Review
hrtapps.com
hrtapps.com
If this solver relies on matrix multiplication and uses the macOS Accelerate framework, you are seeing this speedup because M1 Macs have AMX matrix multiplication co-processors. In single precision GEMM, the M1 is faster than an 8 core Ryzen 3700X and a bit slower than a 12 core Ryzen 5900X. The M1 Pro doubles the GFLOPS of the M1 (due to having AMX co-processors for both performance core clusters). And the M1 Ultra again doubles the GFLOPs (4 performance core clusters, each with an AMX unit).
Single-precision matrix multiplication benchmark results for the Ryzen 3700X/3900X and Apple M1/M1 Pro/M1 Ultra are here:
https://twitter.com/danieldekok/status/1511348597215961093?s...
M1 only has 4 performance cores and they still ran it with 16 threads. So it's not like they are trying to skew the benchmark by not showing the potential of the 5900x.
Also the 5950x is 2 years old and still is an active product.
Also, maybe I can only benchmark on the hardware that I actually have?
In this case that would mean ThreadRipper systems and not Ryzen.
If you want a comparison between a Threadripper system and an M1 Ultra system, I'd be fascinated to read your blog post with your results.
The performance scaling you see between systems pretty much corresponds to the memory bandwidth in those configurations.
Note that on the M1, the CPU can only access a fraction (about 25% iirc) of the total memory bandwidth, you have to use the GPU to really get the full performance of the M1 here.
Also keep in mind that normal x86-64's, even without an IGP only get about 60-65% of peak, even with nothing else sharing the memory bus. I often see this quantified with McCalpin's stream benchmark.
So the M1 Ultra likely has a pretty impressive memory bandwidth of around 440GB/sec, which isn't a large fraction of 800GB/sec, but it still more than any other desktop or server chip I know of. The AMD Epcy maxes out at 8 channels of DDR-3200, which is in the neighborhood of 208GB/sec peak, with an observed bandwidth of 110-120GB/sec.
Correct. The numbers we have are from their M1 Max deep dive, with the M1 Ultra being two M1 Max chips fused together.
For CPU cores:
>Adding a third thread there’s a bit of an imbalance across the clusters, DRAM bandwidth goes to 204GB/s, but a fourth thread lands us at 224GB/s and this appears to be the limit on the SoC fabric that the CPUs are able to achieve, as adding additional cores and threads beyond this point does not increase the bandwidth to DRAM at all. It’s only when the E-cores, which are in their own cluster, are added in, when the bandwidth is able to jump up again, to a maximum of 243GB/s.
https://www.anandtech.com/show/17024/apple-m1-max-performanc...
GPU cores:
>I haven’t seen the GPU use more than 90GB/s (measured via system performance counters). While I’m sure there’s some productivity workload out there where the GPU is able to stretch its legs, we haven’t been able to identify them yet.
Other:
>That leaves everything else which is on the SoC, media engine, NPU, and just workloads that would simply stress all parts of the chip at the same time. The new media engine on the M1 Pro and Max are now able to decode and encode ProRes RAW formats, the above clip is a 5K 12bit sample with a bitrate of 1.59Gbps, and the M1 Max is not only able to play it back in real-time, it’s able to do it at multiple times the speed, with seamless immediate seeking. Doing the same thing on my 5900X machine results in single-digit frames. The SoC DRAM bandwidth while seeking around was at around 40-50GB/s – I imagine that workloads that stress CPU, GPU, media engines all at the same time would be able to take advantage of the full system memory bandwidth, and allow the M1 Max to stretch its legs and differentiate itself more from the M1 Pro and other systems.
Where's the 800GB/s from?
That number for the M1 Ultra (from the OP's post) = 800GB/sec. McCalpin's stream benchmark is often cited as a practical/useful number for usable bandwidth using a straight forward implementation in C or Fortran without trying to play games, much like the vast majority of codes out there.
Also note that the x86-64's in the world use a strict memory model that results in a lower fraction of observed bandwidth vs peak. Arm has a looser memory model which achieves a higher fraction of peak.
The thing that's interesting for me around this is: I hate the on-die memory. I hate the idea of not being able to upgrade after your order, or a few years down the road. It has both practical problems and offends my inner nerd.
But.
This is a useful example of why there's value in it. It seems unlikely that you'd get as good a result with the traditional memory architecture.
Also, if you really need more memory - buy the config with more memory and sell the old config. Resale on macs is usually superb.
I can't remember hearing a friend or gaming buddy say "I upgraded my ram." I haven't upgraded ram in any machine I've owned in twenty years and even before that, it was rare. There was never any point in putting that sort of money into an out of date CPU and memory architecture.
If you owned a trashcan Mac Pro and were a working creative, would you be upgrading its memory this month? Nope...
That said, if the memory is fixed like on the Ultra, then sure, we'll buy what we need (and a bit more).
This is supported by taking the code found in the gist [2] linked from my SO answer and running it on my M1 Pro. Compiling it, we get `dgesv_accelerate` which uses Accelerate to solve a medium-size linear algebra problem, that typically takes ~8s to finish on my M1 Pro. While running, `htop` reports that the process is pegging two cores (analogous to the result in my original SO answer on the M1 pegging one core; this supports the idea that the M1 Pro contains two AMX co-processors). If we run two `dgesv_accelerate` processes in parallel, we see that they take ~15 seconds to finish. So there is some small speedup, but it's very small. And if we run four processes in parallel, we see that they take ~32 seconds to finish.
All in all, the kind of linear scaling shown in the article doesn't map well to the limited number of AMX co-processors available in Apple hardware, as we would expect the M1 Max to contain maybe 8 co-processors at most. This means we should see parallelism step up in 8 steps, rather than 20 steps as was shown in the graph.
Everything I just said is true assuming that a single processor, running well-optimized code can completely saturate an AMX co-processor. That is consistent with the tests that I've run, and I'm assuming that the CFD solver he's running is well-written and making good use of the hardware (it does seem to be doing so from the shape of his graphs!). If this were not the case, one could argue that increasing the number of threads could allow multiple threads to more effectively share the underlying AMX coprocessor and we could get the kind of scaling seen in the article. However, in my experiments, I have found that Accelerate very nicely saturates the AMX resources and there is none left over for future sharing (as shown in the dgesv examples).
Finally, as a last note on performance, we have found that using OpenBLAS to run numerical workloads directly on the Performance cores (and not using the AMX instructions at all) is competitive on larger linear algebra workloads. So it's not too crazy to assume that these results are independent of the AMX's abilities!
[0] https://stackoverflow.com/a/67590869/230778 [1] https://stackoverflow.com/a/69459361/230778 [2] https://gist.github.com/staticfloat/2ca67593a92f77b1568c03ea...
NVIDIA GPUs have much more ubiquitous programming models, from CUDA C++ to OpenMP (both C++ and Fortran), allowing them to be much more useful.
AMD ROCm (a CUDA clone, but not managed too well...) isn't in a great state today, but it definitely provides an OpenMP compiler, for C++ only though, Fortran is not covered.
Intel GPUs have a much better SW story on all fronts for GPGPU than AMD's, and OpenMP is supported for both C++ and Fortran there.
tldr: if you want to run this code on GPUs, the two vendors to look for are NVIDIA and Intel. The others just don't have the software stack to do so. (without more heavyweight porting)
Of course, optimising for GPUs also adds work. That said, an OpenMP version is a very good starting point.
The M1 architecture only allows the CPU to access a fraction of the memory bus, it is around 25% of the total bandwidth or so iirc. The rest has to go to the GPU or NPU or one of the other units, so even if your FP64 code was super shitty emulated with no hardware support, on paper you should be able to quadruple performance by utilizing it.
If you can "pre-digest" it at all, and then push it back from the GPU to the CPU, that might be another viable approach.
Of course none of this will exist out of the box in some 90s FORTRAN code!
That's what's done by Intel for their newer consumer GPUs AFAIK, zero FP64 units present on die, but an emulation mode can be enabled to expose that support. (which is then done by software on the GPU)
> it is around 25%
50%, ~ 200GB/sec on an M1 Max out of ~ 400GB/sec
For CFD it is a different problem. In each time-step you essentially solving a linear system Ax=b, so even if you are looking at element x(100,100,timeperiod2) you need to also know the value of the element x(1,1,timeperiod1).
There are some algorithms that decouple the problem by introducing residuals for each element and then trying to iteratively reduce them, but as you can see, it is not a linear increase in speed as someone would had expected by looking at the GPU specs.
TLDR: Yes you can do CFD with GPUs, but dont expect miracles.
Nobody uses GPUs for "linear speedups by looking at GPU specs". That's not how accelerators are measured.
It is not standard in OpenFoam (only though rapidCFD) or Ansys (only this year they made an announcement for full acceleration).
"GPUs are fantastic for solving linear systems."
*Some* linear systems. They are horrible for sparse algebra.
So this entire thread is kind of pointless with regard to many real-world use cases.
> “If a CPU doesn't run the OS and software I need to run, it might as well be a GPU or …”
> So this entire thread is kind of pointless with regard to many real-world use cases.
Does this kind of reasoning apply to everything that doesn’t suit your unique personal needs?
Personally I think this is an interesting thread, but I don’t use the qualifier of “is it useful to me” to determine whether something is interesting.
Philosophical purity rarely provides top performance. And that's OK, slow and steady works for me, because I'm not doing max performance computing.
That second chart with the "amazing" redline is comparing it to 2010 through 2017 era CPUs. Now compare it to a current generation zen3 based $599 AMD CPU in a $350 motherboard with $450 of RAM.
Or a $599 Intel 12th-gen "core" series.
For $3999 nevermind $7999 you can build one real beast of a workstation that fits in a normal midtower ATX case.
Just because you can afford an absurd $4000-8000 Apple computer doesn't mean you want to waste your money on it. I have a finite amount of money and choose to spend it on other things.
Apple is fine for like, a $1500 laptop. The macbook air is a fine product.
Not for an actual workstation.
[1] hackintoshes need not apply.
They're not vertically integrated, they're sourcing and assembling the same parts you and I could get (and some we can't: they have graphics cards in stock!)
That's for folks that don't need OSX, that like build quality in the form of sturdy, repairable, maintainable tools rather than glossy but glued-together disposable ones, and want Ryzen or Intel performance that will beat a Mac Studio without the hassle.
Edit: I'd say that the Ultra is a much better bang for your buck than the Threadripper, unless you absolutely need humongous amounts of RAM or macOS does not work for you.
Over a few years of usage, that cost is not insignificant, especially if you are using air conditioning part or all of the year.
Of course, the argument can be made — and it is not without merit — that power consumption is secondary to performance for a power or professional user. But the low power consumption of M1 series allows Apple to deliver pretty much unprecedented performance for the given size. Ultra offers performance of large workstation tower in more compact form factor than the smallest HTPC. My M1 Max 16" has the performance of a large workstation laptop with the portability and battery life of an ultrabook — and I can use all that performance while working untethered. It's quite interesting to see all these benchmarks where folks show that the latest ALD laptops can marginally outperform the M1 Pro/Max on the desk, under ideal conditions, while in reality the performance will plummet really fast when you actually want to be mobile. In the meantime I enjoy my desktop-level build times while working on my sofa or on the train.
I've had to take frozen peas out of the freezer and rub them desperately against the bottom of my x86 MBP to quickly cool it down enough to use, in order to make important Zoom meetings on time. It totally ruins the peas, and can't be good for the computer!
Even if they were both the same speed, it would still be so much better, just because of how cool it runs. How fast and long it runs on batteries is just icing on the cake of not getting fucked by speedstep's "kernel_task % CPU 798.6" all the time.
The power requirements of the machine you describe would easily be more than double that of the Mac, but you probably aren't getting anywhere close to double the performance.
Besides, if you care about performance, you’re going to want to do more than just replace the RAM
Unless you live in a cave and using your excrement as manure I could apply the same logic to your lifestyle.
Realistically the only practical difference is virtue signaling (I've seen a bunch of hipsters bring up these kind of irrelevant talking points while discussing their vacation in some exotic destination 5 minutes later).
And like someone else said, Apple devices are the definition of throwaway consumer products designed for a limited shelf life.
I wish I had a 5950X that requires less power. I don't. I am going to choose this CPU over any mac any day because it's faster. A lot of other CPUs consume less electricity, I don't use them neither. End of the story.
PS: I once worked on a project where 10 consultants sat around burning money for over a month because the customer couldn't figure out how to spin up a lab environment. We walked down to the local computer shop across the street, ordered the beefiest Xeon workstation tower we could build, packed it with drives, put VMware ESXi on it, and then the work could finally start. The "research, ordering, build, and troubleshooting" took half a day. It saved hundreds of thousands of dollars of lost time and money.
Docker is definitely a lot slower
I have read of massive performance improvement with Linux on m1pro though so it might be a swap to that when more distros are available.
Question is, has MacOS become bloated or not had attention to performance to make best use of the new hardware?
Huge, and I mean HUGE improvement over my previous 16" Intel, which seemed to make it around 3 hours before calling it a day...
Also, 6-8 hours is a full work day of laptop time for me. If you have more hours in the day you probably will get less days out of the laptop than me.
Plus for PC gaming I’ve been using a cloud server via Paperspace and I play games like Cyberpunk with 4K on ultra settings with pretty good latency (thanks to 1.5 gigabit fibre) and even some VR games using Steam VR + Parsec on a Quest 2.
I really don’t see the point in building a PC anymore. And I was really really close to building a Ryzen one before I discovered Paperspace/https://shadow.tech/
Docker on macOS is busted. I have heard that folks have had much better success with alternative implementation such as nerdctl. M1 virtualisation on itself is very fast and has almost no overhead.
> I have read of massive performance improvement with Linux on m1pro though so it might be a swap to that when more distros are available.
I wouldn't hold my breath. Doubt that M1 distros will ever become more than an impressive tech demo.
> Question is, has MacOS become bloated or not had attention to performance to make best use of the new hardware?
MacOS will always be "more bloated" than a basic Linux installation, it's an opinionated, fully features user OS that runs many more services in the background, e.g. code verification, filesystem event database, full-disk indexing etc. But the CPU/GPU performance is generally excellent. Of course, it boils down to the software you are running, if it is a half-assed port (like Docker on Mac) seems to be, it will eat up any advantage the hardware offers.
https://multipass.run/docs/docker-tutorial
And I second the other commenter's recommendation of VSCode's Remote Container extension.
https://code.visualstudio.com/docs/remote/containers
And possibly using Podman instead of Docker.
https://opensource.com/article/21/7/vs-code-remote-container...
For improving performance, you can enable some of the newer experimental features in latest docker (Virtualization Framework and VirtioFS). The combination works really well except for databases [1] due to the way FS sync is handled. To fix that a setting need to be changed in the linux VM that docker uses. [2]. Hopefully docker will make that a default setting in the future.
[1] https://github.com/docker/roadmap/issues/7#issuecomment-1042... [2] https://github.com/docker/roadmap/issues/7#issuecomment-1044...
Docker is just... faster on Linux. This has been the case for a while, and it's not just kernel-based stuff causing problems: APFS and MacOS' virtualization APIs play a pretty big role in weighing it down.
I'm kinda in the same boat, though. I got a work-issued Macbook that kinda just sits around, most of the time I'll use my desktop or reach for my T460s if I've got the choice. Mostly because I do sysops/devops/Docker stuff, but also because I never really felt like the Mac workflow was all that great. To each their own, I guess.
The downside, of course, is that it's a little annoying to keep your code in the VM if you'd otherwise have it on the host. But I still find it worth it because of just how much faster it is. Now that I think about it, I wonder how well it would work to do your own host mounting using either the Parallels guest folder sharing feature or SMB. I never tried it.
A bigger impact is using the virtiofs (beta) file system mapping/sync. The existing one is horrendously slow to the point of being unusable.
MacOS is somewhat bloated though. There's insane amounts of garbage running in the background.
MacOS is a special beast, because if I run a VM myself inside parallels using Docker (Ubuntu server with the docker snap, for example), I get nearly 5x the performance from that VM than I do just the "regular" docker for mac (including when all of the new mac specific experimental options are enabled for better performance).
It's pretty atrocious, I find it completely unusable and just stick to running my containers in Parallels and forwarding the ports.
I don't get why people still bother with docker for mac...
First, you run virtual machine vs just running natively on Linux.
Second, you emulate x86; and, unlike other x86 that are emulated directly on M1, you emulate it in software through QEMU, because M1 cannot virtualize and run x86 at the same time
So it’s just really slow.
I'm writing this on a cheap core i5 system with intel iris xe graphics. Mediocrity is the name of the game for this type of laptop. Everything is mediocre. The screen, the performance, the touchpad, etc. The only good thing about this system: no blazing fans. That seems to be a thing with most SOC based PCs/laptops. Mediocre performance. Non SOC based solutions exist of course but they suck up a lot more power to the point where laptops start having cooling issues and have to do thermal throttling. I've experienced that with multiple intel based macs. You spend all that money and the system almost immediately starts overheating if you actually bother to use what you paid for.
I actually used to have a wooden plank that I used to insulate my legs from my mac book pro. It was simply too uncomfortable to actually use on my lap.
This guy def knows his math (and working with Apple apps).
So with about half the cores (16 vs 28) and twice the bandwidth (say 420GB/sec vs 180 GB/sec) it manages twice the performance. Looks pretty impressive to me. Looks like the Apple is significantly less memory bottle necked than the 6 channel Xeon W.
The perfect core could be dialed from fanless to water cooling with linear performance, but it doesn't quite exist today. Intel chips have top end, Apple chips have low end, but an i7 fed 3W isn't going to perform and an M1 can't take any more voltage.
The tension is figuring out how to build very large structures to perform useful work with gobs of power, yet still scale down to low power budgets. Imagine a core that can dynamically morph from P core to E core & back on the fly.
But it must be said, I've been really tempted by the Macs. I'm not sure why, the 3 main things I do with my personal computer (game dev, playing games, watching things) are things that linux/windows can do at least as good if not better than a Mac, and yet here I am, holding off for months just trying to be convinced into the apple ecosystem.
I think its probably just the simplicity of it. I really like the idea of replacing a load of bulk (3 monitors, the vesa mount to hold them, big keyboard, mouse, bulky tower and a metric tonne of cables) with just a single laptop that I can pickup and go with at a moments notice, though Im not sure of the practicalities of that, at least for me. I don't even travel much, it just seems nice.
Edit: I know that a lot of people have a mac and a desktop to fill all needs, but that somewhat defeats the simplicity of it for me. One bulky computer is simpler for me than one bulky computer for some things and another much smaller computer for others.
Of the three main things you do, i think only watching things is comparable between macOS and Windows/Linux. Gaming is nonexistent on macOS, unless you stream ( either a cloud service like Stadia/GeforceNow or locally from a PC with Steam in-house streaming/Parsec), and i can't imagine doing game dev somewhere you can't even test.
Absolutely, which is a large part of the reason I'm not sure its practical for me, but its certainly true that one monitor is simpler than three. I suppose I could eventually get used to just having one monitor, it might even have some advantages, but I'm not sure I'd want to.
> A nice middle ground is a laptop with a dock with monitors
This is what I've currently got (laptop is a razer blade). In my case at least it just ended up being more messy and difficult than a desktop. Most docks for example are designed for business use and therefore don't provide "clean" USB power, which I found causes issues with audio interfaces and DACs. I just ended up having a small, poorly cooled, difficult to upgrade desktop.
After that, I decided no more half measures. I either embrace the minimalism of a laptop alone, or return to a desktop. I know a dock works well for some people, sadly not for me.
> Gaming is nonexistent on macOS... i can't imagine doing game dev somewhere you can't even test
This is exactly it (plus game dev for apple sillicon is still somewhat lagging behind), it doesnt seem to be a practical option, yet as I still, here I am.
Of course theres also non-apple laptops as well which would provide similar simplicity. I love my thinkpads, but I'm just not sure they're powerful enough for me to replace a desktop with, and as for the more powerful of the non-apple laptops, the blade has given me a mistrust of them (only had the thing for a year and the battery has bulged beyond being able to fit it in the chassis)
The percentage of all those can change as you change the number of cores, and the different levels of the memory hierarchy get different levels of contention, latency limits, or bandwidth limits. So in an ideal world you can draw a graph and extrapolate, but in the real world you might do significant better or worse. There are cases (admittedly rather rare) where performance increases by more than linear in relation to cores.
Ooh that's interesting, do you have any leads or examples?
It's seen sometime even on multinode jobs. But it's definitely the exception to the rule, but it does happen. It's helped by the fact that each level of cache is approximately 10x faster, so fitting in a cache is a huge win. So running in ram on 1 node might be 20x slower than running in cache on 10 nodes.
It's not the display of the benchmarks that isn't working, but the running of them.
Personally the fact they didn’t produce Linux drivers for these things is baffling. It would just make more people buy them. So what if they don’t run macOS? I would buy a M1 air tomorrow if they published Linux drivers for the full stack.
Followed by
> helping Linux grow doesn't financially benefit them, now or in the future
Is a bit perplexing. Why are you baffled?
Opportunity cost is a very real thing. Apple would probably be better served by developing Windows drivers, if they were going to take developers away from macOS drivers.
Apple, like every company, is limited by the speed at which their supply chain can produce specific parts. And, while they're often the largest customer for a given manufacturer, they're rarely the only.
And why wouldn't you, when it all becomes heat in your room? I'm not saying every single watt matters but, in cases where you're putting the M1 ultra at 40W against a 280W threadripper or Epyc... yeah that hugely matters.
Or you offload it to GCP/AWS but then the wattage of the M1 Ultra is still irrelevant because you could use a Chromebook to do that.
In my experience, a ~500 W gaming PC does not generate any noticeable heat unless you use it in a small room behind a closed door. A 1450 W vacuum cleaner does, so I guess heat becomes significant somewhere around 1 kW of sustained power.
By the way, the M1 Ultra has a 370W power supply, so really the question is do you really care about the difference between 300 and 400 watts in this use case?
yes, many people do, especially since that becomes heat in your room.
> By the way, the M1 Ultra has a 370W power supply, so really the question is do you really care about the difference between 300 and 400 watts in this use case?
Complete non-sequitur, you could put a 1.5kw power supply on an R7 5700G or an i3, that doesn't mean it pulls 1.5kw.
you also should be able to express this without using the flamebait style, that is not appropriate or welcomed on this site. Instead of "by the way...." you can simply say "the M1 Ultra pulls 370W so ...." (or you could, if that were true).
We know that reported power consumption is inaccurate on M1 devices, from Anadtech's testing of the M1 Max, and I doubt it's very different on the Mac Studio.
In any case, the difference is not going to be much.
As far as heating a room, really, I can assure you that you will barely notice 100W. There are screens that will use more power than that (such as, amongst others, the XDR screen). I say this as someone who lived with a 900W power hog of a computer in places where temperature hits 42 in shade.
It looks like the max power consumption of the Mac Studio is 215w. https://support.apple.com/en-us/HT213100
The Studio Ultra has six Thunderbolt 4 ports. Each port is required to deliver at least 15W by the TB4 standard, so that’s already 90W. Add some budget for the USB-A ports and you are near 100W for power delivery.
(Probably not much over that though.)
A 12900K has a max TDP (for the chip, not including ram) of 210 watts, and loses on geekbench5 to the M1 ultra by 1/3rd or so. To get closer you'll need a i9 or threadripper and systems with either of those typically need quite large power supplies and still have a small fraction of the memory bandwidth.
Btw most of that price is for certification for industry software. Does the Mac have that? If not, it is just immediately not an option.
Or even the 440GB/sec of memory bandwidth available to the CPUs.
Sure if you are cache friendly enough and you get enough Zen3 or Intel cores you can win, but you end up spending a fair chunk of change, getting less memory bandwidth, and for a clear win you often need to spend more, like say getting a Lenovo Threadripper (and they have a exclusive rights to the chip for 6 months or something).
"Nano-texture glass" is pretty much just what all screens were like back in the days of CRTs and pre-glossy flat screens. Now Apple are charging $300 for it!
I remember the crackling, if you wiped them, within a few minutes of last being on.
Nowadays, there's lots of folks that think "CRT" is a political hotbutton topic.
Well, who wants big government shoving their noses into how efficient our screens are, and whether they have leaded glass or not. :P