[0]https://www.amd.com/en/products/cpu/amd-ryzen-threadripper-3...
[0]https://www.amd.com/en/products/cpu/amd-ryzen-threadripper-3...
Our "world" build is slightly faster on my M1 Max.
https://twitter.com/kiratpandya/status/1457438725680480257
The 3990x runs a bit faster on the initial compile stage but the linking is single threaded and the M1 Max catches up at that point. I expect the M1 Ultra to crush the 3990x on compile time.
(+ now I see it's rust: how parallel is your build, really?)
Not the OP but I install a lot of Rust projects with Cargo and recently did some benchmarking on DigitalOcean's compute-optimized VMs. Going from 8 cores to 32 cores was a little disappointing:
Bat (~40 crates): 68s -> 61s
Nushell (486 crates): 157s -> 106s
Compilation starts out highly parallel and then quickly drops down to a small number of cores.
If the final x86 production build takes longer it doesn’t matter - that happens on the cloud anyway.
Edit: Rust builds are very parallel until linking. No different than any other LLVM build.
It matters when comparing CPU performance, which is what this benchmark is being used for.
Isn't linking IO-bound?
Does it, though?
I mean, if you read that link you'll notice it boasts the linker's performance by comparing it with cp and how it's "so fast that it is only 2x slower than cp on the same machine."
Is cp supposed to be CPU-bound?
https://llvm.org/devmtg/2017-10/slides/Ueyama-lld.pdf
There is a breakdown in those slides discussing what parts of lld are single threaded and hard to parallelize so I suspect single thread performance plays a big role too. I generally observe one core pegged during linking.
That would mean that these comparisons between Threadripper and the M1 Ultra do not reflect CPU performance but instead showcase whatever choice of SSD they've been using.
Why did you omit the reference to "file system"?
Are we supposed to ignore the fact that a linker's main job is reading object files and write the output to a file?
I find this sort of argument particularly comical given a very old school technique to speed up compilation is to use a RAM drive to store the build's output.
Just that with the same hot caches, the average change-build-test loop that developers do 100+ times a day is just faster on the M1 Max.
Curiosity got the better of me:
Try the same thing with mold.
But you know, I'm still happy to double my current build perf in a small box I can stick in my closet. Ordered one :-)
Also 1st gen threadrippers are getting on a bit now, surely. It's a ~6 year old microarchitecture.
The above statement should also relate to most other C/C++ projects.
I am curious whether there is a real performance difference?
I do lots of computing on high-end workstations. Intel builds used to be extremely expensive if you required ECC. They used that to discriminate prices. Recent AMD offerings helped enormously. I wonder whether these M1 offerings are a significant improvement in terms of performance, making it worthwhile to cope with the hassle of switching architectures?
Threadripper 3990X get about 25k in Geekbench Multicore [1]
M1 Max gets about 12.5k in Geekbench Multicore, so pretty much exactly half [2]
Obviously different tasks will have _vastly_ different performance profiles. For example it's likely that the M1 Ultra will blow the Threadripper out of the water for video stuff, whereas Threadripper is likely to win certain types of compiling.
There's also the upcoming 5995WX which will be even faster: [3]
[1] https://browser.geekbench.com/processors/amd-ryzen-threadrip...
[2] https://browser.geekbench.com/v5/cpu/search?utf8=%E2%9C%93&q...
[3] https://www.amd.com/en/products/cpu/amd-ryzen-threadripper-p...
Maybe the cooling and power delivery difference between laptop formfactors and PC formfactors will be less with these new arm based chips.
If I was to guess, the increased cooling probably helps the Studio sustain similar boost clocks as the laptops, but for longer.
Although it’s possible these are on N4x, which might increase the attainable boost.
24-core scores 20k, 32-core scores 22.3k, and 64-core score 25k. Something isn't scaling there.
At the same time the Threadrippers also have a gargantuan amount of cache that can be accessed at several hundred gigabytes per second per core. Obviously not as nice as being able to hit DRAM at that speed.
Also, not everything fits into cache.
I use a few 32 and 64 core machines for build servers and file servers, and while the 64-core EPYCs are not twice as fast as the 32-core ones due to lower overall frequency, they're 70% or so faster in most of the things I throw at them.
I was under the impression that all of their multi-core tests were "run N independent copies of the single-threaded test", just like SPECrate does.
It sounds pointless to come up with synthetic benchmarks which emulate software that is not able to handle hardware, and then use said synthetic benchmarks to evaluate the hardware performance.
Most consumers are software aware, not hardware aware. They care what they will use the hardware for, not what they can use it for. To that end, benchmarks that correlate with their experience are more useful than a tuned BLAS implementation.
You could still write OpenCL kernels of course. Doesn't mean you can't use it, but not sure if it's all just accessible to CPU-side code.
(or maybe it is? it's still a damn fast piece of hardware either way)
Linking this[1] because TIL that the memory bandwidth number is more about the SoC as a whole. The discussion in the article is interesting because they are actively trying to saturate the memory bandwidth. Maybe the huge bandwidth is a relevant factor for the real-world uses of a machine called "Studio" that retails for over $3,000, but not as much for people running postgres?
1 - https://www.anandtech.com/show/17024/apple-m1-max-performanc...
https://semiaccurate.com/2022/03/08/amd-finally-launches-thr...
The niche for high clocks was arguable with the 2nd-gen products but now you are foregoing v-cache which also improves per-thread performance, so Epyc is relatively speaking even more attractive. And if you take Threadripper you have artificial memory limits, half the memory channels, half the PCIe lanes, etc, plus in some cases it's more expensive than the Epyc chips. It is a lot to pay (not just in cash) just for higher clocks that your 64C workloads probably don't even care about.
AMD moved into rent-seeking mode even before Zen3 came out. Zen2 threadripper clearly beats anything Intel can muster in the segment (unless they wanted to do W-3175X seriously and not as a limited-release thing with $2000 motherboards) and thus AMD had no reason to actually update this segment when they could just coast. Even with this release, they are not refreshing the "mainstream" TRX40 platform but only a limited release for the OEM-only WRX80 platform.
It was obvious when they forced a socket change, and then cranked all the Threadripper 3000 prices (some even to higher-levels than single-socket Epyc "P" skus) what direction things were headed. They have to stay competitive in server, so those prices are aggressive, but Intel doesn't have anything to compete with Threadripper so AMD will coast and raise prices.
And while Milan-X isn't cheap - I doubt these WRX80 chips are going to be cheap either, it would be unsurprising if they're back in the position of Threadripper being more expensive for a chip that's locked-down and cut-down. And being OEM-only you can't shop around or build it yourself, it's take it or leave it.
Don't worry though there will still be room to move the goalposts with "uhhh, but, Apple is designing for high IPC and low clocks, it's totally different and x86 could do it if they wanted to but, uhhh, they don't!".
(I'm personally of the somewhat-controversial opinion that x86 can't really be scaled in the same super-wide-core/super-deep-reorder-buffer fashion that ARM opens up and the IPC gap will persist as a result. The gap is very wide, higher than 3x in floating-point benchmarks, it isn't something that's going to be easy to close.)
Work out the IPC there - the Intel has a 2x thread count advantage, a 17% clock advantage, and Apple comes out 5% ahead. So the IPC gap there is about 2.46x.
It's not a perfect comparison of course, since we're mixing SMT and big/little cores, but in basically every area Intel should (on paper) have more resources available and Apple is coming out on top anyway by sheer IPC.
That's what I'm saying - you can't really do that approach with x86. It's not power-advantageous or transistor-advantageous to go super wide on the decode or reorder buffer like that on x86. And regardless of the tricks x86 uses to mitigate it, you've still got a 2.5x IPC gap at the end of the day. A 2.5x IPC gap will not be closed up by just a single node shrink.
And that's looking at MT, where your task scales perfectly. See where I'm going with this? Intel is using 2x the number of threads, and 3x the number of efficiency cores to get there. Apple can deliver that punch across a much lower number of threads - meaning ST-bottlenecked tasks will scale much much better on Apple.
With a single-threaded test, the M1 is pulling 7W vs 33W for the Alder Lake intel. Obviously that tells us nothing about efficiency, since we'd need to know the scores, but that's the downside, is for normal, poorly-threaded tasks, like surfing the web or editing code, the 12900HK is going to be boosting high to reach the same performance levels the M1 does at 3 GHz. And that's exactly what you see in the power figures there.
In short: you will likely see x86 able to keep up in one metric or another. You can win on performance if you just go nuclear on power. You can match on power on perfectly-threadable tasks that allow the x86 to deploy twice the threads (sharing instruction cache/etc). You can match on single-threaded battery life if you accept lesser performance. But the overall performance of the M1 derives from the massive IPC it generates, and that's something that x86 can't match nearly as easily.
Going ham on a single metric just to claim victory isn't nearly the same thing as the level of all-round performance and efficiency that Apple has achieved there.
(see also, putting a 128-thread Threadripper 3990WX workstation up against a 10-thread M1 Max laptop just to win at rendering... and people here thought that disproved that Apple was great hardware lol)
The reorder buffer size is just a logical consequence of the frontend width.
And yes, scaling an Aarch64 frontend is dead simple compared to x86 due to the fixed instruction width. The disadvantage of x86 is serious, but I don't know if we can count it out quite yet. This is the first time Intel and AMD got any serious pressure on that front. I'm sure they're taking the challenge seriously, and it'll take some years before we'll see the results.
Edit: even AMD themselves call their threadripper lineup workstation chips, not personal.
If the purchase page says to "contact sales" and doesn't list a price then it is not for consumers.
There's no other chip that has the power of an RTX 3090 and more power than an i9-12900K in it - after all, Threadripper doesn't have a lick of graphics power at all. This chip can do 18 8K video streams at once, which Threadripper would get demolished at.
I'm content with giving them the chip crown. Full system? Debatable.
I mean they all are CPUs coming out this year as far as I know.
The performance per watt isn’t in the same universe and that matters.
The M1 Ultra is a workstation part. It goes in machines that start at $4,000. The competition is Xeons, Epycs, and Threadrippers.
I couldn’t give less of a shit about performance-per-watt. The ONLY metric I care about is performance-per-dollar.
A Mac Studio and Threadripper are both boxes that sit on/under my desk. I don’t work from a laptop. I don’t care about energy usage. I even don’t really care about noise. My Threadripper is fine. I would not trade less power for less noise.
One hour of my time is more expensive than an entire month of a computer electricity bill.
Some people just want tasks to perform as fast as possible regardless of power consumption or portability.
Life's short and time is finite.
Every second adds up for repetitive tasks.
I personally stick to the lower wattage ones because I don't generally need high end stuff, so I think Apple is going the right direction here, but it should be noted that Intel has also started down the path of high performance and efficiency cores already. AMD will find itself there too if it turns out that for home use, we just don't need a ton of cores, but instead a small group of fast cores surrounded by a bunch of specialist cores.
Thermal density plays a huge role, the size of the chips is going down faster than the wattage, so thermal density is going up every generation even if you keep the same number of transistors. And everyone is still putting more transistors on their chips as they shrink.
Going forward this is only going to get more complicated - I am very interested to see how the 5800X3D does in terms of thermals with a cache die over the top of the CCD (compute die). But anyway that style of thing seem to be the future - NVIDIA is also rumored to be using a cache die over the top of their Ada/Lovelace architecture. And obviously 60W direct to the IHS is easier to cool than 60W that has to be pulled through a cache die in the middle.
Looking it up though I do see a lot of concerns with the heat they generate. I can only conclude I don't push my chip very hard (which, honestly, I probably don't)
I've been happy with the AMDs I purchased over the past 4 years, we'll see how they hold up and how this next gen comes out. I did see that the recent Intels are quite competitive which is good for everybody.
Yeah, longevity, blah blah, but laptop chips are designed to sit above 90C under load, it's fine.
Just saying that "how hard it is to cool" doesn't solely depend on power consumption anymore. Heat density is making that harder and harder, even if power consumption stays the same.
What does improve though is how much heat it pumps into your room. Yeah, a Rocket Lake at 200W might be roughly as hard to cool as an AMD at 90W or whatever... but one is still putting 200W into your room and the other is still putting 90W. Temperatures are not the same thing as power dissipation either. I don't like having my gaming PC running in my room during the summer, and I'm actually looking at maybe running cables through the walls to have it in the basement instead. I also have a 5700G and some NUCs that are much lower power that I prefer to use for surfing and shitposting.
Sure it does. Reading the rest of your post I think you're more talking about temperature than cooling requirements, but a 200W CPU needs 200W of heat dissipation, while a 60W CPU only needs 60W of heat dissipation. It's literally a 1:1 relationship since CPUs don't do any mechanical work, so power in == heat out.
Keeping temperatures below some arbitrary number does then include things like density, IHS design, etc... But that only matters for something like Intel's "Thermal Velocity Boost" where it's really important to stay under 70C specifically instead of just avoiding thermal throttling.
Power does make a big difference in data centers though - it's often the case that you run out of power before you run out of rack space.
Where power for a computer might make a difference could be in power-constrained (solar/off grid) scenarios.
I don't know if I've ever heard anyone make an argument based on $$$.
https://www.theguardian.com/society/2021/sep/09/transport-no...
As for desktops, watercooling makes computers dead silent.
I personally bought a Ryzen 5950(?) instead of a Threadripper because I figured I'd accidentally spill water all over it or however it works. There are not many watercooled OEM products as far as I know.
Until maybe these M1's (and I'm not entirely convinced) I've not in the 20 years I've been computing seen a reasonably configured desktop (eg not just a laptop on a stick ala iMac but an ACTUAL desktop) ever not smoke the pants off of every single laptop you could put up against it. It's hard to beat the one-two punch of lots of power and room to cool it. If you are sitting at at desk why the heck wouldn't you leverage that?
I still have a proper desk-based working environment hooked up to a docking station though. I really wouldn't want to use a laptop that doesn't have a first-party dock as my primary machine.
I agree that most developers are web/mobile developers who use a laptop. That’s great. I am an increasingly niche developer.
The root comment was a comparison against Threadripper. Normal developers should not waste money on a Threadripper. If someone is a niche developer that warrants a Threadripper then pointing out that most developers don’t need a Threadripper is a waste of time.
So, even if it doesn't quite beat Threadripper in the CPU department - it will absolutely annihilate Threadripper in anything graphics-related.
For this reason, I don't actually have a problem with Apple calling it the fastest. Yes, Threadripper might be marginally faster in real-world work that uses the CPU, but other tasks like video editing, graphics, it won't be anywhere near close.
We all need to take Apple claims with grain of salt as they are always cherrypicked so i wont be surprise if it wont be even 3070 performance in real usage.