First 96-Core AMD Zen 4 Threadripper Tests Show Utter Domination over Intel
extremetech.com
extremetech.com
I'd guess this wins in absolute performance (due to the TDP), but I've heard AMD is pretty competitive on energy efficiency too these days. I'm typing this on an M2, which is simultaneously cooler and faster than my old Intel laptop.
Wondering if it's time to upgrade my decade-old Intel Linux desktop to an AMD yet.
Of note is that Zen 2 (the 3970x's microarchitecture) is 7nm. The new Threadrippers are on a 5nm process. I would expect slightly lower power consumption, but note that the 7000 series of Threadrippers quote a TDP of 350W while the 3000 series quote 280W. So I'm guessing they're using the efficiency for more speed instead of more power savings.
To me, it's fine. I'd always prefer paying more for electricity than to wait longer for something to complete. (But I do respect the desktop chips that can blink my cursor with an efficiency core that uses a few tens of milliwatts.)
These monsters will do better than the 39xx because of the 6mm IO die. But it's still going to have high idle power compared to a monolithic Intel design.
The monolithic AMD dies (laptop/mobile) are competitive in the idle power space.
I don't know if the Zen 4 Threadrippers can be similarly run with a power limit. If they can, I'd expect much better efficiency than whatever Intel offers.
This has eco mode benchmarks: https://youtu.be/W6aKQ-eBFk0?si=o8l1qGUhSQKxmY_c
https://www.notebookcheck.net/R7-7840U-vs-M2_15023_14521.247...
The last benchmark is interesting but unfair given that's an ancient application with Rosetta.
No need to guess: An M2 Ultra will get around 27,000 in Cinebench R23 while this Threadripper will get around 100,000.
The M2 Ultra has a TDP around 90 watts (~ish according to a couple reviews) and this Threadripper has a TDP of 350 watts.
The stock configurations for a lot of desktop chips is to get the maximum performance at the expense of power, especially for things like these Threadripper series which are explicitly aimed at maximum performance. Often even drastically reducing power usage only gives modest performance penalties.
I also run an RTX 3900 GPU at 60% power for 80-90% performance. It's a noticeable decrease in requisite cooling.
Gamers Nexus did a video earlier this year comparing the 7950x at different eco modes and they found that putting the 7950x at a 105 watt eco made drop the performance around 5% but shaved almost 100 watts off of power. heres the link to the video https://www.youtube.com/watch?v=W6aKQ-eBFk0
I know why these companies do it and im lucky electricity isnt too expensive in my area but i rather shave 100 + watts of heat being dumped into my room (especially in the summer months)
Tl;dr, if you want the absolute performance on Desktop, get the Threadripper Pro.
AMD Zen 4 7800X3D, 5nm 5.04Ghz, Single Core GB6 Score @ 2833
Apple A17 Pro, 3nm 3.80Ghz, Single Core GB6 Score @ 2914
Well you can find that with Geekbench but it is very likely worst to publish it as it leads to a wrong conclusion without the basic understanding of Hardware or CPU.
The node and design whether it was tuned for Clockspeed ( Desktop ) or SoC ( Apple ). There are measurable percentage difference in the same core for Risen on desktop and Ryzen SoC. And then you have to factor in Wattage difference in Clockspeed they are operating at. You are expected to know all these basics before the scatterplot would even make sense.
The Desktop 7800X3D will likely use 20W+ in the test, while the Smartphone A17 uses at best 5W.
You will then also have to factor it a lot of these performance test are not sustainable without adequate cooling. i.e You can sustain the above performance for a very long time on 7800X3D, dont expect that on A17 inside a phone.
Many of these are pointing out the obvious. But it seems on HN we are increasingly required to do so otherwise it leads to some sort of flamewar.
For my m1, opening a vm or just plugging in a wireless mouse dongle is all it takes to shave several hours off the lifespan, like from 20 hours to 10.
For the amd laptop, anything waking up the nvidia gpu will do the same, like 10 hours to 3.
When you're talking such low amounts of power, the networking card, sound card, display efficiency, ssd efficiency, and whether the system chooses to run the fan or let the cpu stay warm starts to matter a lot since just like the cpu these are things that are running at the same time. So I guess a fair test would have to use the same ssd and have the monitor off, using an external monitor. Then you're just stuck with the efficiency of the parts of the motherboard you can't choose.
This feels like cheating to me.
I've read it has 12 channels of DDR5.
With smaller sticks, they could populate all the channels while keeping total memory capacity the same as other setups. I don't see why they had to splurge all the way up to 512GB.
Quoting the article:
"To celebrate the much-hyped launch, our sister site PCMag put the flagship CPU through its paces (remotely) in several popular benchmarks."
The people who did the benchmarks didn't splurge... they had no physical access to the remote machine at all. Certainly, it would have been ideal to talk to their remote contact and see if there were any 16GB sticks sitting around to test with instead, but I agree with the other comments that the memory capacity is unlikely to make a significant difference in these benchmarks.
There will certainly be plenty of other benchmarks soon enough, once people get hands on with these processors.
It is not possible to give all the systems the same memory configuration. They don't support the same memory and don't have the same number of channels. And I expect giving all the systems 512 GiB of memory would only widen the lead because larger DIMMs are usually slower.
>Do they expect all of their readers to then audit the code of the benchmarking that they are doing to ensure that it cannot possibly be bound by this?
Cinebench is hardly obscure. The average ExtremeTech reader will already know what it is and what it measures.
you want to max out memory channels on all devices to test max performance
but you also want to use RAM DIMs which are equally good, if possible the same
as AMD has more memory channels this means it will get more "equally good"(1) RAM DIMs and in turn have more memory
and as long as you make sure that no of your benchmarks perform pressure wrt. RAM capacity (e.g. max RAM usage 64GiB) you now have a fair benchmark wrt. max performance
naturally if you don't plan to max out your RAM channels on AMD you do need a different benchmark for non-max performance
(1): "equally good" is also a bit tricky as different vendors have slightly subtle differences in access patterns and DDR5 ram training also can make the same RAM run with different settings on different systems which can make the same ram dim not equally good, but there probably is some RAM which doesn't have such issues
No, the actual amount of RAM isn't directly relevant, but it seems like a possible brown m&m to me.
That's exactly where the differences comes from you want to max out the available memory channels on each device not have the same number. Same for latency etc.
I mean your also not fixing the clock speed to 2GHz no boost or when comparing cars limit things just to the first 3 gears or similar.
Because if you would add such artificial limits you won't compare fairly how fast the _max performance_ can be, it would be a very different comparison and wrt. evaluating max performance very unfair.
Now you still need to e.g. have a fair choice of RAM, so preferably if possible use the same RAM sticks (and also use RAM which works equally well with both Intel and AMD). But what that means is that you _have to_ give the AMD system more RAM because it has more channels. Now you also _have to_ make sure that you benchmarks don't have any RAM pressure wrt. capacity.
Also I'm not sure why you mention cache misses? Because number of cache misses is quite independent of RAM amount? Do you mean disk cache or swap misses instead, but I don't think the benchmark does disk I/O and shouldn't have swap either? But even if the benchmark only uses let's say 32GiB of memory for getting a proper GPU
No. Adding that memory bottleneck would make it an unfair and worthless comparison. If you are testing the maximum CPU performance then you need to eliminate those bottlenecks, not create them.
But the memory is required there to populate the memory channel which push the CPU performance to max.
Something like StackExchange [1] could fit all their 9 server into 1 ( or 2 with one for redundancy ) without any degraded performance.
Nobody believes me, which says a lot about both the hardware and software industries.
I would be happy if Ruby Rails could do it in 2x the rendering time i.e slower with 2x Resources / Server. That is combined 4x difference. Unfortunately even with JIT we are not there yet.
Wouldn't a cheap GPU be better for transcoding tasks?
All i'm seeing is the confirmation that independently parallalizable tasks finish twice as fast on a cpu with twice the cores. cool, amdahl's law is proven.
The handbrake one, all the times seems similar indicating similar single core perf.
This matters for me since single core perf is pretty important for visual studio/other software with alot going on feeling "snappy" even tho the actual compiling/rendering might be faster and scale with moree cores.
https://www.cpubenchmark.net/compare/5726vs5234vs3862/AMD-Ry...
When a pretty big task (compiling, rendering) activates the scheduler and causes the cpu to boost to their respective 5 ghz or w/e, then yeah the numbers would be the same, but they don't be at the beginning when i'm dragging windows around and right clicking on stuff in visual studio. I don't believe this feeling is reflected in the scores.
It takes tens of milliseconds to reach the boost clock: https://stackoverflow.com/a/64254459
The base clock is entirely irrelevant to everything you are talking about. Please reference my earlier comment for the meaning of "base clock": https://news.ycombinator.com/item?id=37980293
And yes it can waste energy you can measure with the your battery monitor (or UPS/PDU) but its also possible to measure the interrupt handling latency/etc, and for some cores it can make a very noticeable difference to desktop latency if you happen to be sensitive to such things.
While I've posted elsewhere about clamping the clock rate to save power, I also do the converse when plugged in by cutting off the bottom 2/3rds of the frequency range. This results in a far more responsive desktop in linux and even windows to a lesser extent.
Ok, so, as I alluded to earlier... make sure to have better-than-minimum-spec cooling to let the processor indefinitely run all cores at the boost clock (or very close to it).
It's been a long time since I've seen a desktop/workstation processor limit itself to the base clock for any reason.
Also, I agree with dralley. Clock speeds are unlikely to be the issue there anyways.
That's why i decided to go with a 5950x + 3090 rtx + gamer ram for my current "workstation" instead of a 5995wx + rtx 6000 + ecc ram for my current machine. even tho its not workstation rated, it seemed like the better way to go for quality of life feel.
When the next generation of cpu/gpu/ram drops in the next couple eyars, i can re-evaluate the gamer vs professional offering to see if it matters.
Once there were a lot of programs in flight it started making a larger and larger difference. Now half my family and friends have TR Pro systems bought refurbished (they watched eBay) because they loved mine so much.
https://www.amd.com/en/products/cpu/amd-ryzen-threadripper-p... vs https://www.amd.com/en/products/cpu/amd-ryzen-9-7900x
Its usually the ones with more cores than the max ryzen cpus that have to make tradeoffs and seeing how it affects day to day use. (the 32, 64, and in this case the 96 core ones).
AFIK it's basically tweaked AMD EPYC™ 9654 with much higher boost and some other changes, so that's probably where the base clock comes from
some situations where the base clock might still matter:
- you perfectly max out CPU utilization (no I/O, hardly cache misses, etc.) and at the same time have exactly only the thermal headroom from the spec and run it for a "long" time
- you want to max optimize for watter/perf the best power efficiency is likely around the base clock (for this one because of my assumption of it being more or less a tweaked server processor, on a typical consumer desktop PC I would say likely quite a bit below the base clock)
I run simulations that are inherently completely single-threaded and I lean on Geekbench which tend to be a good predictor for me. As far as I can tell, Intel still holds the crown in absolute ST perf, granted at horrific efficiencies: https://browser.geekbench.com/processor-benchmarks/ (be sure to ignore the "Top Single-Core Results" which AFAICT is full of BS numbers). Intel's latest 14900K isn't there yet, but should be a smidgen faster: https://www.techspot.com/news/100034-intel-core-i9-14900k-to...
Yes, but if you can squeeze more cores in the same space (and/or power) envelope, and not suffer a performance per core penalty, you're winning.
That never gets old.
> not a GPU
Well, 3970x is able to soft-render DotA in 4k with 8-11 fps...
Thus, the real question is how buggy these new CPUs will be. https://news.ycombinator.com/item?id=37812556
They're hopefully not using a CPU for that - that's all been on the GPU for years.
If you're doing something that follows the same general design as most existing video codecs, absolutely; these days "hardware" encoding is less "fixed function chip" and more "the driver/firmware knows how to stitch together these components so as to comply with the xyz standard" - hence you see support for new encoding standards getting added in driver updates. If you're doing some completely radical out-of-left-field video compression algorithm, maybe not, although even then GPU hardware is pretty general these days and video encoding tends to be well suited to running on it.
For simulation applications there is no such concept as "acceptable", the utility of a CPU scales linearly with how fast it can do the calculations.
CPUbenchmark is giving the new chip about a 100% improvement over the 3990x which is fairly tempting even at the 6-10k price tag I've seen floating around.
Our antenna design folks sends simulation runs to a group of R7525s with 2x EPYC 75F3 CPUs and 2x A100 80GB GPUs.
A small problem might have tens of millions of elements and billions of unknowns to solve and run on one server. Larger problems might be spread across all of the servers, after coordination with everyone else.
Simulations take 8-12 hours to run.
This is not acceptable.
Anything slower than "an amount of time imperceptible to even the most responsive human" is unacceptable.
Everything is too slow.
Networks are too slow, memory is too slow, storage is too slow, GPUs are too slow, and CPUs are definitely too slow.
It's basically a AMD EPYC™ 9654 but with more clock boost speed and some differences wrt. IO (e.g. less but faster memory channels) and I think some differences wrt. enterprise features.
I find this whole thing super fascinating, coming from the perspective of Silicon Valley startups.
[0]: The Problem with Linus Tech Tips: Accuracy, Ethics, & Responsibility
The thing that I find so interesting is that they - for better or worse - essentially live stream their business decisions, challenges, and aspirations. Whether their endeavors are successful or not, I agree with their decision to try and move beyond Youtube and create their own video platform via Floatplane. They've also found success in alternative revenue streams via sponsorships from reputable brands and merchandising and are very open about their metrics.
Lastly, I really respect that they are 100% bootstrapped. I've only ever worked for companies that rely on investor funding for growth, and their influence often feels intrusive and counterproductive.