AMD Introduces New AMD Ryzen Threadripper 7000/Pro 7000 WX
ir.amd.com
ir.amd.com
what I meant to convey in the first place: it doesn't matter how much hardware you throw at an issue if the software can't use it
Speed is actually not that bad either.
How fast is your setup?
Edit: Seems some people are getting 1-2.6 tokens/sec on Ryzen (no GPU acceleration), Llama 70B quantized https://www.reddit.com/r/LocalLLaMA/comments/15rqkuw/llama_2...
Whereas Mac Studio gets 13 tokens/sec https://blog.gopenai.com/how-to-deploy-llama-2-as-api-on-mac...
- you don’t get GPU acceleration just by using unified memory. Llama.cpp still only uses the CPU on Apple Silicon chips.
- the difference in tokens/sec is likely attributable to memory bandwidth. Mac Studios with the base Max chip have 400 GB/s memory bandwidth compared to around 50 GB/s for the Ryzen 5000 series CPUs
mount none /ramdisk -t tmpfsIn Linux it’s part of the kernel, and you can mount with type tempts.
Not sure about OSX or bsds
I heard rumors though that the mainline Ryzen will get 32 cores perhaps next year, so we'll have to see how that does.
Regardless, it's the memory channels and extra PCIE lanes that makes Threadripper shine, oh, and now HEDT gets Registered ECC which is fantastic.
The rumor is that the CCX is going from 8 to 16 cores, at which point a 24 core ryzen would necessarily be 2 partially-enabled CCXs. That wouldn't make much sense as their high end configuration which so far has always been 2 fully enabled compute chiplets. So if that 16 core CCX change is true then 32 full fat cores for the top end Ryzen is the most likely scenario. There wouldn't really be any reason for AMD to just refuse to put 2 fully enabled dies together, especially since as you noted Threadripper & Epyc already offer more benefits that Ryzen just won't get.
Yes and it is Zen 4 Core, so a lot faster than the 3000 days. My only concern is the Quad Memory Channel. Not sure if the 64 Core will be limited by it.
But this is exciting for Desktop Computing.
Did they really do this again? They already have a 12-channel and 6-channel DDR5 socket for EPYC, now they're going to have two completely different sockets just for Threadripper with 4 and 8?
This isn't even market segmentation. Threadripper costs just as much as EPYC. It's pointless.
"AMD also disabled CXL connectivity, as it doesn’t feel there is a need for that capability with workstation/HEDT platforms."
https://www.tomshardware.com/news/amd-announces-threadripper...
(10.17.2023) https://www.eetimes.com/cxl-gets-off-the-drawing-board/
"However, CXL has suffered from a “chicken and egg” problem, according to Srinivas. Last year, there was a great deal of interest from hyperscalers and integrators, but there were no CXL devices ready.. Now that products are sampling, he said, they can work on applications and see benefits. “Now we are seeing practical implementation and how CXL technology can really help provide that additional memory capacity and bandwidth.” ... "CXL 2.0 platforms likely for 2024"
2.)
( 23 hours ago ) "Even a laptop can run RAM externally thanks to a little-known tech called CXL — but it’s not for sale to mere mortals who only want 1TB of RAM"
https://www.techradar.com/pro/even-a-laptop-can-run-ram-exte...
2. Fake news. No laptops support CXL.
FADU CXL 2.0 Switch at FMS 2023
https://www.servethehome.com/fadu-cxl-2-0-switch-and-pcie-ge...
When these 'base clocks' are listed, would a lower core CPU with a higher base clock run more calculations at its max thread count than a higher core cpu? Say 64 threads on a 32 vore 4Ghz CPU vs a 2.5Ghz 96 core monster. For my use case I'm often running a serial computation on a fixed dataset that can be threaded per data point. So if I have 32 data points it's only that parallel. Something between embarrassingly parallel and completely serial.
And second. For these massive caches, if I'm doing a massive serial calculation does an individual core get to use the full cache if it's being thrashed? As in, are there cache benefits to running a computation on these huge CPUs vs some overclocked 8 core beast?
-------
When these 'base clocks' are listed, would a lower core CPU with a higher base clock run more calculations at its max thread count than a higher core cpu? Say 64 threads on a 32 vore 4Ghz CPU vs a 2.5Ghz 96 core monster
-----
You know that GPUs are just CPUs with a massive core count, right?
Like in the 1000s of cores
So depends on what calculations you need done, there is no answer.
It depends on how the program uses the cores, not even only what calculations you need done
Your cache question doesn't really have a simple answer either. E.g. an AMD CPU is split into different CCXs. To simplify somewhat, each core is broken up into several smaller compute units, with their own caches and memory controller. Intel has a completely different ring-based approach that's harder to summarise in once sentence.
Overall though, for the sort of work you're describing the limiting factor is often memory bandwidth, not raw compute. Different platforms have very different membw/core figures, and I suspect if you started measuring that then you'd find it easier to predict your codes performance.
This can be easily noticed when comparing the base clock frequencies, which are more or less proportional with the actual clock frequencies that will be reached in multi-threaded applications. For instance a 7950X has 4.5 GHz versus the 3.2 GHz of 14900K. Similar differences are between Epyc and Xeon and between Threadripper and Xeon W.
In desktop CPUs Intel can hide their very poor multi-threaded performance by allowing a much higher power consumption. However this method does not work for server and workstation CPUs, because these already have the highest TDP that is possible with the current cooling solutions, so in servers and workstations the bad Intel MT performance is much more visible. Intel hopes that this will change in 2024, when they will launch server and workstation CPUs made with the new Intel 3 CMOS process.
In the absence of actual benchmarks, a good proxy for the multi-threaded performance of a CPU is the product between the base clock frequency and the number of cores. For Intel hybrid CPUs, an E-core should be counted as 0.6 cores. For example a Threadripper 7960X should be expected to be (24 cores x 4.2 GHz) / (16 cores x 4.5 GHz) = 1.4 times faster than a 7950X in multi-threaded applications that are limited by the CPU cores (but twice faster in applications that are limited by the memory throughput).
I disagree on this point. I would say this problem is much more critical on Intel's desktop platform than their workstation platform. Xeon Sapphire Rapids is actually very easy to cool, even on air, thanks to the CPU having a much larger surface to dissipate heat than their desktop equivalent.
I have Xeon w9-3495X, and while power consumption is one of its weakest points, it stays under 60°C with water cooling while I pump 500W into it (25°C ambient), of which I see between +30% to +50% gain in multithreaded performance over the default power limit. (Golden Cove needs around ~10W per core, so the default 350W/56c = 6.25W is way below its performance curve.) Noctua has also shown that they're able to achieve ~700W on U12S DX-4677[1] on this platform.
Your metric about clock speed is, I'm afraid to say, so horribly oversimpified as to be flat out wrong. You can't just multiply core count by clock speed like that, as you're failing to take into account all sorts of other scaling factors such as memory bandwidth, cache size, avx support and so on, which matter as much or more than simple IPS.
On AMD chip's (I just dunno about the Intel architecture), each chip is divided up into chiplets with their own set of cores and L3 cache. For these Threadrippers, those will be 8 physical cores, and 32MB of L3 cache. Each core can access the L3 cache within their own chiplet only.
You'll need to dig into the enabled cores+cache arrangements for any particular chip you might be interested in to figure out what will be good for your workloads.
The base clock is only when all cores are loaded, and especially on Threadripper / Ryzen is far from meaningful in practice as they will permanently turbo. So it's very possible that the 96 core threadripper and the 32 core threadripper when asked to run the same 32 threads will actually end up running around the same clockspeed.
See for example this chart on the 5950X from anandtech: https://images.anandtech.com/doci/16214/CoreFreqScale-5950v3...
Note that there's not 2 frequencies, there's a whole range of them and that range varies by the actual workload demands.
The general rule of thumb is that within a generation, higher clock speed yields better performance per core. If you only have 32 data points, then you will probably get better performance with the 4 GHz 32-core CPU.
> if I'm doing a massive serial calculation does an individual core get to use the full cache
Your single core will get to use the entire L3 cache, but L2 and L1 caches are per-core and so your single core doing the work will not have access to those. So yes, there conceivably could be a benefit due to the larger L3 cache.
On a broader note, these kinds of factors (frequency, cache size, parallelism) tend to be extremely workload-specific and unpredictable, so the only real way to find out what's faster is to measure your specific workload.
The 12 and 16 core variants only had the ability to saturate the IO die with enough data to fully utilize 4 channels of RAM due to fewer CCDs being populated. I would bet that will be the case here as well. So those bottom CPUs will be just the budget IO monsters, and be no better than the non-Pro (for RAM bandwidth) even in the Pro boards. The one caveat is that they probably won’t have the speed penalty with 8 RDIMMs that the others probably will.
These chips are great. You get a lot of cores, without having to compromise on frequency (you get high frequency as well).
I wish cloud / dedicated server providers sold these chips. Because most workloads need higher frequency than more cores. And you get the best of both with these.
Threadripper has a significantly higher boost clock, but that only counts when you're not using all the threads, so you only care about this in a workstation running mixed workloads. In a data center you don't use a high core count processor for a single-thread performance-sensitive workload, you put it on something with fewer cores like a Ryzen 7700X because it has a higher boost clock than any Threadripper while costing less and using less power.
https://www.anandtech.com/show/21092/amd-unveils-ryzen-threa...
EPYC 9334 has the same config, but it's clocked lower and has a correspondingly lower TDP.
Whereas the 7965WX has 24 cores from 4 CCDs, so the 7965WX is presumably made of CCDs where 2 of the 8 cores were defective, i.e. it's a 7975WX made of "defective" CCDs. But so is the EPYC 9254. Or they could have used the same four "defective" CCDs to make a pair of the 12-core Ryzen 9 7900, or four of the Ryzen 5 7600.
The EPYC "F" processors are apparently made entirely of "defective" CCDs, presumably because it gives them more L3 cache per core and helps with heat dissipation by spreading the heat load across more CCDs for the same number of cores.
If I were to build a workstation on it right now, I wouldn't buy the most expensive one. My relatively modest compute needs don't come even close to what that monster is able to do.
If you are IO-bound, why spend $3000 more on a CPU if you can spend that in memory, more/faster NVMe or HDDs?
You'd be limited by PCIe lanes (and memory bandwidth, 8 channel DDR5) The CPUs feature 128 PCIe lanes, compared to 7950x's 28.
Although I work with 2-3 monitors all the time, I'm happy with low-end GPUs - I don't do much data visualization or 3D rendering. I mostly develop software, so I stand-up relatively thin ephemeral VMs and containers all the time. For me the core count and amount of memory is the limiting factor. To completely avoid any swapping, 64GB of RAM is needed, and more would allow me to keep a CI pipeline running all the time. If I need more space, I can spare one or two SSDs as bcache for a larger SATA array.
As much as I'd love the biggest Xeon X, EPYC, AmpereOne, MI-300, or Grace Hopper under my desk, I have no need for that much power.
I do more or less the same - I look for tower servers. The hardware is good, the maintenance is easy, plenty of space for storage and memory, and so on.
I’d imagine most people use the horsepower for only a fraction of the day
...Linus might consider an upgrade (he got a Threadripper some time ago) to compile the kernel.
With Threadripper, I can compile as well as play games, all on one machine.
[0]: https://www.anandtech.com/show/21092/amd-unveils-ryzen-threa...
And remember that with "hyper-threading" it appears at 192 cores - which is just insane.
If the platforms are anything like the desktop topology overclocking the memory increases the L3 cache and CCX<->CCX clock speeds as well. Feed the beast.
It's a bit weird (but cool!) that AMD seems to expect to use this chip on ordinary workstations rather than HPC clusters. Hopefully a few of those workstations are bought by the kernel and application developers (who have reasonably parallel workloads like compiling operating systems), otherwise you're likely to have more than an order of magnitude more cores than the average person designing your operating system. Usually, the problem is the opposite - the devs are on new flagship workstation-class devices, and the users on little edge devices (sometimes still with HDDs instead of SSDs!) have a poor experience, but this is kind of the opposite.
This will continue to be true, I think. The vast majority of people buying threadripper machines at threadripper prices are those with serious work to do on them, like compiling huge codebases, hosting piles of VMs, running simulations, or doing certain creative work. I doubt anyone developing 'serious' software that their customers buy 96 core machines to run wouldn't also run similarly powerful machines.
There are some tech fetishists who will buy this premium hardware and underutilize it, but not too many.
The 96 core part has 480MB of cache and 8 channels of DDR5.
Compare that to a 7950X (not the X3D model) which has 16 cores, but only 80MB of cache and 2 channels of DDR5. The 96-core part has 6X as many cores, but 4X as much memory bandwidth and 6X as much cache. The interfaces have scaled similarly to the core count.
Of course, if your workload isn't really very parallel then you won't see much benefit. Such is life with high core counts.
There are still some embarrassingly parallel applications out there that could benefit from this sort of hardware. I'm not sure if it's absolutely necessary to have this number of cores under each desk though. To me this sounds like a luxury/vanity/status symbol.
Also, I think the key point people really miss is that with modern cores you can just say what you want! I really really enjoyed this article where they scaled a AMD 7840HS laptop core from 51W down to 5W. There looks like perfectly acceptable single core performance even down at 10-15W. https://news.ycombinator.com/item?id=37923741
This mirrors what Matthew Dillon of DragonflyBSD was finding 6 years ago with a 2700x. You can just drop the cores target power down from 160W to 85W and it still runs amazingly fast and the efficiency numbers go way way up. https://lists.dragonflybsd.org/pipermail/users/2018-Septembe...
People keep being afraid of high power cores & they just don't know any better. The situation has just gotten better and better and easier and easier to dial in an arbitrary desired power level, and you get it. At some bit to multicore performance, but efficiency usually goes way way up.
The 350W tdp? Its in the article. There's a table of each pro sku and their TDP.
Just wanted to say that you had a good point about modern cores being configurable (for example, on my laptop I set the EPP 100% towards energy efficiency by default and then when I know I'm going to be running some serious long-running computations I'll run a script to make the EPP favor performance). I'll have to wait for these to be released to see how low their power usage can go considering they might have a high floor since I doubt these were designed for low power usage, but this has at least revived my consideration of these CPUs.
That being said, there is one other thing to consider: motherboards. I have a first generation threadripper in my desktop and through a series of dumb moves with very heavy graphics cards, I have broken a number of the PCIe slots. I looked into getting a replacement motherboard since I figured they'd be cheap by now since they're so out of date, but first generation threadripper motherboards are like $700-$1000, while the people who bought AM4 socketed CPUs got multiple generations of CPU upgrades AND their motherboards now cost $100-$200. There are some undeniable advantages to being on their more popular consumer lines instead of the prosumer threadripper.
https://www.pugetsystems.com/solutions/ - can you click through to workflows to specific programs and then look up Hardware Recommendations to get a run down.
While Blender itself is a mostly GPU bound situation, a lot of other media creation software is more mixed.
* Photoshop: mostly still limited to 8 cores (some filters being GPU enabled)
* Premiere: GPU used only to handle specific effects (so it depends on your mix of effects)
* AfterEffects: mostly CPU bound
The 96 core 7995WX: 3.6W/core
The 12 core 7945WX: 29.2W/core
Base frequency:
7995WX: 2.5 GHz
7945WX: 4.7 GHz
96-core: 2.5W/core + 60W I/O-die 12-core: 21.7W/core + 40W I/O-die
Lower TDP settings are still available. With PBO, you might be able to extract more juice out of stock power levels.
it is a half backed bergamo splitted with not all ram channels activated....
The motherboards on the other hand...
I'm sure there are other workflows, especially in the cloud where you are rending out fractions of a CPU. But for individual professionals, what else besides marginal core cost would you consider?
Still, that's good info., I must have missed that part where they are doing their CCD layout like that.
When there was a real insatiable demand for GPUs AMD actually did make money hand over fist
https://ir.amd.com/news-events/press-releases/detail/1146/am...