Threadripper 7000 Storm Peak CPU Surfaces with 64 Zen 4 Cores
tomshardware.com
tomshardware.com
Between Hugging Face, Stable Diffusion, and Whisper, I'm using ML workloads a lot more. Being able to do so:
* with a standard instruction set
* with open-source software
* with my full system RAM
* without having to worry about what is in VRAM versus main RAM
is a big step up. I see about a 10x speed difference between an older 16-core CPU and a hot-off-the-press high-end Ampere card costing 3x as much as the CPU. If 64 core could bring that within 2x, or even 4x, I'd dump the GPU entirely.
My 5950x (measured) flops are ~2 TFLOPS in single-precision, ~1TFLOPS in double precision (obviously, due to half the SIMD vector size). This is a desktop-class 16-core machine.
https://old.reddit.com/r/Amd/comments/9uswbz/how_much_gflops...
https://github.com/Mysticial/Flops/
You can also get a theoretical computation of the Flops, which matches nicely with the experimental measurement. You have to take into account:
- the clock frequency (~3.9 GHz on multithreaded workloads on my machine)
- the number of cores (16)
- the reciprocal throughput of the FMA instruction (~.5, that is, 2 instructions per clock cycle)
- the number of flops per instruction (2 for the FMA instruction, that is, 1 multiply + 1 add)
- the SIMD vector width (4 for double, 8 for float).
Putting it together:
3.9e9 * 16 * 2 * 2 * 4 = 998.4 GFlops (double)
3.9e9 * 16 * 2 * 2 * 8 = 1996.8 GFlops (single)
The measured values on my machine are a bit different, but close (1070 and 2151 respectively).
References:
https://www.agner.org/optimize/instruction_tables.pdf
https://www.agner.org/forum/viewtopic.php?t=56
https://gadgetversus.com/processor/amd-ryzen-9-5950x-gflops-...
The vector flops for the 3090Ti are 33 TFlops for single precision, 0.5 TFlops for double precision. So, 16x faster than the 5950x in single precision, 2x slower for double precision. At almost 3x the price and >4x the power consumption.
Of course, if all you care about is AI, then there's no argument - but then we are not really talking about a general-purpose device any more.
The narrative of GPUs being "hundreds of time" faster than CPUs is vastly blown out of proportion for general-purpose computing.
It is you who barged into the thread with unrelated GPU performance numbers, but whatever :)
Here's the comment I assume you are allegedly trying to "correct":
> with full training you are out of luck with CPUs, the gap is much bigger. 64c TR could only get to roughly 1TFlops
1TFlops is not the main part of that statement, and it is qualified with "roughly" which I suppose is not too far from the truth in the context. And the context is "training ... the gap is much bigger", and in this case "much" is at least 30x even with the updated number.
People might do it just for fun, or maybe to manipulate the share price (make performance better or worse than expected), or maybe even for marketing.
Genuinely asking as I plan to replace my Ryzen 3700X with a 7700X.
Essentially 1024 cores soc server, affordable. Compared to that 64 cores sounds rather unimpressive. IMHO.
Also why even mention that 1024 core CPU instead of upmem. 128 cores per DIMM slot and up to 2560 cores in a single machine and they are fast precisely because they are directly attached to memory with a total memory bandwidth of 2.56TB/s.
It is true that a chip like this probably could render pretty decent 3d in software though. I wonder if combining this with the GPU in a clever way could allow more people to experience real time raytracing?
The whole history of PCs is repeatedly proving otherwise. The NES had hardware sprites. Then Carmack & Romero showed up and proved you can have smooth side scrolling in software, on an underpowered CPU. The whole concept of a PPU was thus rendered obsolete. Repeat for discrete FPUs, discrete sound cards, RAID cards (ZFS), and so on.
Specialised silicon will beat general purpose silicon at the given task, until general purpose silicon + software catches up. You need to keep pouring in proportional R&D effort for the specialised silicon to stay ahead.
What keeps GPUs relevant is that they're in fact much more general than what the "G" originally stood for.
the fun fact, is that if you manually reduce the power limit to 65W the initial single thread results so virtually no loss in ST performance vs 170W, and it appears that the original AMD slides stating 75% more efficient cores at that level not too far off.
So with the power limit set to 65 W or more the single-thread performance was always limited by the maximum turbo frequency (which may depend on the temperature of the CPU) and never by the power limit.
I have not seen yet any published value about the single core power consumption of Zen 4, but it is likely that the single core power is not higher. It is certainly much less than 65 W even at 5.85 GHz.
So the expected behavior is that the single-thread performance does not depend on whether you set in BIOS the steady-state power limit to 170 W, 105 W or 65 W. Only the multi-threaded performance is modified by the power limit, because when the power limit is reached, the clock frequency is decreased until the power consumption matches the limit.
edit: seethe