AMD Ryzen Threadripper 3000 32-core CPU is more bad news for Intel
zdnet.com
zdnet.com
A 32-core chip is still almost certain to show up, since benchmarks have leaked, and folks have also leaked some pics of 3950X's and packaging, but I guess supply/demand have kept everything from coming out yet.
As the Tom's post notes, server and client chips use the same chiplets, so it could be that most of the higher-binned ones are going to server parts; some are higher-margin (7742 is almost $100/core, vs. 3950X under $50/core) and plausible there're some enterprise orders (AWS etc.) getting priority. I wonder if they tightened the binning for the 3950X to respond to the fuss over turbo, too.
I had never really considered everything that goes into getting to that SKU/price list and launch date: you don't really know how much supply you'll have at various perf levels, or what demand there will be for what at what price, and if you don't exactly match them you might end up losing money by underpricing, by having to nerf good silicon to fill highly-demanded lower-end SKUs, or by having a shortage that makes customers go elsewhere. (Plus vendors don't just have the public SKU list, they're working out contracts with OEMs/other big customers. And who knows what the competition will do.) High stakes, and practically speaking no backsies; seeing how upset folks are about turbo clocks, imagine if AMD announced a price hike. Glad it's not my job.
The 7742 is the "halo" chip, though. Or was, rather, until the Epyc 7H12 was announced.
Other Epycs have much lower $/core. For example the 24-core EPYC 7352 is $1350, making it $56/core. The 16-core EPYC 7282 is actually even cheaper than the 3950X at $650, or $40/core.
No doubt bulk orders are going to get the priority, but the margins may not actually be that different depending on what companies are actually bulk-ordering. The $/core drops pretty quickly even just going down slightly in the stack. The 32c 7452 is $65/core.
And we don't entirely know which aspect of the binning is the limiting factor. If it's just functional cores that's going to be different from if they run at the right frequency/voltage. Epyc is all 225 W or lower. Threadripper 3000 AMD could pretty easily slap a 300W TDP on it and ram voltage through the chips that couldn't cut it at Epyc specs. The 2990WX is, after all, a 250W TDP part. The existing socket is already spec'd for more power capability than the top-end Epyc.
FWIW, here's the price-per-core chart for the whole second-gen server line:
https://www.servethehome.com/wp-content/uploads/2019/09/AMD-...
(Doesn't factor in voltage/freq needs for different SKUs, some of them needing fully-working chiplets, etc. Still.)
The low pricing on the 7282 and a few others is interesting, too. Wonder if it's a factor that it can use partially-working chiplets since the server I/O die can take up to eight chiplets and the client die only two, or if that's totally unrelated.
The 3900X has been pretty much sold out since launch, and sells out within hours of inventory showing on newegg/amazon. Been using a 3600 on a new build, waiting for 3950X. And likewise a lot of the aftermarket design rx5700xt cards have been selling out quickly. Took 4 tries to get the gigabyte one. First two times, sold out before I saw the notice, third it sold out while in my card going through checkout. Then I got the order in.
Intel still has the lead in low idle power which is good in laptops.
Ryzen lets different cores have different max frequencies so if your code is single threaded and your operating system isn't the newest that could be a reason to go Intel. Likewise if its single threaded and can take advantage of AVX-512 or the Intel Math Kernel Library.
But otherwise?
Is there any software that's commonly used that has a measurable performance boost with it? Or is it more specialised stuff?
Another disadvantage is that you have to recompile your code to use AVX512 but it seems general enough that compilers will use the instructions (to an extent) without specialized code [1].
[1] https://www.phoronix.com/scan.php?page=news_item&px=GCC-8-AV...
What is especially deceiving is profiling a function in a loop for more than say 50ms, when the normal function execution takes say 0.5ms. Long running functions get the most gain, while short running functions cause the most pain.
That is because downclocking AVX512 lasts 2ms (with a 0.5ms setup). Certain instruction mixes will cause a general slowdown (10% degradation measured by CloudFlare under actual usage) even though the test profiling might predicts a performance gain. Single AVX512 instructions when the CPU is running at full speed have a counterintuitive perverse performance penalty - apparently running 4x slower than when the CPU changes to the slower L1 or L2 clocks.
Sustained AVX512 usage has predictable performance.
“ Intel made more aggressive use of AVX-512 instructions in earlier versions of the icc compiler, but has since removed most use unless the user asks for it with a special command line option.” is a strong indication that you need to be very careful about where you use the instructions.
Running an encoder for 1 second - likely candidate. Occasional 1ms functions or single AVX512 instructions on a web server - likely penalty.
Its a great instruction set, Absolutely great, AVX512 supports gather/scatter, a whole slew of efficient processing instructions, etc. etc.
However, AVX512 has poor implementations right now. Skylake-X is one of the only implementations, and running it drops the clock-rate in ways that are difficult to predict. (One core running AVX512 drops the clock of other cores, slowing down the throughput of the entire server).
Traditionally, the first implementation of these instruction sets always a degree of "emulated". For example, the gather/scatter instructions aren't much faster than load/stores in practice.
So while the AVX512 instruction set could theoretically be efficiently implemented, it seem like Skylake-X's implementation leaves much to be desired. Hopefully future implementations will be better.
------
The other major implementation of AVX512 is Xeon Phi, which has been deemed end-of-life. I like the idea of Xeon Phi, but it just didn't seem to work out in practice.
Independently thermal throttling can occur which would affect all cores, although presumably the CPU is generating heat per numeric operation, so AVX512 is neutral versus other instructions per numeric operation.
intels method makes benchmarking simpler! but may leave performance on the table.
https://blog.cloudflare.com/on-the-dangers-of-intels-frequen...
Basically, if you are interleaving like you suggest, does the processor detect this and reduce the throttling by the "duty cycle" of 512-bit operations?
If not, could there be a way to tell the CPU to do this?
It's actually really, really slow. On newer (I think around Skylake-X, which is when AVX-512 was introduced) CPUs it takes up to 500 microseconds i.e. millions of cycles to activate AVX-512. This can't really be made faster because they actually need to give the voltage regulators time to adjust or the chip literally brows out. During this time AVX-512 instructions execute on the AVX-256 datapath¹.
Once AVX-512 is activated the clock of that core is reduced by about 25% and it starts a 2ms timer which is reset whenever another AVX-512 instruction is issued. AFAIK Intel doesn't say how long it takes to raise the frequency again once the timer expires.
(This is something of a simplification because there are actually two AVX power licenses, the first allowing AVX-256 and a limited set² of AVX-512 instructions, reducing clock by about 15%, and the second allowing everything. Also, executing a single AVX-512 instruction doesn't immediately request a higher power license, you have to execute a certain number of them.)
―
This is actually the better version. On Haswell executing any AVX-256 instruction would reduce the frequency of every core by about 15-20%. But hey, at least it only takes about 150k cycles to activate (not much of a consolation, I know). Beats me how long it stays throttled for.
(I don't know what exactly Broadwell did. I don't think it throttled all cores, but it didn't have the additional power license with reduced throttling that Skylake has.)
―
¹ Or the 128-bit datapath if the core is at the lowest power license (which still lets you use 128-bit SSE instructions, and basic AVX-256 instructions).
² Basically anything that doesn't execute on the floating-point unit, which means no floating point and no integer multiplication (which uses the FPU). This is actually kinda the saving grace of the whole thing, since it means you can vectorize things like memcmp and strlen without requiring the highest power license.
* The processor does not immediately downclock when encountering heavy AVX512 instructions: it will first execute these instructions with reduced performance (say 4x slower) and only when there are many of them will the processor change its frequency. Light 512-bit instructions will move the core to a slightly lower clock. * Downclocking is per core and for a short time after you have used particular instructions (e.g., ~2ms). * The downclocking of a core is based on: the current license level of that core, and also the total number of active cores on the same CPU socket (irrespective of the license level of the other cores).
Both are well optimized for AMD CPUs.
See my other comment on this topic.
By the way, it's interesting to note that Intel has a disclaimer on every MKL documentation page about this; my speculation: this was required by terms of a settlement.
From the above link:
>The Intel CPU dispatcher does not only check the vendor ID string and the instruction sets supported. It also checks for specific processor models. In fact, it will fail to recognize future Intel processors with a family number different from 6. When I mentioned this to the Intel engineers they replied:
> > You mentioned we will not support future Intel processors with non-'6' family designations without a compiler update. Yes, that is correct and intentional. Our compiler produces code which we have high confidence will continue to run in the future. This has the effect of not assuming anything about future Intel or AMD or other processors. You have noted we could be more aggressive. We believe that would not be wise for our customers, who want a level of security that their code (built with our compiler) will continue to run far into the future. Your suggested methods, while they may sound reasonable, are not conservative enough for our highly optimizing compiler. Our experience steers us to issue code conservatively, and update the compiler when we have had a chance to verify functionality with new Intel and new AMD processors. That means there is a lag sometime in our production release support for new processors.
> In other words, they claim that they are optimizing for specific processor models rather than for specific instruction sets. If true, this gives Intel an argument for not supporting AMD processors properly. But it also means that all software developers who use an Intel compiler have to recompile their code and distribute new versions to their customers every time a new Intel processor appears on the market. Now, this was three years ago. What happens if I try to run a program compiled with an old version of Intel's compiler on the newest Intel processors? You guessed it: It still runs the optimal code path. But the reason is more difficult to guess: Intel have manipulated the CPUID family numbers on new processors in such a way that they appear as known models to older Intel software. I have described the technical details elsewhere.
It goes very far back to MMX: https://yro.slashdot.org/comments.pl?sid=155593&cid=13042922
tldr: Intel's compiler doesn't optimize using standardized instructions on non-Intel hardware.
I can’t really blame them. Why support your competitor?
Intel still has superior performance counters and debugging features. Mozilla's rr (Record and Replay framework) only works on Intel for example, and Intel vTune is a very good tool. AVX512 is also an advantage, as you've noted.
There are other instruction set advantages: I think Intel has faster division / modulus operator, and also has single-clock pext / pdep (useful in chess programs).
For most people however, who might not be using those tools, I'd argue that AMD's offerings are superior.
Lets say production is 50% slower than what's tested in staging / developer test cases. Is it the data in production that causes this performance loss? Or is it hardware differences?
If you are using Intel tools to debug performance problems on developer / testing stages, you probably want to keep using those Intel tools in staging / production. There are enough cache differences and instruction-level differences (speed of "division" instruction. PEXT vs PDEP. Cache differenches, branch predictor differences, TLB differences) between the chips.
Intel has interesting optimizations: an Intel Ethernet card drops the data off in L3 cache (bypassing DDR4 RAM entirely). These little differences in the driver / motherboard / CPU can have a huge difference in performance, and complicate performance testing / performance debugging.
If you are deploying to AMD hardware for production, you probably want to be running AMD hardware in testing / developer stages as well. You want all your hardware performing as similarly as possible.
RDMA, the ability to share RAM as if it were local RAM (through a memory-mapped IO mechanism) across Ethernet is not a common setup. The fact that you can perform cache-timing attacks over RDMA + Intel L3 cache is a testament to how efficient the system is if anything.
Consider this interpretation: RDMA + DDIO is so fast, you can perform cache-timing attacks over Gigabit Ethernet(!!). NetCAT (the "vulnerability" you describe) is proof of it.
Cache-timing / side channel attacks aren't exactly the kind of vulnerabilities that most people think of though. Its kinda cool, but its nothing as crazy as Meltdown / Spectre were.
For server/pro market they might be worth considering but again, the huge BOM cut that you get by choosing AMD processor might be worth the performance penalty.
Intel is infamous for severely downclocking the processor for these and other AVX/SSE family instructions, to a point where sometimes using them makes the program slower than it would be otherwise, especially if you're constantly provoking frequency switches between them and regular instructions.
AMD might not have implemented AVX512 specifically yet (there's nothing legally keeping them from doing so however, they have patent sharing agreements with Intel regarding the entire x86/x64 ISA and extensions), but what they currently DO have is all common SIMD extensions implemented (up to SSE4 and AVX2 if I'm not mistaken) without incurring any frequency penalties on clock speeds for using them.
I can live without AVX512 for now, even though I'd be happier to have it. But I would really rather not have it if it came out in the same crap implementation that Intel has.
I think saying that Intel wins in AVX tasks is absolutely fair.
For example, I had a simulation that had to run on the CPU for reasons, but made use of AVX. Intel was consistently faster on any system I tested.
https://www.agner.org/optimize/instruction_tables.pdf
If you want to look at `vmov`s or arithmetic like `vadd` or `vmul`. Particularly glaring is that for moves between memory and (xmm vs ymm) Zen1 has a recirpical throughput of (1 vs 2), ie that on average it is able to complete an xmm-memory move once per cycle, and a 256-bit move one every two cycles. Skylake-X instead has 0.5 for xmm/ymm/zmm-memory. That is, it can move up to 512-bits between a register and memory twice per cycle. That is 8-times the throughput.
Arithmetic isn't as bad, but Zen1's reciprocal throughput goes from 0.5 to 1 on xmm to ymm, while Intel stays at 0.5 independent of vector size.
I haven't seen data on the 7nm Ryzen parts, but their marketing claimed it was supposed to have full width avx2, so I imagine things are different now, and that 7nm Ryzen will do just as well for avx workloads per core and clock as all the Intel parts without avx512.
EDIT: Some instructions on Intel get slower with wider vectors, like vdiv, vsqrtpd, vgather...
For example, the worst 7980XEs do 4.1/3.8/3.6 GHz for each of these respectively. 0.5 GHz down clock isn't too bad. You can change these settings in your bios however you'd like on the unlocked CPUs (ie, the HEDT lineup + W3175X). I do find those down clocks are necessary; I've passed 100C with a 360mm radiator and roaring fans. AVX2 loads don't get anywhere near that hot.
But for all that, I do see a substantial benefit on many workloads from avx512 -- at least 50% better performance than what I'd get from avx2.
I definitely think it's nice to have, especially if you enjoy vectorizing code and looking at assembly. With much bigger performance wins on the line, it's more rewarding and more fun -- and you have more tools to play with, like vector masking (unfortunately, gather/scatter have been disappointing). Fun or not though, if you offer me avx512 on one hand vs twice the cores with full avx2 for the same price on the other, I'd have a hard time rationalizing avx512.
Damn, which radiator? Was there a GPU under load in your water loop?
I was running benchmarks of Intel MKL's zgemm vs zgemm3m because of a Julia PR that recommended replacing the former with the latter. I don't think anything hits a CPU quite as hard as a good BLAS.
I think my thermal paste may be bad, because the CPU idles hot -- nearly 35 C. I ordered a Direct Die-X and MO-RA-420 radiator, so I'm planning on swapping the AIO for an open loop with way more radiator area and flow through the fins.
Running dgemm, that CPU would hit just a tad below 2 teraflops. I'd like to get it just over that (and run much cooler).
oh no, that program will be super fast.
It would make other programs that run concurrently with the AVX one slower.
Can you source that? I can't find it on the Wikipedia page[1] or its homepage[2].
[1]: https://en.wikipedia.org/w/index.php?title=Rr_(debugging)&ol... [2]: https://rr-project.org/
And a head-to-head test of a laptop available in AMD and Intel variants says it has better battery life, although the screen panel could be the reason: https://www.notebookcheck.net/Lenovo-ThinkPad-T495-Review-bu...
I'd like more info, and Intel still probably has a lot of firmware tweaks etc. that AMD has to implement to win microbenchmarks, but to a first order it's not clear Intel has a lead there anymore.
Besides, who spends $15,000 on a mid-high end server to run single threaded applications anyway?
Whether it is a navier-stokes grid/image fluid simulation, arbitrary points in space that work off of nearest neighbors or a combination of both (by rasterizing into a grid and using that to move the particles), there are many straightforward ways to use lots of CPUs.
Fork join parallelism is a start. Sorting particles into a kd-tree is done by recursively partitioning and the partitions can be distributed amount cores. The sorting structure can be read but not written by as many cores as you want, and thus their neighbors can be searched and found by all cores at once.
a) Single instance of application doesn't scale over multiple cores, and
b) Multiple instances of application scales well over multiple independent servers
Can you explain why they are unable to efficiently run multiple instances of the application on the same CPU (with multiple cores)?
The only thing I could think of would be running up against IO/Memory bandwidth limits.
Edit: I was really just responding to "who spends $15,000 on a mid-high end server to run single threaded applications anyway?". I would absolutely consider this a "single threaded application".
This is subjective, but I have a really good feeling about AMD, both on CPU and GPU side.
Only thing I could ask is for more open source (open source CPU firmware), but I think this is only a dream.
I have a 6 core quad channel intel and I've been waiting for this threadripper to come out.
I've been testing a Kubernetes cluster built on 32-core AMD servers and it's unreal the workloads you can throw at them. I'm used to 4 or 8-core Xeon chips and this is a whole different game.
And, mind you, my workstation is 4 year old Xeon with 64 gigs of memory and reasonably fast (but not amazing) SATA SSD.
On a side note, I work more and more from a couple of i3 and i5 laptops and only use the workstation to do heavy lifting tasks, such as replicating these more exotic setups.
https://www.gamersnexus.net/news-pc/3510-hw-news-threadrippe...
More PCIe channels means more NVMe as most servers don't use those channels for GPU.
why? Its not true. Program/data load times change by single digit % between SATA and NVME SSDs.
From 2015 to 2019 we saw a total IPC boost on the AMD side of over 80%, and an increase in max core count in the desktop line of 433%, from 6 to 32.
I'm okay with this progress :)
Just for comparison, here are the numbers from my GeekBench 4 run:
Single-Threaded: 4,746 Multi-Threaded: 34,586
According to a leaked benchmark, the Threadripper 3000's numbers are:
Single-Threaded: 5,519 Multi-Threaded: 68,279
The multithreaded benchmark is 2x, that's a no-brainer since it's likely to have 32 cores vs the 1950x's 16 cores. Now, I will say that TR3000 benchmark is not overclocked (3.6ghz.) But from what I've read, it seems like there's not much room for these latest chips to be overclocked. So, despite being the 3rd iteration, the TR3000 is only 16% faster in single-threaded benchmarks than my 1950X.
Intel's 9900K (overclocked) gets a single-threaded GB4 score of ~7,000. That (or its successor) may be my next machine.
Is the Threadripper 3000 test also from Geekbench V4? I noticed they added V5.
I don't see why it should be so far behind the Ryzen in single thread... unless the boost isn't working properly or something, could be disabled if it's a test chip.
But 15% the last generation is definitely impressive, especially when compared to the 3-5% gains from 14++++ that Intel has been offering.
Right now my poor x201 overheats to death if I use rustc carelessly.
If your development involves compiling code then the 2x larger L3 is also going to result in huge improvements to code compilation speed as seen in the Epyc Rome reviews. Example: https://www.phoronix.com/scan.php?page=article&item=amd-epyc... - compare the 7601 vs. the 7502. Both are 32c/64t parts. Both have close-enough base clocks & turbo clocks. But the new one (7502) absolutely smokes the old one (7601) at both kernel compile in GCC and LLVM compile. Compilers love that L3 it seems.
All signs point to Threadripper 3000 still being sTR4 compatible, so that means you could drop in a CPU that gives you +15% single core performance vs. previous gen's best case, no more NUMA domains bringing more consistent performance, lower power usage, and double the L3 for much faster compiles.
That's pretty awesome particularly given we're only talking a ~2 year gap between the products.
Edit: Found the Zen 2 Epyc marketing materials that describe this. Yes, apparently memory access on a single socket is uniform, in that all memory access is indirected over the IO chiplet[1]! This may hurt best-case access latency for NUMA-aware workloads? Just speculating.
It's not like non-local caches go away with uniform DRAM access latencies — unless L1/L2/L3 are also non-local to the core and indirected behind the IO chiplet. Which would be really surprising.
[1]: https://www.servethehome.com/amd-epyc-7002-series-rome-deliv...
By contrast Threadipper 1 & 2 had multiple memory controllers. It was 2x dual-channel controllers. As such they were full on real NUMA, just like a multi-processor system.
On Zen2 everything goes through the IO die. In a sense this means everything is "far" now, but it is uniform, and performance seems to be very good despite this (perhaps due to the insane amount of cache meaning less need to hit memory as frequently).
The following article calls it a "one-two punch"... wow!
https://www.digitaltrends.com/computing/amd-ryzen-threadripp...
I've read this elsewhere on HN, basically it boils down to having great tech doesn't mean everyone drops all their existing Intel tooling or Intel-optimized source code. If you care a lot about performance, then you'll care about those things. For everyone else, they just want a reasonably priced processor, and most PC manufacturers have large contracts to get those CPUs at a good enough price from Intel to also not want to switch all their motherboards and factory setups and driver in order.
I see it going to it's old peak soon. And it will just be as swift as it was before
> The single-core score of 1,275 is pretty much the same as for the current flagship Threadripper 2990WX, ...
> But when it comes to multi-core, the Threadripper 3000's score of 23,015 absolutely destroys the Threadripper 2990WX's score of 13,400, ...
Another part is the power consumption and resulting higher clocks. The 2990WX was basically power limited, the new chips are on 7nm which will drastically reduce their power consumption and allow higher clocks inside the same power envelope.
Anyway, it's still 32 cores. While eliminating the ultra weird NUMA effect and increasing clocks could explain a significant benefit, I'm still really struggling to intuit 72% higher performance. But it is Windows 10, so your remarks about the Windows scheduler may explain the gap.[2]
[1]: https://cpu.userbenchmark.com/Faq/What-is-single-core-intege...
AMD is advertising the absolute highest that any core on the die can hit under extremely light load for an absolute instant. Sustained single-core clocks will be 100-200 MHz less and sustained all-core will be significantly less. Most cores on a chip are not capable of sustaining the advertised clockrate even under severe voltage and even for instant loads, only the "preferred core" on a chip. There is a significant "binning effect" not just between chips, but between individual cores on a chip.
Intel and previous generations of AMD cores used to advertise the sustained single-core rate, which could be achieved on any of the cores. This is a changeup to how the clockrate has been advertised.
original: https://www.youtube.com/watch?v=DgSoZAdk_E8
followup after a patch: https://www.youtube.com/watch?v=3LesYlfhv3o
As such the clockrates may be significantly different from what you're intuiting based on the advertising.
> Another part is the power consumption and resulting higher clocks. The 2990WX was basically power limited, the new chips are on 7nm which will drastically reduce their power consumption and allow higher clocks inside the same power envelope.
Dwarf Fortress's simulation is all about pointer-indirection and jumping around memory. The CPU doesn't really do much except wait for RAM most of the time. It takes ~50ns to talk to RAM, but the CPU is clocked at 4GHz (0.25 nanoseconds), giving you an idea of scale. The RAM tightening can bring your latency anywhere from 50ns to 200ns depending on how well you tune your RAM parameters, and depending on chips and stuff. (Servers usually have lots of slow LRDIMM RAM over multiple-sockets that can be 200ns latency or worse).
AMD takes their design out of the server-playbook, and seems to have ~100ns main DDR4 RAM Latency. So Dwarf Fortress probably will be faster on Intel i9-9900k (monolithic design with integrated memory controller and 50ns main memory latency).
"Linked list" traverse the integers as follows:
//array is full of 1-billion numbers, randomly sorted
uint32_t idx = 0;
for(int i=0; i<200000000; i++){
idx = array[idx]; // Random traversal
}
Pull out a stopwatch (or use Linux's "time" functionality). Divide the time by 200000000. You've now measured memory latency.Here is a relevant blog post https://lemire.me/blog/2018/11/13/memory-level-parallelism-i...
Google Cache: http://webcache.googleusercontent.com/search?hl=de&ei=ewaGXa...
Results (also in the article): Geekbench v5.0.1 Tryout for Windows x86 (64-bit); 1275 single core, 23015 multi-core for 32C/64T @ 3.59GHz Base Frequency (and with 32GB DDR4, no info on 4 or 8 channel).
Doesn't matter the horsepower the speed limit is the same until we can make software better.
However, if you put this in a machine you use for playing games or checking Facebook then you might think that it’s a waste, and you’d be right. There is plenty of software out there which is still single-threaded, but it’s not dominating our utilization.
Such a shame that all that JS botnet is still single threaded, isn't it ?
Extra valves can increase efficiency, which is probably why the valve count is increasing, can't say I've noticed though.
Extra wheels can improve aerodynamics also.
https://en.m.wikipedia.org/wiki/Tyrrell_P34
I'll stop undermining your point now :)
Multi-core processors can invariably let us do more. I get it - valid point. But we're not doing more. We're just cruising along at the same speed I think.
For developers code compilations benefit from more cores. For hobby 3d artists their renderings benefit from more cores. For hobby video creators their editing & rendering benefits from more cores. For hobby IT usages the extra PCI-E lanes and ECC RAM gets you the capability of doing server-like virtualization without spending server money.
Threadripper is an HEDT platform, same as Skylake-X. This is not targeting mass market usage, so complaining about mass-market viability is rather irrelevant.
In what world is this a valid comparison? No one shopping for a $500 8 core consumer cpu is looking at a $1700 32 core hedt chip. They're entirely different markets. If you can use 32 cores you'd be looking at intels x series at a minimum, if not xeons. This is silly at best, if not intentionally misleading.
Threadripper-users dont care that much about energy efficiency - and its a small market, which gets cannibalized by the 3950X in the new generation.
It runs above 5 GHz.
What are you basing this claim on?
All we know Zen2 targeted beating an hypothetical 10nm Intel CPU which has not shown up.
As it has not shown up, it remains hypothetical. Wild what-ifs.